Image processing method, device, storage medium and program product

By acquiring and replacing the key pixel positions and values ​​in the memory, and combining them with the grayscale image of the current frame, the problems of long time delay, low accuracy and high cost in motion detection in the prior art are solved, and high-accuracy motion information acquisition and low-cost image processing are achieved.

CN116416546BActive Publication Date: 2025-12-12HISILICON (SHANGHAI) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111662733.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-12-12
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

Existing technologies for motion detection suffer from problems such as time delay, low accuracy, high cost, and high power consumption, especially in applications using traditional cameras and event cameras, making it difficult to simultaneously output high-precision images and event streams.

Method used

By obtaining the key pixel positions and values ​​of the previous frame image stored in memory, combining them with the grayscale image of the current frame image, the events of the current frame are calculated, and the data of the previous frame in memory is replaced, thereby reducing storage and processing costs and improving the accuracy of motion information.

Benefits of technology

It enables the acquisition of high-accuracy motion information of objects in images at a low cost, reduces data processing and storage costs, and ensures the synchronous output of events and images. It is applicable to sensors and processors and improves the practicality of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116416546B_ABST
    Figure CN116416546B_ABST
Patent Text Reader

Abstract

The application relates to an image processing method and device, a storage medium and a program product. The method comprises the following steps: acquiring key pixel positions and key pixel values of a previous frame image stored in a storage; obtaining key pixel positions of a current frame image according to a gray image corresponding to the current frame image and the key pixel positions of the previous frame image; obtaining events of the current frame according to the gray image corresponding to the current frame image, the key pixel positions and the key pixel values of the previous frame image; obtaining key pixel values of the current frame image according to the key pixel positions of the current frame image, and replacing the key pixel values and the key pixel positions of the previous frame image stored in the storage with the key pixel values and the key pixel positions of the current frame image. According to the image processing method, higher-accuracy motion information of an object in an image can be obtained at a lower cost and time delay.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, and in particular, to an image processing method and device, a storage medium, and a program product. BACKGROUND

[0002] With the continuous development of information technology and the gradual improvement of people's requirements for image processing technology, the motion detection technology for images is widely used. The motion detection technology is a kind of motion detection method with a wide range of applications, which obtains the motion information of objects in images by analyzing video frame images. In order to provide more reliable motion detection results to users, how to process the video frame images to extract the motion information of objects in the images so that the extracted motion information is more accurate and the cost required for image processing is smaller has become a key problem in the field of image processing. SUMMARY

[0003] Therefore, an image processing method, device, storage medium, and program product are provided. According to the image processing method, higher-accuracy motion information of objects in images can be obtained at a lower cost and time delay.

[0004] In a first aspect, an embodiment of the present application provides an image processing method, which includes: obtaining a key pixel position of a previous frame image and a key pixel value of the previous frame image stored in a storage, the key pixel position being a position of a pixel point representing a feature of an object in a frame image, and the key pixel value being a value of the pixel point at the key pixel position; obtaining a key pixel position of a current frame image according to a gray image corresponding to the current frame image and the key pixel position of the previous frame image; obtaining an event of the current frame according to the gray image corresponding to the current frame image, the key pixel position of the previous frame image, and the key pixel value of the previous frame image, the event of the current frame indicating motion information of an object in the current frame image compared with the previous frame image; obtaining a key pixel value of the current frame image according to the key pixel position of the current frame image, and replacing the key pixel value of the previous frame image and the key pixel position of the previous frame image stored in the storage with the key pixel value of the current frame image and the key pixel position of the current frame image.

[0005] According to the image processing method provided in the embodiments of the present application, the event of the current frame is calculated by combining the gray image of the current frame image with the key pixel position of the previous frame image and the key pixel value of the previous frame image. Since the key pixel positions of adjacent two frame images in the time domain are closer when the frame rate is higher, the key pixel position of the previous frame image and the key pixel value at the key position, and the key pixel position and the key pixel value of the current frame image are in a strong correlation relationship, so that the event of the current frame can be accurately obtained. The key pixel position of the current frame image is obtained by the gray image corresponding to the current frame image and the key pixel position of the previous frame image, the key pixel value of the current frame image is obtained according to the key pixel position of the current frame image, and the key pixel value of the current frame image and the key pixel position of the current frame image are used to replace the key pixel value of the previous frame image and the key pixel position of the previous frame image stored in the storage, so that only the key pixel position and the key pixel value of one frame image are stored in the storage, the amount of data stored in the storage is low, and the cost can be reduced. The event of the current frame indicates the motion information of the object in the current frame image compared with the previous frame image, so that the high-accuracy motion information of the object in the image can be obtained at a low cost.

[0006] According to the first aspect, in a first possible implementation manner of the image processing method, the event of the current frame is obtained according to the gray image corresponding to the current frame image, the key pixel position of the previous frame image, and the key pixel value of the previous frame image, including: obtaining the initial key pixel value of the current frame image according to the gray image corresponding to the current frame image and the key pixel position of the previous frame image; obtaining the event of the current frame according to the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image.

[0007] The pixel value is equal to the value of the pixel point, and the event of the current frame is obtained by the difference between the pixel values, so that the event of the current frame is associated with the change degree of the value of the pixel point, and the accuracy of the event of the current frame can be ensured.

[0008] In a second possible implementation of the image processing method according to the first aspect, the events of the current frame include common events and disappearing events, the events of the current frame are obtained according to the difference between the initial key pixel values of the current frame image and the key pixel values of the previous frame image, including: taking the pixel points common to the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image as common positions, and taking the common events of the pixel points at each of the common positions when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than a first threshold value or less than a second threshold value; taking the pixel points other than the common positions in the key pixel positions of the previous frame image as disappearing positions, and taking the disappearing events of the pixel points at each of the disappearing positions when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than the first threshold value or less than the second threshold value.

[0009] In this way, the common events and the disappearing events of the current frame can be obtained. Since the common events are obtained at the common positions which are the pixel points common to the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image, it can be indicated that the pixel points at the common positions appear motion between the capture time of the current frame image and the capture time of the previous frame image. Since the disappearing events are obtained at the disappearing positions which are the pixel points other than the common positions in the key pixel positions of the previous frame image, it can be indicated that the pixel points at the disappearing positions appear motion between the capture time of the current frame image and the capture time of the previous frame image. Thus, the motion information of the object in the current frame image can be represented by the common events and the disappearing events.

[0010] In a third possible implementation of the image processing method according to the first aspect, or any of the above possible implementation of the first aspect, the events of the current frame are obtained according to the gray image corresponding to the current frame image, the key pixel positions of the previous frame image, and the key pixel values of the previous frame image, including: performing feature extraction (such as edge detection) on the gray image corresponding to the current frame image to obtain the initial key pixel positions of the current frame image; and obtaining the events of the current frame according to the similarity between the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image.

[0011] The key pixel positions are the positions of the pixel points representing the features of the object, and the initial key pixel positions are obtained by feature extraction and also include the pixel points representing the features of the object. Therefore, the events of the current frame are obtained according to the similarity between the initial key pixel positions of the current frame and the key pixel positions of the previous frame, so that the events of the current frame are associated with the similarity between the features of the object in the current frame image and the previous frame image, and the accuracy of the events of the current frame can be ensured.

[0012] In a fourth possible implementation of the image processing method according to the third possible implementation of the first aspect, the events of the current frame include new events, and the events of the current frame are obtained according to a similarity between the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image, including: determining, for each region divided from the gray image corresponding to the current frame image, a ratio of a number of pixel points in an intersection of the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image to a number of pixel points in a union of the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image; and obtaining, when the ratio is less than a third threshold, a new event corresponding to a pixel point of the initial key pixel position of the current frame image in the region.

[0013] In this way, the new events of the current frame can be obtained. Since the new events are obtained according to the ratio of the intersection and the union of the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image in each region, the pixel point of the initial key pixel position of the current frame image in the region can represent that the pixel point appears motion between the capture time of the current frame image and the capture time of the previous frame image. Thus, the motion information of the object in the current frame image can be represented by the new events.

[0014] In a fifth possible implementation of the image processing method according to the first aspect or any possible implementation of the first aspect, the key pixel positions of the current frame image are obtained according to the gray image corresponding to the current frame image and the key pixel positions of the previous frame image, including: performing feature extraction (such as edge detection) on the gray image corresponding to the current frame image to obtain initial key pixel positions of the current frame image; and performing time domain smoothing processing on the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image to obtain the key pixel positions of the current frame image.

[0015] In this way, the initial key pixel positions of the previous frame image are obtained by feature extraction (such as edge detection), and then the initial key pixel positions and the key pixel positions of the previous frame image are used to perform time domain smoothing processing to obtain the key pixel positions of the current frame image, so that the case that a large error is caused when the initial key pixel positions of the current frame image are directly used as the key pixel positions of the current frame image due to the existence of noise in the gray image corresponding to the current frame image is avoided, and the obtained key pixel positions of the current frame image are more accurate.

[0016] In a sixth possible implementation of the image processing method according to the first aspect or any possible implementation of the first aspect, the method further includes: obtaining the gray image corresponding to the current frame image according to the current frame image, and a resolution of the gray image corresponding to the current frame image is less than or equal to a resolution of the current frame image.

[0017] In this way, the gray image corresponding to the current frame image can be obtained. The resolution of the gray image corresponding to the current frame image is equal to the resolution of the current frame image, so that the resolution of the event can be ensured without adjusting the resolution of the current frame image when the step of obtaining the gray image corresponding to the current frame image is performed. The resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, so that the data processing cost of the step of obtaining the event of the current frame and the key pixel position of the current frame image according to the gray image corresponding to the current frame image can be reduced.

[0018] According to a sixth possible implementation of the first aspect, in the seventh possible implementation of the image processing method, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the gray image corresponding to the current frame image is obtained according to the current frame image, including: performing one or more of filtering processing, down-sampling processing, interpolation processing, and non-linear transformation processing on the current frame image to obtain the gray image corresponding to the current frame image.

[0019] In the case that the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, when one or more of filtering processing, down-sampling processing, interpolation processing, and non-linear transformation processing are performed on the current frame image to obtain the gray image corresponding to the current frame image, the down-sampling processing, interpolation processing, and filtering processing make the resolution of the obtained gray image corresponding to the current frame image lower, and further reduce the data processing cost of the image processing method. The filtering processing reduces the interference of noise, so that the generated event is more accurate and has less noise. The non-linear transformation processing makes it easier to select the threshold (the first threshold and the second threshold) of the generated event, and the accuracy of the obtained event is higher.

[0020] According to the sixth or seventh possible implementation of the first aspect, in the eighth possible implementation of the image processing method, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the feature extraction is performed on the gray image corresponding to the current frame image to obtain the initial key pixel position of the current frame image, including: determining the position correspondence relationship between the pixel points in the gray image corresponding to the current frame image and the pixel points in the current frame image according to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image; and obtaining the initial key pixel position of the current frame image according to the position correspondence relationship and the position of the pixel points included in the extracted feature in the gray image corresponding to the current frame image.

[0021] In this way, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the initial key position of the current frame image can be obtained through the correspondence between the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image, so that the initial key position and the key pixel position of the previous frame image are in the same resolution, and it is possible to perform time domain smoothing processing on the initial key position of the current frame image and the key pixel position of the previous frame image.

[0022] According to the first aspect, or any one of the possible implementation ways of the first aspect, in a ninth possible implementation way of the image processing method, the key pixel position of the current frame image is stored in the memory in the form of a binary image or in the form of a lossless compressed data packet.

[0023] In this way, the data storage cost of the memory can be reduced.

[0024] According to the first aspect, or any one of the possible implementation ways of the first aspect, in a tenth possible implementation way of the image processing method, the method is applied to a sensor or a processor connected to the sensor, and the current frame image includes a RAW image in a raw format captured by the sensor or a three-channel RGB image in a color format generated by the processor according to the image captured by the sensor, wherein when the method is applied to the sensor and the current frame image includes the RAW image in the raw format captured by the sensor, the sensor outputs the event of the current frame while outputting the current frame image, or the sensor outputs the event of the current frame.

[0025] When the image processing method of the present application is applied to the sensor, the sensor outputs the event of the current frame while outputting the current frame image, which can ensure the synchronous output of the event of the current frame and the current frame image in time, and can ensure the accuracy when the event of the current frame and the current frame image are used to further perform other tasks. The sensor only outputs the event of the current frame, which can reduce the data transmission cost of the sensor. When the image processing method of the present application is applied to the processor, the data processing cost and power consumption of the image sensor can be reduced. In this way, the image processing method of the present application is more practical.

[0026] In a second aspect, embodiments of the present application provide an image processing apparatus, the apparatus comprising: a first obtaining module configured to obtain key pixel positions of a previous frame image and key pixel values of the previous frame image stored in a memory, the key pixel positions being positions of pixel points representing features of an object in the previous frame image, and the key pixel values being values of the pixel points at the key pixel positions; a second obtaining module configured to obtain key pixel positions of a current frame image according to a gray image corresponding to the current frame image and the key pixel positions of the previous frame image; a third obtaining module configured to obtain events of the current frame according to the gray image corresponding to the current frame image, the key pixel positions of the previous frame image, and the key pixel values of the previous frame image, the events of the current frame indicating motion information of the object in the current frame image compared with the previous frame image; and a first replacing module configured to obtain key pixel values of the current frame image according to the key pixel positions of the current frame image, and replace the key pixel values of the previous frame image and the key pixel positions of the previous frame image stored in the memory with the key pixel values of the current frame image and the key pixel positions of the current frame image.

[0027] According to a first possible implementation of the second aspect, in the second possible implementation of the image processing apparatus, the events of the current frame include common events and disappearance events, the events of the current frame are obtained according to differences between initial key pixel values of the current frame image and the key pixel values of the previous frame image, including: pixel points common to the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image as common positions, a common event of each pixel point at the common positions is obtained when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than a first threshold or less than a second threshold; and pixel points other than the common positions in the key pixel positions of the previous frame image as disappearance positions, a disappearance event of each pixel point at the disappearance positions is obtained when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than the first threshold or less than the second threshold.

[0028] According to the second aspect, or any possible implementation of the second aspect above, in a third possible implementation of the image processing apparatus, the events of the current frame are obtained according to the gray image corresponding to the current frame image, the key pixel positions of the previous frame image, and the key pixel values of the previous frame image, including: performing feature extraction on the gray image corresponding to the current frame image to obtain initial key pixel positions of the current frame image; and obtaining the events of the current frame according to a similarity degree between the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image.

[0029] According to a third possible implementation of the second aspect, in the fourth possible implementation of the image processing apparatus, the event of the current frame includes an added event, the event of the current frame is obtained according to a similarity degree between the initial key pixel position of the current frame image and the key pixel position of the previous frame image, including: determining, for each region divided by the gray image corresponding to the current frame image, a ratio of a number of pixel points in an intersection of the initial key pixel position of the current frame image and the key pixel position of the previous frame image to a number of pixel points in a union of the initial key pixel position of the current frame image and the key pixel position of the previous frame image; when the ratio is less than a third threshold value, obtaining an added event corresponding to the pixel point of the initial key pixel position of the current frame image in the region.

[0030] According to the second aspect, or any possible implementation of the second aspect above, in a fifth possible implementation of the image processing apparatus, the key pixel position of the current frame image is obtained according to the gray image corresponding to the current frame image and the key pixel position of the previous frame image, including: performing feature extraction on the gray image corresponding to the current frame image to obtain an initial key pixel position of the current frame image; and performing time domain smoothing processing on the initial key pixel position of the current frame image and the key pixel position of the previous frame image to obtain the key pixel position of the current frame image.

[0031] According to the second aspect, or any possible implementation of the second aspect above, in a sixth possible implementation of the image processing apparatus, the apparatus further includes a fourth obtaining module configured to obtain, according to the current frame image, a gray image corresponding to the current frame image, the resolution of the gray image corresponding to the current frame image being less than or equal to the resolution of the current frame image.

[0032] According to the sixth possible implementation of the second aspect, in a seventh possible implementation of the image processing apparatus, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the gray image corresponding to the current frame image is obtained according to the current frame image, including: performing one or more of filtering processing, downsampling processing, interpolation processing, and nonlinear transformation processing on the current frame image to obtain the gray image corresponding to the current frame image.

[0033] In an eighth possible implementation manner of the image processing apparatus according to the sixth or seventh possible implementation manner of the second aspect, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the feature extraction on the gray image corresponding to the current frame image to obtain the initial key pixel position of the current frame image comprises: determining a position correspondence between a pixel point in the gray image corresponding to the current frame image and a pixel point in the current frame image according to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image; and obtaining the initial key pixel position of the current frame image according to the position correspondence and the position of the pixel point included in the extracted feature in the gray image corresponding to the current frame image.

[0034] In a ninth possible implementation manner of the image processing apparatus according to the second aspect, or any one of the possible implementation manners of the second aspect, the key pixel position of the current frame image is stored in the memory in the form of a binary image or in the form of a lossless compressed data packet.

[0035] In a tenth possible implementation manner of the image processing apparatus according to the second aspect, or any one of the possible implementation manners of the second aspect, the apparatus is applied to a sensor or a processor connected to the sensor, and the current frame image comprises a RAW image in a raw format captured by the sensor or a three-channel RGB image generated by the processor according to an image captured by the sensor, wherein when the apparatus is applied to the sensor and the current frame image comprises the RAW image in the raw format captured by the sensor, the sensor outputs the event of the current frame while outputting the current frame image, or the sensor outputs the event of the current frame.

[0036] In a third aspect, the embodiments of the present application provide an image processing apparatus, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the image processing method of the first aspect or one or more of the possible implementation manners of the first aspect.

[0037] In a fourth aspect, the embodiments of the present application provide a non-volatile computer readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the image processing method of the first aspect or one or more of the possible implementation manners of the first aspect.

[0038] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer readable code, or a non-volatile computer readable storage medium carrying the computer readable code, when the computer readable code is run in an electronic device, a processor in the electronic device performs the image processing method in the first aspect or one or more of the possible implementation manners of the first aspect.

[0039] These and other aspects of the application will become more fully understood from the following (detailed description). BRIEF DESCRIPTION OF DRAWINGS

[0040] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate the exemplary embodiments, features and aspects of the application and, together with the description, further serve to explain the principles of the application.

[0041] Figure 1 An example of events generated by a dynamic vision sensor of the prior art is shown.

[0042] Figure 2 A method schematic diagram of a time domain frame difference operation of the prior art is shown.

[0043] Figure 3 A flow schematic diagram of a motion detection technique based on a binarized image of the prior art is shown.

[0044] Figure 4a An exemplary application scenario of the image processing method according to an embodiment of the present application is shown.

[0045] Figure 4b Another exemplary application scenario of the image processing method according to an embodiment of the present application is shown.

[0046] Figure 5 An exemplary work flow of the image processing method according to an embodiment of the present application is shown.

[0047] Figure 6 An example of initial key pixel positions of a current frame image and key pixel positions of a previous frame image according to an embodiment of the present application is shown.

[0048] Figure 7 An example of the image data processing method according to an embodiment of the present application in actual application is shown.

[0049] Figure 8a An exemplary schematic diagram of events obtained according to the prior art is shown.

[0050] Figure 8b An exemplary schematic diagram of events of a current frame obtained according to the image processing method according to an embodiment of the present application is shown.

[0051] Figure 9 An exemplary schematic diagram showing a pixel point corresponding to the information stored in the memory according to an embodiment of the present application is shown.

[0052] Figure 10 An exemplary structural schematic diagram of an image processing apparatus according to an embodiment of the present application is shown.

[0053] Figure 11 An exemplary structural schematic diagram of an image processing apparatus according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0054] Various exemplary embodiments, features, and aspects of the present application will be described herein below with reference to the accompanying drawings. The same reference numbers in different drawings indicate the same or similar elements. Although various aspects of the embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.

[0055] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any implementation described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other implementations.

[0056] In addition, for the purpose of convenience and brevity, detailed descriptions of well-known functions and structures incorporated in the application are omitted. It will be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described herein, embody the principles of the application and fall within the spirit and scope of the application.

[0057] The following describes the motion detection technology in the prior art.

[0058] Most of the existing motion detection technologies are based on image signal processors (ISPs) to capture and process color images obtained at a fixed frame rate, and use traditional methods or neural network methods to extract motion information. After using such methods to obtain a region of interest (ROI) representing motion information, the ROI can be used for tasks such as snapshot. However, this method has problems such as long time delay and low accuracy.

[0059] With the continuous development of advanced processes and long-term efforts in the industry, complementary metal-oxide-semiconductor (CMOS) image sensors have significantly improved the imaging quality of images, but such sensors usually only output high-quality image frames and cannot output events (i.e., motion information).

[0060] Based on this, the prior art proposes an event camera. The event camera is an image sensor inspired by the biological retina. Unlike a traditional camera (i.e., an image signal processor) that captures a complete image at a fixed frame rate, each pixel of the event camera detects the change in relative illumination alone and generates an event representing motion information only when there is a change in relative illumination and outputs an event stream, which is usually sparse and low-redundant. The event stream includes consecutive events, each event including the time of event occurrence, the pixel coordinates (x, y) corresponding to the event, and the event polarity. The time of event occurrence and the pixel coordinates corresponding to the event uniquely determine the position of the event in the space-time domain. The event polarity indicates whether the pixel brightness at the pixel coordinates (x, y) corresponding to the event increases or decreases by more than a certain threshold. When the brightness increases by more than a certain threshold, the event polarity is 1. When the brightness decreases by more than a certain threshold, the event polarity is -1. Compared with a traditional camera, the event camera has obvious advantages: high temporal resolution (the temporal resolution of the event camera is usually in the order of μs), high dynamic range (the dynamic range of the event camera can reach 140 dB, and the dynamic range of the traditional camera is usually 60 dB), low power consumption, and high pixel bandwidth (the pixel bandwidth of the event camera is usually in the order of kHz). Therefore, the event camera has great potential in the fields of intelligent robots, computer vision, and the like, especially in scenes that are very challenging for traditional cameras, such as high-speed motion scenes, long-endurance scenes, and the like.

[0061] The sensor based on differential vision sampling is the mainstream in the event camera, and the most commonly used are dynamic vision sensor (DVS), asynchronous time-based image sensor (ATIS), and dynamic and active pixel vision sensor (DAVIS). The above sensors usually adopt a logarithmic differential model, i.e., the photoelectric current and the voltage are in a logarithmic mapping relationship. When the change in the photoelectric current intensity of a certain pixel causes the voltage change to exceed a preset threshold, the pixel generates an event, for example, generates a pulse signal. Figure 1 An example of an event generated by a dynamic vision sensor of the prior art is shown. As shown in Figure 1 t represents the time of event occurrence, <x, y> represents the pixel coordinates corresponding to the event, DI(x, y) / dt represents the difference between the pixel brightness at time t and the pixel brightness at the previous time, and sign(DI(x, y) / dt) represents the event polarity.

[0062] Dynamic vision sensor usually only outputs event stream, and the output events are asynchronous output, so it is necessary to accumulate enough events to analyze and process the motion of the object in the image. Asynchronous time-based image sensor and dynamic active pixel vision sensor are derived from dynamic vision sensor, and can output event stream and image at the same time, wherein the output events are also asynchronous output. Taking dynamic active pixel vision sensor as an example, dynamic active pixel vision sensor is derived from the original dynamic vision sensor by additionally introducing an active pixel sensor (APS), and the two sensors share a photosensitive sensor. Dynamic vision sensor is used to monitor the change of light to generate events, and active pixel sensor is used to measure brightness to generate images, so as to realize the output of event stream and image at the same time.

[0063] Such an event camera that can output event stream and image at the same time usually takes event stream output as the main output, and the power consumption of the output event stream is relatively large. Limited by the requirement of low power consumption, the quality of the image may need to be sacrificed. Therefore, the image output by the event camera has the problems of low resolution, motion blur, low dynamic range and the like. For example, limited by power consumption, the maximum resolution of the active pixel sensor on the market is 1280*800, which is relatively low. The active pixel sensor usually cannot achieve high frame rate, so the image generated by the active pixel sensor still has problems such as motion blur, especially in high-speed motion scenes. The active pixel sensor cannot introduce the design to improve the dynamic range, resulting in generally low dynamic range of the image. The sampling speed of the active pixel sensor is far lower than that of the dynamic vision sensor, and the images and event streams generated by the two sensors cannot be precisely synchronized. This makes the adaptability of the event camera poor in the application scene that requires high-precision image and event stream at the same time. Moreover, the event camera needs to accumulate enough events before executing a task that needs to use events (such as motion wake-up snapshot, exposure time control, etc.), resulting in high latency.

[0064] The prior art also proposes a motion detection technology based on target tracking. This technology detects the target in the image by acquiring the motion information of the target (object) through real-time image processing of each frame of image collected by the sensor. Many sensors in the prior art can integrate the device that can realize the above image processing function in the sensor.

[0065] The principle of this technical solution is to find the target centroid to find the target to be tracked. Frame difference operation can effectively obtain the centroid. Figure 2 A method schematic diagram of the time domain frame difference operation of the prior art is shown.

[0066] As Figure 2As shown, in order to implement the frame difference operation in the time domain, a memory needs to be set to buffer the pixel values of a frame of image. After the current frame of image is pre-processed (for example, filtering, etc.) to remove noise, the pixel values of the current frame of image are obtained, and the pixel values of the previous frame of image stored in the memory are differentiated to obtain the motion information of the current frame. The pixel values of the current frame of image can be used to replace the pixel values of the previous frame of image in the memory, and the pixel values of the next frame of image are used when the motion information of the next frame of image is obtained.

[0067] The memory can be, for example, a static random access memory (SRAM). In some sensors, the memory is usually set inside the pixel. The disadvantage of this scheme is that the memory corresponding to each pixel needs to be set to buffer the pixel values of each pixel point. When the image resolution is large, the number of pixels is also large, and a large number of memories need to be set, and the cost and power consumption increase greatly due to storage.

[0068] The prior art also proposes a motion detection technology based on a binary image. Figure 3 A flowchart of the prior art motion detection technology based on a binary image is shown. The principle of this technology scheme is as follows: a memory is set to store a frame of binary image in a three-channel color RGB space and a frame of binary image in an HSV space including hue, saturation, and value information. The current frame of image in the RGB space collected by the sensor is converted to the HSV space to obtain the current frame of image in the HSV space, the current frame of image in the RGB space and the current frame of image in the HSV space are respectively binary processed to obtain the binary image of the current frame in the RGB space and the binary image of the current frame in the HSV space, and the binary image of the current frame in the RGB space and the binary image of the previous frame in the RGB space, and the binary image of the current frame in the HSV space and the binary image of the previous frame in the HSV space are used to perform a difference operation to obtain the region of interest (i.e., the motion information). Here, the acquisition of the motion information completely depends on the binary image, and the region of interest finally obtained is usually the centroid or other features of the region where the object is located in the image. The disadvantage of this scheme is that the motion information is of coarse granularity, and is not detailed and accurate enough. Moreover, the binaryzation is too absolute, and in scenes such as dark light and high dynamic range, there is a large error between the binary image after binaryzation and the binary image before binaryzation, and the region of the object cannot be completely and accurately extracted from the binary image, so the practicability of this scheme is very low.

[0069] Therefore, an image processing method, device, storage medium, and program product are provided. The image processing method according to the embodiments of the present application can obtain high-accuracy motion information of an object in an image at a low cost.

[0070] Figure 4a This illustrates an exemplary application scenario of the image processing method according to an embodiment of this application.

[0071] like Figure 4a As shown, this application scenario may include an optical module (e.g., a lens) and an image sensor connected to the optical module. Optionally, it may also include an image processor connected to the image sensor. The natural scene includes moving objects. The image processing method of this application embodiment can be executed by an image processing device integrated in the image sensor. The image sensor may be, for example, a complementary metal-oxide-semiconductor image sensor. The image sensor may be disposed, for example, on a terminal device. The terminal device may be a smartphone, netbook, tablet computer, laptop (see the example above for implementation), wearable electronic device (such as a smart bracelet, smartwatch, etc.), TV, virtual reality device, etc. The image processor may be disposed on the same terminal device as the image sensor, or it may be disposed separately on another terminal device; this application does not limit this.

[0072] Image sensors may also include, for example, existing photosensitive circuits, correlated double sampling (CDS) circuits, gain amplification circuits, and analog-to-digital converters (ADCs). When an image sensor acquires an image, the received photons are converted into electrical signals through the photosensitive circuit, correlated double sampling circuit, gain amplification circuit, and ADC circuit, thereby obtaining a frame of sensor RAW image (i.e., raw format RAW image). The sensor RAW image obtained by the image sensor can be an image in a Bayer image pattern arranged in RGGB, BGGR, etc., or it can be an image with any arrangement of 4*4 pixel window units; this application does not impose any limitations on this.

[0073] The image sensor can take the sensor RAW image as a current frame image, and the image processing apparatus executes the image processing method of the embodiments of the present application to first obtain a grayscale image of the current frame image according to the current frame image, and then obtain events of the current frame and key pixel positions and key pixel values of the current frame image according to the grayscale image of the current frame image and the key pixel positions and key pixel values of the previous frame image obtained from the memory (not shown). The memory can be integrated in the image sensor, connected with the image processing apparatus, or arranged outside the image sensor and connected with the image sensor. The key pixel positions can be positions of pixel points (such as edges) representing object features in the image, and the key pixel values can be values of the pixel points at the key pixel positions. After the image processing apparatus obtains the key pixel positions and key pixel values of the current frame image, the image processing apparatus can write the key pixel positions and key pixel values of the current frame image into the memory to overwrite the key pixel positions and key pixel values of the previous frame image stored in the memory. The key pixel positions and key pixel values of the current frame image can be used when the image processing apparatus processes the grayscale image of the next frame image. The image sensor can synchronously output the events of the current frame and the current frame image, for example, to the image processor.

[0074] Figure 4b Another exemplary application scenario of the image processing method according to the embodiments of the present application is shown.

[0075] In the application scenario, Figure 4b In the application scenario, the image processing apparatus executing the image processing method of the embodiments of the present application can be arranged on an image processor connected with the image sensor, and correspondingly, the memory can also be arranged on the image processor or connected with the image processor. The image sensor can transmit the collected sensor RAW image to the image processor. The image processor can directly take the sensor RAW image as a current frame image, or obtain a three-channel color RGB image from the sensor RAW image based on the prior art to take the RGB image as a current frame image. Then, the image processing method of the embodiments of the present application is executed according to the current frame image. The image processor can store or output the events of the current frame and the current frame image to other devices or apparatuses for subsequent execution of tasks such as target detection and recognition (such as face detection and face recognition).

[0076] Those skilled in the art should understand that the specific arrangement positions of the image processing apparatus and the memory are not limited by the present application as long as the image processing apparatus and the memory are arranged to enable the image processing apparatus to obtain the grayscale image of the current frame image and the information stored in the memory.

[0077] Figure 5An exemplary workflow of the image processing method according to an embodiment of the present application is shown.

[0078] As shown in Figure 5 The present application provides an image processing method, which comprises steps S1-S4:

[0079] In step S1, the key pixel position of a previous frame image and the key pixel value of the previous frame image stored in a memory are acquired, the key pixel position is the position of a pixel point representing a feature of an object in a frame image, and the key pixel value is the value of the pixel point at the key pixel position.

[0080] The feature of the object in the image can be an edge of the object in the image or the like which can distinguish the object from the image. The value of the pixel point can be a gray value included in a YUV image, or a gray value calculated from three-channel data included in an RGB image, or a value calculated from RGGB data included in a sensor RAW image. The present application does not limit the specific type of the value of the pixel point. The key pixel position of the previous frame image and the key pixel value of the previous frame image can be acquired and stored according to a gray image of the previous frame image.

[0081] In step S2, the key pixel position of the current frame image is obtained according to a gray image corresponding to the current frame image and the key pixel position of the previous frame image.

[0082] The gray image corresponding to the current frame image can be obtained according to the current frame image, and its exemplary implementation manner can be referred to the related description of the following and Figure 7 Step S2 can be to obtain the initial key pixel position of the current frame image according to the gray image corresponding to the current frame image, and then to obtain the key pixel position of the current frame image according to the initial key pixel position of the current frame image and the key pixel position of the previous frame image. Its exemplary implementation manner can be referred to the related description of the following and Figure 7 .

[0083] In step S3, the event of the current frame is obtained according to the gray image corresponding to the current frame image, the key pixel position of the previous frame image and the key pixel value of the previous frame image, the event of the current frame indicating the motion information of the object in the current frame image compared with the previous frame image.

[0084] Step S3 can be obtaining the initial key pixel position of the current frame image according to the gray image corresponding to the current frame image, and then obtaining the event of the current frame according to the initial key pixel position of the current frame image, the key pixel position of the previous frame image and the key pixel value of the previous frame image. Step S3 can be executed simultaneously with step S2, or can be executed after step S2, or can be executed before step S2. The application does not limit the execution order of step S3 and step S2. When the initial key pixel position of the current frame image has been obtained before step S2 or step S3 is executed, step S2 or step S3 can directly use the initial key pixel position of the current frame image to save data processing cost and improve data processing efficiency. The method for obtaining the initial key pixel position of the current frame image can be referred to the related description of the application in the following and Figure 9

[0085] Step S4, obtaining the key pixel value of the current frame image according to the key pixel position of the current frame image, and replacing the key pixel value of the previous frame image and the key pixel position of the previous frame image stored in the memory with the key pixel value of the current frame image and the key pixel position of the current frame image.

[0086] The key pixel value of the current frame image and the key pixel position of the current frame image can be used for processing the next frame image. That is, the key pixel value and the key pixel position of one frame image can be stored in the memory.

[0087] According to the image processing method of the application, the event of the current frame is calculated by combining the gray image corresponding to the current frame image, the key pixel position of the previous frame image and the key pixel value of the previous frame image. Since the key pixel positions of adjacent two frame images in the time domain are closer when the frame rate is higher, the key pixel position of the previous frame image and the key pixel value at the key position are strongly related to the key pixel position and the key pixel value of the current frame image, so that the event of the current frame can be accurately obtained. The key pixel position of the current frame image is obtained according to the gray image corresponding to the current frame image and the key pixel position of the previous frame image, the key pixel value of the current frame image is obtained according to the key pixel position of the current frame image, and the key pixel value of the previous frame image and the key pixel position of the previous frame image stored in the memory are replaced with the key pixel value of the current frame image and the key pixel position of the current frame image, so that only the key pixel position and the key pixel value of one frame image are stored in the memory, the amount of data stored in the memory is low, and the cost can be reduced. The event of the current frame indicates the motion information of the object in the current frame image compared with the previous frame image, so that the high-accuracy motion information of the object in the image can be obtained at a low cost.

[0088] ​In a possible implementation, the method is applied to a sensor or a processor connected to the sensor, the current frame image comprises a raw format RAW image captured by the sensor or a three-channel color RGB image generated by the processor according to the image captured by the sensor, and when the method is applied to the sensor and the current frame image comprises the raw format RAW image captured by the sensor, the sensor outputs the event of the current frame while outputting the current frame image, or the sensor outputs the event of the current frame.

[0089] The method of the embodiments of the present application can be applied to a sensor, for example, the image sensor in the above Figure 4a and related description. The image sensor capturing a raw format RAW image can refer to the above Figure 4a and related description. The sensor outputs the event of the current frame while outputting the current frame image, so that the current frame image and the event of the current frame are synchronously transmitted to a device or apparatus for further processing of the current frame image and the event of the current frame, thus ensuring that the correspondence between the current frame image and the event of the current frame is more accurate, and the subsequent application layer program or module is more friendly to the use of the event and the image, that is, the accuracy of the application layer program or module in further performing a task using the event and the image is higher. When the subsequent task does not need to use the current frame image, the sensor can output only the event of the current frame, so as to reduce the data transmission cost of the sensor.

[0090] The method of the embodiments of the present application can be applied to a processor, for example, the image processor in the above Figure 4b and related description. The image processor generating a three-channel color RGB image can refer to the above Figure 4b and related description.

[0091] When the image processing method of the present application is applied to a sensor, the sensor outputs the event of the current frame while outputting the current frame image, which can ensure the synchronous output of the event of the current frame and the current frame image in time, and can ensure the accuracy of further performing other tasks using the event of the current frame and the current frame image. The sensor outputs only the event of the current frame, which can reduce the data transmission cost of the sensor. When the image processing method of the present application is applied to a processor, the data processing cost and power consumption of the image sensor can be reduced. In this way, the image processing method of the present application is more practical.

[0092] In a possible implementation, the method further comprises:

[0093] obtaining a gray image corresponding to the current frame image according to the current frame image, and the resolution of the gray image corresponding to the current frame image is less than or equal to the resolution of the current frame image.

[0094] In this way, the grayscale image corresponding to the current frame image can be obtained. Furthermore, by setting the resolution of the grayscale image corresponding to the current frame image to be equal to the resolution of the current frame image, it is unnecessary to adjust the resolution of the current frame image when performing the step of obtaining the grayscale image corresponding to the current frame image, thus ensuring event resolution. Setting the resolution of the grayscale image corresponding to the current frame image to be less than the resolution of the current frame image reduces the data processing cost of performing the steps of obtaining the event of the current frame and the key pixel positions of the current frame image from the grayscale image corresponding to the current frame image.

[0095] For example, from the above text Figure 4a and Figure 4b As described in the relevant description, the current frame image is either an image acquired by an image sensor or an image generated by an image signal processor (ISP) based on an image acquired by an image sensor. When the current frame image is acquired by an image sensor, it may be a high-resolution Bayer format RAW image, in which case a grayscale image cannot be directly obtained. When the current frame image is generated by an image processor, its resolution may also be high, and it may not be a grayscale image. In this application, processing is performed using a grayscale image and information stored in memory (i.e., steps S2 and S3 mentioned above). Therefore, before executing step S2 or step S3, it is necessary to perform the step of obtaining the grayscale image corresponding to the current frame image based on the current frame image.

[0096] In one possible implementation, when the resolution of the grayscale image corresponding to the current frame image is less than the resolution of the current frame image, obtaining the grayscale image corresponding to the current frame image based on the current frame image includes:

[0097] The current frame image is subjected to one or more of the following processing methods: filtering, downsampling, interpolation, and nonlinear transformation, to obtain a grayscale image corresponding to the current frame image.

[0098] Depending on the type of the current frame image (e.g., RAW, RGB, or YUV image), and the different cases where the resolution of the corresponding grayscale image is less than or equal to the resolution of the current frame image, there can be different ways to obtain the grayscale image corresponding to the current frame image. Several exemplary implementations of obtaining the grayscale image corresponding to the current frame image are given below.

[0099] In one possible implementation, when the resolution of the grayscale image corresponding to the current frame image is equal to the resolution of the current frame image, obtaining the grayscale image corresponding to the current frame image based on the current frame image can include:

[0100] the current frame image is taken as a gray image of the current frame image, or

[0101] the current frame image is filtered to obtain a filtered image, and the filtered image is taken as a gray image of the current frame image, or

[0102] the current frame image is subjected to nonlinear transformation to obtain a nonlinearly transformed image, and the nonlinearly transformed image is taken as a gray image of the current frame image, or

[0103] the current frame image is filtered to obtain a filtered image, the filtered image is subjected to nonlinear transformation to obtain a nonlinearly transformed image, and the nonlinearly transformed image is taken as a gray image of the current frame image.

[0104] For example, when the resolution of the current frame image is 1080*1080, the resolution of the gray image of the current frame image is also 1080*1080, for example. The current frame image can be an image that already includes gray information, such as a YUV format image, and thus the current frame image can be directly taken as the gray image of the current frame image.

[0105] For another example, the current frame image can be a linear domain image that already includes gray information, and the current frame image can be subjected to nonlinear transformation to obtain a nonlinearly transformed image, and the nonlinearly transformed image is taken as the gray image of the current frame image. The nonlinear transformation can normalize the variation of brightness in dark and bright environments, and can be implemented in a manner known in the art, such as gamma transformation or logarithmic transformation, etc. For example, when the resolution of the current frame image is 1080*1080, the resolution of the nonlinearly transformed image can be 1080*1080, for example, and the pixel value range of the nonlinearly transformed image can be 0-255, for example. The pixel value range of the nonlinearly transformed image in a dark region can be transformed from 0-25 to 0-89, and the pixel value range of the nonlinearly transformed image in a bright region can be transformed from 178-255 to 216-255. Since the pixel value of the gray image of the current frame image is also used in step S3 to set the first threshold value and the second threshold value, which are further used to determine whether an event is obtained (see the description of step S3 below for an example), using the nonlinearly transformed image as the gray image of the current frame image can enable a small brightness variation in a dark environment to be detected to obtain a corresponding event, and a larger brightness variation in a bright environment is required to obtain a corresponding event, which can make the selection of the first threshold value and the second threshold value in step S3 more accurate, and thus reduce the data processing cost when steps S2 and S3 are executed, while the accuracy of the event obtained in step S3 is higher.

[0106] For example, the current frame image can be an image without including grayscale information, such as an infrared RAW image or an RGB image. The current frame image can be filtered to obtain a filtered image, and the filtered image can be used as the grayscale image of the current frame image. The filtering can be implemented in a manner known in the art, such as median filtering or mean filtering, etc. For example, when the resolution of the current frame image is 1080*1080, the resolution of the filtered image is also 1080*1080, for example. The operation of using the filtered image as the grayscale image of the current frame image can reduce the noise in the image, and thus improve the accuracy of the image processing method. It should be understood by those skilled in the art that when the current frame image already includes grayscale information, the filtering can be performed to further reduce the noise, and the present application does not limit the specific type of the current frame image that can be filtered.

[0107] For example, the current frame image can be an image without including grayscale information, and the current frame image can be filtered to obtain a filtered image, and the filtered image can be nonlinearly transformed to obtain a nonlinearly transformed image, and the nonlinearly transformed image can be used as the grayscale image of the current frame image. The exemplary manner of filtering and the exemplary manner of nonlinear transformation can be referred to the description above, and will not be described herein. By using the filtering and the nonlinear transformation, the noise in the grayscale image of the current frame image can be reduced, and the accuracy of the event obtained using the grayscale image of the current frame image can be improved. It should be understood by those skilled in the art that when the current frame image already includes grayscale information, the filtering can be performed to further reduce the noise, and the nonlinear transformation can be performed to further improve the accuracy of the event obtained using the grayscale image of the current frame image, and the present application does not limit the specific type of the current frame image that can be filtered and nonlinearly transformed.

[0108] In a possible implementation, when the resolution of the grayscale image corresponding to the current frame image is less than the resolution of the current frame image, obtaining the grayscale image corresponding to the current frame image according to the current frame image can include:

[0109] filtering the current frame image to obtain a filtered image, and using the filtered image as the grayscale image of the current frame image, or

[0110] filtering the current frame image to obtain a filtered image, nonlinearly transforming the filtered image to obtain a nonlinearly transformed image, and using the nonlinearly transformed image as the grayscale image of the current frame image, or

[0111] down-sampling or interpolation processing on the current frame image to obtain a down-sampled image or an interpolated image, and taking the down-sampled image or the interpolated image as the gray scale image corresponding to the current frame image, or

[0112] down-sampling or interpolation processing on the current frame image to obtain a down-sampled image or an interpolated image, and taking the down-sampled image or the interpolated image as the gray scale image corresponding to the current frame image, or

[0113] down-sampling or interpolation processing on the current frame image to obtain a down-sampled image or an interpolated image, and taking the down-sampled image or the interpolated image as the gray scale image corresponding to the current frame image, or

[0114] down-sampling or interpolation processing on the current frame image to obtain a down-sampled image or an interpolated image, and taking the down-sampled image or the interpolated image as the gray scale image corresponding to the current frame image, or

[0115] For example, when the resolution of the current frame image is 1080*1080, the resolution of the gray scale image corresponding to the current frame image can be, for example, 540*540. The current frame image can be a RAW image, and the filtered image can be directly obtained by filtering the current frame image, and the filtered image can be taken as the gray scale image of the current frame image. The exemplary manner of filtering can be referred to the description above, and when the RAW image is 2cell / 4cell, the stride of filtering can be set as 2 / 4, and the filter of 4*4 or 8*8 can be selected. The filtered image obtained can be an image with reduced noise and resolution. This operation is relatively simple, and the complexity of the image processing method is greatly reduced.

[0116] For example, when the resolution of the current frame image is 1080*1080, the resolution of the gray scale image corresponding to the current frame image can be, for example, 540*540. The current frame image can be a RAW image, and the filtered image can be directly obtained by filtering the current frame image, and the filtered image can be taken as the gray scale image of the current frame image. The exemplary manner of filtering can be referred to the description above, and when the RAW image is 2cell / 4cell, the stride of filtering can be set as 2 / 4, and the filter of 4*4 or 8*8 can be selected. The filtered image obtained can be an image with reduced noise and resolution. This operation is relatively simple, and the complexity of the image processing method is greatly reduced.

[0117] For another example, the current frame image can be an RGB image containing grayscale information, and the current frame image can not be filtered, but can be directly down-sampled or interpolated to obtain a down-sampled image or an interpolated image, and the down-sampled image or the interpolated image can be taken as a grayscale image corresponding to the current frame image. The down-sampling and the interpolation can reduce the resolution of the image, and can be implemented in a manner known in the art. For example, the down-sampling can be mean down-sampling, and the interpolation can be bilinear interpolation, nearest neighbor interpolation, etc. Taking the down-sampling as an example, when the resolution of the current frame image is 1080*1080, the down-sampling parameter is 2*2 (i.e., the pixels in a 2*2 window of the image before down-sampling are changed into one pixel), and the resolution of the down-sampled image can be, for example, 540*540. Since the resolution of the grayscale image corresponding to the current frame image is lower, the data processing cost when processing the grayscale image and the information stored in the memory (i.e., steps S2 and S3 described above) is also reduced.

[0118] For another example, the current frame image can be an RGB image containing grayscale information, and the current frame image can not be filtered, but can be directly down-sampled or interpolated to obtain a down-sampled image or an interpolated image, and the down-sampled image or the interpolated image can be taken as a grayscale image corresponding to the current frame image. The down-sampling and the interpolation can reduce the resolution of the image, and can be implemented in a manner known in the art. For example, the down-sampling can be mean down-sampling, and the interpolation can be bilinear interpolation, nearest neighbor interpolation, etc. Taking the down-sampling as an example, when the resolution of the current frame image is 1080*1080, the down-sampling parameter is 2*2 (i.e., the pixels in a 2*2 window of the image before down-sampling are changed into one pixel), and the resolution of the down-sampled image can be, for example, 540*540. Since the resolution of the grayscale image corresponding to the current frame image is lower, the data processing cost when processing the grayscale image and the information stored in the memory (i.e., steps S2 and S3 described above) is also reduced.

[0119] For another example, the current frame image can be an RGB image containing grayscale information, and the current frame image can not be filtered, but can be directly down-sampled or interpolated to obtain a down-sampled image or an interpolated image, and the down-sampled image or the interpolated image can be taken as a grayscale image corresponding to the current frame image. The down-sampling and the interpolation can reduce the resolution of the image, and can be implemented in a manner known in the art. For example, the down-sampling can be mean down-sampling, and the interpolation can be bilinear interpolation, nearest neighbor interpolation, etc. Taking the down-sampling as an example, when the resolution of the current frame image is 1080*1080, the down-sampling parameter is 2*2 (i.e., the pixels in a 2*2 window of the image before down-sampling are changed into one pixel), and the resolution of the down-sampled image can be, for example, 540*540. Since the resolution of the grayscale image corresponding to the current frame image is lower, the data processing cost when processing the grayscale image and the information stored in the memory (i.e., steps S2 and S3 described above) is also reduced.

[0120] Those skilled in the art should understand that when the current frame image is a RAW image, filtering processing can also be performed to reduce noise, reduce resolution, and then further reduce resolution through downsampling processing or interpolation processing, and then further improve the accuracy of the event obtained using the grayscale image of the current frame image through nonlinear transformation processing. The present application does not limit the specific type of current frame image that can perform filtering processing, downsampling processing, or interpolation processing and nonlinear transformation processing.

[0121] In the case where the resolution of the grayscale image corresponding to the current frame image is less than the resolution of the current frame image, one or more of filtering processing, downsampling processing, interpolation processing, and nonlinear transformation processing are performed on the current frame image to obtain the grayscale image corresponding to the current frame image. Downsampling processing, interpolation processing, and filtering processing result in a lower resolution of the obtained grayscale image corresponding to the current frame image, thereby reducing the data processing cost of the image processing method. Filtering processing reduces the interference of noise, resulting in more accurate and less noisy events. Nonlinear transformation processing makes it easier to select the threshold (first threshold, second threshold) of the generated event, resulting in higher accuracy of the event.

[0122] Optionally, one or more of filtering processing, downsampling processing, interpolation processing, and nonlinear transformation processing can be performed on the current frame image to obtain a grayscale image corresponding to the current frame image with a resolution equal to that of the current frame image, and a grayscale image corresponding to the current frame image with a resolution less than that of the current frame image. In step S3, when the event is obtained based on the grayscale image corresponding to the current frame image, the grayscale image corresponding to the current frame image with a resolution equal to that of the current frame image can be used, so that the event can be obtained based on a more data-rich grayscale image, further improving the accuracy of the obtained event. In step S2, when the key pixel position of the current frame image is obtained based on the grayscale image corresponding to the current frame image and the key pixel position of the previous frame image, the grayscale image corresponding to the current frame image with a resolution less than that of the current frame image can be used, so that the data processing cost of the image processing method can be reduced.

[0123] Those skilled in the art should understand that when the grayscale image corresponding to the current frame image is obtained based on the current frame image, the execution order of the multiple processing means should not be limited to the above examples, and the present application does not limit the execution order of each processing means. In addition to the processing means of filtering processing, downsampling processing, interpolation processing, and nonlinear transformation processing, more processing means can be included when the grayscale image corresponding to the current frame image is obtained based on the current frame image, and the present application does not limit the specific processing means used to obtain the grayscale image corresponding to the current frame image.

[0124] After obtaining the gray image corresponding to the current frame image, and the key pixel position of the previous frame image and the key pixel value of the previous frame image have been obtained in step S1, step S2 and step S3 can be performed to obtain the key pixel position of the current frame image and the event of the current frame, respectively. The exemplary method of step S2 to obtain the key pixel position of the current frame image is introduced first.

[0125] In a possible implementation, step S2 comprises:

[0126] performing feature extraction on the gray image corresponding to the current frame image, and obtaining the initial key pixel position of the current frame image according to the positions of the pixel points included in the extracted features in the gray image corresponding to the current frame image;

[0127] performing time domain smoothing processing on the initial key pixel position of the current frame image and the key pixel position of the previous frame image to obtain the key pixel position of the current frame image.

[0128] The initial key pixel position of the previous frame image is obtained by the feature extraction method first, and then the initial key pixel position and the key pixel position of the previous frame image are used to perform time domain smoothing processing to obtain the key pixel position of the current frame image, which avoids the case that a large error is caused when the initial key pixel position of the current frame image is directly used as the key pixel position of the current frame image due to the noise in the gray image corresponding to the current frame image, and makes the obtained key pixel position of the current frame image more accurate.

[0129] For example, the key pixel position of the current frame image is the position of the pixel point representing the object feature in the current frame image. As can be known from the related description above, Figure 5 the object feature in the image can be the edge of the object in the image or other features that can distinguish the object from the image, and therefore, the position of the pixel point representing the object feature in the current frame image can be extracted by using the low-level feature extraction method such as edge detection and / or histogram of oriented gradient (HOG) detection. The specific method of implementing the feature extraction is not limited in the present application.

[0130] The position of the pixel point included in the extracted feature in the gray scale image corresponding to the current frame image can be used as the initial key pixel position of the current frame image, and then time domain smoothing processing is performed according to the initial key pixel position of the current frame image and the key pixel position of the previous frame image to obtain the key pixel position of the current frame image. The time domain smoothing can be implemented based on the prior art, for example, two smoothing parameters a and b are pre-set, for example, a+b=1, the key pixel position of the previous frame image is represented by a first binary image (where the pixel value of the pixel point of the key pixel position is represented by 1, and the pixel value of the pixel point of the non-key pixel position is represented by 0), and the initial key pixel position of the current frame image is represented by a second binary image (where the pixel value of the pixel point of the initial key pixel position is represented by 1, and the pixel value of the pixel point of the non-initial key pixel position is represented by 0). For each same pixel point in the two binary images, the sum of the product of the pixel value (0 or 1) in the first binary image at the pixel point and the smoothing parameter a and the product of the pixel value (0 or 1) in the second binary image and the smoothing parameter b is close to "1" (for example, the absolute value of the difference from 1 is less than a certain pre-set threshold value), it can be determined that the pixel point is the pixel point of the key pixel position of the current frame image. Those skilled in the art should understand that the time domain smoothing should not be limited to the above exemplary implementation, and the specific method adopted by the embodiment of the present application for time domain smoothing is not limited. The time domain smoothing reduces the noise of the position of the pixel point representing the object feature, that is, makes the key pixel position of the current frame image more accurate.

[0131] In the following, the way of feature extraction is exemplarily described by taking the feature extraction of the gray scale image corresponding to the current frame image by edge detection as an example.

[0132] In a possible implementation, the feature extraction of the gray scale image corresponding to the current frame image comprises:

[0133] Obtaining the gradient of each pixel point in the gray scale image corresponding to the current frame image in at least one direction;

[0134] Obtaining the edge feature of the gray scale image corresponding to the current frame image according to the pixel point with the gradient greater than the fourth threshold value.

[0135] In this way, the pixel point indicating the feature of the object in the image can be obtained. The edge feature can be obtained according to the gradient in multiple directions, so that the obtained feature is more accurate.

[0136] The gradient of each pixel in the gray image corresponding to the current frame image in at least one direction can include the gradient of each pixel in the horizontal direction and / or the vertical direction. The gradient can be calculated according to the prior art, for example, the gradient in the horizontal direction can be equal to the rate of change of the gray value of a pixel and the adjacent pixel of the pixel in the horizontal direction in the image. Alternatively, the gradient can also include the gradient in other directions, for example, the diagonal direction of the image, and the like, which are not limited in the present application. The edge features of the gray image corresponding to the current frame image are obtained according to the pixel corresponding to the gradient greater than the fourth threshold value in all the obtained gradients. Those skilled in the art should understand that the fourth threshold value can be a pre-set fixed value, or can be determined according to the calculated gradient, for example, the minimum value of the top 10% of the calculated gradient is taken as the fourth threshold value. Those skilled in the art should understand that 10% is only an example, and the present application does not limit the specific acquisition method of the fourth threshold value.

[0137] After the edge features of the gray image corresponding to the current frame image are extracted, the initial key pixel position of the current frame image can be obtained according to the edge features of the gray image corresponding to the current frame image. Since the initial key pixel position of the current frame image needs to be time-domain smoothed with the key pixel position of the previous frame image, the initial key pixel position of the current frame image and the key pixel position of the previous frame image need to be positions in the same resolution. In the present application, it can be set that each frame image obtained by the image sensor or the image processor can be an image of the same resolution (for example, a first resolution), and it can be set that the key pixel position of the first frame image is the corresponding position in the resolution of the first frame image. That is, when the initial key pixel position of the current frame image becomes a position in the resolution of the current frame image, it satisfies the condition that the initial key pixel position of the current frame image and the key pixel position of the previous frame image are positions in the same resolution.

[0138] For different cases where the resolution of the gray image corresponding to the current frame image is less than or equal to the resolution of the current frame image, the feature extraction is performed on the gray image corresponding to the current frame image, and the initial key pixel position of the current frame image is obtained, so that the initial key pixel position of the current frame image becomes a position in the resolution of the current frame image, which can have different implementation manners.

[0139] For example, if the resolution of the gray image corresponding to the current frame image is equal to the resolution of the current frame image, the position correspondence between the pixel points in the gray image corresponding to the current frame image and the pixel points in the current frame image can be a one-to-one correspondence according to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image. When the initial key pixel positions of the current frame image are obtained according to the position correspondence and the positions of the pixel points included in the extracted feature in the gray image corresponding to the current frame image, the positions of the pixel points in the extracted feature can be directly used as the initial key pixel positions of the current frame image. For example, the resolution of the current frame image can be 1080*1080, the resolution of the gray image corresponding to the current frame image can be 1080*1080, and the feature obtained by performing feature extraction in step S2 can include two pixel points (10, 10) and (100, 100). Therefore, the initial key pixel positions of the current frame image can include (10, 10) and (100, 100).

[0140] In a possible implementation, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the feature extraction is performed on the gray image corresponding to the current frame image to obtain the initial key pixel positions of the current frame image, including:

[0141] determining the position correspondence between the pixel points in the gray image corresponding to the current frame image and the pixel points in the current frame image according to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image;

[0142] obtaining the initial key pixel positions of the current frame image according to the position correspondence and the positions of the pixel points included in the extracted feature in the gray image corresponding to the current frame image.

[0143] For example, if the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the position correspondence relationship between the pixel points in the gray image corresponding to the current frame image and the pixel points in the current frame image determined according to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image can be a multiple mapping relationship. In the multiple mapping relationship, the multiple can be the ratio of the resolution of the current frame image to the resolution of the gray image corresponding to the current frame image. When the initial key pixel position of the current frame image is obtained according to the position correspondence relationship and the position of the pixel points included in the feature in the gray image corresponding to the current frame image, the pixel point positions in the feature can be mapped to the corresponding positions in the current frame image at the resolution or the preset resolution according to the multiple in the multiple mapping relationship. The mapping can be realized by, for example, upsampling, 8-neighbor connectivity and other existing technologies. The mapped pixel point positions can be used as the initial key pixel positions of the current frame image. For example, the resolution of the current frame image can be 1080*1080, the resolution of the gray image corresponding to the current frame image can be 540*540, and the multiple in the multiple mapping relationship can be (1080*1080) / (540*540)=2*2. The feature obtained by performing feature extraction on the gray image corresponding to the current frame image can include two pixel points (10, 10) and (100, 100). According to the multiple 2*2 in the multiple mapping relationship, the pixel points (10, 10) and (100, 100) in the feature can be mapped to pixel points (20, 20) and (200, 200), and the initial key pixel positions of the current frame can include (20, 20) and (200, 200).

[0144] In this way, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the initial key position of the current frame image can be obtained through the correspondence relationship between the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image, so that the initial key position and the key pixel position of the previous frame image are at the same resolution, and it is possible to perform time domain smoothing processing on the initial key position of the current frame image and the key pixel position of the previous frame image.

[0145] The key pixel position of the first frame image is the corresponding position under the resolution of the first frame image as an example. Those skilled in the art should understand that when the key pixel position of the first frame image is set to a preset position different from the corresponding position under the resolution (for example, the second resolution) of the first frame image, the position correspondence between the pixel points in the gray image corresponding to the current frame image and the pixel points in the image under the preset resolution corresponding to the current frame image can be determined according to the resolution of the gray image corresponding to the current frame image and the preset resolution, and the implementation manner can refer to the example of determining the position correspondence between the pixel points in the gray image corresponding to the current frame image and the pixel points in the current frame image according to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, wherein the selection of the multiple can be the ratio of the preset resolution to the resolution of the gray image corresponding to the current frame image. According to the obtained position correspondence and the position of the pixel points included in the extracted feature in the gray image corresponding to the current frame image, the initial key pixel position of the current frame image is obtained, and at this time, the initial key pixel position of the current frame image can be the position under the preset resolution. In this way, the initial key position of the current frame image and the key pixel position of the previous frame image are also positions under the same resolution. The preset resolution can be set to be less than the resolution of the first frame image, and in this case, the data storage cost of the key pixel position of one frame image stored in the memory is lower.

[0146] The following describes an exemplary method of obtaining the event of the current frame in step S3.

[0147] In a possible implementation, step S3 comprises:

[0148] According to the gray image corresponding to the current frame image and the key pixel position of the previous frame image, the initial key pixel value of the current frame image is obtained.

[0149] According to the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image, the event of the current frame is obtained.

[0150] The pixel value is equal to the value of the pixel point, and the event of the current frame is obtained by the difference between the pixel values, so that the event of the current frame is associated with the change degree of the pixel value, and the accuracy of the event of the current frame can be ensured.

[0151] For example, the key pixel positions of the previous frame image can be obtained from the memory, and the key pixel values of the previous frame image corresponding to the key pixel positions of the previous frame image are also stored in the memory. The values of the pixel points in the gray image corresponding to the current frame image at the key pixel positions of the previous frame image, i.e., the initial key pixel values of the current frame image, are obtained. In this way, each pixel point at the key pixel positions of the previous frame image corresponds to two values of adjacent frames, and the two values can be subtracted to determine whether an event occurs at the pixel point according to the difference.

[0152] According to the resolution of the key pixel positions of the previous frame image and the resolution of the gray image corresponding to the current frame image, different methods can be used to obtain the initial key pixel values of the current frame image in step S3.

[0153] For example, when the key pixel positions of the previous frame image are positions at the first resolution (or the second resolution) and the gray image corresponding to the current frame image is an image at the first resolution (or the second resolution), the values of the pixel points in the gray image corresponding to the current frame image at the key pixel positions of the previous frame image can be directly used as the initial key pixel values of the current frame image.

[0154] When the key pixel positions of the previous frame image are positions at the second resolution and the gray image corresponding to the current frame image is an image at the first resolution, the key pixel positions of the previous frame image can be mapped so that the mapped key pixel positions of the previous frame image become positions at the first resolution. The mapping method can refer to the example of determining the position correspondence between the pixel points in the gray image corresponding to the current frame image and the pixel points in the current frame image according to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image. Then, the values of the pixel points in the gray image corresponding to the current frame image at the mapped key pixel positions of the previous frame image are used as the initial key pixel values of the current frame image.

[0155] The following describes an example method of obtaining the events of the current frame according to the difference between the initial key pixel values of the current frame image and the key pixel values of the previous frame image.

[0156] In a possible implementation, the events of the current frame include common events and disappearance events, and the events of the current frame are obtained according to the difference between the initial key pixel values of the current frame image and the key pixel values of the previous frame image, including:

[0157] a pixel point common to the initial key pixel position of the current frame image and the key pixel position of the previous frame image as a common position, and obtaining a common event of the pixel point when a difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image at each pixel point of the common position is greater than the first threshold value or smaller than the second threshold value;

[0158] a pixel point other than the common position in the key pixel position of the previous frame image as a disappearance position, and obtaining a disappearance event of the pixel point when a difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image at each pixel point of the disappearance position is greater than the first threshold value or smaller than the second threshold value.

[0159] In this way, the common event and the disappearance event of the current frame can be obtained. Since the common event is obtained at the common position which is a pixel point common to the initial key pixel position of the current frame image and the key pixel position of the previous frame image, it can be indicated that the pixel point of the common position appears motion between the acquisition time of the current frame image and the acquisition time of the previous frame image. Since the disappearance event is obtained at the disappearance position which is a pixel point other than the common position in the key pixel position of the previous frame image, it can be indicated that the pixel point of the disappearance position appears motion between the acquisition time of the current frame image and the acquisition time of the previous frame image. Thus, the motion information of the object in the current frame image can be represented by the common event and the disappearance event.

[0160] For example, when the object in the natural scene appears motion between the acquisition time of the current frame image and the acquisition time of the previous frame image, the initial key pixel position of the current frame image and the key pixel position of the previous frame image are usually different, Figure 6An example of the initial key pixel position of the current frame image and the key pixel position of the previous frame image according to the embodiment of the present application is shown. Wherein, the pixel point shared by the initial key pixel position of the current frame image and the key pixel position of the previous frame image is the shared position, which can represent the pixel point position that still embodies the object feature after the object motion. The pixel point other than the shared position in the key pixel position of the previous frame image is the disappeared position, which can represent the pixel point position that no longer embodies the object feature after the object motion. And when the object in the image moves, the value of each pixel point in the gray image corresponding to the current frame image will change compared with the value of the corresponding pixel point in the gray image corresponding to the previous frame image, so it can be considered that when the change degree of the value of the same pixel point in the two frames is greater than a certain threshold, the event can be obtained at the pixel point. If the change degree of the value of the same pixel point in the two frames is less than or equal to a certain threshold, it can be considered that it is the noise that can occur in the image acquisition process, such as the noise caused by the time consistency problem, etc., and no event is obtained at the pixel point. On this basis, based on the pixel point of the shared position and the pixel point of the disappeared position, the shared event and the disappeared event can be obtained respectively.

[0161] Wherein, at each pixel point of the shared position, when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than the first threshold or less than the second threshold, the shared event of the pixel point is obtained. The shared event obtained at a certain pixel point can include the time when the event occurs, the pixel coordinates corresponding to the event, and the attribute of the event, wherein the attribute of the shared event can be the identification of the shared event.

[0162] At each pixel point of the disappeared position, when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than the first threshold or less than the second threshold, the disappeared event of the pixel point is obtained. The disappeared event obtained at a certain pixel point can include the time when the event occurs, the pixel coordinates corresponding to the event, and the attribute of the event, wherein the attribute of the disappeared event can be the identification of the disappeared event.

[0163] The first threshold value and the second threshold value can be preset according to a pixel value range in the gray image corresponding to the current frame image. Since the values of the first threshold value and the second threshold value are directly used as the basis for event acquisition, when the pixel value range in the gray image corresponding to the current frame image is 0-255, the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image can be a value between [-255, 255], and when the pixel value range in the gray image corresponding to the current frame image is 0-15, the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image can be a value between [-15, 15]. It can be seen that it is easier to determine the first threshold value and the second threshold value for [-15, 15] than for [-255, 255]. The specific values of the first threshold value and the second threshold value are not limited in the embodiments of the present application.

[0164] Optionally, the common event and the disappearing event can further include a polarity of the event. If the event is obtained when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than the first threshold value, the event can be a positive event, and the event polarity can be represented by "1". If the event is obtained when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is less than the second threshold value, the event can be a negative event, and the event polarity can be represented by "-1".

[0165] The following describes an exemplary method of obtaining the event of the current frame according to the initial key pixel position of the current frame image and the initial key pixel position of the current frame image.

[0166] In a possible implementation, the step S3 includes:

[0167] performing feature extraction on the gray image corresponding to the current frame image to obtain the initial key pixel position of the current frame image;

[0168] obtaining the event of the current frame according to the similarity between the initial key pixel position of the current frame image and the key pixel position of the previous frame image.

[0169] The implementation of the step of performing feature extraction on the gray image corresponding to the current frame image to obtain the initial key pixel position of the current frame image can refer to the foregoing step S2 and the example of obtaining the gray image of the current frame image according to the current frame image. If the initial key pixel position of the current frame image has been obtained before the feature extraction step is performed, the obtained initial key pixel position of the current frame image can be directly used, and the feature extraction step does not need to be performed again.

[0170] After the initial key pixel position of the current frame image is obtained, the event of the current frame can be obtained according to the similarity between the initial key pixel position of the current frame image and the key pixel position of the previous frame image. The similarity represents the coincidence degree of the pixel points included in the initial key pixel position of the current frame image and the key pixel position of the previous frame image.

[0171] The key pixel position is the position of the pixel point representing the object feature, the initial key pixel position is obtained by feature extraction, and also includes the pixel point representing the object feature. Therefore, the event of the current frame is obtained by the similarity between the key pixel position of the previous frame image and the initial key pixel position of the current frame, so that the event of the current frame is associated with the similarity of the object features of the current frame image and the previous frame image, and the accuracy of the event of the current frame can be ensured.

[0172] In a possible implementation, the event of the current frame includes a new event, and the event of the current frame is obtained according to the similarity between the initial key pixel position of the current frame image and the key pixel position of the previous frame image, including:

[0173] For each region of the gray image corresponding to the current frame image, the ratio of the number of pixel points in the intersection of the initial key pixel position of the current frame image and the key pixel position of the previous frame image to the number of pixel points in the union of the initial key pixel position of the current frame image and the key pixel position of the previous frame image is determined.

[0174] When the ratio is less than a third threshold, a new event corresponding to the pixel point in the initial key pixel position of the current frame image in the region is obtained.

[0175] For example, the initial key pixel position of the current frame image is obtained according to the gray image corresponding to the current frame image, and therefore, if the gray image corresponding to the current frame image is divided into multiple regions, the pixel points of the initial key pixel position can be included in each region. For each region after the gray image corresponding to the current frame image is divided, the intersection of the initial key pixel position of the current frame image and the key pixel position of the previous frame image in the region can be determined, for example, by the number of common pixel points. And the union of the initial key pixel position of the current frame image and the key pixel position of the previous frame image in the region can be determined, for example, by the sum of the number of pixel points of the initial key pixel position of the current frame image and the key pixel position of the previous frame image in the region minus the number of common pixel points. In theory, the closer the ratio of the number of pixel points in the intersection of the initial key pixel position of the current frame image and the key pixel position of the previous frame image in the region to the number of pixel points in the union is to 1, the smaller the motion amplitude of the object in the region within the capture time of the current frame image and the capture time of the previous frame image. Therefore, a third threshold value can be set in advance, when the above ratio is less than the third threshold value, it is considered that the motion amplitude of the object in the region is larger, and the event can be obtained at the initial key pixel position of the current frame image. When the above ratio is greater than or equal to the third threshold value, it is considered that noise can be generated at the pixel point, and no event is obtained at the pixel point. In practical applications, the third threshold value can be set by the ISO sensitivity of the current frame image. Since the initial key pixel position of the current frame image corresponds to the current frame image, the new event corresponding to the pixel point of the initial key pixel position of the current frame image in the region is obtained.

[0176] In this way, the new event of the current frame can be obtained. Since the new event is obtained by the ratio of the intersection and the union of the initial key pixel position of the current frame image and the key pixel position of the previous frame image in each region, it can be indicated that the pixel point of the initial key pixel position of the current frame image in the region appears motion between the capture time of the current frame image and the capture time of the previous frame image. Thus, the motion information of the object in the current frame image can be represented by the new event.

[0177] After the common event, the disappearing event and the new event are obtained, the key pixel position of the previous frame image stored in the memory and the key pixel value of the previous frame image are used up, and thus the value of each pixel point at the key pixel position of the current frame image, i.e., the key pixel value of the current frame image, can be obtained according to the gray image corresponding to the current frame image, and the key pixel position of the current frame image and the key pixel value of the current frame image are used to replace the key pixel position of the previous frame image and the key pixel value of the previous frame image stored in the memory. In a possible implementation, the key pixel position of the current frame image is stored in the memory in the form of a binary image or in the form of a lossless compressed data packet.

[0178] For example, the key pixel position of the current frame image can be directly stored in the form of a binary image, and the storage cost of the pixel value of each pixel point in the binary image is 1 bit, for example, the pixel value of the pixel point of the key pixel position of the current frame image is set to 1, and the pixel value of the pixel point of the key pixel position of the previous frame image is set to 0. Alternatively, a lossless compression method can be used to compress the key pixel position of the current frame image (optionally, the key pixel value of the current frame image is also included) to obtain a lossless compressed data packet, and the compressed data packet is stored in the memory. When the next frame image is processed, the stored data packet is first decompressed to obtain the key pixel position of the current frame image (optionally, the key pixel value of the current frame image is also included) from the memory.

[0179] In this way, the data storage cost of the memory can be reduced.

[0180] After the event of the current frame is obtained, the event of the current frame can be output, for example, in the form of information coding. The event of the current frame can be encoded in a fixed bit coding manner of the prior art, or in an encoding manner of a conventional event camera, i.e., address-event-representation (AER). In the fixed bit coding, the resolution of the image is W*H, and 2-bit data (0, 1, 2) is used to represent the motion information for each pixel in the image. 0, 1 and 2 can represent no event, disappearing event, new event and common event respectively. Alternatively, when the image processing method of the embodiments of the present application is applied to an image sensor, the encoding can be completed by the image sensor to reduce the transmission bandwidth of the image sensor to the image processor. It should be understood by those skilled in the art that there can be more ways of encoding, which are not limited in the present application.

[0181] Figure 7 An example of the image data processing method according to the embodiments of the present application in actual application is shown.

[0182] AsFigure 7 As shown, the current frame image can be a sensor RAW image with a resolution of 346*260, and the current frame image can be processed in a spatial domain first. For example, pre-processing is performed according to the current frame image, and the pre-processing can include the filtering processing and the down-sampling processing described above. The filtering processing of the current frame image can obtain a filtered image (a first gray-scale image), and the first gray-scale image can be an image with a resolution of 346*260. The down-sampling processing of the filtered image can obtain a down-sampled image (a second gray-scale image), and the second gray-scale image can be an image with a resolution of 173*130. The key pixel positions of the previous frame image and the key pixel values of the previous frame image can be obtained from the memory. Alternatively, when the key pixel positions of the encoded previous frame image and the key pixel values of the previous frame image are stored in the memory, the information stored in the memory can be decoded first and then obtained. According to the first gray-scale image and the key pixel positions of the previous frame image, the initial key pixel values of the current frame image can be obtained (for details, refer to the example of obtaining the initial key pixel values of the current frame image according to the gray-scale image corresponding to the current frame image and the key pixel positions of the previous frame image in step S3 described above). According to the second gray-scale image, the initial key pixel positions of the current frame image can be obtained (for details, refer to the example of obtaining the initial key pixel positions of the current frame image by performing feature extraction on the gray-scale image corresponding to the current frame image in steps S2 / S3 described above).

[0183] The initial key pixel values of the current frame image and the initial key pixel positions of the current frame image can be processed in a space-time domain. For example, the difference between the initial key pixel values of the current frame image and the key pixel values of the previous frame image is obtained, and the events of the current frame are obtained according to the difference (for details, refer to the example of obtaining the events of the current frame according to the difference between the initial key pixel values of the current frame image and the key pixel values of the previous frame image in step S3 described above). The key pixel positions of the current frame image are obtained by performing time-domain smoothing processing on the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image (for details, refer to the example of obtaining the key pixel positions of the current frame image by performing time-domain smoothing processing on the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image in step S2 described above). The key pixel positions of the current frame image and the key pixel values of the current frame image, i.e., the key pixel values of the current frame image, are written into the memory to replace the information of the previous frame stored in the memory. Alternatively, the key pixel positions of the current frame image and the key pixel values of the current frame image can be written after being encoded online. The events of the current frame can be output to other devices or equipment. Alternatively, the events of the current frame can be output after being encoded.

[0184] Figure 8a and Figure 8bExemplary schematic diagrams are shown respectively of events obtained according to the prior art and events of the current frame obtained according to the image processing method of the embodiments of this application. Figure 8a As shown in the diagram, in the prior art, the number of events obtained directly based on whole-frame differential is 14814. However, in this embodiment, the object undergoes the same motion as in the prior art, and the number of events occurring within the same time interval is 8071, while the object's features remain relatively clear. This allows for noise suppression and ensures the accuracy of generated events while reducing the number of generated events, further reducing output bandwidth consumption. Based on this, the present embodiment can achieve more efficient tracking of moving objects, making it particularly suitable for scenes with rapid motion.

[0185] The image processing method of this application can reduce hardware storage overhead. For example, the value range of a pixel is 0-255, thus requiring 8 bits of data storage. The current frame image is 346*260 pixels in size, and the existing whole-frame differential scheme requires 346*260*8bit = 89960 bytes of storage space. Figure 9 An exemplary schematic diagram showing pixels corresponding to information stored in the memory according to an embodiment of this application is provided. Figure 9 As shown, this embodiment of the application needs to store 6644 key pixel positions of the current frame image. As described above, the key pixel positions of the current frame image only require 1 bit of data storage. Therefore, the storage space required by this invention is 1 bit * 346 * 260 (space occupied by the key pixel positions of the current frame image) + 6644 * 8 bits (space occupied by the key pixel values ​​of the current frame image) = 17889 bytes. Compared with the storage space required by the prior art of 89960 bytes, the storage space required by this embodiment of the application is reduced by 80%.

[0186] When the image processing method of this application is applied to an image sensor or image processor, it can obtain high-quality imaging images (high resolution, high dynamic range, etc.) while retaining traditional photosensitive circuitry, and also acquire high-quality motion information (events). The output motion information (events) corresponds to the motion information of each frame of the image, and can be directly and efficiently used for application tasks (such as motion wake-up capture, exposure time control), greatly reducing latency. When this application is applied to an image sensor, it can ensure the synchronization of image and motion information, thereby avoiding accuracy errors caused by the asynchrony between image and motion information in related applications (such as motion wake-up capture, exposure time control).

[0187] The embodiment of the present application can effectively weaken and eliminate the influence of noise in a dark light scene by processing a current frame image to obtain a gray image corresponding to the current frame image in a simple and efficient manner and by using a more robust key pixel position extraction manner of the current frame image, and can effectively extract high-quality events in various scenes. The events obtained by the embodiment of the present application are not coarse-grained regions of interest, but can contain more effective information, and thus can be used in more application scenarios, such as pose estimation, accurate eye tracking, doorbell monitoring, smart home, smart sound box, and other scenarios requiring low power consumption and high endurance, thereby realizing real-time continuous, low-power monitoring, intelligent wake-up, and real-time accurate positioning tracking in high-speed motion scenarios in the above scenarios.

[0188] Figure 10 An exemplary structural schematic diagram of an image processing apparatus according to an embodiment of the present application is shown.

[0189] As Figure 10 shown, the embodiment of the present application provides an image processing apparatus, which comprises:

[0190] The first acquisition module 101 is configured to acquire a key pixel position of a previous frame image and a key pixel value of the previous frame image stored in a memory, the key pixel position being a position of a pixel point representing a feature of an object in a frame image, and the key pixel value being a value of the pixel point at the key pixel position.

[0191] The second acquisition module 102 is configured to acquire a key pixel position of a current frame image according to a gray image corresponding to the current frame image and the key pixel position of the previous frame image.

[0192] The third acquisition module 103 is configured to acquire an event of the current frame according to the gray image corresponding to the current frame image, the key pixel position of the previous frame image, and the key pixel value of the previous frame image, the event of the current frame indicating motion information of an object in the current frame image compared with the previous frame image.

[0193] The first replacement module 104 is configured to acquire a key pixel value of the current frame image according to the key pixel position of the current frame image, and replace the key pixel value of the previous frame image and the key pixel position of the previous frame image stored in the memory with the key pixel value of the current frame image and the key pixel position of the current frame image.

[0194] In a possible implementation, the events of the current frame include common events and disappearance events, and the events of the current frame are obtained according to a difference between initial key pixel values of the current frame image and key pixel values of the previous frame image, including: taking, as common positions, pixel points that are common to the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image, and taking, as disappearance positions, pixel points that are not in the common positions among the key pixel positions of the previous frame image, and taking, as a common event of each pixel point in the common positions, a case where the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than a first threshold value or less than a second threshold value, and taking, as a disappearance event of each pixel point in the disappearance positions, a case where the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than the first threshold value or less than the second threshold value.

[0195] In a possible implementation, the events of the current frame are obtained according to a corresponding gray image of the current frame image, the key pixel positions of the previous frame image, and the key pixel values of the previous frame image, including: performing feature extraction on the corresponding gray image of the current frame image to obtain initial key pixel positions of the current frame image; and obtaining the events of the current frame according to a similarity degree between the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image.

[0196] In a possible implementation, the events of the current frame include addition events, and the events of the current frame are obtained according to a similarity degree between the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image, including: determining, for each region of the corresponding gray image of the current frame image after division, a ratio of a number of pixel points in an intersection to a number of pixel points in a union of the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image in the region; and obtaining an addition event of a pixel point corresponding to the initial key pixel position of the current frame image in the region when the ratio is less than a third threshold value.

[0197] In a possible implementation, the key pixel positions of the current frame image are obtained according to the corresponding gray image of the current frame image and the key pixel positions of the previous frame image, including: performing feature extraction on the corresponding gray image of the current frame image to obtain initial key pixel positions of the current frame image; and performing time domain smoothing processing on the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image to obtain the key pixel positions of the current frame image.

[0198] In a possible implementation, the apparatus further includes a fourth obtaining module configured to obtain a corresponding gray image of the current frame image according to the current frame image, and a resolution of the corresponding gray image of the current frame image is less than or equal to a resolution of the current frame image.

[0199] In a possible implementation, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the gray image corresponding to the current frame image is obtained according to the current frame image, including: performing one or more of filtering processing, down-sampling processing, interpolation processing, and non-linear transformation processing on the current frame image to obtain the gray image corresponding to the current frame image.

[0200] In a possible implementation, when the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the initial key pixel position of the current frame image is obtained by performing feature extraction on the gray image corresponding to the current frame image, including: determining a position correspondence relationship between a pixel point in the gray image corresponding to the current frame image and a pixel point in the current frame image according to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image; and obtaining the initial key pixel position of the current frame image according to the position correspondence relationship and the position of the pixel point included in the extracted feature in the gray image corresponding to the current frame image.

[0201] In a possible implementation, the key pixel position of the current frame image is stored in the memory in the form of a binary image or in the form of a lossless compressed data packet.

[0202] In a possible implementation, the device is applied to a sensor or a processor connected to the sensor, and the current frame image includes a RAW image in a raw format captured by the sensor or a three-channel RGB image generated by the processor according to an image captured by the sensor, where, when the device is applied to the sensor and the current frame image includes the RAW image in the raw format captured by the sensor, the sensor outputs the event of the current frame while outputting the current frame image, or the sensor outputs the event of the current frame.

[0203] Figure 11 An exemplary structural schematic diagram of an image processing device according to an embodiment of the present application is shown.

[0204] As Figure 11 shown, an embodiment of the present application provides an image processing device, including: a processor and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions.

[0205] The image processing apparatus can be disposed in an electronic device, which can include at least one of a mobile phone, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, a smart speaker, an ultra-mobile personal computer (UMPC), a netbook, an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a drone, an in-vehicle device, a smart home device, or a smart city device. The embodiments of the present application do not specially limit the specific type of the electronic device to which the image processing apparatus belongs. The electronic device can include Figure 4a an image sensor in an example of the electronic device, or Figure 4b an image processor in an example of the electronic device.

[0206] The image processing apparatus can include a processor 110, an internal memory 121, a communication module 160, and the like.

[0207] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), and the like. Different processing units can be independent devices or integrated in one or more processors. For example, the processor 110 can implement the image processing method of the embodiments of the present application.

[0208] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 can be a cache memory. The memory can store instructions or data that have been used or used frequently by the processor 110, such as the key pixel positions of the previous frame image, the key pixel values of the previous frame image, and the like in the embodiments of the present application. If the processor 110 needs to use the instructions or data, it can directly call from the memory. Avoiding repeated access reduces the waiting time of the processor 110, thus improving the efficiency of the system.

[0209] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, a universal asynchronous receiver / transmitter (UART) interface, a general-purpose input / output (GPIO) interface, and the like. The processor 110 can connect modules (not shown) such as an external memory or other processors through at least one of the above interfaces.

[0210] The memory 121 can be used to store computer executable program codes including instructions. The memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required by a function (such as an application for feature extraction, etc.), and the like. The data storage area can store data created during the use of the image processing apparatus (such as a gray scale image of a current frame image, key pixel positions of a current frame image, key pixel values of a current frame image, etc.). In addition, the memory 121 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), and the like. The processor 110 executes various functional methods or data processing of the image processing apparatus by running instructions stored in the memory 121 and / or instructions stored in a memory disposed in the processor.

[0211] The communication module 160 can be used to receive data from or send data to other apparatuses or devices through wired or wireless communication. For example, the electronic device includes Figure 4a When the image sensor is included in the example of the electronic device and the image sensor and the image processor are disposed on different devices, the other apparatuses or devices can be devices including the image processor. For example, a wireless communication solution including a WLAN (such as a Wi-Fi network), Bluetooth (BT), a global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, and the like can be provided and applied to the image processing apparatus. When the image processor apparatus is connected to other apparatuses or devices, the communication module 160 can also use a wired communication solution.

[0212] It can be appreciated that the structure illustrated by the embodiments of the present application does not constitute a specific limitation to the computing device. In other embodiments of the present application, the computing device can include more or fewer components than illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0213] Embodiments of the present application provide a non-volatile computer readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the above method.

[0214] Embodiments of the present application provide a computer program product comprising computer readable code, or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is run in a processor of an electronic device, the processor in the electronic device performs the above method.

[0215] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an electrically programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital video disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch cards or a groove having instructions recorded thereon, and any suitable combination of the above. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se.

[0216] Computer readable program instructions or code for carry out the operations described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adaptation card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium in the respective computing / processing device.

[0217] Computer readable program instructions for carrying out operations of the present application can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or any combination of source code or object code in any combination of one or more programming languages including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on a user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.

[0218] Various aspects of the present application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0219] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer readable storage medium having no data, programs, program modules, e.g., instructions for operation, or digital content stored thereon or therein for a brief period of time while the machine is in an

[0220] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0221] The flow diagrams and block diagrams in the attached figures illustrate the architecture, functionality, and operation of possible implementations of apparatuses, systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical functions (‘instruction(s)’). In some alternative implementations, the functions noted in the block can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by hardware, e.g., circuitry or ASICs (Application Specific Integrated Circuit), or by a combination of hardware and software, e.g., firmware or the like.

[0222] It is also noted that each of the blocks of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by hardware, e.g., circuitry or ASICs (Application Specific Integrated Circuit), or by a combination of hardware and software, e.g., firmware or the like.

[0223] Although the application has been described in connection with various embodiments, it will be understood that the application is capable of further modifications. These and other changes, along with the apparent alternatives and equivalents, fall within the scope of the claimed application. The description herein is intended to be illustrative only and is presented to enable any person skilled in the art to make and use the application. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope of the described application. The scope of the described application is not to be limited by the specific illustrative embodiments contained herein but only by the scope of the appended claims, which follow this disclosure.

[0224] Various embodiments of the application have been described in connection with the embodiments described above. The description is intended to be illustrative only and not limiting of the application. Many modifications and variations of the described embodiments are possible in light of the above teachings without departing from the scope of the described embodiments. The scope of the described embodiments is not to be limited by the specific illustrative embodiments described above but only by the scope of the claims that follow this disclosure.

Claims

1. An image processing method, characterized by, The method comprises: obtaining key pixel positions of a previous frame image and key pixel values of the previous frame image stored in a memory, the key pixel positions being positions of pixel points representing object features in a frame image, and the key pixel values being values of the pixel points at the key pixel positions; obtaining key pixel positions of a current frame image according to a gray image corresponding to the current frame image and the key pixel positions of the previous frame image; obtaining events of the current frame according to the gray image corresponding to the current frame image, the key pixel positions of the previous frame image and the key pixel values of the previous frame image, the events of the current frame indicating motion information of an object in the current frame image compared with the previous frame image; obtaining key pixel values of the current frame image according to the key pixel positions of the current frame image, and replacing the key pixel values of the previous frame image and the key pixel positions of the previous frame image stored in the memory by using the key pixel values of the current frame image and the key pixel positions of the current frame image.

2. The method of claim 1, wherein, The events of the current frame are obtained according to the gray image corresponding to the current frame image, the key pixel positions of the previous frame image and the key pixel values of the previous frame image, and the events of the current frame include: obtaining initial key pixel values of the current frame image according to the gray image corresponding to the current frame image and the key pixel positions of the previous frame image; obtaining the events of the current frame according to differences between the initial key pixel values of the current frame image and the key pixel values of the previous frame image.

3. The method of claim 2, wherein, The events of the current frame include common events and disappearance events, and the events of the current frame are obtained according to the differences between the initial key pixel values of the current frame image and the key pixel values of the previous frame image, and the events of the current frame include: pixel points common to the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image are taken as common positions, and a common event of each pixel point at the common positions is obtained when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than a first threshold value or less than a second threshold value; pixel points other than the common positions in the key pixel positions of the previous frame image are taken as disappearance positions, and a disappearance event of each pixel point at the disappearance positions is obtained when the difference between the initial key pixel value of the current frame image and the key pixel value of the previous frame image is greater than the first threshold value or less than the second threshold value.

4. The method according to any one of claims 1 to 3, characterized in that, The events of the current frame are obtained according to the gray image corresponding to the current frame image, the key pixel positions of the previous frame image and the key pixel values of the previous frame image, and the events of the current frame include: feature extraction is performed on the gray image corresponding to the current frame image to obtain initial key pixel positions of the current frame image; the events of the current frame are obtained according to a similarity degree between the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image.

5. The method of claim 4, wherein, The events of the current frame include new events, and the events of the current frame are obtained according to the similarity degree between the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image, and the events of the current frame include: for each region divided from the gray image corresponding to the current frame image, a ratio of a number of pixel points in an intersection of the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image to a number of pixel points in a union of the initial key pixel positions of the current frame image and the key pixel positions of the previous frame image is determined. When the ratio is less than a third threshold value, a new event of a pixel point corresponding to the initial key pixel position of the current frame image in the region is obtained.

6. The method according to any one of claims 1-3, characterized in that, The key pixel position of the current frame image is obtained according to the key pixel position of the previous frame image and the gray image corresponding to the current frame image, and the method comprises: feature extraction is performed on the gray image corresponding to the current frame image to obtain the initial key pixel position of the current frame image; temporal smoothing processing is performed on the initial key pixel position of the current frame image and the key pixel position of the previous frame image to obtain the key pixel position of the current frame image.

7. The method according to any one of claims 1-3, characterized in that, The method further comprises: The gray image corresponding to the current frame image is obtained according to the current frame image, and the resolution of the gray image corresponding to the current frame image is less than or equal to the resolution of the current frame image.

8. The method of claim 7, wherein, When the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the gray image corresponding to the current frame image is obtained according to the current frame image, and the method comprises: one or more of filtering processing, downsampling processing, interpolation processing and nonlinear transformation processing is performed on the current frame image to obtain the gray image corresponding to the current frame image.

9. The method of claim 7, wherein, When the resolution of the gray image corresponding to the current frame image is less than the resolution of the current frame image, the initial key pixel position of the current frame image is obtained by performing feature extraction on the gray image corresponding to the current frame image, and the method comprises: According to the resolution of the gray image corresponding to the current frame image and the resolution of the current frame image, the position correspondence relationship between the pixel points in the gray image corresponding to the current frame image and the pixel points in the current frame image is determined; According to the position correspondence relationship and the position of the pixel points included in the extracted features in the gray image corresponding to the current frame image, the initial key pixel position of the current frame image is obtained.

10. The method of any one of claims 1-3, wherein, The key pixel position of the current frame image is stored in the memory in the form of a binary image or in the form of a lossless compressed data packet.

11. The method of any one of claims 1-3, wherein, The method is applied to a sensor or a processor connected to the sensor, and the current frame image comprises a RAW image in a raw format collected by the sensor or a three-channel RGB image generated by the processor according to the image collected by the sensor, When the method is applied to a sensor, and the current frame image comprises a RAW image in a raw format collected by the sensor, the sensor outputs the event of the current frame while outputting the current frame image, or the sensor outputs the event of the current frame.

12. An image processing apparatus characterized by comprising: The device comprises: The first acquisition module is configured to acquire the key pixel position and the key pixel value of the previous frame image stored in the memory, wherein the key pixel position is the position of a pixel point representing an object feature in a frame image, and the key pixel value is the value of the pixel point at the key pixel position. The second acquisition module is configured to obtain the key pixel position of the current frame image according to the key pixel position of the previous frame image and the gray image corresponding to the current frame image. A third obtaining module, configured to obtain an event of a current frame according to a gray image corresponding to the current frame image, key pixel positions of a previous frame image, and key pixel values of the previous frame image, the event of the current frame indicating motion information of an object in the current frame image compared with the previous frame image; A first replacing module, configured to obtain key pixel values of the current frame image according to the key pixel positions of the current frame image, and replace the key pixel values and the key pixel positions of the previous frame image stored in the memory with the key pixel values and the key pixel positions of the current frame image.

13. An image processing apparatus characterized by comprising: Comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the method of any one of claims 1-11 when executing the instructions.

14. A non-transitory computer readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by the processor, implement the method of any one of claims 1-11.

15. A computer program product comprising computer readable code, or a non-transitory computer readable storage medium having computer readable code embodied thereon, the computer readable code comprising instructions for causing a computer to perform the method of any one of claims 1 to 14. When the computer readable code runs in the electronic device, the processor in the electronic device executes the method of any one of claims 1-11.

Citation Information

Patent Citations

  • Motion detection method based on edge detection and frame difference

    CN102307274A

  • Object identification method, device and storage medium

    CN108596128A