Image processing method, related device and system

By performing deep feature extraction of blurred images and event streams, clear images at any moment during the exposure time are restored, problems of difficulty in information loss and noise suppression in the prior art are solved, and image quality and analysis capabilities are improved.

CN114841870BActive Publication Date: 2025-05-06HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210334494.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-05-06
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

The prior art can only restore information at a specific point in time when restoring clear images in motion blur images, resulting in the loss of important scene information at a certain moment, and cannot effectively suppress the space-time noise of events, affecting image quality and analysis of motion scenes.

Method used

By extracting the blurred image and event stream respectively, preliminary features are obtained, and then deep feature extraction is performed to obtain the defuzzed image features and event stream features after suppressing event noise. Based on these features, the recovery of clear images at any moment during the exposure time is achieved.

Benefits of technology

It realizes the recovery of clear images at any moment in the exposure time, suppresses event feature noise, improves image quality and analysis capabilities of motion scenes, and avoids information loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841870B_ABST
    Figure CN114841870B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an image processing method, related devices and systems. The method includes: extracting features from a blurred image and an event stream respectively to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a change in scene brightness within an exposure time corresponding to the blurred image; performing deep feature extraction on the first feature of the blurred image and the first feature of the event stream to obtain a second feature of the blurred image and a second feature of the event stream; obtaining a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, wherein the target moment is any moment within the exposure time. By adopting this method, it is possible to suppress event feature noise and deblur image features, and to restore a clear image at any moment within the exposure time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, related devices and systems. Background Art

[0002] Motion deblurring is a traditional and important problem in the field of computer vision and photography. Motion blurring is the aliasing of information of different spatial positions caused by camera motion or scene changes during the imaging process, which leads to the loss of important high-frequency information in the image. However, motion blurred images usually contain complete information of the moving scene during the exposure time. Therefore, it is entirely possible to restore the motion state of an object and clear object texture from a blurred image. On the one hand, motion deblurring allows the restored image / video to be presented with better visual effects, allowing us to understand the useful information hidden in the blurred image and improve the visual experience of the end user. On the other hand, it helps to solve the problem of low-quality images being difficult to analyze, which has a positive role in promoting both the analysis of people and the improvement of the performance of other computer vision algorithms such as detection and tracking.

[0003] The prior art designs a CNN network, and uses a single motion blurred image and event data as input, and learns multiple clear video sequences through the supervision of synthetic data. In order to guide the network to perform semi-supervised training to improve the generalization ability of the model on real data, another CNN branch is designed to use the event stream as input and output optical flow information as the motion information of the scene. Through the physical blur formation process, based on the estimated clear image and motion information, a new blurred image is re-rendered, and the new blurred image and the blurred image as input establish a blur consistency cost function to guide the training of the CNN network in a self-supervised manner. In addition, the estimated clear image forms a new clear image through the transmission of motion information, and the new clear image and the clear image estimated by the network establish a self-supervised cost function of optical measurement consistency to further guide the training of the network. In the process of establishing these two cost functions, the piecewise linear motion model is used to estimate the high nonlinearity of motion, so that the network outputs accurate and high-density motion flow.

[0004] However, it can only restore the clear image corresponding to a specific time point, which will lose important scene information at a certain moment, thus affecting the analysis and judgment of the entire scene. Moreover, the spatiotemporal noise of the event cannot be well suppressed, and the event noise will misguide the network training, resulting in noise in the restored image where there is no blur; in addition, the noise will affect the estimation of the brightness change of the moving scene, resulting in wrong timestamps and the number of event points, which will also affect the network's estimation of the motion. Summary of the invention

[0005] The present application discloses an image processing method, related device and system, which can restore a clear image at any time within the exposure time and suppress the spatiotemporal noise of the event.

[0006] In a first aspect, an embodiment of the present application provides an image processing method, comprising: performing feature extraction on a blurred image and an event stream respectively to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a change in scene brightness within an exposure time corresponding to the blurred image; performing deep feature extraction on the first feature of the blurred image and the first feature of the event stream to obtain a second feature of the blurred image and a second feature of the event stream; and obtaining a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, wherein the target moment is any moment within the exposure time.

[0007] In the embodiment of the present application, preliminary feature extraction is performed on the blurred image and the event stream respectively to obtain the first feature of the blurred image and the first feature of the event stream, and then deep feature extraction is performed on the first feature of the blurred image and the first feature of the event stream to obtain the image feature after deblurring (i.e., the second feature of the blurred image) and the event stream feature after event noise suppression (i.e., the second feature of the event stream); then a clear image of the target moment is obtained based on the second feature of the blurred image and the second feature of the event stream. By adopting this method, event feature noise suppression and image feature deblurring can be achieved. By restoring a clear image at any time within the exposure time, the entire motion scene corresponding to the blurred image can be more completely understood and analyzed without missing any information at important moments.

[0008] In a possible implementation, the performing of depth feature extraction on the first feature of the blurred image and the first feature of the event stream to obtain the second feature of the blurred image and the second feature of the event stream includes: obtaining depth feature parameters of the blurred image and the depth feature parameters of the event stream according to the first feature of the blurred image and the first feature of the event stream, respectively; and interactively processing the depth feature parameters of the blurred image and the depth feature parameters of the event stream to obtain the second feature of the blurred image and the second feature of the event stream.

[0009] This embodiment utilizes the ultra-high temporal resolution of event information in a complementary manner to solve the image deblurring problem at any time within the exposure time, and at the same time utilizes the information smoothness of the blurred image to suppress noise in the event, thereby achieving a more accurate estimation of the scene dynamics within the exposure time.

[0010] In a possible implementation, the method of obtaining a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, where the target moment is any moment within the exposure time, includes: encoding the target moment to obtain a time vector; obtaining an image feature that fuses time information according to the time vector, the second feature of the blurred image, and the second feature of the event stream; and decoding the image feature that fuses time information to obtain a clear image at the target moment.

[0011] This embodiment encodes continuous time signals to achieve the fusion of any continuous time information and depth features; by adopting the MLP network structure of continuous time decoding, a clear image corresponding to any moment within the exposure time is decoded from the deep features of the fused time information, so that the entire motion scene corresponding to the blurred image can be more completely understood and analyzed without missing any information of important moments.

[0012] In a second aspect, an embodiment of the present application provides an image processing device, comprising: a first extraction module, used to perform feature extraction on a blurred image and an event stream respectively, so as to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a change in scene brightness within an exposure time corresponding to the blurred image; a second extraction module, used to perform deep feature extraction on the first feature of the blurred image and the first feature of the event stream, so as to obtain a second feature of the blurred image and a second feature of the event stream; a processing module, used to obtain a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, wherein the target moment is any moment within the exposure time.

[0013] In a possible implementation, the second extraction module is used to: obtain depth feature parameters of the blurred image and depth feature parameters of the event stream according to the first feature of the blurred image and the first feature of the event stream, respectively; and interactively process the depth feature parameters of the blurred image and the depth feature parameters of the event stream to obtain second features of the blurred image and second features of the event stream.

[0014] In one possible implementation, the processing module is used to: encode the target moment to obtain a time vector; obtain image features that fuse time information based on the time vector, the second feature of the blurred image, and the second feature of the event stream; and decode the image features that fuse time information to obtain a clear image of the target moment.

[0015] In a third aspect, the present application provides an image processing device, comprising a processor and a memory; wherein the memory is used to store program code, and the processor is used to call the program code to execute a method provided in any possible implementation manner of the first aspect.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium, comprising computer instructions, which, when executed on an electronic device, enables the electronic device to execute a method provided in any possible implementation manner of the first aspect.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute a method provided in any possible implementation manner of the first aspect.

[0018] It can be understood that the device described in the second aspect, the device described in the third aspect, the computer-readable storage medium described in the fourth aspect, or the computer program product described in the fifth aspect provided above are all used to execute any method provided in the first aspect. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The following is an introduction to the drawings used in the embodiments of the present application.

[0020] Figure 1 is a schematic diagram of the architecture of an image processing system provided in an embodiment of the present application;

[0021] Figure 2 It is a flowchart of an image processing method provided in an embodiment of the present application;

[0022] Figure 3a is a flowchart of another image processing method provided in an embodiment of the present application;

[0023] Figure 3b It is a schematic diagram of a model processing provided in an embodiment of the present application;

[0024] Figure 3c is a schematic diagram of a network structure of a dual attention with an implicit structure provided in an embodiment of the present application;

[0025] Figure 4a is a schematic diagram of another image processing method provided in an embodiment of the present application;

[0026] Figure 4b It is a schematic diagram of a network structure provided in an embodiment of the present application;

[0027] Figure 4cThis is another schematic diagram of a network structure provided in an embodiment of the present application;

[0028] Figure 5 is a structural schematic diagram of an image processing device provided in an embodiment of the present application;

[0029] Figure 6 It is a structural schematic diagram of another image processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method part of the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0031] Reference Figure 1 , is a schematic diagram of the architecture of an image processing system provided by an embodiment of the present application. The system may include an image processing device and an autonomous driving vehicle / robot. The image processing device is used to process a blurred image and an event stream triggered by scene brightness changes during the exposure time corresponding to the blurred image, and obtain a clear image corresponding to any time during the exposure time. Based on multiple clear images corresponding to any time, a clear video sequence can be obtained, which can then be processed in an autonomous driving vehicle / robot to achieve purposes such as face recognition, semantic segmentation, object detection, or 3D reconstruction.

[0032] The image processing device may be a server or any terminal device, for example, it may be located inside the autonomous driving vehicle or robot, or connected to the outside of the autonomous driving vehicle or robot, for performing image processing.

[0033] This embodiment is only described by taking an autonomous driving vehicle or a robot as an example, and it may also be other equipment, etc., and this solution does not make specific limitations on this.

[0034] Reference Figure 2 As shown in FIG. 1 , it is a flow chart of an image processing method provided by an embodiment of the present application. Figure 2 As shown, the method includes steps 201-203, which are specifically as follows:

[0035] 201. Perform feature extraction on a blurred image and an event stream respectively to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a brightness change of a scene within an exposure time corresponding to the blurred image;

[0036] The blurred image can be understood as a blurred image obtained due to motion. For example, when an object is moving, a blurred image is obtained when the object is photographed.

[0037] The feature extraction may be preliminary processing of the blurred image and the event stream respectively, for example, it may be a graph feature.

[0038] In a possible implementation, the feature extraction of the blurred image and the event stream may be performed by inputting both the blurred image and the event stream into a preset neural network model for processing, thereby obtaining the first feature of the blurred image and the first feature of the event stream. Other means may also be used, which are not specifically limited in this solution.

[0039] 202. Perform deep feature extraction on the first feature of the blurred image and the first feature of the event stream to obtain a second feature of the blurred image and a second feature of the event stream;

[0040] The second feature of the blurred image can be understood as a deblurred image feature obtained by processing the first feature of the blurred image.

[0041] The second feature of the event stream can be understood as the event stream feature after suppressing event noise obtained by processing the first feature of the event stream.

[0042] In a possible implementation, the deep feature extraction of the first feature of the blurred image and the first feature of the event stream may be performed by inputting the first feature of the blurred image and the first feature of the event stream into a preset neural network model for processing, thereby obtaining the second feature of the blurred image and the second feature of the event stream. Other means may also be used, which are not specifically limited in this solution.

[0043] 203. Obtain a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, where the target moment is any moment within the exposure time.

[0044] In one possible implementation, based on the deblurred image features (i.e., the second features of the blurred image) and the event stream features after event noise suppression (i.e., the second features of the event stream) obtained above, time information is added to the second features of the blurred image and the second features of the event stream, and then a clear image of the target moment is obtained from the features with time information.

[0045] In the embodiment of the present application, preliminary feature extraction is performed on the blurred image and the event stream respectively to obtain the first feature of the blurred image and the first feature of the event stream, and then deep feature extraction is performed on the first feature of the blurred image and the first feature of the event stream to obtain the deblurred image feature (i.e., the second feature of the blurred image) and the event stream feature after event noise suppression (i.e., the second feature of the event stream); then a clear image of the target moment is obtained based on the second feature of the blurred image and the second feature of the event stream. By adopting this method, event feature noise suppression and image feature deblurring can be achieved. By restoring a clear image at any time within the exposure time, the entire motion scene corresponding to the blurred image can be more completely understood and analyzed without missing any information at important moments. This solution can restore more image details with higher clarity. This solution can greatly promote the performance of other high-level computer vision algorithms.

[0046] Reference Figure 3a As shown in FIG. 1 , it is a flow chart of another image processing method provided by an embodiment of the present application. Figure 3a As shown, the method includes steps 301-304, which are specifically as follows:

[0047] 301. Perform feature extraction on a blurred image and an event stream respectively to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a brightness change of a scene within an exposure time corresponding to the blurred image;

[0048] The blurred image can be understood as a blurred image obtained due to motion. For example, when an object is moving, a blurred image is obtained when the object is photographed.

[0049] The feature extraction may be preliminary processing of the blurred image and the event stream respectively, for example, it may be a graph feature.

[0050] In a possible implementation, the feature extraction is performed on the blurred image and the event stream respectively, and both the blurred image and the event stream are input into a preset neural network model for processing, thereby obtaining a first feature of the blurred image and a first feature of the event stream.

[0051] Specifically, Figure 3b As shown in FIG. 1 , a schematic diagram of a shallow feature extraction network model processing provided in an embodiment of the present application is shown. Figure 3b As shown, the shallow feature extraction network model SFE includes a pixel shuffle operation reshf layer and two convolution layers, wherein the observable blurred image B and the high temporal resolution event stream ε are respectively input into the shallow feature extraction network model, and the two signals are respectively passed through two convolution layers to obtain the shallow feature B feat(i.e. the first characteristic of the blurred image) and E feat (That is, the first characteristic of event flow).

[0052] This shallow feature extraction network model can make the subsequent transformer network training more stable when obtaining the second feature of the blurred image and the second feature of the event stream; at the same time, it can make inputs of different scales produce features of the same scale, which is conducive to the subsequent transformer feature embedding.

[0053] Other means may also be used, and this solution does not specifically limit this.

[0054] 302. Obtain a depth feature parameter of the blurred image and a depth feature parameter of the event stream respectively according to the first feature of the blurred image and the first feature of the event stream;

[0055] In one possible implementation, Figure 3c As shown, the embodiment of the present application also provides a network structure with a double attention with an implicit structure. Figure 3c As shown in Figure 1, this network structure enables the two input signals to interact with each other, thereby suppressing event noise and deblurring image features. The network structure is composed of multiple stacked DALS blocks, each of which takes two features extracted by a shallow feature extraction network or the output of the previous DALS as input.

[0056] Specifically, the network structure firstly analyzes the first feature B of the blurred image feat and the first feature E of the event stream feat The features are embedded through the residual dense block (RDB) respectively, and then the Query (Q), Key (K) and Value (V) required to calculate the respective attention are encoded through the linear layer. The attention of the two signals is calculated by the following methods:

[0057]

[0058] where d k is the dimension of a single feature in Q, K, V, and W-MSA(Q,K,V) is the feature representation in the middle of the DALS block.

[0059] Based on this processing, the depth feature parameters of the blurred image and the depth feature parameters of the event stream can be obtained. This embodiment is described by taking the depth feature parameters as the attention parameters as an example.

[0060] It may also be other parameters obtained through processing by other network models, and this solution does not make any specific limitation on this.

[0061] 303. Interactively process the depth feature parameter of the blurred image and the depth feature parameter of the event stream to obtain a second feature of the blurred image and a second feature of the event stream;

[0062] After obtaining the attention of each signal, the two signals are interactively processed in the following two ways.

[0063] 1) If Figure 3c As shown, the event attention suppresses noise by correcting the attention of the blurred image, and then outputs deep features through MLP:

[0064] Attn E ←Attn E +Attn B

[0065]

[0066]

[0067] Attn E and Attn B are the attention of the event signal and the blurred image respectively, V· is the value calculation operator, and E feat Refers to the deep features (i.e., the second features) of the event signal.

[0068] 2) In order to eliminate the ambiguity of image features, the image depth feature B feat (i.e., the second feature) is calculated from the features of the two signals:

[0069]

[0070]

[0071] This scheme uses the ultra-high temporal resolution of event information in a complementary way to solve the image deblurring problem at any time within the exposure time, and at the same time uses the information smoothness of the blurred image to suppress noise in the event, thereby achieving a more accurate estimation of the scene dynamics within the exposure time.

[0072] Based on the above processing, the deblurred image features (ie, the second features of the blurred image) and the event stream features after event noise suppression (ie, the second features of the event stream) can be obtained.

[0073] 304. Obtain a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, where the target moment is any moment within the exposure time.

[0074] In one possible implementation, based on the deblurred image features (i.e., the second features of the blurred image) and the event stream features after event noise suppression (i.e., the second features of the event stream) obtained above, time information is added to the second features of the blurred image and the second features of the event stream, and then a clear image of the target moment is obtained from the features with time information.

[0075] In the embodiment of the present application, preliminary feature extraction is performed on the blurred image and the event stream respectively to obtain the first feature of the blurred image and the first feature of the event stream, and then the first feature of the blurred image and the first feature of the event stream are input into the dual feature extraction model for deep feature extraction to obtain the deblurred image feature (i.e., the second feature of the blurred image) and the event stream feature after event noise suppression (i.e., the second feature of the event stream); then a clear image of the target moment is obtained based on the second feature of the blurred image and the second feature of the event stream. By adopting this method, event feature noise suppression and image feature deblurring can be achieved. By restoring a clear image at any time within the exposure time, the entire motion scene corresponding to the blurred image can be more completely understood and analyzed without missing any information at important moments.

[0076] Reference Figure 4a FIG. 4 is a schematic diagram of another image processing method provided by an embodiment of the present application. The method includes steps 401-406, which are as follows:

[0077] 401. Perform feature extraction on a blurred image and an event stream respectively to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a scene brightness change within an exposure time corresponding to the blurred image;

[0078] For an introduction to this section, see Figure 4b The description of the aforementioned embodiments will not be repeated here.

[0079] 402. Obtain a depth feature parameter of the blurred image and a depth feature parameter of the event stream respectively according to the first feature of the blurred image and the first feature of the event stream;

[0080] like Figure 4a Dual feature extraction network model in f γ , through the dual feature extraction network model f γ The deep features of the scene clear image and motion state within the exposure time are extracted from the input observable blurred image B and the event stream ε.

[0081] For an introduction to this section, see Figure 4b The description of the aforementioned embodiments will not be repeated here.

[0082] 403. Interactively process the depth feature parameter of the blurred image and the depth feature parameter of the event stream to obtain a second feature of the blurred image and a second feature of the event stream;

[0083] For the introduction of this part, please refer to the aforementioned embodiment, which will not be described again here.

[0084] Based on the above processing, the deblurred image features (ie, the second features of the blurred image) and the event stream features after event noise suppression (ie, the second features of the event stream) can be obtained.

[0085] 404. Encode the target time to obtain a time vector;

[0086] like Figure 4a As shown, It is a clear image decoding network that integrates continuous time information. By inputting the extracted deep features and the corresponding time t (or the encoding information of t) into the clear image decoding network that integrates continuous time information, a clear image of the scene corresponding to time t is decoded, where t is the time information of any moment in the exposure event. In the process of network training, as the continuous time t changes continuously, the clear image obtained by the network also has the corresponding time sequence information of t, so that in the network inference stage, the image of the corresponding clear scene can be output according to any input time t (within the exposure time).

[0087] Should Figure 4a The corresponding processing methods are as follows:

[0088] I(t)=B+φ θ (t;f γ (B, ε τ ))

[0089]

[0090] Where x and p represent the position and polarity of the event trigger respectively.

[0091] In one possible implementation, Figure 4c As shown, the network structure for clear image decoding that integrates continuous time information provided by the embodiment of the present application is shown. Figure 4c As shown in the figure, the deep features extracted by DALS are processed by multiple convolutional layers, and the encoded time signal is integrated into the deep features output by the convolutional layer, and then a clear image at time t is decoded by an MLP.

[0092] In a possible implementation, the exposure time is first normalized. For example, if the exposure time is 0.1 s, it is normalized so that it is 1 after normalization. Based on this operation, the normalized time of any time in the exposure time can be obtained.

[0093] Then, the target time t can be encoded into a time vector using Fourier encoding. The time vector can be a multi-dimensional time vector, for example, a 2L-dimensional vector, where L can be any positive integer, which is not specifically limited in this solution.

[0094] For example, the multidimensional time vector can be expressed as:

[0095] n(t)=(cos(2 0 πt), sin(2 0 πt),…,cos(2 L-1 πt), sin(2 L-1 πt)).

[0096] 405. Obtain an image feature that integrates time information according to the time vector, the second feature of the blurred image, and the second feature of the event stream;

[0097] like Figure 4c As shown, the encoded time vector and the obtained deblurred image features (i.e., the second features of the blurred image) and the event stream features after event noise suppression (i.e., the second features of the event stream) are cascaded to achieve pixel-by-pixel cascading of the time information into the above features.

[0098] 406. Decode the image features fused with the time information to obtain a clear image at the target moment.

[0099] Among them, the clear image at time t can be obtained by decoding the clear image at that moment through the pixel-by-pixel MLP network structure.

[0100] In a possible implementation, multiple clear images at any time can also be obtained, which is not specifically limited in this solution.

[0101] In the embodiment of the present application, preliminary feature extraction is performed on the blurred image and the event stream respectively to obtain the first feature of the blurred image and the first feature of the event stream, and then the first feature of the blurred image and the first feature of the event stream are input into the dual feature extraction model for deep feature extraction to obtain the deblurred image feature (i.e., the second feature of the blurred image) and the event stream feature after event noise suppression (i.e., the second feature of the event stream); then the obtained second feature of the blurred image and the second feature of the event stream are input into the clear image decoding network that integrates continuous time information, thereby obtaining a clear image at the target moment. By adopting this method, event feature noise suppression and image feature deblurring can be achieved. By restoring a clear image at any time within the exposure time, the entire motion scene corresponding to the blurred image can be more completely understood and analyzed without missing any information at important moments.

[0102] On the basis of the foregoing embodiments, the embodiments of the present application provide a model training method. Among them, the model describes the relationship between the input blurred image, the event stream and the clear image of the scene at any time during the exposure time by learning an implicit neural representation INR. The learned INR can be called an implicit video function (IVF). First, the deep features of the clear image and motion state of the scene during the exposure time are extracted from the input observable blurred image B and the event stream ε through a dual feature extraction network, and then the extracted deep features and the corresponding time t (or the encoding information of t) are input into the continuous time decoding network to decode the clear image of the corresponding scene at time t. During the network training process, as the continuous time t continues to change, the clear image obtained by the network also has the corresponding time sequence information of t, so that the corresponding clear scene can be output according to the input arbitrary time t (within the exposure time) in the network reasoning stage.

[0103] This embodiment jointly trains the dual feature extraction network f in an end-to-end manner. γ and MLP continuous-time decoding network Because this embodiment only sparsely samples a limited number of high-definition images within the exposure time as the correct labeled ground truth for supervision, the cost function at the sampling moment and the non-sampling moment are constructed differently.

[0104] The cost function at the sampling time directly generates a clear image I(t i ) and ground The difference between:

[0105]

[0106] Since the ground truth at the non-sampling time is unknown, the optical flow estimation network EV-Flow(·) is first used to estimate the optical flow from the event information. Then calculate the ground truth at any time. The specific process is shown in the following formula:

[0107]

[0108] After obtaining the clear image at the required time, construct the cost function of the one-norm in the same way:

[0109]

[0110] Combining the sampling time and non-sampling time, the total cost function can be expressed as:

[0111]

[0112] This embodiment is only introduced by taking the above implementation method as an example, and it can also be other methods, which are not specifically limited in this solution.

[0113] Reference Figure 5 As shown in FIG. 1 , an image processing device is provided in an embodiment of the present application. Figure 5 As shown, the device includes a first extraction module 501, a second extraction module 502 and a processing module 503, wherein:

[0114] A first extraction module 501 is used to perform feature extraction on the blurred image and the event stream respectively to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a scene brightness change within an exposure time corresponding to the blurred image;

[0115] A second extraction module 502, configured to perform deep feature extraction on the first feature of the blurred image and the first feature of the event stream to obtain a second feature of the blurred image and a second feature of the event stream;

[0116] The processing module 503 is used to obtain a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, where the target moment is any moment within the exposure time.

[0117] Optionally, the second extraction module 502 is used to:

[0118] Obtaining a depth feature parameter of the blurred image and a depth feature parameter of the event stream respectively according to the first feature of the blurred image and the first feature of the event stream;

[0119] The depth feature parameters of the blurred image and the depth feature parameters of the event stream are interactively processed to obtain a second feature of the blurred image and a second feature of the event stream.

[0120] Optionally, the processing module 503 is used to:

[0121] Encoding the target time to obtain a time vector;

[0122] Obtaining an image feature that integrates time information according to the time vector, the second feature of the blurred image, and the second feature of the event stream;

[0123] The image features fused with the time information are decoded to obtain a clear image at the target moment.

[0124] The specific functional implementation of the image processing device can refer to the description of the above-mentioned image processing method, which will not be repeated here. Each unit or module in the device can be separately or completely combined into one or several other units or modules to constitute, or one (some) of the units or modules can be further divided into multiple functionally smaller units or modules to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above-mentioned units or modules are divided based on logical functions. In actual applications, the function of a unit (or module) can also be implemented by multiple units (or modules), or the functions of multiple units (or modules) can be implemented by one unit (or module).

[0125] Based on the description of the above method embodiment and device embodiment, an embodiment of the present invention further provides an image processing device. Figure 6 A schematic structural diagram of an image processing device provided in an embodiment of the invention. Figure 6 The image processing device 600 shown includes a memory 601 , a processor 602 , a communication interface 603 and a bus 604 . The memory 601 , the processor 602 , and the communication interface 603 are connected to each other through the bus 604 .

[0126] The memory 601 may be a read-only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM).

[0127] The memory 601 can store programs. When the program stored in the memory 601 is executed by the processor 602, the processor 602 executes each step of the image processing method of the embodiment of the present application through the communication interface 603.

[0128] The processor 602 can adopt a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits to execute relevant programs to implement the functions required to be performed by the units in the image processing device of the embodiment of the present application, or to execute the image processing method of the method embodiment of the present application.

[0129] The processor 602 may also be an integrated circuit chip having signal processing capabilities. In the implementation process, each step of the image processing method of the present application may be completed by an integrated logic circuit of hardware in the processor 602 or by instructions in the form of software. The above-mentioned processor 602 may also be a CPU, a digital signal processor (Digital Signal Processing, DSP), an ASIC, a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application may be directly embodied as being executed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in a decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 601, and the processor 602 reads the information in the memory 601, and combines its hardware to complete the functions required to be performed by the units included in the image processing device of the embodiment of the present application, or executes the image processing method of the method embodiment of the present application.

[0130] The communication interface 603 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the image processing device 600 and other devices or a communication network. For example, data can be acquired through the communication interface 603 .

[0131] The bus 604 may include a path for transmitting information between various components of the image processing apparatus 600 (eg, the memory 601 , the processor 602 , and the communication interface 603 ).

[0132] It should be noted that although Figure 6 The image processing device 600 shown in the figure only shows a memory, a processor, and a communication interface. However, in the specific implementation process, those skilled in the art should understand that the image processing device 600 also includes other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the image processing device 600 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the image processing device 600 may also only include the devices necessary to implement the embodiments of the present application, and does not necessarily include Figure 6 All devices shown in .

[0133] An embodiment of the present application further provides a chip, which includes a processor and a data interface. The processor reads instructions stored in a memory through the data interface to implement the image processing method.

[0134] In a possible implementation, the chip may further include a memory, in which instructions are stored, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the image processing method.

[0135] An embodiment of the present application also provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is executed on a computer or a processor, the computer or the processor executes one or more steps in any of the above methods.

[0136] The embodiment of the present application further provides a computer program product including instructions. When the computer program product is executed on a computer or a processor, the computer or the processor executes one or more steps in any of the above methods.

[0137] Those skilled in the art will appreciate that the functions described in conjunction with the various illustrative logic blocks, modules, and algorithm steps disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions described in the various illustrative logic blocks, modules, and steps can be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to tangible media, such as data storage media, or includes any media (e.g., based on a communication protocol) that facilitates the transfer of a computer program from one place to another. In this manner, computer-readable media may generally correspond to (1) non-temporary tangible computer-readable storage media, or (2) communication media, such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in this application. A computer program product may include computer-readable media.

[0138] As an example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other media that can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio and microwaves, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio and microwaves are included in the definition of media. However, it should be understood that the computer-readable storage media and data storage media do not include connections, carriers, signals or other temporary media, but are actually directed to non-temporary tangible storage media. As used herein, disks and optical discs include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media. Instructions may be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the aforementioned structures or any other structures suitable for implementing the techniques described herein. In addition, in some aspects, the functions described by the various illustrative logic blocks, modules, and steps described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Moreover, the techniques may be fully implemented in one or more circuits or logic elements.

[0139] The techniques of the present application may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in the present application to emphasize the functional aspects of the devices for performing the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. In fact, as described above, the various units may be combined in an encoded hardware unit in conjunction with appropriate software and / or firmware, or provided by interoperating hardware units (including one or more processors as described above).

[0140] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the specific descriptions of the corresponding steps in the aforementioned method embodiments, and will not be repeated here.

[0141] It should be understood that in the description of the present application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; wherein A and B can be singular or plural. Also, in the description of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, wherein a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second", etc. are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not limit them to be necessarily different. Meanwhile, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.

[0142] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the division of the unit is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling, direct coupling, or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0143] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0144] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from a website site, a computer, a server or a data center to another website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, DSL) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integration. The available medium can be ROM, or RAM, or magnetic media, for example, a floppy disk, a hard disk, a tape, a disk, or an optical medium, for example, a digital versatile disc DVD, or a semiconductor medium, for example, a solid state disk (SSD), etc.

[0145] The above is only a specific implementation of the embodiment of the present application, but the protection scope of the embodiment of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed in the embodiment of the present application should be included in the protection scope of the embodiment of the present application. Therefore, the protection scope of the embodiment of the present application should be based on the protection scope of the claims.

Claims

1. An image processing method, characterized in that: include: Performing feature extraction on the blurred image and the event stream respectively to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a brightness change of the scene within an exposure time corresponding to the blurred image; Performing deep feature extraction on the first feature of the blurred image and the first feature of the event stream to obtain a second feature of the blurred image and a second feature of the event stream; A clear image at a target moment is obtained according to the second feature of the blurred image and the second feature of the event stream, where the target moment is any moment within the exposure time.

2. The method according to claim 1, characterized in that The performing deep feature extraction on the first feature of the blurred image and the first feature of the event stream to obtain the second feature of the blurred image and the second feature of the event stream includes: Obtaining a depth feature parameter of the blurred image and a depth feature parameter of the event stream respectively according to the first feature of the blurred image and the first feature of the event stream; Interactively processing the depth feature parameters of the blurred image and the depth feature parameters of the event stream to obtain a second feature of the blurred image and a second feature of the event stream; The interactive processing of the depth feature parameter of the blurred image and the depth feature parameter of the event stream to obtain the second feature of the blurred image and the second feature of the event stream includes: Correcting the depth feature parameters of the event stream based on the depth feature parameters of the blurred image to obtain processed depth feature parameters of the event stream; Inputting the processed deep feature parameters of the event stream into a multi-layer perceptron MLP for processing to obtain a second feature of the event stream; Processing is performed based on the processed deep feature parameters of the event stream and the deep feature parameters of the blurred image and the multi-layer perceptron MLP to obtain a second feature of the blurred image.

3. The method according to claim 1 or 2, characterized in that: The obtaining a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, wherein the target moment is any moment within the exposure time, comprises: Encoding the target time to obtain a time vector; Obtaining an image feature that integrates time information according to the time vector, the second feature of the blurred image, and the second feature of the event stream; The image features fused with the time information are decoded to obtain a clear image at the target moment.

4. An image processing device, characterized in that: include: A first extraction module is used to perform feature extraction on the blurred image and the event stream respectively to obtain a first feature of the blurred image and a first feature of the event stream, wherein the event stream is triggered by a scene brightness change within an exposure time corresponding to the blurred image; A second extraction module, configured to perform deep feature extraction on the first feature of the blurred image and the first feature of the event stream to obtain a second feature of the blurred image and a second feature of the event stream; A processing module is used to obtain a clear image at a target moment according to the second feature of the blurred image and the second feature of the event stream, wherein the target moment is any moment within the exposure time.

5. The device according to claim 4, characterized in that The second extraction module is used to: Obtaining a depth feature parameter of the blurred image and a depth feature parameter of the event stream respectively according to the first feature of the blurred image and the first feature of the event stream; Interactively processing the depth feature parameters of the blurred image and the depth feature parameters of the event stream to obtain a second feature of the blurred image and a second feature of the event stream; The interactive processing of the depth feature parameter of the blurred image and the depth feature parameter of the event stream to obtain the second feature of the blurred image and the second feature of the event stream includes: Correcting the depth feature parameters of the event stream based on the depth feature parameters of the blurred image to obtain processed depth feature parameters of the event stream; Inputting the processed deep feature parameters of the event stream into a multi-layer perceptron MLP for processing to obtain a second feature of the event stream; Processing is performed based on the processed deep feature parameters of the event stream and the deep feature parameters of the blurred image and the multi-layer perceptron MLP to obtain a second feature of the blurred image.

6. The device according to claim 4 or 5, characterized in that The processing module is used for: Encoding the target time to obtain a time vector; Obtaining an image feature that integrates time information according to the time vector, the second feature of the blurred image, and the second feature of the event stream; The image features fused with the time information are decoded to obtain a clear image at the target moment.

7. An image processing device, characterized in that: The method comprises a processor and a memory; wherein the memory is used to store program codes, and the processor is used to call the program codes to execute the method according to any one of claims 1 to 3.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 3.

9. A computer program product, characterized in that When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Monocular image depth estimation method and system

    CN107204010A

  • Infrared image enhanced road airborne scene semantic segmentation method

    CN113012175A