Image processing method and device, electronic equipment and storage medium

By fusing information from frame cameras and event cameras, and using the Bezier curve algorithm and feature stitching to generate motion trajectories, the problem of image quality degradation in high-speed motion scenes by frame cameras is solved, and efficient reconstruction of complex motion scenes is achieved.

CN120147936BActive Publication Date: 2025-10-24JILIN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510615291.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-10-24
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Frame cameras have limitations in capturing subtle changes in high-speed motion, and event cameras lack description of absolute brightness information and spatial details, resulting in degraded image quality and limiting their application in motion analysis and image reconstruction.

Method used

By acquiring image frames and event information from frame cameras and event cameras, performing feature extraction and stitching, and using the Bezier curve algorithm to generate the motion trajectory of physical points, combined with event directed graphs and voxel processing, accurate and efficient reconstruction of complex motion scenes can be achieved.

Benefits of technology

Multiple target image frames that are continuous and clear from a visual perspective are generated. By making full use of the complementary advantages of frame cameras and event cameras, the limitations of a single sensor under high-speed motion and extreme lighting conditions are overcome, thereby improving the accuracy and computational efficiency of motion trajectory generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147936B_ABST
    Figure CN120147936B_ABST
Patent Text Reader

Abstract

The application provides an image processing method and device, electronic equipment and storage medium. The method acquires image frames collected by a frame camera and event information collected by an event camera in a target scene in a target period, which can ensure that the image information and the event information are consistent in space and time. Sampling image features matched with each physical pixel from an image feature map of the image frames, and splicing the pixel features of each physical pixel with the matched image features can fully utilize the complementary advantages of the frame camera and the event camera, enhance the space-time representation capability of the fused features, and overcome the limitations of a single sensor in feature acquisition. Using the fused features and a Bezier curve algorithm, a motion trajectory of at least one physical point is generated, and discrete event information is converted into a smooth motion path. Based on the motion trajectory and a background image region on the image frames, precise and efficient reconstruction of a complex motion scene is realized, and multiple image frames that are continuous and clear in a visual angle are generated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electronic devices, and more particularly, to an image processing method and device, an electronic device, and a storage medium. BACKGROUND

[0002] In recent years, with the rapid development of computer vision, image processing and artificial intelligence technology, motion analysis and image reconstruction have attracted widespread attention in the fields of robot navigation, augmented reality, autonomous driving, etc.

[0003] Currently, frame cameras rely on fixed frame rate and exposure time, which makes frame cameras have significant limitations when capturing subtle changes in high-speed motion. When the moving speed of objects in the environment exceeds the sampling rate of the frame camera, some image details in the frame images collected by the frame camera are lost. In addition, in a high dynamic range environment, due to limited exposure, frame cameras cannot clearly capture scenes with a large range of brightness at the same time, resulting in a decline in image quality. The above problems limit the application of frame images in motion analysis and image reconstruction.

[0004] In addition, as a new type of sensor, event cameras, unlike the above-mentioned frame cameras, can capture the polarity and time of changes in environmental brightness with a time resolution of microseconds, that is, event cameras can focus on subtle changes in high-speed motion. However, it lacks the description of absolute brightness information and spatial details.

[0005] Therefore, how to fuse frame images and event information collected by event cameras, make full use of the complementary advantages of the two, and realize accurate and efficient reconstruction of complex motion scenes has become a key technical problem to be solved. SUMMARY

[0006] The present application provides an image processing method and device, an electronic device and a storage medium, which can make full use of the complementary advantages of frame cameras and event cameras, realize accurate and efficient reconstruction of complex motion scenes, and generate multiple target image frames that are continuous and clear in visual angle.

[0007] In a first aspect, an image processing method is provided, which includes: obtaining, in a target period, an image frame captured by a frame camera and event information captured by an event camera in a target scene, the event information being information of a plurality of physical pixels in the event camera that have a change in luminance; performing feature extraction on the image frame to obtain an image feature map, sampling image features matched with each physical pixel in the event information from the image feature map, and splicing pixel features of each physical pixel and the matched image features to obtain each target feature; generating a motion trajectory of at least one physical point based on each target feature and a Bezier curve algorithm, and generating a target image frame based on the motion trajectory of the at least one physical point and a background image region on the image frame, the plurality of physical pixels being used to reflect a change in position of at least one physical point in the target scene in the target period.

[0008] In the above technical solution, the image frame captured by the frame camera and the event information captured by the event camera in the target period in the same target scene are obtained, which can ensure that the image information (image frame) and the event information are consistent in space and time. At the same time, the image frame can reflect rich static texture information in the target scene, and the event information can focus on subtle changes in the target scene. Further, the image frame is subjected to feature extraction to obtain an image feature map, and image features matched with each physical pixel in the event information are sampled from the image feature map, and pixel features of each physical pixel are spliced with the matched image features to obtain target features (fusion features). This can simultaneously focus on static texture information and subtle changes in the target scene, that is, the complementary advantages of the frame camera and the event camera can be fully utilized to enhance the spatiotemporal representation capability of the target features and overcome the limitations of a single sensor in feature acquisition in some scenes (such as high-speed motion or extreme illumination). The motion trajectory of at least one physical point is generated by using each target feature and the Bezier curve algorithm, which can convert discrete event information into a smooth motion path. Based on the motion trajectory and the background image region on the image frame, precise and efficient reconstruction of a complex motion scene can be achieved, and a plurality of target image frames that are continuous and clear in a visual angle are generated. The joint use of the frame camera and the event camera can meet the replacement demand for high-quality cameras.

[0009] In some possible implementation manners, the method further includes: constructing an event directed graph based on the pixel positions and corresponding times of the physical pixels in the event information, wherein a node position of each event node in the event directed graph is related to a pixel position, and a direction of a directed edge between two event nodes is related to corresponding times of the two event nodes; determining a node feature of each event node based on the node position of the event node, a change polarity of the physical pixel corresponding to the event node, and an edge feature of the directed edge corresponding to the event node, N being a positive integer greater than 2; and sampling, from the image feature map, an image feature matching each physical pixel in the event information, and splicing the pixel feature of each physical pixel and the matching image feature to obtain each target feature, including: sampling, from the image feature map, an image feature matching the node position and corresponding time of each event node; and splicing the node feature of each event node and the matching image feature to obtain each target feature.

[0010] In the technical solution, the event directed graph is constructed based on the pixel positions and corresponding times of the physical pixels, the causality (such as an object moving path) of event propagation is encoded by using the time directionality (such as a timestamp sequence) between two event nodes, and the edge feature is combined, which can reflect the correlation of multiple physical pixels in space-time. The node feature is dynamically generated based on the node position of the event node, the change polarity of the corresponding physical pixel, and the edge feature of the corresponding directed edge, which can accurately determine the node feature. Further, after sampling by space-time alignment, the node feature of the event node is spliced with the frame image feature (the matching image feature), which can fuse the high-frequency response capability of the event camera and the rich texture information of the frame camera, and overcome the information limitation of a single sensor. The scheme can significantly improve the accuracy of motion trajectory generation in a dynamic scene by using the fusion mechanism of the directed graph structure and the cross-modal features.

[0011] In a possible implementation of the first aspect and the above implementation, based on the node position of each event node in the N event nodes, the change polarity of the physical pixel corresponding to the event node, and the edge feature of the directed edge corresponding to the event node, the node feature of each event node is determined, including: determining the initial feature of each event node based on the node position of each event node and the change polarity of the physical pixel corresponding to the event node; taking any event node as a target event node, performing linear transformation on the initial feature of the target event node to obtain a first feature; determining at least one weight based on the edge feature of the directed edge between each adjacent event node in the at least one adjacent event node of the target event node and the target event node, and performing weighted summation on the initial features of the at least one adjacent event node based on the at least one weight to obtain a second feature; and determining the sum of the first feature and the second feature as the node feature of the target event node to obtain the node feature of each event node.

[0012] In a possible implementation of the first aspect and the above implementation, the method further includes: voxelizing the N event nodes, clustering at least one event node in each voxel of M voxels to obtain the node feature, node position, and corresponding time of each target node of M target nodes, one voxel corresponding to one target node, M being a positive integer greater than 1 and less than or equal to N; and sampling image features matching each physical pixel in the event information from the image feature map, and splicing the pixel feature of each physical pixel with the matching image feature to obtain each target feature, including: sampling image features matching the node position and corresponding time of each target node from the image feature map; and splicing the node feature of each target node with the matching image feature to obtain each target feature.

[0013] In the above technical solution, voxelizing the N event nodes can reduce the number of event nodes to M target nodes, that is, can significantly reduce the number of event nodes. Further, after obtaining the node feature, node position, and corresponding time of each target node, sampling image features matching the node position and corresponding time of each target node from the image feature map can reduce the number of sampling times in the image feature map. Finally, splicing the node feature of each target node with the matching image feature to obtain each target feature can deeply fuse the high temporal resolution advantage of the event camera and the rich texture information of the frame camera, and enhance the representation breadth of the target feature. At the same time, voxelization and clustering can also greatly reduce the computational complexity of the electronic device.

[0014] In a possible implementation manner of the first aspect, the node feature, the node position and the corresponding time of each target node of the M target nodes are obtained by clustering the at least one event node in each voxel of the M voxels, including: taking any voxel as a target voxel, and for at least one event node in the target voxel, determining the maximum feature in the node feature of the at least one event node as the node feature of a target node corresponding to the target voxel; determining the node position of the target node based on a first coordinate mean value of the node position of the at least one event node in a horizontal direction and a second coordinate mean value of the node position of the at least one event node in a vertical direction; and determining the corresponding time of the target node as a maximum time corresponding to the at least one event node.

[0015] In the technical solution, the maximum pooling strategy is used to extract the maximum feature of the event node in the target voxel, and the maximum feature is determined as the node feature of the target node in the target voxel, which can retain the key feature of the event node and enhance the robustness when the node feature is expressed. The node feature of the target node is determined by the coordinate mean value in the horizontal direction and the coordinate mean value in the vertical direction, which can consider the node features of the event nodes in the target voxel and determine a more reasonable comprehensive feature. Finally, the maximum time stamp is selected as the corresponding time of the target node, which can ensure the causal consistency of the event time sequence. Meanwhile, the method can realize data dimension reduction through voxel-level clustering, and greatly reduce the computational complexity.

[0016] In a possible implementation manner of the first aspect, the motion trajectory of the at least one physical point is generated based on the target features and the Bezier curve algorithm, including: classifying the target features to obtain a plurality of groups of features, each group of features corresponding to one physical point, and a plurality of features in each group of features reflecting changes of the corresponding physical point in the target time period; determining a plurality of control positions for fitting the motion trajectory of each physical point in the plurality of physical points based on each group of features in the plurality of groups of features; and generating the motion trajectory of each physical point based on the plurality of control positions and the Bezier curve algorithm.

[0017] In the technical solution, the feature grouping and physical point mapping mechanism can effectively decouple the dynamic changes of different physical points, avoid interference caused by motion coupling of multiple physical points, and improve the accuracy of fitting the motion trajectories of different physical points. The control points (control positions) corresponding to the Bezier curve are inversely deduced by using the time sequence changes of each group of features in the target time period, and the abstract features are converted into control position parameters in the geometric space, which can enhance the physical interpretability when the motion trajectory is generated. Further, by using the Bezier curve and the plurality of control positions, the motion anomaly of a specific physical point can be locally corrected while ensuring the smoothness and continuity of the trajectory, and the global stability and dynamic adjustment capability are considered.

[0018] With reference to the first aspect and the above implementation manners, in some possible implementation manners, based on each group of features in the multiple groups of features, determining the multiple control positions used for fitting the motion trajectory of each physical point in the multiple physical points comprises: taking any group of features as target group of features, for each pair of adjacent features in the target group of features, determining a correlation matrix between each pair of adjacent features; determining an initial control position used for generating a motion trajectory of a target physical point corresponding to the target group of features; and based on the initial control position and the multiple correlation matrices corresponding to the target group of features, determining the multiple control positions used for generating the motion trajectory of the target physical point.

[0019] In the above technical solution, by calculating the correlation matrix between adjacent features in the target group of features, the coupling relationship between the target features can be accurately captured, and more fine-grained data support is provided for the generation of the multiple control positions (control points). After the initial control position is determined, based on the initial control position and the multiple correlation matrices corresponding to the target group of features, the multiple control positions of the motion trajectory of the target physical point are determined, that is, the correlation between the high-dimensional target features is converted into geometric constraints on the control points, which can ensure the accuracy of the multiple control positions. At the same time, the motion trajectory generated based on the multiple control positions can also be smoother.

[0020] With reference to the first aspect and the above implementation manners, in some possible implementation manners, based on the multiple control positions and the Bezier curve algorithm, the motion trajectory of each physical point is generated, comprising: taking any physical point as a target physical point, based on a preset order corresponding to the Bezier curve algorithm, performing formula transformation on the Bezier curve algorithm to obtain a target expression after transformation; based on a plurality of preset times, an initial position of the target physical point, the multiple control positions and the target expression, determining a plurality of predicted positions to generate the motion trajectory of the target physical point.

[0021] In the above technical solution, the Bezier curve algorithm is subjected to formula transformation based on a preset order, which can obtain a target expression to determine an expression that can accurately fit the motion trajectory, and avoid the underfitting problem caused by a low-order curve. Furthermore, in combination with the plurality of preset times and the initial position, the multiple control positions (multiple control points) can be mapped to predicted positions under the spatiotemporal constraints to generate the motion trajectory of the target physical point, and the continuity and smooth transition of the motion trajectory are ensured.

[0022] In a second aspect, an image processing apparatus is provided, which comprises: an acquisition module configured to acquire, in a target time period, an image frame captured by a frame camera and event information captured by an event camera in a target scene, the event information being information of a plurality of physical pixels in the event camera that have undergone a change in luminance; a determination module configured to perform feature extraction on the image frame to obtain an image feature map, sample image features matching each physical pixel in the event information from the image feature map, and splice pixel features of each physical pixel and the matching image features to obtain each target feature; and a generation module configured to generate a motion trajectory of at least one physical point based on each target feature and a Bezier curve algorithm, and generate a target image frame based on the motion trajectory of the at least one physical point and a background image region on the image frame, the plurality of physical pixels being used to reflect a change in position of at least one physical point in the target scene in the target time period.

[0023] With reference to the second aspect, in some possible implementation manners, the apparatus further comprises: a construction module configured to construct an event directed graph based on pixel positions of each physical pixel in the event information and corresponding times, a node position of each event node in the event directed graph being related to a pixel position, and a direction of a directed edge between two event nodes being related to corresponding times of the two event nodes; the determination module is specifically configured to determine a node feature of each event node based on the node position of each event node in N event nodes, a change polarity of the physical pixel corresponding to the event node, and an edge feature of the directed edge corresponding to the event node, N being a positive integer greater than 2; and the determination module is specifically further configured to sample image features matching the node position of each event node and the corresponding time from the image feature map, and splice the node feature of each event node and the matching image features to obtain each target feature.

[0024] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the determination module is specifically further configured to: determine an initial feature of each event node based on the node position of each event node and the change polarity of the physical pixel corresponding to the event node; perform linear transformation on the initial feature of a target event node to obtain a first feature, the target event node being any event node; determine at least one weight based on an edge feature of a directed edge between each adjacent event node in at least one adjacent event node of the target event node and the target event node, and perform weighted summation on the initial features of the at least one adjacent event node based on the at least one weight to obtain a second feature; and determine a sum of the first feature and the second feature as the node feature of the target event node, to obtain the node features of each event node.

[0025] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the apparatus further includes: a voxel processing module, configured to perform voxelization processing on the N event nodes, cluster at least one event node in each voxel of M voxels, and obtain node features, node positions and corresponding times of each target node of M target nodes, one voxel corresponding to one target node, M being a positive integer greater than 1 and less than or equal to N; and the determination module is specifically further configured to: sample, from the image feature map, image features matching the node positions and corresponding times of each target node; and splice the node features of each target node and the matched image features to obtain each target feature.

[0026] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the voxel processing module is specifically configured to: take any voxel as a target voxel, for at least one event node in the target voxel, determine, as the node features of a target node corresponding to the target voxel, a maximum feature in the node features of the at least one event node; determine, as the node position of the target node, a first coordinate mean value in a horizontal direction and a second coordinate mean value in a vertical direction of the node positions of the at least one event node; and determine, as the corresponding time of the target node, a maximum time corresponding to the at least one event node.

[0027] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the determination module is specifically further configured to: classify the plurality of target features to obtain a plurality of groups of features, each group of features corresponding to one physical point, and a plurality of features in each group of features being used to reflect changes of the corresponding physical point in the target time period; determine, based on each group of features in the plurality of groups of features, a plurality of control positions used to fit motion trajectories of each physical point in the plurality of physical points; and the generation module is specifically configured to generate the motion trajectories of each physical point based on the plurality of control positions and the Bezier curve algorithm.

[0028] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the determination module is specifically further configured to: take any group of features as a target group of features, for a plurality of pairs of adjacent features in the target group of features, determine a correlation matrix between each pair of adjacent features; determine an initial control position used to generate a motion trajectory of a target physical point corresponding to the target group of features; and determine, based on the initial control position and a plurality of correlation matrices corresponding to the target group of features, a plurality of control positions used to generate the motion trajectory of the target physical point.

[0029] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the determining module is specifically further configured to take any physical point as a target physical point, perform formula transformation on the Bezier curve algorithm based on a preset order corresponding to the Bezier curve algorithm, and obtain a target expression after the formula transformation; and the generating module is specifically further configured to determine a plurality of predicted positions based on a plurality of preset times, an initial position of the target physical point, the plurality of control positions, and the target expression, to generate a motion trajectory of the target physical point.

[0030] In a third aspect, an electronic device is provided, including a memory and a processor. The memory is configured to store executable program code, and the processor is configured to invoke and run the executable program code from the memory, so that the electronic device executes the method in the first aspect or any possible implementation manner of the first aspect.

[0031] In a fourth aspect, a computer-readable storage medium is provided, which stores executable program code. When the executable program code is run on a computer, the computer executes the method in the first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 is a schematic diagram of a framework of image reconstruction provided by an embodiment of the present application;

[0033] Figure 2 is a schematic flowchart of an image processing method provided by an embodiment of the present application;

[0034] Figure 3 is a schematic block diagram of determining a target image frame provided by an embodiment of the present application;

[0035] Figure 4 is a schematic diagram of constructing an event directed graph provided by an embodiment of the present application;

[0036] Figure 5 is a schematic flowchart of another image processing method provided by an embodiment of the present application;

[0037] Figure 6 is a schematic diagram of constructing another event directed graph provided by an embodiment of the present application;

[0038] Figure 7 is a schematic flowchart of another image processing method provided by an embodiment of the present application;

[0039] Figure 8 is a schematic diagram of determining a target feature provided by an embodiment of the present application;

[0040] Figure 9Fig. 1 is a structural schematic diagram of an image processing device provided by an embodiment of the present application.

[0041] Figure 10 Fig. 2 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0042] The technical solutions in the present application will be described clearly and exhaustively below with reference to the drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B: "and / or" in the text is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0043] Hereinafter, the terms "first" and "second" are only for descriptive purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features.

[0044] Figure 1 Fig. 3 is a schematic diagram of a framework of image reconstruction provided by an embodiment of the present application.

[0045] As shown in Fig. 1, for example, the frame camera can acquire images under the target scene within the target period to obtain a plurality of image frames; the event camera can capture the change polarity and time of the environmental brightness under the target scene within the target period to obtain event information. Figure 1

[0046] Since the frame camera relies on a fixed frame rate and exposure time, this makes the frame camera have significant limitations in capturing subtle changes in high-speed motion. In the above target scene, due to limited exposure, each frame camera cannot clearly capture the scene with a large brightness range at the same time, resulting in a decrease in image quality, for example, object A with a large brightness range and the details of the lower right corner of the chair in the first image frame cannot be clearly displayed. The above problem can limit the application of multiple frame images in motion analysis and image reconstruction. The event camera is different from the above frame camera, which can capture the change polarity and time of the environmental brightness under the target scene with a time resolution of microseconds (i.e., the event camera can focus on subtle changes in high-speed motion), but lacks description of absolute brightness information and spatial details. Therefore, the above problem can limit the application of event information in motion analysis and image reconstruction.

[0047] ​Based on the respective advantages of the frame camera and the event camera, the electronic device can fuse the multiple frame images and the event information to obtain multiple fused images. Each image in the multiple fused images can clearly display the details of each object with a large luminance range, and can display the change polarity, time, absolute luminance information, and spatial detail description of the environmental luminance. For example, in the first fused image, object A with a large luminance range and the details of the lower right corner of the chair can be clearly displayed, and the change polarity (brightening), time, absolute luminance information (target luminance), and spatial detail (spatial position) of the luminance of the street lamp can be displayed.

[0048] However, how to fuse the frame images and the event information, fully utilize the complementary advantages of the two, and realize accurate and efficient reconstruction of a complex motion scene, becomes a key technical problem to be solved at present. In order to solve the above problems, the present application provides an image processing method, and the specific implementation steps are as follows Figure 2 .

[0049] Figure 2 is a schematic flowchart of an image processing method provided by an embodiment of the present application.

[0050] It should be understood that the image processing method provided by the embodiments of the present application can be applied to an electronic device as shown in Figure 1 .

[0051] For example, as shown in Figure 2 , the method 200 includes the following steps 201-203.

[0052] Step 201: In a target time period, an image frame captured by a frame camera and event information captured by an event camera under a target scene are obtained, and the event information is information of multiple physical pixels in the event camera that have a luminance change.

[0053] It should be understood that the "image frame" in the above step 201 is captured for a target scene and in a target time period. In order to improve the fusion efficiency of the frame image and the event information, the number of images corresponding to the multiple image frames should be as large as possible, that is, the time interval of each image frame captured in the target time period is short. Optionally, the time interval is 50 milliseconds.

[0054] It should also be understood that the "event information" in the above step 201 is also collected for the target scene and in the target time period. An event camera is an event-based vision sensor with a pixel array. Each physical pixel has a dedicated circuit inside to monitor the change of ambient brightness in real time. If the change of brightness exceeds a preset threshold, an event output (i.e., an event point) is triggered, which is a physical pixel that has a change in brightness that reaches the preset threshold. The pixel position, time, and change polarity of the physical pixel that outputs the change in brightness are output. The pixel position is represented by coordinate values, specifically (coordinate value of the physical pixel in the horizontal direction, coordinate value of the physical pixel in the vertical direction), and the change polarity p includes increase (+) or decrease (-). That is, the above information includes the pixel position, time, and change polarity.

[0055] It should be emphasized that the "event information" in the above step 201 is actually the information of multiple physical pixels that have changes in brightness in the target time period, i.e., in this case, the information of the same physical pixel (the same pixel position or the same identified physical pixel is considered the same physical pixel) can appear multiple times. Alternatively, the event information includes the information of physical pixel PX 12 , the information of physical pixel PX 34 , the information of physical pixel PX 24 , the information of physical pixel PX 69 , the information of physical pixel PX 23 , the information of physical pixel PX 12 , and the information of physical pixel PX 69 .

[0056] It should also be emphasized that each of the above image frames is collected for the target scene, and the above event information is also collected for the target scene, i.e., each physical pixel in the event camera has a mapping relationship with an object (specifically a physical point on the object) in the target scene, and each physical pixel in the event camera has a mapping relationship with an image pixel on the image frame.

[0057] It should also be noted that each physical pixel in the event camera detects changes in the brightness of the environment it receives independently, and the change in brightness of the object in the target scene can be directly mapped to the polarity event of the event camera. Alternatively, the change in brightness of the street lamp is projected as a change in brightness of multiple physical pixels in a certain pixel area.

[0058] In some embodiments, the acquiring, in step 201, the multiple image frames captured by the frame camera and the event information captured by the event camera under the target scene within the target time period comprises: acquiring, within the target time period, the image frames captured by the frame camera under the target scene at a preset time interval, to obtain the multiple image frames; and acquiring the event information triggered when the luminance change of the physical pixels in the event camera under the target scene exceeds a preset threshold.

[0059] In step 202, the image features are extracted from the image frames to obtain an image feature map, the image features matched with each physical pixel in the event information are sampled from the image feature map, the pixel features of each physical pixel are spliced with the matched image features to obtain each target feature.

[0060] It should be understood that the "sampling the image features matched with each physical pixel in the event information from the image feature map" in step 202 above refers to sampling the image features matched with each physical pixel in position and time from the image feature map, i.e., the matched physical pixel and image feature at the position are used to describe the luminance change of the same object (specifically, the physical point on the object) at the time.

[0061] It should also be understood that the "pixel feature of the physical pixel" in step 202 above is determined by the pixel position of the physical pixel and the change polarity of the luminance change of the physical pixel.

[0062] It should also be understood that the "splicing the pixel feature of each physical pixel with the matched image feature" in step 202 above refers to splicing the pixel feature of any physical pixel with the matched image feature in the channel dimension. The size of the feature is represented by X*Y*Z, X is used to represent the width, Y is used to represent the height, and Z is used to represent the channel dimension. Alternatively, the size of the pixel feature of the first physical pixel is X1*Y1*Z1, the size of the matched image feature is X1*Y1*Z2, and the size of the spliced target feature is X1*Y1*(Z1+Z2). Alternatively, the size of the pixel feature of the first physical pixel is X2*Y2*Z1, the size of the matched image feature is X3*Y2*Z2, X2>X3, and the size of the spliced target feature is X2*Y2*(Z1+Z2). Wherein, a part of the feature elements in the spliced target feature are 0 elements.

[0063] In some embodiments, the feature extraction on the image frames in step 202 to obtain the image feature map comprises: performing convolution processing on the image frames through a preset convolution layer to obtain a third feature; and performing residual processing on the third feature through a preset residual layer to obtain the image feature map.

[0064] In some embodiments, the sampling, from the image feature map, an image feature matching each physical pixel in the event information, and splicing a pixel feature of each physical pixel with the matching image feature to obtain a target feature in step 202 comprises: taking any physical pixel as a target physical pixel, determining a target feature corresponding to the target physical pixel based on the following formula (1) and formula (2).

[0065] g r = g (I G , po r , t r ) (1)

[0066]

[0067] wherein formula (1) can be understood as sampling, from an image feature map I G , an image feature g r matching a target physical pixel r at a position po r and a time t r , and formula (2) can be understood as splicing a pixel feature f r of the target physical pixel with the matching image feature g r to obtain a target feature

[0068] In some embodiments, the splicing, in step 202, of a pixel feature of each physical pixel with a matching image feature to obtain a target feature comprises: taking any physical pixel as a target physical pixel, determining whether a width or height of a pixel feature of the target physical pixel is the same as a width or height of a matching image feature; in a case where the width or height of the pixel feature of the target physical pixel is the same as the width or height of the matching image feature, splicing the pixel feature of the target physical pixel with the matching image feature in a channel dimension to obtain a target feature corresponding to the target physical pixel; in a case where the width and / or height of the pixel feature of the target physical pixel is different from the width and / or height of the matching image feature, adjusting the pixel feature of the target physical pixel to obtain a candidate pixel feature of the target physical pixel, and adjusting the matching image feature to obtain a candidate image feature of the matching image feature, the width and height of the candidate pixel feature being the same as the width and height of the candidate image feature; splicing the candidate pixel feature with the candidate image feature in the channel dimension to obtain the target feature corresponding to the target physical pixel.

[0069] In some embodiments, in the case that the width and / or height of the pixel feature of the target physical pixel is different from the width and / or height of the matched image feature, the pixel feature of the target physical pixel is adjusted to obtain a candidate pixel feature of the target physical pixel, and the matched image feature is adjusted to obtain a candidate image feature of the matched image feature, including: in the case that the width of the pixel feature of the target physical pixel is less than the width of the matched image feature, the pixel feature of the target physical pixel is zero-padded in the width direction, the first feature after zero padding is determined as the candidate pixel feature, and the matched image feature is determined as the candidate image feature; in the case that the width of the pixel feature of the target physical pixel is greater than the width of the matched image feature, the matched image feature is zero-padded in the width direction, the second feature after zero padding is determined as the candidate image feature, and the pixel feature of the target physical pixel is determined as the candidate pixel feature; in the case that the height of the pixel feature of the target physical pixel is less than the height of the matched image feature, the pixel feature of the target physical pixel is zero-padded in the height direction, the third feature after zero padding is determined as the candidate pixel feature, and the matched image feature is determined as the candidate image feature; in the case that the height of the pixel feature of the target physical pixel is greater than the height of the matched image feature, the matched image feature is zero-padded in the height direction, the fourth feature after zero padding is determined as the candidate image feature, and the pixel feature of the target physical pixel is determined as the candidate pixel feature.

[0070] In step 203, a motion trajectory of at least one physical point is generated based on each target feature and a Bezier curve algorithm, and a target image frame is generated based on the motion trajectory of the at least one physical point and a background image region on the image frame, and the plurality of physical pixels are used to reflect changes of at least one physical point in the target scene within the target time period.

[0071] It should be understood that the "Bezier curve algorithm" in the above step 203 is commonly used to fit the motion trajectory of a certain physical point, and the motion trajectory is generated by a plurality of predicted positions, and the core is to define a smooth motion trajectory by control points. The Bezier curve algorithm corresponds to an n-order Bezier curve generated by n+1 control points, and the corresponding expression is as follows formula (3);

[0072]

[0073] wherein P i is a plurality of control points (a plurality of control positions), P i is (x i , y i ), t is time, is a combination number. The coordinate value of each control point in the plurality of control points in the horizontal direction and the coordinate value in the vertical direction are determined by each target feature in the plurality of target features, B(t) is the predicted position of a certain physical point at time t, n is the order corresponding to the Bezier curve algorithm, which is used to control the type of trajectory line of the motion trajectory. When n is 1, the type is a straight line, and when n is 3, the type is an S-shaped curve.

[0074] It should also be understood that each target feature in the plurality of target features is obtained by fusing the pixel features of the physical pixels and the matched image features. In the case where a plurality of physical pixels are used to reflect the position change of at least one physical point in the target scene within the target period, the plurality of target features are specifically used to indicate the characteristics of the at least one physical point. One physical point corresponds to part of the plurality of target features.

[0075] It should be noted that the "background image region" in the above step 203 refers to the static background region in the image frame. In addition, the "target image frame" in the above step 203 refers to a plurality of image frames. Since the Bezier curve algorithm (corresponding to the expression with time t) is combined in step 203, a plurality of predicted positions (motion trajectories) that are continuous in time can be generated. Further, based on the motion trajectory of the at least one physical point and the background image region on the image frame, a plurality of target image frames that are continuous and clear in the visual angle, i.e., a plurality of target image frames that are continuous and clear in time, can be generated. The joint use of the frame camera and the event camera can meet the replacement demand for high-quality cameras.

[0076] In some embodiments, the step 203 of generating a target image frame based on the motion trajectory of the at least one physical point and the background image region on the image frame comprises: filling the corresponding physical points on the blank image frame based on the predicted positions of different physical points at a first time in a plurality of times, and filling the corresponding image content of the background image region in the remaining region on the blank image frame to generate a first image frame in the target image frame corresponding to the first time. Similarly, filling the corresponding physical points on the blank image frame based on the predicted positions of different physical points at a Gth time in a plurality of times, and filling the corresponding image content of the background image region in the remaining region on the blank image frame to generate a Gth image frame in the target image frame corresponding to the Gth time, G is a positive integer greater than 1.

[0077] It should be understood that the "plurality of times" in the above scheme refers to the times substituted into formula (3), including the first time t1, the second time t2,..., the Gth time t G .

[0078] In a possible implementation, the step of generating the motion trajectory of each physical point based on the target features and the Bezier curve algorithm in step 203 includes: classifying the target features to obtain a plurality of groups of features, each group of features corresponding to a physical point, and the target features in each group of features being used to reflect the change of the corresponding physical point in the target time period; determining, based on each group of features in the plurality of groups of features, a plurality of control positions used to fit the motion trajectory of each physical point in the plurality of physical points; and generating the motion trajectory of each physical point based on the plurality of control positions and the Bezier curve algorithm.

[0079] It should be understood that the "plurality of target features" in the above scheme is the fused features obtained after tracking at least one physical point in the target time period, and therefore, the plurality of target features needs to be classified to obtain a plurality of groups of features, each group of features corresponding to a physical point, that is, each group of features is the fused features obtained after tracking the corresponding physical point.

[0080] In the above technical scheme, based on the feature grouping and physical point mapping mechanism, the dynamic changes of different physical points can be effectively decoupled, the interference caused by the motion coupling of multiple physical points can be avoided, and the accuracy of fitting the motion trajectories of different physical points can be improved. The control points (control positions) corresponding to the Bezier curve are inversely deduced by using the time sequence changes of each group of features in the target time period, the abstract features are converted into control position parameters in the geometric space, and the physical interpretability in generating the motion trajectory can be enhanced. Further, by using the Bezier curve and the plurality of control positions, the motion anomaly of a specific physical point can be locally corrected while ensuring the smoothness and continuity of the trajectory, and the global stability and dynamic adjustment capability are taken into account.

[0081] In some embodiments, classifying the plurality of target features to obtain a plurality of groups of features includes: determining the similarity between a plurality of pairs of target features at different times in the plurality of target features, and determining the target features with a similarity greater than a preset similarity as a group of features.

[0082] It should be understood that in the above scheme, the target features at the same time cannot be classified into a group of features.

[0083] For example, for the target features at the first time Target features and target features Target features at the second time and target features Target features at the third time and target features

[0084] The similarity S1 between and is determined, and The similarity S2 between and The similarity between them is S3, and The similarity between them is S4, and The similarity between them is S5, and The similarity between S6, and The similarity between S7, and The similarity between them is S8, and The similarity between S9, and The similarity S between 10 , and The similarity S between 11 , and The similarity S between 12 , and The similarity S between 13 , and The similarity S between 14 , and The similarity S between 15 , and The similarity S between 16 .

[0085] In S1, S3 and S 13 When both are greater than the preset similarity, and Classify into a group of features, when S6 is greater than the preset similarity, and into a set of features, and in S 12 When the similarity is greater than the preset value, and grouped into a set of features.

[0086] In one possible implementation, based on each group of features in the multiple groups of features, multiple control positions for fitting the motion trajectory of each physical point in the multiple physical points are determined, including: taking any group of features as the target group features, and for multiple pairs of adjacent features in the target group features, determining the correlation matrix between each pair of adjacent features; determining an initial control position for generating the motion trajectory of a target physical point, the target physical point corresponding to the target group features; and determining multiple control positions for generating the motion trajectory of the target physical point based on the initial control position and multiple correlation matrices corresponding to the target group features.

[0087] It should be understood that when the number of features corresponding to the "target group features" in the above solution is Q, the number of correlation matrices corresponding to them is Q-1. Optionally, the features in the target group features include and When and Determine the correlation matrix C 12 ,Depend on and Determine the correlation matrix C 23 and by and Determine the correlation matrix C 34 .

[0088] In the above technical solution, by calculating the correlation matrix between adjacent features within the target group features, the coupling relationship between the target features can be accurately captured, providing more fine-grained data support for the generation of multiple control positions (control points). After determining the initial control position, based on the multiple correlation matrices corresponding to the initial control position and the target group features, multiple control positions of the motion trajectory of the target physical point are determined, that is, the correlation between the high-dimensional target features is converted into geometric constraints on the control points, which can ensure the accuracy of the multiple control positions. At the same time, it can also make the motion trajectory generated based on multiple control positions smoother.

[0089] In some embodiments, determining a correlation matrix between each pair of adjacent features includes: determining a correlation matrix between each pair of adjacent features based on the following formula (4);

[0090]

[0091] Among them, C i is the first target feature in the i-th pair of adjacent features and the second target feature The correlation matrix between The first target feature The u-th characteristic element in The second target feature the u-th feature element in the N-th target feature the number of feature elements in the N-th target feature the first target feature the second target feature

[0092] In some embodiments, determining the initial control position comprises: determining the initial control position based on the following formula (5);

[0093] P0 = (x min ,y min )(5)

[0094] wherein P0 is the initial control position represented by coordinates, specifically the pixel coordinates of the physical pixel in which the luminance change occurs at the minimum time.

[0095] In some embodiments, determining the plurality of control positions of the motion trajectory of the target physical point based on the initial control position and the plurality of correlation matrices corresponding to the target group of features comprises: determining the plurality of control positions of the motion trajectory of the target physical point based on the following formula (6) and formula (7);

[0096] ΔP i-1 = U(C i ,P i-1 ), i = 1, …, M (6)

[0097] P i = P i-1 + ΔP i-1 (7)

[0098] wherein ΔP i-1 is used to reflect the position change amount of the control position, U(C i ,P i-1 ) is a preset algorithm used to determine the change of the control position, M is the number of adjacent features, P i is the i-th control position in the plurality of control positions.

[0099] In a possible implementation, generating the motion trajectory of each physical point based on the plurality of control positions and the Bezier curve algorithm comprises: taking any physical point as a target physical point, performing formula transformation on the Bezier curve algorithm based on a preset order corresponding to the Bezier curve algorithm to obtain a transformed target expression; determining a plurality of predicted positions based on a plurality of preset times, the initial position of the target physical point, the plurality of control positions and the target expression, to generate the motion trajectory of the target physical point.

[0100] It should be understood that the "formula transformation is performed on the Bezier curve algorithm based on the preset order corresponding to the Bezier curve algorithm to obtain the target expression" in the above scheme refers to expanding the expression corresponding to the Bezier curve algorithm based on the preset order to generate the target expression.

[0101] In the above technical solution, the formula transformation is performed on the Bezier curve algorithm based on the preset order, which can obtain the target expression to determine the expression that can accurately fit the motion trajectory and avoid the underfitting problem caused by the low-order curve. Furthermore, in combination with the preset multiple times and the initial position, the multiple control positions (multiple control points) can be mapped to the predicted positions under the space-time constraint to generate the motion trajectory of the target physical point, ensuring the continuity and smooth transition of the motion trajectory.

[0102] Optionally, when the preset order is 1, the expression corresponding to the first-order Bezier curve is B(t) = (1-t)P0+tP1; when the preset order is 2, the expression corresponding to the second-order Bezier curve is B(t) = (1-t) 2 P0+2(1-t)tP1+t 2 P2; when the preset order is 3, the expression corresponding to the third-order Bezier curve is B(t) = (1-t) 3 P0+3(1-t) 2 tP1+3(1-t)t 2 P2+t 3 P3.

[0103] In some embodiments, based on the preset multiple times, the initial position of the target physical point, the multiple control positions, and the target expression, the multiple predicted positions are determined, including: determining a target number of target control positions corresponding to the preset order from the multiple control positions; substituting the initial position of the target physical point and the target control position into the target expression to obtain a first expression; substituting each time in the preset multiple times into the first expression in turn to obtain the multiple predicted positions.

[0104] It should be understood that the number of control positions determined by formula (5) to formula (7) is greater than the target number (i.e., the required number) corresponding to the preset order. For example, the preset order is 3, the target number corresponding to the preset order is 4, that is, (P0, P1, P2, and P3) are included, and the number of control positions determined by formula (5) to formula (7) is greater than 4. Therefore, it is necessary to determine a target number of target control positions corresponding to the preset order from the multiple control positions. In addition, the target number of target control positions can be determined by clustering the multiple control positions.

[0105] Figure 3This is a schematic block diagram of a method for determining a target image frame provided in an embodiment of the present application.

[0106] For example, Figure 3 As shown, image frames captured by a frame camera in a target scene within a target time period are obtained, and features are extracted from the image frames to obtain an image feature map. Event information captured by an event camera in the target scene within the target time period is obtained. The event information is information about multiple physical pixels in the event camera that experience brightness changes. Image features that match each physical pixel in the event information are sampled from the image feature map, that is, image features that match the pixel features of each physical pixel are obtained. The image features and pixel features have a matching relationship. The pixel features of each physical pixel are spliced ​​with the matching image features to obtain each target feature. Based on the multiple target features, multiple control positions are determined. Then, combined with the initial positions (the multiple control positions and the initial positions are referred to as multiple control points) and a Bezier curve algorithm, multiple predicted positions are generated, that is, the motion trajectory of at least one physical point. Based on the multiple predicted positions and the background image area on the image frame, a target image frame is generated. The multiple physical pixels are used to reflect the position change of at least one physical point in the target scene within the target time period.

[0107] The second process of "obtaining each target feature" is described below.

[0108] In one possible implementation, the method 200 further includes: constructing an event directed graph based on the pixel position and corresponding time of each physical pixel in the event information, wherein the node position of each event node in the event directed graph is related to the pixel position, and the direction of the directed edge between two event nodes is related to the time corresponding to each of the two event nodes; determining the node features of each event node based on the node position of each event node in N event nodes, the change polarity of the physical pixel corresponding to the event node, and the edge features of the directed edge corresponding to the event node, where N is a positive integer greater than 2; and, in step 202, sampling image features matching each physical pixel in the event information from the image feature map, splicing the pixel features of each physical pixel with the matching image features to obtain each target feature, including: sampling image features matching the node position and corresponding time of each event node from the image feature map; splicing the node features of each event node with the matching image features to obtain each target feature.

[0109] It should be understood that the "event directed graph" in the above scheme is actually a directed graph, which is composed of multiple event nodes and multiple directed edges. Each event node has corresponding node features, and each directed edge has corresponding edge features.

[0110] It should also be understood that the "sampling, from the image feature map, image features matching the node positions and corresponding times of the respective event nodes" in the above technical solution refers to sampling, from the image feature map, image features matching the respective event nodes in terms of node position and corresponding time, i.e., both the physical pixels and the image features matching the node position and corresponding time are used to describe the brightness change of the same object (specifically, a physical point on the object) at the corresponding time.

[0111] In the above technical solution, the event directed graph is constructed based on the pixel positions and corresponding times of the physical pixels, the causality of event propagation (such as the object movement path) is encoded using the time directionality (such as the timestamp order) between two event nodes, and the edge features are combined, which can reflect the correlation of multiple physical pixels in space-time. In combination with the node position of the event node, the change polarity of the corresponding physical pixel, and the edge feature of the corresponding directed edge, the node feature is dynamically generated, which can accurately determine the node feature. Further, after sampling through space-time alignment, the node feature of the event node is spliced with the frame image feature (matching image feature), which can fuse the high-frequency response capability of the event camera and the rich texture information of the frame camera, and overcome the information limitation of a single sensor. The scheme can significantly improve the accuracy of motion trajectory generation in a dynamic scene through the fusion mechanism of directed graph structure and cross-modal features.

[0112] In some embodiments, based on the pixel positions and corresponding times of the respective physical pixels in the event information, the event directed graph is constructed, including: based on the respective physical pixels, constructing respective event nodes in the event directed graph, determining a ratio between a coordinate value of the pixel position of each physical pixel in the horizontal direction and the width of the image frame as a first coordinate value, and determining a ratio between a coordinate value of the pixel position of each physical pixel in the vertical direction and the height of the image frame as a second coordinate value; determining the first coordinate value as a coordinate value of the node position of the corresponding event node in the horizontal direction, and determining the second coordinate value as a coordinate value of the node position of the corresponding event node in the vertical direction; determining the distance between each pair of event nodes in the plurality of event nodes based on the node position of the event node; determining that there is a directed edge between each pair of event nodes corresponding to the distance less than the preset distance, and determining the direction of the directed edge as from the event node with the smaller time to the event node with the larger time.

[0113] Figure 4 is a schematic diagram of constructing an event directed graph provided by an embodiment of the present application.

[0114] For example, as shown in Figure 4 , the event information includes the pixel positions and corresponding times of the respective physical pixels in the 5 physical pixels, and the pixel positions of the 5 physical pixels include the physical pixels PX 12pixel position (1, 2), physical pixel PX 34 pixel position (1, 3), physical pixel PX 24 pixel position (2, 2), physical pixel PX 69 pixel position (3, 1), and physical pixel PX 23 pixel position (3, 3), 5 physical pixels corresponding to the time including physical pixel PX 12 corresponding to the first time 1.0 ms, physical pixel PX 34 corresponding to the second time 1.2 ms, physical pixel PX 24 corresponding to the third time 1.5 ms, physical pixel PX 69 corresponding to the fourth time 1.8 ms, and physical pixel PX 23 corresponding to the fifth time 2.0 ms, and, taking the preset distance √5 as an example, the construction process of the event directed graph is described.

[0115] physical pixel PX 12 corresponding event node A 12 , event node A 12 The node position is (1, 2); physical pixel PX 34 corresponding event node A 34 , event node A 34 The node position is (1, 3); physical pixel PX 24 corresponding event node A 24 , event node A 24 The node position is (2, 2); physical pixel PX 69 corresponding event node A 69 , event node A 69 The node position is (3, 1); physical pixel PX 23 corresponding event node A 23 , event node A 23 The node position is (3, 3).

[0116] The distance between pixel position (1, 2) and pixel position (1, 3) is 1, the distance between pixel position (1, 2) and pixel position (2, 2) is 1, the distance between pixel position (1, 2) and pixel position (3, 1) is √5, and the distance between pixel position (1, 2) and pixel position (3, 3) is The distance between pixel position (1, 3) and pixel position (2, 2) is √2, and the distance between pixel position (1, 3) and pixel position (3, 1) is The distance between pixel position (1, 3) and pixel position (3, 3) is 2; the distance between pixel position (2, 2) and pixel position (3, 1) is The distance between pixel position (2, 2) and pixel position (3, 3) is determined to be 1. The distance between pixel position (3, 1) and pixel position (3, 3) is determined to be 2.

[0117] Since the above 1, 2 and are all less than the preset distance Therefore, correspondingly, the above physical pixels PX 12 and physical pixel PX 34 There is a directed edge between physical pixel PX 12 and physical pixel PX 24 There is a directed edge between physical pixel PX 34 and physical pixel PX 24 There is a directed edge between physical pixel PX 34 and physical pixel PX 23 There is a directed edge between physical pixel PX 24 and physical pixel PX 69 There is a directed edge between physical pixel PX 24 and physical pixel PX 23 There is a directed edge between physical pixel PX 69 and physical pixel PX 23 There is a directed edge between physical pixel PX In addition, since the first time 1.0 ms is less than the second time 1.2 ms, the second time 1.2 ms is less than the third time 1.5 ms, the third time 1.5 ms is less than the fourth time 1.8 ms, and the fourth time 1.8 ms is less than the fifth time 2.0 ms, an event directed graph as shown in Figure 4 can be constructed.

[0118] In some embodiments, the edge feature of the directed edge corresponding to the event node includes: determining the edge feature of the directed edge corresponding to the event node based on the following formula (8);

[0119]

[0120] wherein, is the edge feature of the directed edge between event node o1 and event node o2, i.e. the edge feature of the directed edge corresponding to event node o1, is the initial feature of event node o1, is the initial feature of event node o2.

[0121] In a possible implementation, the node feature of each event node is determined based on the node position of each event node in the N event nodes, the change polarity of the physical pixel corresponding to the event node, and the edge feature of the directed edge corresponding to the event node, including: determining the initial feature of each event node based on the node position of each event node and the change polarity of the physical pixel corresponding to the event node; taking any event node as a target event node, performing linear transformation on the initial feature of the target event node to obtain a first feature; in the case that the target event node has adjacent event nodes, determining at least one weight based on the edge feature of the directed edge between each adjacent event node in the at least one adjacent event node and the target event node, and performing weighted summation on the initial features of the at least one adjacent event node based on the at least one weight to obtain a second feature; determining the sum of the first feature and the second feature as the node feature of the target event node, to obtain the node features of the event nodes.

[0122] It should be understood that the "initial feature of the event node" in the above scheme is determined only by the node position of the event node and the change polarity of the corresponding physical pixel; the "node feature of the event node" is also related to the corresponding directed edge, that is, the node feature of the event node is determined by the node position of the event node, the change polarity of the corresponding physical pixel, and the edge feature of the corresponding directed edge.

[0123] In some embodiments, the initial feature of each event node is determined based on the node position of each event node and the change polarity of the physical pixel corresponding to the event node, including: determining the initial feature of each event node based on the following formula (9);

[0124]

[0125] wherein the initial feature of the event node o is which can be represented as n(x o ,y o ,p o ), (x o ,y o ) is the node position of the event node o, x o is the coordinate value of the node position in the horizontal direction, y o is the coordinate value of the node position in the vertical direction, and p o is the change polarity of the physical pixel corresponding to the event node o.

[0126] It should be understood that the "physical pixel corresponding to the event node" in the above scheme refers to the physical pixel at the same pixel position as the node position of the event node.

[0127] In some embodiments, determining the sum of the first feature and the second feature as the node feature of the target event node includes: determining the node feature of the target event node based on the following formula (10);

[0128]

[0129] in, For the target event node The node characteristics of is the first feature, E is the event directed graph, For the target event node Adjacent event nodes The node characteristics of Based on adjacent event nodes With the target event node Edge characteristics of directed edges between Determine the weight, is the second feature. Wherein W is a preset mapping matrix used to map the target event node i * The initial features of the edge are linearly transformed, and W() can be dynamically generated through the spline basis function and the preset learnable coefficient matrix to encode the spatial relationship of the edge features. The spline basis function refers to a smooth basis function defined on the space of edge features, which is used to map the edge features to the weight space.

[0130] It should be noted that the process of determining the node features of the target event node by using formula (10) in the above scheme can be regarded as determining the node features by aggregating the node information of adjacent nodes using spline convolution.

[0131] In the above technical solution, the initial features of the event node are constructed by combining the node position of the event node with the change polarity of the corresponding physical pixel, which can effectively encode spatial attributes and enhance the perception of the physical meaning of the event. By linearly transforming the initial features through a learnable mapping matrix (obtaining the first feature), nonlinear expression capabilities can be introduced, and at least one weight is dynamically generated using the spline basis function and the learnable coefficient matrix to encode the edge features as spatial relationship constraints, and by weighted aggregation of node information of adjacent nodes (the second feature), context-aware feature enhancement can be achieved. The sum of the first feature and the second feature is determined as the node feature of the event node, which can integrate the correlation between event nodes while retaining the independent characteristics of the nodes to determine more accurate node features.

[0132] It should be noted that the above is another way to determine the characteristics of each target. In combination with the second way to determine the characteristics of each target, another process for generating the motion trajectory of at least one physical point is specifically given below in steps 501 to 505.

[0133] Step 501, in a target time period, acquiring an image frame captured by a frame camera and event information captured by an event camera in a target scene, the event information being information of a plurality of physical pixels in the event camera that have brightness changes.

[0134] It should be understood that the meaning of "image frame" in the above step 501 and the corresponding embodiments and the meaning of "event information" and the corresponding embodiments are the same as the meaning of "image frame" and the corresponding embodiments and the meaning of "event information" and the corresponding embodiments in the aforementioned step 201, which will not be repeated here.

[0135] Step 502, performing feature extraction on the image frame to obtain an image feature map, and constructing an event directed graph based on the pixel positions and corresponding times of each physical pixel in the event information, the node positions of each event node in the event directed graph being related to the pixel positions, and the direction of the directed edge between two event nodes being related to the corresponding times of the two event nodes.

[0136] It should be understood that the meaning of "event directed graph" in the above step 502 and the corresponding embodiments are the same as the meaning of "event directed graph" and the corresponding embodiments in the aforementioned scheme, which will not be repeated here.

[0137] Step 503, determining the node feature of each event node based on the node position of each event node in the N event nodes, the change polarity of the physical pixel corresponding to the event node, and the edge feature of the directed edge corresponding to the event node, N being a positive integer greater than 2.

[0138] Step 504, sampling image features matching the node position and corresponding time of each event node from the image feature map; and splicing the node feature of each event node with the matched image features to obtain each target feature.

[0139] Step 505, generating a motion trajectory of at least one physical point based on each target feature and a Bezier curve algorithm, and generating a target image frame based on the motion trajectory of the at least one physical point and a background image region on the image frame, the plurality of physical pixels being used to reflect changes of at least one physical point in the target scene in the target time period.

[0140] It should be understood that the respective specific embodiments of the above steps 504 and 505 are the same as the principles of the respective specific embodiments of the aforementioned steps 202 and 203, which will not be repeated here.

[0141] The third process of "obtaining each target feature" is described as follows.

[0142] In a possible implementation manner, the method 200 further includes: voxelizing the N event nodes, clustering at least one event node in each voxel of M voxels, to obtain node features, node positions, and corresponding times of each target node of M target nodes, one voxel corresponds to one target node, M is a positive integer greater than 1, and M is less than or equal to N; and sampling, from the image feature map, image features matching each physical pixel in the event information, splicing pixel features of each physical pixel with the matched image features, to obtain each target feature, including: sampling, from the image feature map, image features matching the node positions and corresponding times of each target node; and splicing the node features of each target node with the matched image features, to obtain each target feature.

[0143] It should be understood that the "voxelization processing" in the above solution refers to dividing the N event nodes by a voxel grid with a certain voxel size g x ×g y ×g t in space, so that at least one event node is included in one voxel, where g t is the time dimension.

[0144] It should be noted that the process of voxelizing the N event nodes and clustering at least one event node in each voxel of M voxels in the above solution can be regarded as a kind of pooling operation.

[0145] In the above technical solution, by voxelizing the N event nodes, the number of event nodes can be reduced to M target nodes, that is, the number of event nodes can be significantly reduced. Further, after obtaining the node features, node positions, and corresponding times of each target node, sampling, from the image feature map, image features matching the node positions and corresponding times of each target node can reduce the sampling times in the image feature map. Finally, splicing the node features of each target node with the matched image features to obtain each target feature can deeply fuse the high time resolution advantage of the event camera and the rich texture information of the frame camera, and enhance the representation breadth of the target feature. At the same time, by voxelizing and clustering to reduce the sampling times, the computational complexity of the electronic device can also be greatly reduced.

[0146] In some embodiments, voxelizing the N event nodes includes: dividing the N event nodes into M voxels according to a preset voxel size.

[0147] An example is that 10 event nodes include event node 1, event node 2, event node 3, event node 4, event node 5, event node 6, event node 7, event node 8, event node 9 and event node 10, and the preset voxel size g x ×g y ×g t Specifically, 2×2×0.2, the node feature corresponding to the event node 1 is 2, the node position is (1, 2) and the corresponding event is 0.1 ms, the node feature corresponding to the event node 2 is 2.3, the node position is (1.2, 2) and the corresponding event is 0.2 ms, the node feature corresponding to the event node 3 is 3, the node position is (0.5, 1.2) and the corresponding event is 0.2 ms, the node feature corresponding to the event node 4 is 2.6, the node position is (1, 1.3) and the corresponding event is 0.16 ms, the node feature corresponding to the event node 5 is 3.6, the node position is (2.2, 2.3) and the corresponding event is 0.3 ms, the node feature corresponding to the event node 6 is 2.3, the node position is (3, 2.6) and the corresponding event is 0.32 ms, the node feature corresponding to the event node 7 is 3, the node position is (2.7, 3.8) and the corresponding event is 0.4 ms, the node feature corresponding to the event node 8 is 4, the node position is (4.2, 5) and the corresponding event is 0.5 ms, the node feature corresponding to the event node 9 is 2.2, the node position is (4.6, 4.9) and the corresponding event is 0.45 ms, and the node feature corresponding to the event node 10 is 2.6, the node position is (7.2, 6.3) and the corresponding event is 0.62 ms, which are taken as examples to describe the process of voxelizing the 10 event nodes.

[0148] The 10 event nodes are divided into the voxel 1 including the event node 1, the event node 2, the event node 3 and the event node 4, the voxel 2 including the event node 5, the event node 6 and the event node 7, the voxel 3 including the event node 8 and the event node 9, and the voxel 4 including the event node 10.

[0149] In a possible implementation, the at least one event node in each voxel of the M voxels is clustered to obtain the node feature, the node position and the corresponding time of each target node of the M target nodes, including: taking any voxel as a target voxel, for at least one event node in the target voxel, determining the maximum feature in the node feature of the at least one event node as the node feature of the target node corresponding to the target voxel; determining the node position of the target node based on the first coordinate mean value in the horizontal direction and the second coordinate mean value in the vertical direction of the node position of the at least one event node; and determining the maximum time corresponding to the at least one event node as the time corresponding to the target node.

[0150] In the technical solution, the maximum pooling strategy is used to extract the maximum feature of the event nodes in the target voxel, and the maximum feature is determined as the node feature of the target node in the target voxel, which can retain the key feature of the event node and enhance the robustness when expressing the node feature. The node feature of the target node is determined by the average values of the coordinates in the horizontal direction and the vertical direction, which can consider the node features of all event nodes in the target voxel and determine a more reasonable comprehensive feature. Finally, the maximum time stamp is selected as the time corresponding to the target node, which can ensure the causal consistency of the event time sequence. At the same time, the above method can realize data dimension reduction through voxel-level clustering, which can greatly reduce the computational complexity.

[0151] In some embodiments, based on the first coordinate average value in the horizontal direction and the second coordinate average value in the vertical direction of the node position of the at least one event node, the node position of the target node is determined, including any one of the following: determining the first coordinate average value as the coordinate value in the horizontal direction of the node position of the target node, and determining the second coordinate average value as the coordinate value in the vertical direction of the node position of the target node; rounding the first coordinate average value to obtain a third coordinate value, and rounding the second coordinate average value to obtain a fourth coordinate value; determining the third coordinate value as the coordinate value in the horizontal direction of the node position of the target node, and determining the fourth coordinate value as the coordinate value in the vertical direction of the node position of the target node.

[0152] For example, the above 10 event nodes include event node 1, event node 2, event node 3, event node 4, event node 5, event node 6, event node 7, event node 8, event node 9 and event node 10, and the process of clustering at least one event node in each voxel is described.

[0153] For the event nodes 1, 2, 3 and 4 included in the above-mentioned voxel 1, the maximum feature 3 is determined from the node features 2, 2.3, 3 and 2.6 corresponding to the event nodes 1, 2, 3 and 4, respectively, as the node feature of the target node corresponding to the voxel 1; based on the average values of (1, 2), (1.2, 2), (0.5.1.2) and (1, 1.3) on the horizontal axis and the vertical axis, the node position of the target node corresponding to the voxel 1 is determined as (0.925, 1.625); the maximum time of 0.1ms, 0.2ms, 0.2ms and 0.16ms is determined as the time corresponding to the target node corresponding to the voxel 1 as 0.2ms.

[0154] For the event node 5, the event node 6 and the event node 7 included in the voxel 2, the maximum feature 3.6 is determined as the node feature of the target node corresponding to the voxel 2 from the node feature corresponding to the event node 5, the node feature corresponding to the event node 6 and the node feature corresponding to the event node 7; the node position of the target node corresponding to the voxel 2 is determined as (2.63, 2.9) based on the average value on the horizontal axis and the average value on the vertical axis of (2.2, 2.3), (3, 2.6) and (2.7, 3.8); the maximum time of 0.3ms, 0.32ms and 0.4ms is determined as the time corresponding to the target node corresponding to the voxel 2, which is 0.4ms.

[0155] For the event node 8 and the event node 9 included in the voxel 3, the maximum feature 4 is determined as the node feature of the target node corresponding to the voxel 3 from the node feature corresponding to the event node 8 and the node feature corresponding to the event node 9; the node position of the target node corresponding to the voxel 3 is determined as (4.4, 3.7) based on the average value on the horizontal axis and the average value on the vertical axis of (4.2, 5) and (4.6, 4.9); the maximum time of 0.5ms and 0.45ms is determined as the time corresponding to the target node corresponding to the voxel 3, which is 0.5ms.

[0156] For the event node 10 included in the voxel 4, the node feature 2.6 corresponding to the event node 10 is determined as the node feature of the target node corresponding to the voxel 4; the node position (7.2, 6.3) of the event node 10 is determined as the node position of the target node corresponding to the voxel 4; 0.62ms is determined as the time corresponding to the target node corresponding to the voxel 4.

[0157] In some embodiments, determining the maximum feature in the node features of the at least one event node as the node feature of the target node corresponding to the target voxel comprises: determining the node feature of the target node corresponding to the target voxel based on the following formula (11);

[0158]

[0159] Wherein, n” o is the node feature of the target node corresponding to the target voxel T * is the number of event nodes in the target voxel T * is the node feature of the event node o in the target voxel T * .

[0160] In some embodiments, determining the node position of the target node based on the average value on the horizontal axis and the average value on the vertical axis of the node positions of the at least one event node comprises: determining the node position of the target node based on the following formula (12);​

[0161]

[0162] Among them, p″ o (x,y) is the node position of the target node, p' o (x,y) is the node position of the event node.

[0163] In some embodiments, in order to maintain resolution consistency, the horizontal coordinate value of the target node position can be divided by H, and the vertical coordinate value of the target node position can be divided by W. * , the time corresponding to the target node can be multiplied by β. Among them, H and W * are the height and width of the image frame respectively, and β is the temporal normalization factor.

[0164] It should be understood that directed edges in an event directed graph follow a temporal order, pointing only from event nodes with a shorter timeframe to event nodes with a longer timeframe. However, after voxelizing the N event nodes in the event directed graph, this temporal order constraint may be partially lost. To address this issue, the following solution is proposed.

[0165] Specifically, time information is used to constrain directed edges, ensuring that directed edges in the newly constructed event directed graph are retained only when the time corresponding to the source event node is less than the time corresponding to the target node. This condition effectively filters the number of directed edges after voxelization, helping to reduce redundant computation.

[0166] For example, continue with Figure 4 The event directed graph shown is used to illustrate this. Assume that event node A 12 and event node A 24 are divided into the same voxel, then Figure 4 Specifically, event node A 12 and event node A 24 Within the same voxel, based on Figure 4 The directed edges in the event directed graph in can be generated as follows Figure 6 The result shown in (a) is shown in the figure. 12 and event node A 24 The respective node positions and corresponding times, after determining that the node position of the target node A* is (1.5, 2) and the corresponding time is 1.5ms, can be generated as follows Figure 6 The result shown in (b) is shown in the figure. 34There are two directed edges between the target node A* and the event node A. Based on the principle that a directed edge can only be retained if the time corresponding to the source event node is less than the time corresponding to the target node, the directed edge from the target node A* to the event node A 34 is deleted, i.e., the event directed graph shown in (c) in Figure 6 is obtained.

[0167] It should be noted that the above is another way of determining each target feature, and another generation process of the motion trajectory of at least one physical point is specifically given in steps 801 to 806 in combination with the third way of determining each target feature.

[0168] Step 701: In a target time period, image frames captured by a frame camera and event information captured by an event camera in a target scene are obtained, the event information being information of a plurality of physical pixels in the event camera that have a change in luminance.

[0169] It should be understood that the meaning of "image frames" in the above step 701 and the corresponding embodiments and the meaning of "event information" and the corresponding embodiments are the same as the meaning of "image frames" in the aforementioned step 201 and the corresponding embodiments and the meaning of "event information" and the corresponding embodiments, which will not be repeated here.

[0170] Step 702: Feature extraction is performed on the image frames to obtain an image feature map, and an event directed graph is constructed based on the pixel positions and corresponding times of each physical pixel in the event information, the node positions of each event node in the event directed graph being related to the pixel positions, and the direction of the directed edge between two event nodes being related to the times corresponding to the two event nodes.

[0171] It should be understood that the meaning of "event directed graph" in the above step 702 and the corresponding embodiments is the same as the meaning of "event directed graph" in the aforementioned scheme and the corresponding embodiments, which will not be repeated here.

[0172] Step 703: Based on the node position of each event node in the N event nodes, the change polarity of the physical pixel corresponding to the event node, and the edge feature of the directed edge corresponding to the event node, the node feature of each event node is determined, N being a positive integer greater than 2.

[0173] Step 704: The N event nodes are voxelized, at least one event node in each voxel of M voxels is clustered, and the node feature, node position, and corresponding time of each target node of the M target nodes are obtained, one voxel corresponding to one target node, M being a positive integer greater than 1, and M being less than or equal to N.

[0174] It should be understood that the meaning of "voxelization processing" in the above step 704 and the corresponding implementation are the same as the meaning of "voxelization processing" in the foregoing scheme and the corresponding implementation, which will not be repeated here.

[0175] Step 705, sampling the image features matching the node positions and corresponding times of the respective target nodes from the image feature map; splicing the node features of the respective target nodes with the matched image features to obtain the respective target features.

[0176] Step 706, generating a motion trajectory of at least one physical point based on the respective target features and the Bezier curve algorithm, and generating a target image frame based on the motion trajectory of the at least one physical point and the background image region on the image frame, the plurality of physical pixels being used to reflect the changes of at least one physical point in the target scene within the target time period.

[0177] It should be understood that the respective specific implementations of the above steps 705 and 706 are the same as the principles of the respective specific implementations of the foregoing steps 202 and 203, which will not be repeated here.

[0178] Figure 8 is a schematic diagram for determining a target feature provided by an embodiment of the present application.

[0179] As shown in Figure 8 , the image frames collected by the frame camera under the target scene within the target time period are obtained, the image frames are first subjected to convolution processing to obtain first features, and the first features are then subjected to residual processing to obtain second features (image feature map). The event information collected by the event camera under the target scene within the target time period is obtained. Based on the pixel positions and corresponding times of the respective physical pixels in the event information, an event directed graph is constructed, the node positions of the respective event nodes in the event directed graph are related to the pixel positions, and the directions of the directed edges between two event nodes are related to the respective corresponding times of the two event nodes. In a spline convolution manner, based on the node position of each event node in the N event nodes, the change polarity of the physical pixel corresponding to the event node, and the edge feature of the directed edge corresponding to the event node, the node feature of each event node is determined. The N event nodes are subjected to voxelization processing, at least one event node in each voxel of the M voxels is clustered (i.e., a pooling operation) to obtain the node feature, node position and corresponding time of each target node in the M target nodes, and one voxel corresponds to one target node. The image features matching the node positions and corresponding times of the respective target nodes are sampled from the image feature map; the node features of the respective target nodes are spliced with the matched image features to obtain the respective target features.

[0180] Furthermore, the spline convolution process and pooling operation can be performed on each target feature obtained above to obtain an updated node feature, and residual processing can be further performed on the image feature map for the image frame to obtain a processed image feature map, and updated image features that match the updated node features are sampled from the processed image feature map. Finally, each updated node feature is spliced ​​with the matched updated image feature to obtain each updated target feature;

[0181] Of course, each target feature can be updated again according to the above process. Figure 8 Only three processes of determining target features are illustrated.

[0182] Figure 9 It is a structural diagram of an image processing device provided in an embodiment of the present application.

[0183] For example, Figure 9 As shown, the apparatus 900 includes:

[0184] An acquisition module 901 is configured to acquire image frames captured by a frame camera and event information captured by an event camera in a target scene within a target time period. The event information is information about a plurality of physical pixels in the event camera that have brightness changes.

[0185] Determination module 902 is configured to extract features from the image frame to obtain an image feature map, sample image features that match each physical pixel in the event information from the image feature map, and concatenate the pixel features of each physical pixel with the matching image features to obtain each target feature;

[0186] Generation module 903 is used to generate a motion trajectory of at least one physical point based on various target features and a Bezier curve algorithm, and to generate a target image frame based on the motion trajectory of the at least one physical point and the background image area on the image frame. The multiple physical pixels are used to reflect the changes of at least one physical point in the target scene within the target time period.

[0187] Optionally, the apparatus 900 further includes: a constructing module configured to construct an event directed graph based on the pixel position of each physical pixel in the event information and the corresponding time, wherein the node position of each event node in the event directed graph is related to the pixel position, and the direction of the directed edge between two event nodes is related to the corresponding time of the two event nodes; the determining module 902 is specifically configured to determine the node feature of each event node based on the node position of each event node in the N event nodes, the change polarity of the physical pixel corresponding to the event node, and the edge feature of the directed edge corresponding to the event node, N being a positive integer greater than 2; and the determining module 902 is specifically further configured to: sample the image feature matching the node position and the corresponding time of each event node from the image feature map; and splice the node feature of each event node with the matched image feature to obtain each target feature.

[0188] Optionally, the determining module 902 is specifically further configured to: determine the initial feature of each event node based on the node position of the event node and the change polarity of the physical pixel corresponding to the event node; perform linear transformation on the initial feature of any event node as a target event node to obtain a first feature; in the case that there is a neighboring event node of the target event node, determine at least one weight based on the edge feature of the directed edge between each neighboring event node in the at least one neighboring event node and the target event node, and perform weighted summation on the initial feature of the at least one neighboring event node based on the at least one weight to obtain a second feature; and determine the sum of the first feature and the second feature as the node feature of the target event node to obtain the node feature of each event node.

[0189] Optionally, the apparatus 900 further includes: a voxel processing module configured to perform voxelization processing on the N event nodes, and cluster at least one event node in each voxel of M voxels to obtain the node feature, the node position and the corresponding time of each target node of M target nodes, one voxel corresponding to one target node, M being a positive integer greater than 1 and M being less than or equal to N; and the determining module 902 is specifically further configured to: sample the image feature matching the node position and the corresponding time of each target node from the image feature map; and splice the node feature of each target node with the matched image feature to obtain each target feature.

[0190] Optionally, the voxel processing module is specifically configured to: take any voxel as a target voxel, for at least one event node in the target voxel, determine a maximum feature in the node features of the at least one event node as a node feature of a target node corresponding to the target voxel, determine a node position of the target node based on a first coordinate mean value of the node position of the at least one event node in a horizontal direction and a second coordinate mean value of the node position of the at least one event node in a vertical direction, and determine a maximum time corresponding to the at least one event node as a time corresponding to the target node.

[0191] Optionally, the determining module 902 is further configured to: classify the plurality of target features to obtain a plurality of groups of features, each group of features corresponding to a physical point, and a plurality of features in each group of features being used to reflect a change of the corresponding physical point in the target time period; determine, based on each group of features in the plurality of groups of features, a plurality of control positions used to fit a motion trajectory of each physical point in the plurality of physical points; and the generating module 903 is configured to generate the motion trajectory of each physical point based on the plurality of control positions and the Bezier curve algorithm.

[0192] Optionally, the determining module 902 is further configured to: take any group of features as a target group of features, for a plurality of pairs of adjacent features in the target group of features, determine a correlation matrix between each pair of adjacent features, and determine an initial control position used to generate a motion trajectory of a target physical point corresponding to the target group of features based on the plurality of correlation matrices corresponding to the target group of features and the initial control position; and the generating module 903 is configured to determine a plurality of control positions used to generate the motion trajectory of the target physical point based on the initial control position and the plurality of correlation matrices corresponding to the target group of features.

[0193] Optionally, the determining module 902 is further configured to: take any physical point as a target physical point, perform formula transformation on the Bezier curve algorithm based on a preset order corresponding to the Bezier curve algorithm to obtain a target expression after transformation; and the generating module is further configured to determine a plurality of predicted positions based on a plurality of preset times, an initial position of the target physical point, the plurality of control positions, and the target expression, to generate the motion trajectory of the target physical point.

[0194] Figure 10 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application.

[0195] As shown in FIG. 1, the electronic device 1000 includes a memory 1001 and a processor 1002, wherein the memory 1001 stores executable program code 1003, and the processor 1002 is configured to invoke and execute the executable program code 1003 to perform an image processing method. Figure 10

[0196] ​In addition, the embodiment of the present application also protects a device, which can include a memory and a processor, wherein the memory stores executable program code, and the processor is configured to invoke and execute the executable program code to perform the image processing method provided by the embodiment of the present application.

[0197] The embodiment can divide the device into functional modules according to the above method examples. For example, each functional module can be provided, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware. It should be noted that the division of the modules in the embodiment is illustrative, and is only a logical function division. In actual implementation, another division mode can be used.

[0198] When each functional module is divided according to each function, the device can further include an acquisition module, a determination module, a generation module, a construction module, a voxel processing module, and the like. It should be noted that all related contents involved in the above method embodiments can be referred to the function description of the corresponding functional module, and will not be repeated here.

[0199] It should be understood that the device provided by the embodiment is used to perform the above image processing method, and thus the same effect as the above implementation method can be achieved.

[0200] When the integrated unit is used, the device can include a processing module and a storage module. When the device is applied to an electronic device, the processing module can be used to control and manage the actions of the electronic device. The storage module can be used to support the electronic device to execute related executable program codes and the like.

[0201] The processing module can be a processor or a controller, which can implement or execute various exemplary logical blocks, modules and circuits shown in combination with the disclosure of the present application. The processor can also be a combination of computing functions, such as one or more microprocessor combinations, digital signal processing (DSP) and microprocessor combinations, and the like. The storage module can be a memory.

[0202] In addition, the device provided by the embodiment of the present application can be a chip, an assembly or a module. The chip can include a connected processor and a memory. The memory is used to store instructions, and when the processor invokes and executes the instructions, the chip can perform the image processing method provided by the above embodiment.

[0203] The embodiment also provides a computer readable storage medium, which stores executable program code. When the executable program code runs on the computer, the computer executes the above related method steps to implement the image processing method provided by the above embodiment.

[0204] The embodiment further provides a computer program product, which, when running on a computer, causes the computer to execute the above related steps to implement the image processing method provided by the above embodiment.

[0205] Wherein, the apparatus, the computer readable storage medium, the computer program product or the chip provided by the embodiment are all used to execute the corresponding method provided above, thus the beneficial effects that can be achieved are referable to the beneficial effects in the corresponding method provided above, which will not be repeated here.

[0206] Through the above description of the implementation mode, those skilled in the art can understand that, for the convenience and brevity of description, only the above division of functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above.

[0207] In the embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiment described above is only illustrative, for example, the division of modules or units is only a logical function division, and in actual implementation, there can be another division manner, for example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, and can be electrical, mechanical or other forms.

[0208] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image processing method, characterized by, The method comprises: acquiring, in a target period, an image frame collected by a frame camera and event information collected by an event camera in a target scene, the event information being information of a plurality of physical pixels in the event camera that have brightness changes; performing feature extraction on the image frame to obtain an image feature map, sampling image features matching each physical pixel in the event information from the image feature map, splicing pixel features of each physical pixel with the matching image features to obtain each target feature; generating a motion trajectory of at least one physical point based on each target feature and a Bezier curve algorithm, and generating a target image frame based on the motion trajectory of the at least one physical point and a background image region on the image frame, the plurality of physical pixels being used to reflect changes of at least one physical point in the target scene in the target period; wherein the generating of the motion trajectory of the at least one physical point based on each target feature and the Bezier curve algorithm comprises: classifying a plurality of target features to obtain a plurality of groups of features, each group of features corresponding to one physical point, and a plurality of features in each group of features being used to reflect changes of the corresponding physical point in the target period; determining a plurality of control positions for fitting a motion trajectory of each physical point in the plurality of physical points based on each group of features in the plurality of groups of features; generating the motion trajectory of each physical point based on the plurality of control positions and the Bezier curve algorithm.

2. The method of claim 1, wherein, The method further comprises: constructing an event directed graph based on pixel positions and corresponding times of each physical pixel in the event information, a node position of each event node in the event directed graph being related to a pixel position, and a direction of a directed edge between two event nodes being related to corresponding times of the two event nodes; determining a node feature of each event node based on a node position of each event node in N event nodes, a change polarity of a physical pixel corresponding to the event node, and an edge feature of a directed edge corresponding to the event node, N being a positive integer greater than 2; and the sampling of image features matching each physical pixel in the event information from the image feature map, the splicing of pixel features of each physical pixel with the matching image features to obtain each target feature, comprises: sampling image features matching a node position and a corresponding time of each event node from the image feature map; splicing the node feature of each event node with the matching image features to obtain each target feature.

3. The method of claim 2, wherein, The determining of the node feature of each event node based on the node position of each event node in the N event nodes, the change polarity of the physical pixel corresponding to the event node, and the edge feature of the directed edge corresponding to the event node comprises: determining an initial feature of each event node based on the node position of each event node and the change polarity of the physical pixel corresponding to the event node; performing linear transformation on the initial feature of the target event node to obtain a first feature. determine at least one weight based on an edge feature of a directed edge between each of at least one adjacent event node of the target event node and the target event node, and perform a weighted sum on initial features of the at least one adjacent event node based on the at least one weight to obtain a second feature; determine a node feature of the target event node as a sum of the first feature and the second feature to obtain the node feature of each of the event nodes.

4. The method of claim 2, wherein, The method further comprises: perform voxelization processing on the N event nodes to cluster at least one event node in each of M voxels to obtain a node feature, a node position, and a corresponding time of each of M target nodes, one voxel corresponding to one target node, M being a positive integer greater than 1 and less than or equal to N; and the sampling, from the image feature map, an image feature matching each of the physical pixels in the event information, splicing a pixel feature of each of the physical pixels with the matching image feature to obtain each of the target features comprises: sampling, from the image feature map, an image feature matching the node position and the corresponding time of each of the target nodes; splicing the node feature of each of the target nodes with the matching image feature to obtain each of the target features.

5. The method of claim 4, wherein, The clustering, of at least one event node in each of M voxels, to obtain a node feature, a node position, and a corresponding time of each of M target nodes comprises: taking any voxel as a target voxel, determining a maximum feature in a node feature of at least one event node in the target voxel as a node feature of a target node corresponding to the target voxel; determining the node position of the target node based on a first coordinate mean value of the node position of the at least one event node in a horizontal direction and a second coordinate mean value in a vertical direction; determining a maximum time corresponding to the at least one event node as a corresponding time of the target node.

6. The method of claim 1, wherein, The determining, of a plurality of control positions for fitting a motion trajectory of each of a plurality of physical points, based on each of the plurality of groups of features comprises: taking any group of features as a target group of features, determining a correlation matrix between each pair of adjacent features in the target group of features; determining an initial control position for generating a motion trajectory of a target physical point corresponding to the target group of features; determining a plurality of control positions for generating a motion trajectory of the target physical point based on the initial control position and a plurality of correlation matrices corresponding to the target group of features.

7. The method of claim 1, wherein, The generating, of a motion trajectory of each of the physical points based on the plurality of control positions and the Bezier curve algorithm comprises: taking any physical point as a target physical point, performing formula transformation on the Bezier curve algorithm to obtain a transformed target expression based on a preset order corresponding to the Bezier curve algorithm; Based on the preset plurality of times, the initial position of the target physical point, the plurality of control positions and the target expression, a plurality of predicted positions are determined to generate a motion trajectory of the target physical point.

8. An image processing apparatus characterized by comprising: The device comprises: An acquisition module is configured to acquire an image frame captured by a frame camera and event information captured by an event camera in a target scene within a target period, the event information being information of a plurality of physical pixels in the event camera that have brightness changes; A determination module is configured to perform feature extraction on the image frame to obtain an image feature map, sample image features matched with each physical pixel in the event information from the image feature map, and splice pixel features of each physical pixel with the matched image features to obtain each target feature; A generation module is configured to generate a motion trajectory of at least one physical point based on each target feature and a Bezier curve algorithm, generate a target image frame based on the motion trajectory of the at least one physical point and a background image region on the image frame, and use the plurality of physical pixels to reflect position changes of at least one physical point in the target scene within the target period; The generation module is specifically configured to: classify a plurality of target features to obtain a plurality of groups of features, each group of features corresponding to one physical point, and a plurality of features in each group of features being used to reflect changes of the corresponding physical point within the target period; determine a plurality of control positions used to fit a motion trajectory of each physical point in the plurality of physical points based on each group of features in the plurality of groups of features; and generate a motion trajectory of each physical point based on the plurality of control positions and the Bezier curve algorithm.

9. An electronic device, comprising: The electronic device comprises: a memory configured to store executable program code; a processor configured to call and run the executable program code from the memory, so that the electronic device performs the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores executable program code, when the executable program code is executed, the method of any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Machine vision monitoring image processing algorithm and system

    CN118154622A