Visual bionic imaging method, device and equipment based on single lens and storage medium
By combining vision sensors and dynamic differential sensors in a single lens, image data is acquired and processed, and target images are generated using imaging fusion models, the image quality problem of traditional cameras in high dynamic range and fast moving target scenes is solved, and high-quality images are captured and aligned.
Patent Information
- Application Number
- CN202510078424.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional cameras are difficult to capture high-quality images in high dynamic range and fast moving target scenarios, and the image alignment is difficult in the prior art, resulting in low image quality.
A single lens is used to combine vision sensors and dynamic differential sensors to collect event stream data and image data to be imaged, and an event voxel grid is generated through voxel encoding processing, and an imaging fusion model is input to generate a target image.
The alignment error of binocular stereoscopic vision devices or spectroscopic mirror systems is avoided, the image quality is improved, and the capture ability of fast moving targets is enhanced.
Smart Images

Figure CN120017981A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a single-lens-based visual simulation imaging method, device, equipment and storage medium. Background Art
[0002] When capturing images, traditional cameras have a small high dynamic range and are prone to overexposure in complex lighting scenes such as high reflections. In addition, traditional cameras cannot effectively capture fast-moving targets, which can easily lead to dynamic blur in the image. Based on this, the current event camera (dynamic differential sensor) has the characteristics of fast response, high dynamic range, and the ability to quickly capture moving targets. It can make up for the shortcomings of traditional cameras by asynchronously outputting event stream data. However, in the process of using event cameras (dynamic differential sensors) to obtain event data streams, there is a problem that the rich image detail information collected by traditional cameras cannot be obtained.
[0003] In the prior art, the event data stream collected by the event camera (dynamic differential sensor) is fused with the image data collected by the traditional camera to obtain high-quality image frames. However, the event data stream and the image data need to be aligned before fusion. Currently, data alignment is achieved by binocular stereo vision devices and binocular systems based on spectroscopes.
[0004] However, since binocular stereo vision devices are based on the principle of parallax to achieve depth perception, they need to be precisely calibrated and have large alignment errors. The binocular system based on beam splitters needs to share an optical path, which has problems with lens distortion correction, hardware interchangeability, and convenience. Based on this, using existing technologies, it is still impossible to obtain high-quality images. Summary of the invention
[0005] Based on this, it is necessary to provide a visual simulation imaging method, device, equipment and storage medium based on a single lens for the above technical problems. Among them, the single lens includes a visual sensor and a dynamic differential sensor, and the event stream data within a preset time length is collected by the dynamic differential sensor, and the image data to be imaged corresponding to the event stream data is collected by the visual sensor. The event stream data is voxel encoded to obtain an event voxel grid corresponding to the event stream data. The event voxel grid and the image data to be imaged are input into a trained imaging fusion model to obtain a target image output by the imaging fusion model. In this way, a single lens including a visual sensor and a dynamic differential sensor is used to collect event stream data and image data to be imaged, thereby avoiding the difficulty in image alignment processing when using a binocular stereo vision device or a binocular system based on a spectroscope for image alignment in the prior art, and the imaging fusion model is used to obtain the target image, thereby improving the quality of the target image.
[0006] In a first aspect, an embodiment of the present invention provides a visual simulation imaging method based on a single lens, wherein the single lens includes a visual sensor and a dynamic differential sensor, and the method includes:
[0007] The dynamic differential sensor is used to collect event stream data within a preset time period, and the visual sensor is used to collect image data to be imaged corresponding to the event stream data;
[0008] Performing voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data;
[0009] The event voxel grid and the image data to be imaged are input into a trained imaging fusion model to obtain a target image output by the imaging fusion model.
[0010] In one embodiment, the image data to be imaged includes first image data and second image data, and the image data to be imaged corresponding to the event stream data collected by the visual sensor includes:
[0011] At the start time of the preset time period, collecting the first image data by the visual sensor;
[0012] At the end of the preset time period, the second image data is collected by the visual sensor.
[0013] In one embodiment, performing voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data includes:
[0014] Acquire multiple event data within the preset time length, wherein each event data is represented by e=(x, y, p, t), where x represents the abscissa of the event, y represents the ordinate of the event, p represents the polarity of the event, and t represents the timestamp of the event;
[0015] According to the abscissa, the ordinate, the timestamp and the polarity of the plurality of event data, voxel encoding processing is performed on the event stream data to obtain the corresponding event voxel grid.
[0016] In one embodiment, performing voxel encoding processing on the event stream data according to the abscissa, the ordinate and the polarity of the plurality of event data to obtain the corresponding event voxel grid includes:
[0017] Determine polarity values of the polarities corresponding to the plurality of event data respectively;
[0018] Determine, in a preset voxel grid, target filling positions corresponding to the plurality of event data respectively according to the abscissas, the ordinates, and the timestamps corresponding to the plurality of event data respectively;
[0019] According to the target filling positions respectively corresponding to the plurality of event data, the polarity values respectively corresponding to the plurality of event data are filled into the preset voxel grid to obtain the event voxel grid.
[0020] In one embodiment, before performing voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data, the method further includes:
[0021] The event stream data and the image data to be imaged are corrected by a checkerboard calibration method.
[0022] In one embodiment, the imaging fusion model includes: an input layer, a first fusion network, a second fusion network and an output layer, inputting the event voxel grid and the image data to be imaged into the trained imaging fusion model, and obtaining the target image output by the imaging fusion model, including:
[0023] Inputting the event voxel grid and the image data to be imaged into the first fusion network through the input layer to obtain a first fused image and a second fused image;
[0024] The first fused image and the second fused image are input into the second fusion network for fusion, and a target image is obtained through the output layer.
[0025] In one embodiment, inputting the event voxel grid and the image data to be imaged into the first fusion network to obtain a first fused image and a second fused image includes:
[0026] fusing the event voxel grid and the first image data through the first fusion network to obtain the first fused image;
[0027] The event voxel grid and the second image data are fused through the first fusion network to obtain the second fused image.
[0028] In a second aspect, an embodiment of the present invention provides a visual simulation imaging device based on a single lens, wherein the single lens includes a visual sensor and a dynamic differential sensor, and the device includes:
[0029] A collection module, used to collect event stream data within a preset time period through the dynamic differential sensor, and to collect image data to be imaged corresponding to the event stream data through the visual sensor;
[0030] An event voxel grid acquisition module, used for performing voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data;
[0031] The target image acquisition module is used to input the event voxel grid and the image data to be imaged into the trained imaging fusion model to acquire the target image output by the imaging fusion model.
[0032] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the single-lens based visual simulation imaging method described in the first aspect when executing the computer program.
[0033] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the single-lens based visual simulation imaging method described in the first aspect.
[0034] The technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art:
[0035] The embodiment of the present invention provides a visual simulation imaging method based on a single lens, wherein the single lens includes a visual sensor and a dynamic differential sensor. In this way, the dynamic differential sensor is used to collect event stream data within a preset time length, and the image data to be imaged corresponding to the event stream data is collected by the visual sensor. The event stream data is subjected to voxel encoding processing to obtain an event voxel grid corresponding to the event stream data. The event voxel grid and the image data to be imaged are input into a trained imaging fusion model to obtain a target image output by the imaging fusion model. In this way, the event stream data and the image data to be imaged are collected using a single lens including a visual sensor and a dynamic differential sensor, thereby avoiding the difficulty in image alignment processing when using a binocular stereo vision device or a binocular system based on a spectroscope for image alignment in the prior art, and the imaging fusion model is used to obtain the target image, thereby improving the quality of the target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0038] Figure 1 A schematic flow chart of a single-lens-based visual simulation imaging method provided by an embodiment of the present invention;
[0039] Figure 2 A schematic diagram of the structure of an imaging fusion model provided by an embodiment of the present invention;
[0040] Figure 3 A schematic structural diagram of a single-lens visual simulation imaging device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
[0042] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present invention, rather than all of the embodiments.
[0043] When capturing images, traditional cameras have a small high dynamic range and are prone to overexposure in complex lighting scenes such as high reflections. In addition, traditional cameras cannot effectively capture fast-moving targets, which can easily lead to dynamic blur in images. Based on this, the current event camera (dynamic differential sensor) has the characteristics of fast response, high dynamic range, and the ability to quickly capture moving targets. By asynchronously outputting event stream data, it can make up for the shortcomings of traditional cameras. However, in the process of using event cameras (dynamic differential sensors) to obtain event data streams, there is a problem that the rich image detail information collected by traditional cameras cannot be obtained.
[0044] In the prior art, the event data stream collected by the event camera (dynamic differential sensor) is fused with the image data collected by the traditional camera to obtain high-quality image frames. However, the event data stream and the image data need to be aligned before fusion. Currently, data alignment is achieved by binocular stereo vision devices and binocular systems based on spectroscopes.
[0045] However, since binocular stereo vision devices are based on the principle of parallax to achieve depth perception, they need to be precisely calibrated and have large alignment errors. The binocular system based on beam splitters needs to share an optical path, which has problems with lens distortion correction, hardware interchangeability, and convenience. Based on this, using existing technologies, it is still impossible to obtain high-quality images.
[0046] Therefore, the present invention provides a visual simulation imaging method based on a single lens, wherein the single lens includes a visual sensor and a dynamic differential sensor, and the event stream data within a preset time length is collected by the dynamic differential sensor, and the image data to be imaged corresponding to the event stream data is collected by the visual sensor. The event stream data is subjected to voxel encoding processing to obtain an event voxel grid corresponding to the event stream data. The event voxel grid and the image data to be imaged are input into a trained imaging fusion model to obtain a target image output by the imaging fusion model. In this way, the event stream data and the image data to be imaged are collected by a single lens including a visual sensor and a dynamic differential sensor, thereby avoiding the difficulty in image alignment processing when using a binocular stereo vision device or a binocular system based on a spectroscope for image alignment in the prior art, and the target image is obtained by using the imaging fusion model, thereby improving the quality of the target image.
[0047] In one embodiment, Figure 1 As shown, Figure 1 A flow chart of a single-lens visual bionic imaging method provided by an embodiment of the present invention includes: a visual sensor and a dynamic differential sensor, wherein the visual sensor refers to a method based on picture frame imaging, i.e., used to obtain traditional image frames, and the dynamic differential sensor refers to a method based on the event-driven principle, using a differential circuit to output event data greater than a preset threshold, i.e., to obtain event stream data with a high dynamic range. Based on this, the following steps are specifically included:
[0048] S10: collecting event stream data within a preset time period through a dynamic differential sensor, and collecting image data to be imaged corresponding to the event stream data through a visual sensor.
[0049] Among them, the preset duration refers to the duration set for collecting the corresponding event data stream when acquiring the event feature graph through the dynamic differential sensor, such as 30ms, that is, collecting 30ms of event data stream through the dynamic differential sensor, but it is not limited to this. The present invention is not specifically limited, and those skilled in the art can set it according to actual conditions.
[0050] The image data to be imaged corresponding to the above event stream data refers to the image data collected by the visual sensor under preset conditions such as the same shooting object, shooting position, shooting time and shooting angle, etc. However, it is not limited to this, the present invention is not specifically limited, and those skilled in the art can set it according to actual conditions.
[0051] Specifically, for the dynamic differential sensor arranged in a single lens, a preset time length is set in advance, and the event data stream within the preset time length is collected by the dynamic differential sensor, and the image data to be imaged corresponding to the event stream data one by one is collected by the visual sensor arranged in the single lens.
[0052] Optionally, based on the above embodiment, in some embodiments of the present invention, the image data to be imaged includes first image data and second image data. Based on this, an implementation method of collecting the image data to be imaged corresponding to the event stream data through a visual sensor may be:
[0053] S101: At the start time of a preset time period, first image data is collected by a visual sensor.
[0054] Specifically, at the starting moment of the preset time length, when the dynamic differential sensor starts to collect the event data stream within the preset time length, the visual sensor is controlled to collect the first image data corresponding to the event data stream, that is, to collect the first image data at the starting moment of the preset time length.
[0055] Optionally, based on the above embodiments, in some embodiments of the present invention, the single lens also includes a timestamp alignment module. Assuming that the starting time of the preset duration is T1, when the dynamic differential sensor starts to collect the event data stream at time T1, the timestamp alignment module uses the waveform signal generating function to control the visual sensor to collect the first image data corresponding to the event data stream, that is, to collect the first image data at time T1.
[0056] S102: At the end of the preset time period, second image data is collected by a visual sensor.
[0057] Specifically, at the end of the preset duration, when the dynamic differential sensor collects the event data stream at the last moment within the preset duration, the visual sensor is controlled to collect the second image data corresponding to the event data stream, that is, to collect the second image data at the end of the preset duration.
[0058] Optionally, following the above embodiment, assuming that the end time of the preset duration is T2, when the event data stream is collected at time T2 by the dynamic differential sensor, the timestamp alignment module is used to utilize the waveform signal generating function to control the visual sensor to collect the second image data corresponding to the event data stream, that is, to collect the second image data at time T2.
[0059] In this way, this embodiment can collect the first image data at the start time of the preset time length through the visual sensor, and collect the second image data at the end time of the preset time length through the visual sensor, so that the event data stream is aligned in time with the image data to be imaged.
[0060] S11: Perform voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data.
[0061] Voxel is the abbreviation of volume element. A volume containing voxels can be expressed by stereo rendering or extracting polygonal isosurfaces with given threshold contours. Voxel is the smallest unit of digital data in three-dimensional space segmentation, similar to the smallest unit pixel in two-dimensional space, that is, voxel is a pixel in three-dimensional space. The above-mentioned voxel grid refers to the one composed of voxels.
[0062] Specifically, after the event data stream within a preset time length is collected by the dynamic differential sensor, voxel encoding is performed on the event stream data to obtain an event voxel grid corresponding to the event stream data.
[0063] Optionally, based on the above embodiment, in some embodiments of the present invention, an implementation manner of S11 may be:
[0064] S111: Acquire multiple event data within a preset time period.
[0065] Each event data is represented by e=(x, y, p, t), where x represents the horizontal coordinate of the event, y represents the vertical coordinate of the event, p represents the polarity of the event, and t represents the timestamp of the event.
[0066] Multiple event data refers to multiple events included in a preset duration. Following the above embodiment, the preset duration is set to 30ms, and 30ms includes multiple events, that is, it can be understood that event data of multiple events can be collected within 30ms, that is, the horizontal coordinate of the event, the vertical coordinate of the event, the polarity of the event, and the timestamp of the event. However, it is not limited to this, and the present invention is not specifically limited, and those skilled in the art can set it according to actual conditions.
[0067] It should be noted that the polarity of the event reflects the trend of light intensity change at the current moment. When the light intensity increases, the event polarity is determined to be 1, indicating that the event is an ON event (i.e., a positive polarity event). When the light intensity decreases, the event polarity is determined to be -1, indicating that the event is an OFF event (i.e., a negative polarity event). It can be understood that when there is a change in light intensity and exceeds a certain threshold, that is, an event occurs, the polarity value of the polarity of the event is determined to be 1, otherwise, the polarity value of the polarity of the event is determined to be 0. Based on this, since the dynamic differential sensor only responds to changes in light intensity, compared with the visual sensor, it does not require energy integration process and AD conversion process, and has the characteristics of fast imaging speed and large dynamic range.
[0068] S112: Perform voxel encoding processing on the event stream data according to the horizontal coordinates, vertical coordinates, timestamps and polarities of the multiple event data to obtain a corresponding event voxel grid.
[0069] Specifically, multiple event data within a preset time length are collected by a dynamic differential sensor. Furthermore, the event stream data is voxel encoded according to the horizontal coordinates, vertical coordinates, timestamps and polarities corresponding to the multiple event data to obtain an event voxel grid corresponding to the event stream data.
[0070] Optionally, based on the above embodiment, in some embodiments of the present invention, an implementation manner of S112 may be:
[0071] S1121: Determine polarity values corresponding to polarities of a plurality of event data.
[0072] Specifically, for a plurality of event data, a polarity value corresponding to each event is determined.
[0073] Exemplarily, following the above embodiment, when the light intensity changes and exceeds a certain threshold, an event occurs, and the polarity value of the polarity of the current event is determined to be 1, otherwise, the polarity value is 0, so as to obtain the polarity value corresponding to each event. However, this is not limited to this, and the present invention is not specifically limited, and those skilled in the art can set it according to actual conditions.
[0074] S1122: Determine target filling positions corresponding to the plurality of event data in a preset voxel grid according to the abscissas, ordinates, and timestamps corresponding to the plurality of event data.
[0075] Among them, the preset voxel grid is used to obtain the event voxel grid corresponding to the event stream data. The size of the preset voxel grid can be determined according to the event stream data and the data to be imaged. The present invention does not specifically limit it, and those skilled in the art can set it according to actual conditions.
[0076] Specifically, according to the abscissa, ordinate, and time stamp corresponding to each event data, a target filling position of the polarity value of each event data is determined in a preset voxel grid.
[0077] Exemplarily, assuming that a vertex of the preset voxel grid is the coordinate origin, and the three sides are the coordinate axes x-axis, y-axis and time t-axis, then the target filling position of the polarity value of each event data is determined in the preset voxel grid according to the horizontal coordinate, vertical coordinate and timestamp corresponding to each event data. However, it is not limited to this, the present invention is not specifically limited, and those skilled in the art can set it according to actual conditions.
[0078] S1123: According to the target filling positions respectively corresponding to the plurality of event data, the polarity values respectively corresponding to the plurality of event data are filled into a preset voxel grid to obtain an event voxel grid.
[0079] Specifically, after obtaining the target filling positions and polarity values respectively corresponding to the multiple event data, the polarity values respectively corresponding to the multiple event data are filled into the preset voxel grid to obtain the event voxel grid corresponding to the event stream data.
[0080] In this way, this embodiment can obtain the event voxel grid corresponding to the event stream data by performing voxel encoding processing on the event stream data, ensuring that the event stream data can be subsequently learned and features extracted through the imaging fusion model to obtain a high-quality target image.
[0081] S12: Input the event voxel grid and the image data to be imaged into the trained imaging fusion model to obtain the target image output by the imaging fusion model.
[0082] The imaging fusion model is obtained by training based on a sample training set, and the sample training set includes: sample event voxel grids corresponding to sample event stream data within a plurality of preset time lengths, and sample image data corresponding to the sample event voxel grids.
[0083] Specifically, the event voxel grid corresponding to the event stream data and the image data to be imaged are input into the trained imaging fusion model to obtain the target image output by the imaging fusion model.
[0084] Thus, the present embodiment provides a single-lens-based visual simulation imaging method, wherein the single lens includes a visual sensor and a dynamic differential sensor, and the event stream data within a preset time length is collected by the dynamic differential sensor, and the image data to be imaged corresponding to the event stream data is collected by the visual sensor. The event stream data is subjected to voxel encoding processing to obtain an event voxel grid corresponding to the event stream data. The event voxel grid and the image data to be imaged are input into a trained imaging fusion model to obtain a target image output by the imaging fusion model. In this way, a single lens including a visual sensor and a dynamic differential sensor is used to collect event stream data and image data to be imaged, thereby avoiding the difficulty in image alignment processing when using a binocular stereo vision device or a binocular system based on a spectroscope for image alignment in the prior art, and the imaging fusion model is used to obtain the target image, thereby improving the quality of the target image.
[0085] Optionally, based on the above embodiments, in some embodiments of the present invention, reference is made to Figure 2 As shown, Figure 2 A schematic diagram of the structure of an imaging fusion model provided by an embodiment of the present invention, the imaging fusion model includes an input layer 20, a first fusion network 21, a second fusion network 22, and an output layer 23.
[0086] Among them, the first fusion network 21 can be a U-Net network, using the U-Net network, through the multi-layer convolution and pooling layers contained in the U-Net network, the event voxel grid and the image data to be imaged are convoluted and pooled, so that the event voxel grid and the high-level feature information of the image data to be imaged can be extracted. Through the decoding stage of the U-Net network, the spatial resolution of the feature image can be restored by upsampling operation. And the U-Net network can directly pass the feature information of each layer in the encoding stage to the corresponding layer in the decoding stage by means of jump connection, thereby retaining the event voxel grid and the low-level feature information of the image data to be imaged, so as to obtain a target image with rich detail information and improve the quality of the target image. The above-mentioned second fusion network 22 can be, for example, a residual network, but is not limited to this. The present invention is not specifically limited, and those skilled in the art can set it according to actual conditions. Based on this, an implementation method of S12 can be:
[0087] S121: Inputting the event voxel grid and the image data to be imaged into the first fusion network through the input layer to obtain a first fused image and a second fused image.
[0088] Specifically, the event voxel grid corresponding to the event stream data and the image data to be imaged are input into the first fusion network through the input layer, and the first fusion network is used to perform feature extraction and fusion on the event voxel grid and the image data to be imaged, so as to obtain the first fused image and the second fused image output by the first fusion network.
[0089] Optionally, based on the above embodiment, in some embodiments of the present invention, an implementation manner of S121 may be:
[0090] S1211: Fusing the event voxel grid and the first image data through a first fusion network to obtain a first fused image.
[0091] Specifically, the input event voxel grid and the first image data are feature extracted and fused through the first fusion network to obtain a first fused image output by the first fusion network.
[0092] S1212: Fusing the event voxel grid and the second image data through the first fusion network to obtain a second fused image.
[0093] Specifically, the input event voxel grid and the second image data are feature extracted and fused through the first fusion network to obtain a second fused image output by the first fusion network.
[0094] S122: Input the first fused image and the second fused image into the second fusion network for fusion, and obtain the target image through the output layer.
[0095] Specifically, after the first fusion image and the second fusion image are obtained through the first fusion network, the first fusion image and the second fusion image are input into the second fusion network, the first fusion image and the second fusion image are fused using the second fusion network, and the target image is obtained by outputting through the output layer.
[0096] In this way, the present invention first extracts features from the event voxel grid and the first image data and the second image data respectively through the first fusion network included in the imaging fusion model, so as to obtain a first fused image and a second fused image with high dynamics, high resolution, and rich detail information, and further fuses the first fused image and the second fused image through the second fusion network to obtain a high-quality target image.
[0097] Optionally, based on the above embodiment, in some embodiments of the present invention, before executing S11, the following steps are further included:
[0098] S21: Correcting the event stream data and the image data to be imaged by using a checkerboard calibration method.
[0099] Among them, the checkerboard calibration method is a commonly used calibration method. By using a known checkerboard pattern to calibrate the intrinsic and extrinsic parameters of the camera, the distortion problem of the event stream data and the image data to be imaged can be avoided.
[0100] Specifically, before voxel encoding is performed on the event stream data to obtain an event voxel grid corresponding to the event stream data, the event stream data and the image data to be imaged are further corrected by a checkerboard calibration method.
[0101] In this way, the present invention corrects the event stream data and the image data to be imaged by using the checkerboard calibration method, so as to improve the image quality of the target image subsequently acquired based on the event stream data and the image data to be imaged.
[0102] It should be understood that although Figure 1 to Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 to Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0103] In one embodiment, Figure 3 As shown, a visual simulation imaging device based on a single lens is provided. The single lens includes a visual sensor and a dynamic differential sensor. The device includes: an acquisition module 10, an event voxel grid acquisition module 11, and a target image acquisition module 12.
[0104] The acquisition module 10 is used to acquire event stream data within a preset time period through the dynamic differential sensor, and to acquire image data to be imaged corresponding to the event stream data through the visual sensor;
[0105] An event voxel grid acquisition module 11 is used to perform voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data;
[0106] The target image acquisition module 12 is used to input the event voxel grid and the image data to be imaged into the trained imaging fusion model to acquire the target image output by the imaging fusion model.
[0107] In the above embodiment, the acquisition module acquires event stream data within a preset time length through a dynamic differential sensor, and acquires image data to be imaged corresponding to the event stream data through a visual sensor. The event voxel grid acquisition module performs voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data. The target image acquisition module inputs the event voxel grid and the image data to be imaged into a trained imaging fusion model to obtain a target image output by the imaging fusion model. In this way, a single lens including a visual sensor and a dynamic differential sensor is used to acquire event stream data and image data to be imaged, thereby avoiding the difficulty in image alignment processing when using a binocular stereo vision device or a binocular system based on a spectroscope for image alignment in the prior art, and using an imaging fusion model to acquire a target image, thereby improving the quality of the target image.
[0108] For the specific definition of the visual bionic imaging device based on a single lens, please refer to the definition of the visual bionic imaging method based on a single lens above, which will not be repeated here. Each module in the above-mentioned server can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0109] An embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the visual simulation imaging method based on a single lens provided in the embodiment of the present invention can be implemented. For example, when the processor executes the computer program, Figure 1 to Figure 2The technical solution of any of the method embodiments shown has similar implementation principles and technical effects, which will not be repeated here.
[0110] The embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. The computer program is executed by a processor to implement the visual simulation imaging method based on a single lens provided by the embodiment of the present invention. For example, when the computer program is executed by the processor, the computer program can be implemented. Figure 1 to Figure 2 The technical solution of any of the method embodiments shown has similar implementation principles and technical effects, which will not be repeated here.
[0111] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be accomplished by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to a memory, a database or other medium used in the embodiments provided by the present invention can include at least one of a non-volatile and a volatile memory. The non-volatile memory can include a read-only memory (ROM), a magnetic tape, a floppy disk, a flash memory or an optical memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as a static random access memory (SRAM) and a dynamic random access memory (DRAM).
[0112] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0113] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A visual bionic imaging method based on a single lens, characterized in that: The single lens includes a visual sensor and a dynamic differential sensor, and the method includes: The dynamic differential sensor is used to collect event stream data within a preset time period, and the visual sensor is used to collect image data to be imaged corresponding to the event stream data; Performing voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data; The event voxel grid and the image data to be imaged are input into a trained imaging fusion model to obtain a target image output by the imaging fusion model.
2. The method according to claim 1, characterized in that The image data to be imaged includes first image data and second image data, and the image data to be imaged corresponding to the event stream data collected by the visual sensor includes: At the start time of the preset time period, collecting the first image data by the visual sensor; At the end of the preset time period, the second image data is collected by the visual sensor.
3. The method according to claim 1, characterized in that The performing voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data includes: Acquire multiple event data within the preset time length, wherein each event data is represented by e=(x, y, p, t), where x represents the abscissa of the event, y represents the ordinate of the event, p represents the polarity of the event, and t represents the timestamp of the event; According to the abscissa, the ordinate, the timestamp and the polarity of the plurality of event data, voxel encoding processing is performed on the event stream data to obtain the corresponding event voxel grid.
4. The method according to claim 3, characterized in that The step of performing voxel encoding processing on the event stream data according to the abscissa, the ordinate and the polarity of the plurality of event data to obtain the corresponding event voxel grid comprises: Determine polarity values of the polarities corresponding to the plurality of event data respectively; Determine, in a preset voxel grid, target filling positions corresponding to the plurality of event data respectively according to the abscissas, the ordinates, and the timestamps corresponding to the plurality of event data respectively; According to the target filling positions respectively corresponding to the plurality of event data, the polarity values respectively corresponding to the plurality of event data are filled into the preset voxel grid to obtain the event voxel grid.
5. The method according to claim 1, characterized in that: Before performing voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data, the method further includes: The event stream data and the image data to be imaged are corrected by a checkerboard calibration method.
6. The method according to claim 2, characterized in that The imaging fusion model includes: an input layer, a first fusion network, a second fusion network and an output layer. The event voxel grid and the image data to be imaged are input into the trained imaging fusion model to obtain a target image output by the imaging fusion model, including: Inputting the event voxel grid and the image data to be imaged into the first fusion network through the input layer to obtain a first fused image and a second fused image; The first fused image and the second fused image are input into the second fusion network for fusion, and a target image is obtained through the output layer.
7. The method according to claim 6, characterized in that The step of inputting the event voxel grid and the image data to be imaged into the first fusion network to obtain a first fused image and a second fused image includes: fusing the event voxel grid and the first image data through the first fusion network to obtain the first fused image; The event voxel grid and the second image data are fused through the first fusion network to obtain the second fused image.
8. A visual bionic imaging device based on a single lens, characterized in that: The single lens includes a visual sensor and a dynamic differential sensor, and the device includes: A collection module, used to collect event stream data within a preset time period through the dynamic differential sensor, and to collect image data to be imaged corresponding to the event stream data through the visual sensor; An event voxel grid acquisition module, used for performing voxel encoding processing on the event stream data to obtain an event voxel grid corresponding to the event stream data; The target image acquisition module is used to input the event voxel grid and the image data to be imaged into the trained imaging fusion model to acquire the target image output by the imaging fusion model.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the single-lens based visual simulation imaging method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the single-lens based visual simulation imaging method described in any one of claims 1 to 7 are implemented.