Methods, apparatuses, and computer readable media for imaging
By acquiring images with different resolutions and bit widths, and using image fusion algorithms and convolutional neural networks to perform image fusion, the low efficiency and real-time performance issues of ultra-large-scale pixel imaging in existing technologies have been solved, achieving fast and high-definition image acquisition and transmission.
Patent Information
- Application Number
- CN202110522201.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-13
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2041-05-13
AI Technical Summary
Existing ultra-large-scale pixel imaging technology suffers from problems such as overly smooth or false textures in images, large system size and weight, complex circuit design, slow reading speed, large data volume, and inability to acquire and transmit video in real time.
By acquiring images with different resolutions and bit widths, image fusion algorithms and convolutional neural networks are used for image fusion. Combining the sparsity characteristics of differential images, compression encoding is performed, regions of interest are identified and targeted output is performed, and run-length encoding and Huffman encoding are used to optimize data transmission.
It enables fast and efficient acquisition of ultra-large-scale pixel images and real-time image transmission, improving image accuracy and recognizability, reducing transmission bandwidth, and meeting the requirements of real-time imaging.
Smart Images

Figure CN115345777B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a method, device and computer readable medium for imaging, and more particularly, to a method, device and computer readable medium for real-time imaging of ultra-large scale pixels. BACKGROUND
[0002] There are several existing imaging solutions that can achieve ultra-large scale pixels (e.g., typically referring to tens of millions of pixels): (1) through super-resolution technology, which enlarges a low-resolution image by a certain multiple through a traditional interpolation algorithm or a neural network-based image restoration algorithm to increase the resolution of the image, however, due to the loss of information in the acquisition process, the restored image is too smooth or has false textures (especially when the magnification is too large), resulting in poor final imaging results; (2) by shooting multiple low-resolution partial images and stitching them into a complete image of ultra-large scale pixels, such a system usually has the characteristics of large volume, heavy weight and complex circuit design; (3) single-chip ultra-large scale pixel imaging, which can shoot an ultra-large scale pixel image through a single sensor chip, but has the problems of slow reading speed and large data volume, and cannot realize real-time acquisition and transmission of video. The above methods directly acquire information in the image domain, and due to the large amount of related data, the acquisition and transmission time is long, and real-time shooting of video cannot be realized. SUMMARY
[0003] TECHNICAL PROBLEM
[0004] The purpose of the present application is to provide a method, device and computer readable medium for quickly or even real-time shooting of ultra-large scale pixel images.
[0005] TECHNICAL SCHEME
[0006] In order to overcome the technical problems in the above technical background and develop a shooting and processing technology for ultra-large scale pixel images, according to one aspect of the present application, a new image acquisition and fusion technology is proposed. This technology acquires two images of different resolutions and different bit widths (e.g., a high-resolution low-bit-width difference image and a low-resolution high-bit-width image) for a target field of view, and uses an image fusion algorithm to fuse the regions of interest in the two images to update the complete image of the target field of view, thereby quickly and efficiently obtaining a clear complete image of the target field of view with ultra-large scale pixels.
[0007] According to one embodiment, the sparse characteristics of the difference image (relative to the intra-frame) can be used to compress and encode the acquired high-resolution low-bit-width difference image, thereby greatly reducing the data volume of the image, improving the reading speed, reducing the transmission bandwidth, and helping to realize real-time acquisition of ultra-large scale pixel images.
[0008] According to an embodiment, a high-resolution low-bit-width difference image and a low-resolution high-bit-width image can also be fused and reconstructed using a convolutional neural network-based image fusion algorithm.
[0009] According to an embodiment, an image recognition method can also be used for the collected image, for example, to determine and identify a region of interest included in the collected image based on morphological and / or dynamic characteristics, and to selectively output the morphological and dynamic characteristics of the region of interest.
[0010] According to an aspect of the present application, the present application proposes a method for imaging, further comprising: photographing and quantizing a target field of view at a first resolution to obtain a first image having a first bit width; photographing and differentiating the target field of view at a second resolution to obtain a second image having a second bit width, wherein the differentiation includes: for a pixel point photographed at the second resolution, quantizing a difference between the pixel point and a neighboring or nearby pixel point of the pixel point to obtain a quantized difference as a value of a corresponding pixel point in the second image; and fusing the first image and the second image to obtain a third image, wherein the first resolution is lower than the second resolution, and the first bit width is higher than the second bit width.
[0011] According to an embodiment, fusing the first image and the second image to obtain the third image further includes: using an image recognition method to determine a region of interest of the target field of view in the second image; obtaining a corresponding region in the first image corresponding to the region of interest; and fusing the corresponding region of the first image and the region of interest of the second image to obtain the third image.
[0012] According to an embodiment, quantizing the difference between the pixel point and the neighboring or nearby pixel point of the pixel point for the pixel point photographed at the second resolution includes: quantizing the difference to a selected array.
[0013] According to an embodiment, the array is {-1, 0, +1}. In this way, the image data can be stored, processed and transmitted with the smallest number of bits, i.e. the most economical data resources.
[0014] According to an embodiment, the value of the pixel point can also be divided into five levels, for example, using two thresholds to compare the pixel point value in two positive and negative directions. For example, the array is {-2, -1, 0, +1, +2}. This five-value array scheme is more conducive to identification and fusion than the three-value array scheme, especially more conducive to real-time and accurate identification and positioning of the region of interest.
[0015] According to an embodiment, the values of the pixels can also be divided into four levels, for example using two thresholds, and the values of the pixels are compared in both positive and negative directions. For example, the array {-2, -1, 0, +1} or {-1, 0, +1, +2}. Using this four-value array scheme, only 2 bits are needed to represent the value of a pixel.
[0016] According to an embodiment of the present application, the target field of view corresponding to the first image is the same as the target field of view corresponding to the second image. In practice, there is usually a time and / or spatial difference between the two. This is because the locations and times at which the two images are taken can be slightly different, so that the two target fields of view do not necessarily completely coincide. However, as long as the area in which the two coincide is the main part of the image in practice, or both cover the area of interest, the two can be considered to be the same or equivalent fields of view. That is, fields of view that differ in time and / or space by an amount that is within a tolerable range can be considered to be the same field of view.
[0017] According to an embodiment, the method further comprises, before the fusing, encoding the second image for transmission, and decoding the second image after the transmission for the fusing.
[0018] According to an embodiment, the second image is encoded using run-length encoding, in which a bit sequence of the second image is encoded into a count sequence L recording the number of times of repeating the data and a data sequence D recording the data itself that is repeated.
[0019] According to an embodiment, the count sequence L is recorded using Huffman encoding, and the data sequence D is recorded using fixed-length encoding.
[0020] According to an embodiment, the fusing the first image with the second image comprises fusing the first image with the second image using a convolutional neural network.
[0021] According to an embodiment, the determining the area of interest in the target field of view in the second image using an image recognition method comprises determining the area of interest in combination with images taken for the target field of view within a certain time range or previously stored images taken for the target field of view.
[0022] According to an embodiment, the method further comprises training the image recognition method using deep learning, wherein the image recognition method is trained using deep learning based on a selected target, a spatial condition at the time of taking, and manual annotation.
[0023] According to an embodiment, the fusing the area of interest with the corresponding area to obtain the third image further comprises fusing the area of interest with the corresponding area to obtain the third image only when the area of interest includes a certain object.
[0024] According to an embodiment, the method further comprises outputting the third image after the fusing, or outputting the updated full image after updating the full image with the third image.
[0025] According to an embodiment, a camera or camcorder implementing the method of the present application can be manufactured.
[0026] For the discovery and reporting of regions of interest, real-time image reporting can be difficult if limited by the capability of the system, especially the hardware. As a beneficial improvement, the camera or camcorder of the present application can extract and report a region of interest when it discovers a region of interest including a specific object, to better meet the requirement of real-time.
[0027] According to another aspect of the present application, the present application further provides an apparatus for imaging, comprising: an image capturing component configured to capture and quantize a target field of view at a first resolution to obtain a first image having a first bit width, and capture and differentially process the target field of view at a second resolution to obtain a second image having a second bit width, wherein the differentially processing comprises: for a pixel point captured at the second resolution, quantizing a difference between the pixel point and a neighboring or nearby pixel point of the pixel point to obtain a quantized difference as a value of a corresponding pixel point in the second image; and a data processing component coupled to the image capturing component and configured to fuse the first image and the second image to obtain a third image, wherein the first resolution is lower than the second resolution, and the first bit width is higher than the second bit width.
[0028] According to an embodiment, fusing the first image and the second image to obtain the third image further comprises: using an image recognition method to determine a region of interest of the target field of view in the second image; obtaining a corresponding region in the first image corresponding to the region of interest; and fusing the corresponding region of the first image and the region of interest of the second image to obtain the third image.
[0029] According to an embodiment, quantizing, for a pixel point captured at the second resolution, a difference between the pixel point and a neighboring or nearby pixel point of the pixel point comprises: quantizing the difference to a selected array.
[0030] According to an embodiment, the array is {-1, 0, +1} or {-2, -1, 0, +1, +2} or {-2, -1, 0, +1} or {-1, 0, +1, +2}.
[0031] According to one embodiment, the apparatus further comprises an encoding component coupled to the image capturing component and configured to encode the second image before the fusing; and a transmitting component coupled to the encoding component and the data processing component and configured to transmit the encoded second image to the data processing component, and the data processing component is further configured to decode the encoded second image after receiving the encoded second image for the fusing.
[0032] According to one embodiment, the encoding component is further configured to encode the second image using run-length encoding, wherein a bit sequence of the second image is encoded into a count sequence L recording a number of repeated data and a data sequence D recording the repeated data itself.
[0033] According to one embodiment, the encoding component is further configured to record the count sequence L using Huffman encoding and record the data sequence D using fixed-length encoding.
[0034] According to one embodiment, the fusing the first image with the second image comprises fusing the first image with the second image using a convolutional neural network.
[0035] According to one embodiment, the determining the region of interest of the target field of view in the second image using an image recognition method comprises determining the region of interest in combination with images taken for the target field of view within a certain time range or previously stored images taken for the target field of view.
[0036] According to one embodiment, the data processing component is further configured to train the image recognition method using deep learning, and wherein the image recognition method is trained using deep learning based on a selected target, a spatial condition at the time of taking, and manual annotation.
[0037] According to yet another aspect of the present application, a non-transitory computer readable medium having program code recorded thereon is also provided, the program code, when executed by a computer, performs the method as described above.
[0038] Using the technical solution of the present application, by reducing the bit number of image acquisition, the reading time is shortened, which helps to realize real-time readout of super-large-scale pixel difference images; in the image fusion process, by performing convolutional neural network operation, the accuracy and recognizability of the image can be effectively improved, and a clear super-large-scale pixel image is obtained. In addition, using the on-chip compression coding scheme optimized by the present application, the image data can be compressed specifically, which greatly reduces the transmission bandwidth, and real-time or even high-speed image transmission can be realized. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figures 1A-1B is a schematic diagram of a flow of a method for imaging according to an embodiment of the application.
[0040] Figures 2A-2B is a schematic diagram of another flow of a method for imaging according to an embodiment of the application.
[0041] Figure 3 is a block diagram of a convolutional neural network based image fusion algorithm according to an embodiment of the application.
[0042] Figure 4A is a structural block diagram of an apparatus for imaging according to an embodiment of the application, Figure 4B is an example implementation of an apparatus for imaging according to an embodiment of the application.
[0043] Figure 5 is a structural block diagram of a real-time imaging apparatus according to an embodiment of the application.
[0044] Figure 6 is a schematic diagram of an optional pixel cell of a pixel array module of a real-time imaging apparatus according to an embodiment of the application.
[0045] Figure 7 is a schematic diagram of a pixel cell used by a pixel array module of a real-time imaging apparatus according to an embodiment of the application.
[0046] Figure 8 is a schematic diagram of an optional architecture of a pixel array module of a real-time imaging apparatus according to an embodiment of the application.
[0047] Figure 9 is a multi-level shift circuit used by a row / column drive module of a real-time imaging apparatus according to an embodiment of the application.
[0048] Figure 10 is a schematic diagram of an optional timing logic scheme of a row / column decode module of a real-time imaging apparatus according to an embodiment of the application.
[0049] Figure 11 is a schematic diagram of an optional combinational logic scheme of a row / column decode module of a real-time imaging apparatus according to an embodiment of the application.
[0050] Figure 12 is a schematic diagram of a readout module for a low resolution high bit width raw image of a real-time imaging apparatus according to an embodiment of the application.
[0051] Figure 13 is a schematic diagram of a readout module for a high resolution low bit width differential image of a real-time imaging apparatus according to an embodiment of the application.
[0052] Figure 14This is a schematic diagram of a readout module for a real-time imaging apparatus according to an embodiment of the present invention, for low-resolution high-bit-width raw images and high-resolution low-bit-width differential images. Detailed Implementation
[0053] The methods, apparatus, and computer-readable media of the present invention will now be described by way of example with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure typically described and shown in the drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the drawings is not intended to limit the scope of the claimed disclosure, but merely to illustrate selected embodiments of the disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0054] It should be noted that similar numbers and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0055] Figures 1A-1B This is a schematic diagram of the flow of an imaging method 100 according to an embodiment of the present invention. This method 100 can be used for image acquisition, processing, and output.
[0056] like Figure 1A As shown, in method 100, two images (e.g., two images differing in pixels and bit width) are acquired (or captured) separately for the target field of view in a parallel manner. Specifically, the first image (e.g., a low-resolution, high-bit-width original image) is acquired using the process or channel shown in the right half of the figure (step S10), and the second image (e.g., a high-resolution, low-bit-width differential image (obtained by differential processing of the corresponding original image)) is acquired using the process or channel shown in the left half of the figure (step S20). If one set of image acquisition components is used, the acquisition of the two types of images described above is performed sequentially. If two sets of image acquisition components are used, the acquisition of the two types of images can be performed simultaneously, or concurrently. This invention does not impose any limitation on the number of image acquisition components used.
[0057] Then, the acquired first image and second image can be fused (step S30). According to an embodiment, the first image can be acquired at a first resolution and a first bit width, and the second image can be acquired at a second resolution and a second bit width, wherein the second image can be a differential image processed by difference. According to an embodiment, wherein the first resolution can be lower than the second resolution, and the first bit width can be higher than the second bit width, for example, the first resolution can be much lower than the second resolution (e.g., 10 times or more lower), the first bit width can be 8 bits, and the second bit width can be 2 bits. In this case, relatively, the first image of low resolution and high bit width can be used to obtain rough information (e.g., background) of the target field of view, and the second image of high resolution and low bit width can be used to obtain fine information (e.g., vehicles appearing in the target field of view) of the target field of view, and the first image contains a smaller amount of data, and the second image contains a larger amount of data.
[0058] According to an embodiment of the present application, the second image can be processed by difference, including: for a pixel point (e.g., each pixel point in the second image taken at the second resolution), quantizing the difference between the pixel point and its adjacent or nearby pixel points to obtain a quantized difference as the value of the corresponding pixel point in the second image. Thus, the second image with the second bit width (such as a low bit width of 2 bits) is obtained.
[0059] In the foregoing and hereinafter descriptions, the present application does not make any limitation on the so-called first or second resolution (or high / low resolution) and the first or second bit width (or high / low bit width), and the meanings thereof can be relative, and a person skilled in the art can set the specific values thereof according to the needs to obtain advantageous technical effects.
[0060] Optionally, as Figure 1BAs shown, after the acquisition of the first and second images (step S10 and step S20) is completed, to improve the transmission efficiency, the acquired images can be encoded (step S101 and step S201) so as to improve the transmission efficiency when subsequently transmitted (step S102 and step S202). After the transmitted encoded images are received by the receiving end, they are decoded (step S103 and step S203) and image fused (step S30), and finally the third image (e.g., a high-resolution high-bit-width image of the field of view of interest) is obtained. Here, the third image is a complete image for the target field of view, and can contain the coarse information of the first image and the fine information of the second image. Here, the complete image can refer to an image (e.g., a high-resolution high-bit-width image) for the complete region of the field of view. At the receiving end, one receiver can be used to alternately receive the encoded data of the above two types of images, or two receivers can be used to separately and simultaneously receive the encoded data of the above two types of images. The present application does not impose any limitation on the receiving mode here.
[0061] Figures 2A-2B is a schematic diagram of another flow of the method 200 for imaging according to an embodiment of the present application. The method 200 can be regarded as an improvement of the method 100. As shown, Figure 2A and 2B similarly to Figures 1A-1B , using the method 200, two images can also be acquired in parallel, i.e., the first image (e.g., a low-resolution high-bit-width original image) is acquired by the process or channel shown in the right half of the figure (step S10), and the second image (e.g., a high-resolution low-bit-width difference image) is acquired by the process or channel shown in the left half of the figure (step S20). The description of the same or similar steps is omitted here.
[0062] Unlike Figures 1A-1B , in Figure 2A and 2BIn some embodiments, prior to image fusion (step S30), an image recognition method can be used to identify a region of interest in the second image (and optionally, the first image, as shown by the dashed line in the figure) that can contain fine information (step S40). For example, image recognition (e.g., object-specific) and motion feature detection can be performed on a target of interest in the field of view based on data from one image (e.g., previously taken) or a few consecutive images or a number of images taken within a predetermined time range, according to empirical data (e.g., morphological characteristics of different targets). The image recognition method can include only identifying an image region where a particular object can exist, or can or can not include identifying a particular object (e.g., a vehicle) included in the region of interest, but the present application does not limit the specific image recognition method, which can be any method or algorithm capable of extracting a region of interest (or a motion region involving motion) from an image.
[0063] In addition, the object in the field of view can also be identified and dynamically determined based on empirical data or information about the morphology of the target of interest, by manual visual inspection or computer scanning, before or after image fusion.
[0064] According to embodiments of the present application, the first image and the second image can be directly fused (to obtain a third image as a complete image of the field of view), or a portion of the first image and a corresponding portion of the second image can be fused (to obtain a third image as a partial image of the field of view). For example, fusing the first image and the second image can include using an image recognition method to determine a region of interest of the target field of view in the second image, obtaining a corresponding region in the first image corresponding to the region of interest, and fusing the corresponding region of the first image and the region of interest of the second image to obtain the third image. For example, the region of interest and the corresponding region can also be fused to obtain the third image only when the region of interest includes a particular object (e.g., when the region of interest is determined to include a particular object such as a car or a person by computer recognition or manual recognition).
[0065] In addition, the third image can be output after fusion, or the updated complete image can be output after the complete image is updated with the third image.
[0066] According to embodiments of the present application, a method (e.g., method 100) for ultra-large-scale pixel real-time imaging can include:
[0067] Step 1 (or Step S10 and Step S20): Taking pictures and quantifying for the target field of view to get, for example, high-resolution low-bit-width (e.g., 2-bit) difference images and low-resolution high-bit-width (e.g., 8-bit) raw images. The acquisition speed can be improved by reducing the bit number of the acquired (i.e., difference-processed raw images at the time of taking pictures) difference images to facilitate real-time imaging. According to embodiments of the present application, for example, 2-bit ternary bit-width (e.g., {-1, 0, 1}) difference images can be used for image fusion here to get a clear fused image. The acquisition of low-resolution high-bit-width raw images can be realized using ordinary commercial sensor chips.
[0068] Step 2 (or Step S101 and / or Step S201): Compressively encoding the high-resolution low-bit-width difference images using the sparse characteristics of the data of the difference images to reduce the data transmission bandwidth to facilitate real-time transmission. Optionally, the low-resolution high-bit-width raw images can also be compressively encoded as needed, and generally the data size of the low-resolution high-bit-width raw images is much smaller than that of the high-resolution low-bit-width difference images.
[0069] For example, the compressive encoding method can consist of two parts of optimized run-length encoding and Huffman encoding. The run-length encoding is to encode the original bit sequence into a count sequence L recording the number of repetitions of data and a data sequence D recording the repeated data itself. On the one hand, for the count sequence L recording the number of repetitions, the numerical distribution is very uneven, so Huffman variable-length encoding can be used instead of fixed-length code word encoding scheme to further reduce the data amount. On the other hand, for the data sequence D recording the repeated data itself, theoretically, the three different values of the difference image need to be represented by 2 bits, however, since the two adjacent data in the data sequence D of run-length encoding are necessarily unequal, then the next data of the current data only has two possibilities. Therefore, the larger one of the two possible values is recorded as 1, and the smaller one is recorded as 0. Based on the above principle, the data amount can be further reduced to facilitate the condition of real-time transmission of the transmission bandwidth (Step S102 or Step S202).
[0070] Step 3 (or Step S103 or Step S203): Decoding the transmitted compressively encoded high-resolution low-bit-width difference images (and optionally, the transmitted compressively encoded low-resolution high-bit-width raw images) in Step 2.
[0071] Step 4 (or Step S30): Fusing the (decoded) high-resolution low-bit-width difference images with the (decoded) low-resolution high-bit-width raw images. In the present application, a convolutional neural network-based image fusion algorithm can be used to fuse and reconstruct the two.
[0072] According to the embodiments of the present application, when acquiring the low-resolution high-bit-width original image, the field of view of the captured image is the target field of view. When acquiring the high-resolution low-bit-width difference image, ideally the same field of view as the original image is used. However, in practice, the two operations are not necessarily performed simultaneously, or even by the same device, so the fields of view of the two images are not necessarily exactly the same. According to the spirit of the present application, as long as the field of view of the first image (e.g., the low-resolution high-bit-width original image) is substantially the same as the field of view of the second image (e.g., the high-resolution low-bit-width difference image), the third image generated according to the method and device of the present application can be obtained through subsequent image processing (e.g., image fusion of the same field of view of the two images), to update the complete image or the region of interest of the field of view. In other words, according to the embodiments of the present application, the two fields of view corresponding to the first image and the second image can be referred to as the target field of view, which can be defined as a field of view at a specific time and / or a specific space, and a certain error in time and space is allowed.
[0073] According to the embodiments of the present application, the method of acquiring the high-resolution low-bit-width difference image can be, for example, capturing a high-resolution original image in the same field of view as the original image, and performing a difference comparison between a pixel point and its adjacent or nearby pixel points in the high-resolution original image. For example, the comparison result (the difference between the two) can be quantized to one of {-1, 0, +1} according to the judgment of {less than, equal to, greater than} (which can be referred to as "three-interval method"). The above operation (difference comparison and quantization) can be repeated for multiple pixel points of the high-resolution original image, and finally a high-resolution low-bit-width difference image is generated. For example, the above operation can be repeated for each pixel point, or for pixel points or a plurality of representative pixel points of the field of view selected in other ways at a fixed interval (e.g., every two adjacent rows / columns) or a variable interval, and the present application does not make any limitation in this regard. Herein, the process of difference comparison and quantization can be referred to as "difference processing".
[0074] For example, the difference processing can be a comparison of the pixel point values of two adjacent pixel points, i.e., a comparison of the pixel point values of the (n+1)th pixel and the nth pixel in a certain direction. In a similar manner, a comparison of the pixel point values of the (n+2)th or (n+3)th pixel and the nth pixel can also be performed, i.e., a comparison of the pixel point values of the (n+i)th pixel (i=1, 2,...) and the nth pixel. This is particularly suitable for general occasions where spatial high resolution is not required, such as routine screening (e.g., whether there is a new object in the field of view), which can be used to detect whether an abnormal situation occurs. For example, the difference comparison can be performed every two rows or every two columns.
[0075] As mentioned above, when quantizing the pixel point value to one of {-1, 0, +1}, the pixel point value can be, for example, a luminance value or other value representing color, which can be, for example, one of 256 values and can be represented by 8 binary bits. When performing the differential comparison of the pixel point values, the decision made can be based only on the direct difference between the two, i.e. a decision of -1 or 1 is made if the two are different, even if the difference corresponds to only one of the 256 values (e.g. a difference of 1). Correspondingly, if the difference is 0, a decision of 0 is made. Alternatively, the decision criterion can be set to a difference between the two greater than or equal to a threshold value, such as 4 of the 256 values, and a decision of different is made and correspondingly -1 or 1 is outputted when the decision criterion is met. This processing method can be used, for example, to qualitatively reveal details at the edges of a high-contrast image.
[0076] As another embodiment of the present application, a more complex quantization scheme as described below can also be used.
[0077] For example, with a five-interval method, two threshold values are used, i.e. threshold values C1 and C2, where C1 and C2 are both greater than 0 and C2 > C1,
[0078] It can be determined whether the comparison result (difference) of the pixel point values satisfies the following conditions: <-C2, <-C1, in the interval close to a reference value (e.g. 0), >C1, >C2. In other words, it is determined which interval the comparison result is in, and the comparison result of the pixel point values is assigned values of {-2, -1, 0, 1, 2} respectively, i.e.
[0079]
[0080] where I represents the result before quantization (e.g. the result of the differential comparison, i.e. the difference as described above). In this way, the pixel point values can be measured in a wide range. However, this five-interval method requires 3 bits to represent the comparison result of the pixel point values, which occupies more resources.
[0081] Slightly more resource-saving than the above-mentioned five-interval method is a four-interval method, which also uses the same two threshold values, i.e. threshold values C1 and C2, where C1 and C2 are both greater than 0 and C2 > C1. It can be determined whether the comparison result of the pixel point values satisfies the following conditions: <-C2, <-C1, in the interval close to a reference value (e.g. 0), >C1. In other words, it is determined which interval the comparison result is in, and the comparison result of the pixel point values is assigned values of {-2, -1, 0, 1} respectively, i.e.
[0082]
[0083] Alternatively, it can be determined whether the pixel point value satisfies the following condition: <-C1, in the interval close to a reference value (e.g., 0), >C1, >C2. In other words, it is determined in which interval the comparison result is, and the comparison result of the pixel point value is assigned as {-1, 0, 1, 2}, respectively, that is:
[0084]
[0085] where I represents the result before quantization (e.g., the result of the difference comparison, i.e., the difference value as described above). In this way, only 2 bits are needed to represent the comparison result of the pixel point value.
[0086] Different bit widths of the difference image can be obtained based on any one of the three-interval (three-value) method, the four-interval (four-value) method, and the five-interval (five-value) method as described above. The present application will be described in detail below in combination with the drawings of the present application and the embodiments of the three-value case. The following embodiments are merely exemplary and not limiting. It should be noted that the principles of the four-value or five-value embodiments are the same as the following three-value embodiments.
[0087] Example 1:
[0088] In this embodiment, a method (e.g., method 100) for imaging is designed based on a high-resolution low-bit-width (three-value, {-1, 0, 1}) difference image and a low-resolution high-bit-width (e.g., 8 bits) grayscale image. The specific steps include:
[0089] Step 1-1 (or steps S10 and S20): photographing and quantization are performed for a target field of view to obtain a low-resolution high-bit-width grayscale image and a high-resolution low-bit-width difference image. For example, the resolution of the low-resolution high-bit-width grayscale image can be The high-resolution low-bit-width difference image can be based on the difference value of every two adjacent columns in its corresponding original image, and its default quantization interval can be, for example, [-255, -4), [-4, 4], and (4, 255]. In this way, the three-interval method can be used to assign the difference value as {-1, 0, 1} according to the quantization interval (i.e., the three intervals correspond to -1, 0, and 1, respectively). By reducing the number of bits for acquiring the difference image, the acquisition speed can be improved, which helps to achieve real-time imaging.
[0090] Step 1-2 (or steps S101, S102, and / or steps S201, S202): taking advantage of the sparsity of the difference image, the high-resolution low-bit-width difference image is compressed and encoded to reduce the data transmission bandwidth, thereby helping to achieve real-time transmission.
[0091] For example, the compression encoding method can be composed of two parts of optimized run-length encoding and Huffman encoding. Specifically, the run-length encoding is to encode the original bit sequence into a count sequence L recording the repetition times of data and a data sequence D recording the repeated data itself. On one hand, for the count sequence L recording the repetition times, the value distribution is very uneven, so Huffman encoding can be used instead of fixed-length code word encoding scheme to further reduce the data amount. On the other hand, for the data sequence D recording the repeated data itself, theoretically, the two adjacent data must be unequal. In this case, the value of the difference image has 3 possibilities, while the next data of the current data has only 2 possibilities, so in the data sequence D, only the first data has 3 possibilities, and each of the following data has only 2 possibilities. Therefore, the first data can be encoded using 2 bits, and each of the remaining data can be encoded using 1 bit.
[0092] Step 1-3 (or step S103 and / or S203): Decoding the encoded high-resolution low-bit-width difference image at the receiving end.
[0093] Step 1-4 (or step S30): Fusing the high-resolution low-bit-width difference image and the low-resolution high-bit-width gray image.
[0094] In the embodiment, the two images can be optimally fused and reconstructed using the convolutional neural network-based image fusion algorithm. In the embodiment, the input high-resolution low-bit-width difference image and the low-resolution high-bit-width image both have a channel of 1, and the output high-resolution high-bit-width image has a resolution same as that of the input high-resolution low-bit-width difference image and a channel number of 1.
[0095] Figure 3 is a block diagram of the convolutional neural network-based image fusion algorithm according to the embodiment of the application. The following will be described in combination with Figure 3 The network structure of the convolutional neural network will be introduced.
[0096] According to the embodiment of the application, the two input images (for example, the high-resolution low-bit-width image and the low-resolution high-bit-width image) can have a difference in resolution, for example, the resolution of one image is 1 / 64 of that of the other image. To deal with the problem of resolution mismatch, a multi-scale feature fusion network can be used to fuse the high-frequency information of the high-resolution low-bit-width difference image and the low-frequency information of the low-resolution high-bit-width image at different scales.
[0097] As Figure 3As shown, this multi-scale feature fusion network can be divided into three different branches, namely one super-resolution branch and two difference branches. The input of the super-resolution branch is a low-resolution high-bit-width image, and the output is a corresponding (for example, 8*8 times) high-resolution high-bit-width image, and in the latter half of the super-resolution branch, the high-frequency components from the two difference branches are fused using feature fusion connections for the synthesis of a high-resolution high-bit-width image with clear details. The input of the two difference branches is a high-resolution low-bit-width difference image, since in this embodiment the difference image only has 1 x direction (i.e., difference comparison between columns and columns), the input of the two branches is the same, which is the difference image in the x direction. In order to better fuse, the two difference branches respectively output high-resolution high-bit-width difference images in the x direction (i.e., difference comparison between columns and columns) and the y direction (i.e., difference comparison between rows and rows), in order to achieve this purpose, the supervision maps used by the two difference branches during training are high-resolution high-bit-width difference images in the x direction and the y direction, respectively. Both difference branches fuse low-frequency components from the super-resolution branch at different scales to guide the fusion of high-resolution high-bit-width difference images.
[0098] The details of each branch structure are introduced below, including feature fusion connections, loss functions, and training strategies.
[0099] Reference Figure 3 The super-resolution branch can be divided into two parts, the first part completes the (for example, 8*8 times) super-resolution processing of the input low-resolution high-bit-width image, and at the same time obtains feature maps at different scales, which can be used to guide the fusion reconstruction of the difference branch; the second part fuses the feature maps from the difference branch to complete the final fusion reconstruction of the high-resolution high-bit-width image. The super-resolution branch can use a progressive super-resolution algorithm, for example, 3 times 2*2 upsampling layers can be used to complete 8*8 times super-resolution, and the upsampling layer uses deconvolution to achieve. In the latter half of the network, the super-resolution branch fuses the features of the two difference branches to complete the final fusion reconstruction.
[0100] Reference Figure 3The structures of the two difference branches are the same, but their supervision graphs can be different and can not share parameters. The first half of the difference branch can use a U-shaped structure similar to the U-net, but unlike the U-net, the RRDB (Residual-in-Residual Dense Block) can also be used instead of the basic convolution layer. The maximum pooling layer can be used as the down-sampling layer and the de-convolution layer can be used as the up-sampling layer. The RRDB and the maximum pooling layer can be combined as a basic down-sampling module, and in the down-sampling process, 3 times of 2*2 down-sampling modules can be used to obtain feature maps at 4 different scales (for example, the original scale, 1 / 2*1 / 2 times resolution, 1 / 4*1 / 4 times resolution, and 1 / 8*1 / 8 times resolution), which will be fused with the feature maps in the up-sampling process in the manner as shown in FIG. 1. Figure 3 The RRDB and the de-convolution layer can be combined as a basic 2*2 up-sampling module, and 3 times of 2*2 up-sampling modules can be used to restore the feature maps to the original size, and in the up-sampling process, the feature maps obtained at different scales in the down-sampling of the same difference branch and the feature maps in the up-sampling process of the super-resolution branch are fused.
[0101] According to the embodiment of the present application, the algorithm of the feature fusion connection can include three kinds of feature connections: high-resolution difference image-to-high-resolution difference image (HRD-to-HRD) feature fusion connection, low-resolution original image-to-high-resolution difference image (LRI-to-HRD) feature fusion connection, and high-resolution difference image-to-low-resolution original image (HRD-to-LRI) feature fusion connection, which are located at different positions of the network and play different roles, and each feature connection uses the operation (Concat) of feature map splicing. In the first half of the entire network, the three branches are a multi-scale structure, in this part, the feature maps of the super-resolution branch can be fused into the two difference branches in the up-sampling part (LRI-to-HRD feature fusion connection), and at the same time, the feature maps generated in the down-sampling process of the same difference network can be fused in the up-sampling part (HRD-to-HRD feature fusion connection); in the second half of the entire network, the feature maps generated by the two difference branches can be fused into the super-resolution branch (HRD-to-LRI feature fusion connection), to complete the final fusion. In this network, except for the layers containing feature fusion connections, the number of feature maps of each layer is set to 16, and in the layers containing feature fusion connections, the number of feature maps is an integer multiple of 16, and the multiple is the number of different branches fused, such as the number of feature maps being 32 when 2 branches are fused.
[0102] According to the embodiment of the present application, the minimum mean square error between the clear high-resolution high-bit-width image and the fusion result can be used as the loss function (MSELoss) of the super-resolution branch. Meanwhile, the minimum mean square error between the high-resolution high-bit-width difference image in the x direction and the output of the two difference branches can also be used as the loss function of the two difference branches. Therefore, the total loss function can be represented by the following equation:
[0103]
[0104] According to the embodiment of the present application, in terms of the training method, the DIV2K super-resolution data set can be used to make the training set. For example, during training, the convolution kernel size is set to 3, the Adam algorithm is used as the optimizer, the hyperparameters β and γ in the loss function are set to 0.1, the learning rate is set to 1x10 -4 , and after every 20K iterations, the learning rate is multiplied by the decay factor 0.5. A total of 100K iterations are trained, and the Batchsize is set to 16.
[0105] Example 2:
[0106] The embodiment proposes an imaging method (for example, method 200) which increases image recognition based on difference images. Unlike directly performing image recognition on the original image without difference processing, the embodiment directly performs image recognition on the difference image, and only fuses the region of interest after identifying the region of interest.
[0107] Step 2-1: same as step 1-1 of embodiment 1;
[0108] Step 2-2: same as step 1-2 of embodiment 1;
[0109] Step 2-3: same as step 1-3 of embodiment 1;
[0110] Step 2-4: construct a training data set, and use the generated training set to train the image recognition method (for example, YOLOv3) to obtain a trained image recognition method. The training data set of the embodiment can be generated from a public image recognition data set, or can be generated by manual annotation. The specific methods of the two ways are as follows:
[0111] 1) generated from a public data set: after downloading the public image recognition data set, the labels are not processed, but the original image (for example, a high-resolution unquantized original image) in the data set is subjected to the difference processing as described above according to the quantization interval set in step 2-1 to obtain the corresponding difference image (for example, a high-resolution low-bit-width difference image), and then the existing labels in the data set are combined to generate a pair of training sets.
[0112] 2) Artificial annotation generation: The regions of interest in the collected difference images are annotated by human, and the data pairs are formed to construct the training dataset.
[0113] Wherein, steps 2-4 can be preformed, i.e. the existing trained image recognition method is used in the implementation of the present method. Generally, after obtaining the high-resolution low-bit-width difference image in step 2-3, it is directly jumped to step 2-5.
[0114] Step 2-5: The trained image recognition method (YOLOv3 obtained by step 2-4) can be used to apply the image recognition method to the high-resolution low-bit-width difference image to identify the region of interest.
[0115] Step 2-6: Obtain the corresponding region in the low-resolution high-bit-width image corresponding to the identified region of interest (for example, determined by the positioning of the region of interest in the target field of view).
[0116] Step 2-7: The identified region of interest is fused with the identified corresponding region to obtain a fused image (third image) for the region of interest, wherein the fusion method can be the same as step 1-4 of embodiment 1. For example, the fused image can be used to update the complete image of the field of view to obtain an updated complete image. The complete image can be updated (for example, with the fused image) only when changes occur in the field of view (for example, the appearance of a specific object or the movement of an original object, etc.). Alternatively, different update rates can be set, for example, the first image and the second image (images containing information of the entire field of view) are used to update the complete image at a first update rate, and the fused image is used to update the complete image (of the corresponding part that needs to be updated or has changed) at a second update rate, wherein the first update rate can be less than the second update rate (for example, 1 fps and 30 fps, respectively), and the present application does not make any limitation on the specific values. In this way, the complete image can be updated with the least data.
[0117] Example 3:
[0118] The present embodiment proposes a design scheme of an imaging method based on high-resolution low-bit-width (for example, ternary) difference images and low-resolution high-bit-width (for example, 8-bit) grayscale images, which can adaptively adjust the quantization interval.
[0119] Step 3-1: Constructing a code rate-quantization interval-fusion quality database: using existing public datasets to construct a code rate-quantization interval-fusion quality database.
[0120] Step 3-2: Determine the quantization interval for collecting the low resolution gray scale image and the high resolution difference image. Set the initial code rate and fusion quality, for example, the resolution of the low resolution gray scale image is set to be 1 / 64 of the high resolution difference image, and the high resolution difference image is set to be the difference value of every two adjacent columns based on its corresponding original image. Based on the set initial system code rate and fusion quality, query the database constructed in step 3-1 to determine the corresponding quantization interval for completing the collection of the low resolution gray scale image and the high resolution difference image with different bit widths. The code rate and fusion quality can be manually adjusted in real time, and the corresponding quantization interval in the database will also change accordingly.
[0121] Then, steps 1-1 to 1-4 can be implemented, or steps 2-1 to 2-7 can be implemented.
[0122] Example 4:
[0123] The present embodiment proposes an imaging method (e.g., method 200) that increases motion detection based on difference images. Unlike the conventional motion detection on the fused image (e.g., the third image), the method of the present embodiment performs motion detection (e.g., image recognition as shown in step S40 of Figures 2A-2B ) on the image before fusion (e.g., the low resolution high bit width image or the first image, and / or the high resolution low bit width image or the second image as described above). After the motion region (or region of interest) is identified, only the motion region (or region of interest) is fused (e.g., image fusion as shown in step S30 of Figures 2A-2B ). The method can be implemented with reference to Figures 2A-2B . The specific steps of the method are described below with the high resolution low bit width difference image as an example.
[0124] Step 4-1: same as step 1-1 of embodiment 1;
[0125] Step 4-2: same as step 1-2 of embodiment 1;
[0126] Step 4-3: same as step 1-3 of embodiment 1;
[0127] Step 4-4: use the inter-frame method to calculate the difference between the adjacent two frames of high resolution low bit width difference images, and extract the region with difference change (e.g., the difference value meets a certain threshold condition) to obtain the motion region (e.g., in the form of coordinates), i.e., the region of interest.
[0128] Step 4-5: same as steps 2-6 to 2-7 of embodiment 2, to obtain the fused motion region;
[0129] Step 4-6: update the full image of the target field of view previously taken with the fused moving region, and the region other than the moving region in the full image is not updated.
[0130] Step 4-7: repeat steps 4-1 to 4-6.
[0131] Example 5:
[0132] The embodiment proposes an imaging method based on a low-resolution high-bit-width (for example, 8-bit) RGB image and a high-resolution low-bit-width (for example, three-value) difference image.
[0133] Step 5-1: acquire a low-resolution high-bit-width RGB image and a high-resolution low-bit-width difference image. In an example, the resolution of the low-resolution high-bit-width RGB image is 1 / 64 of the high-resolution low-bit-width difference image, and the high-resolution low-bit-width difference image can be based on the difference value of every two adjacent columns of its corresponding original gray image, and the default quantization interval is (-255, -4), [-4, 4] and (4, 255), and the corresponding quantized values are -1, 0 and 1 respectively. By reducing the number of bits of the acquired difference image, the acquisition speed is improved to help achieve real-time imaging effect.
[0134] Step 5-2: same as step 1-2 of embodiment 1;
[0135] Step 5-3: same as step 1-3 of embodiment 1;
[0136] Step 5-4: fuse with the acquired low-resolution high-bit-width RGB color image. In this embodiment, a convolutional neural network-based image fusion algorithm is used to optimize the fusion reconstruction of the two. In this embodiment, for the convolutional neural network, the channel numbers of the input high-resolution low-bit-width difference image and the low-resolution high-bit-width RGB image are 1 and 3 respectively, and the resolution of the output high-resolution high-bit-width image is the same as that of the input high-resolution low-bit-width difference image, and the channel number is 3, that is, the output high-resolution high-bit-width RGB image.
[0137] Embodiments of a device for real-time imaging according to the present application
[0138] Figure 4A is a structural block diagram of the device 400 for imaging according to an embodiment of the present application, Figure 4B is an example implementation of the device 400 for imaging according to an embodiment of the present application.
[0139] As Figure 4AAs shown, the device 400 may include an image capturing component 401, a data processing component 402, an encoding component 403, and a transmission component 404. According to an embodiment, the image capturing component 401 may be configured to perform the operations of acquiring (including capturing and quantizing) images as described above. The data processing component 402 may be configured to perform operations associated with data processing, including operations such as image recognition, image fusion, and decoding. The encoding component 403 may be configured to encode the data of the acquired image for transmission to, for example, the data processing component 402 or other data processing devices. The transmission component 404 may be configured to transmit various types of data, for example, transmitting encoded image data from the encoding component 403 to the data processing component 402 for further processing. The various components included in the device 400 may perform various operations, for example, under the control of a control component (not shown).
[0140] like Figure 4B As shown, the device according to the present invention (e.g., device 400) can be implemented as comprising a real-time imaging device development board 410 (which can be used as a combination of an image capturing component 401 and an encoding component 403) and a host computer 412 (which can be used as a data processing component 402), which can be electrically connected via an Ethernet interface 414 (which can be used as a transmission component 404). The real-time imaging device development board 410 may include a real-time imaging device 411, an FPGA control unit 415, and a system peripheral chip 413. The host computer 412 uses the Ethernet interface 414 to configure the real-time imaging device development board 410. The FPGA control unit 415 and the system peripheral chip 413 on the real-time imaging device development board 410 provide the timing signals and control voltages required for the operation of the real-time imaging device 411 according to the configuration information. Finally, the acquired image data is transmitted back to the host computer 412 via the Ethernet interface 414, completing one acquisition task.
[0141] Figure 5 This is a structural block diagram of a real-time imaging device (e.g., real-time imaging device 411) according to an embodiment of the present invention. Figure 5 As shown, the real-time imaging device may include a pixel array module, a row / column driving module, a row / column decoding module, a readout module, and an I / O module. The FPGA control unit (e.g., FPGA control unit 415) controls the row / column decoding module through the I / O interface module, and can provide operating voltage to the row / column driving module. At the same time, it controls the pixel array module to sense the target scene, and finally reads out the image data through the readout module.
[0142] Figure 6 This is a schematic diagram of optional pixel units of a pixel array module of a real-time imaging apparatus (e.g., real-time imaging apparatus 411) according to an embodiment of the present invention. Figure 6As shown, the pixel array module can include two levels of pixel cell and pixel array. At the pixel cell level, there are multiple options based on planar silicon process manufacturing technology, such as photodiode (PD), phototriode (PT), charge-coupled device (CCD), active pixel sensor (APS), etc.
[0143] Figure 7 is a schematic diagram of a pixel cell used by a pixel array module of a real-time imaging device (e.g., real-time imaging device 411) according to an embodiment of the present application. Figure 8 is a schematic diagram of an alternative architecture of a pixel array module of a real-time imaging device (e.g., real-time imaging device 411) according to an embodiment of the present application. In the present application, as shown in Figure 6 In one preferred embodiment of the pixel cell, as shown, a dual-transistor photosensitive detector in CN201210442007.X can be used. Figure 7 In one preferred embodiment of the pixel cell, as shown, a dual-transistor photosensitive detector in CN201210442007.X can be used. Figure 8 As shown, the pixel array level can choose a NAND architecture or a NOR architecture, and in one preferred embodiment, a NOR architecture can be used.
[0144] Figure 9 is a schematic diagram of a multi-level shift circuit used by a row / column drive module of a real-time imaging device (e.g., real-time imaging device 411) according to an embodiment of the present application. As an example, the row / column drive module can use a multi-level shift circuit in CN202010384765.5. The module needs to input a pre-shift positive voltage signal VVPP, a pre-shift negative voltage signal VVPN, a shift positive voltage signal VPHV, a shift negative voltage signal VNHV, and a pre-shift control signal VIN. The shift voltage output signal VO outputs the shift positive voltage signal VPHV or the shift negative voltage signal VNHV under the control of the pre-shift control signal VIN, providing driving voltage for the pixel array module. Figure 9 According to an embodiment of the present application, the row / column decoding module can use a timing logic scheme of a shift register, or a combinational logic scheme of a decoder.
[0145] is a schematic diagram of an alternative timing logic scheme of a row / column decoding module of a real-time imaging device (e.g., real-time imaging device 411) according to an embodiment of the present application. The embodiment uses a combinational logic scheme described in Verilog HDL. Figure 10 is a schematic diagram of an alternative combinational logic scheme of a row / column decoding module of a real-time imaging device (e.g., real-time imaging device 411) according to an embodiment of the present application. As an example, a 1024bit timing logic scheme ( Figure 11 ) and an 8bit combinational logic scheme ( Figure 10 ) can be used. Figure 11
[0146] Figures 12-14 are schematic diagrams of different embodiments of the readout module of the real-time imaging device (e.g., the real-time imaging device 411) according to the embodiments of the present application. The readout module can be implemented in two schemes, that is, the readout circuit can be designed separately for the high-resolution low-bit-width differential image and the low-resolution high-bit-width original image, or the readout circuits for the high-resolution low-bit-width differential image and the low-resolution high-bit-width original image can be combined. For the first scheme, the readout of the low-resolution high-bit-width original image can be performed using the readout circuit in the existing patent CN201911257219.9 (as shown in Figure 12 ), and the readout of the high-resolution low-bit-width differential image can be performed using the current subtraction circuit in the existing patent CN202010697791.3 (as shown in Figure 13 ). For the second scheme, the readout circuit in the patent CN201911257219.9 (as shown in Figure 12 ) can be continued to be used, and when the readout of the low-resolution high-bit-width original image is performed, the up-down counter is configured in the up-counting mode, and when the readout of the high-resolution low-bit-width differential image is performed, the up-down counter is first configured in the up-counting mode, and then configured in the down-counting mode, so as to realize the differential readout. The present application uses a new type of readout circuit (as shown in Figure 14 ), in the case of operating on the low-resolution high-bit-width original image, DIR is 0, the current mirror CM1 works, CM2 does not work, the current of BLN is used to discharge the capacitor C, the counter works all the time before the voltage on the capacitor is discharged to below the reference voltage VP of the comparator CMP1, and after the voltage on the capacitor is discharged below the reference voltage VP, the counter stops working, and the quantization result is sent out through the parallel-serial conversion module; in the case of operating on the high-resolution low-bit-width differential image, DIR is 1, the current mirrors CM1 and CM2 work at the same time, the current of BLN-BLN+1 is used to discharge the capacitor C, and after the comparators CMP1 and CMP2, a 2-bit quantization result is obtained, which is sent out through the parallel-serial conversion module. The I / O interface module can use a general input / output interface with ESD protection function from any manufacturer.
[0147] It should be noted that each of the embodiments in the present specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same and similar parts between the embodiments can be mutually referred to.
[0148] In several embodiments provided in the present application, it should be understood that each block in a flowchart or block diagram can represent a module, a segment or a portion of code which includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or combinations thereof, or can be implemented by a combination of dedicated hardware and computer instructions.
[0149] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present disclosure essentially or the parts of the prior art that make contributions or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0150] It should also be noted that the aforementioned description is merely specific implementation manners of the present application and not intended to limit the protection scope of the present application. Any variations or replacements of the present application, which are obvious to persons skilled in the art without creative efforts, within the idea and technical scope of the present application shall fall into the protection scope of the present application. Furthermore, it should be noted that in the present application, the terms "comprise", "contain" or any other variants are intended to cover non-exclusive inclusion, so that a process, a method, an article or an apparatus including a series of elements not only includes those elements, but also further includes other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus including the element.
[0151] Those skilled in the art can reasonably understand the details not described in the above disclosure of the present application, the method and device of the present application and its embodiments without difficulty. The above description is within the protection scope of the present application.
Claims
1. A method for imaging, comprising: The target field of view is captured and quantized at a first resolution to obtain a first image with a first width. The target field of view is captured at a second resolution and differentially processed to obtain a second image with a second bit width. The differential processing includes: quantizing the difference between a pixel captured at the second resolution and its adjacent or nearby pixels to obtain the quantized difference as the value of the corresponding pixel in the second image; and The first image and the second image are merged to obtain the third image. The first resolution is lower than the second resolution, and the first bit width is higher than the second bit width.
2. The method as described in claim 1, wherein, The process of fusing the first image with the second image to obtain the third image also includes: Image recognition methods are used to determine the region of interest of the target field of view in the second image; Obtain the corresponding region in the first image that corresponds to the region of interest; and The corresponding region of the first image is fused with the region of interest of the second image to obtain the third image.
3. The method of claim 1, further comprising encoding the second image for transmission prior to fusion, and decoding the second image for fusion after transmission.
4. The method of claim 3, wherein, The second image is encoded using run-length encoding, wherein the bit sequence of the second image is encoded into a count sequence L that records the number of repetitions of the repeating data and a data sequence D that records the repeating data itself.
5. The method of claim 4, wherein, The counting sequence L is recorded using Huffman coding, and the data sequence D is recorded using fixed-length coding.
6. The method of claim 1, wherein, Fusing the first image and the second image includes using a convolutional neural network to fuse the first image and the second image.
7. The method of claim 2, wherein, Determining the region of interest (ROI) of the target field of view in the second image using an image recognition method includes: combining images taken within a specific time range for the target field of view or previously stored images taken for the target field of view to determine the ROI.
8. The method of claim 2 further includes training the image recognition method using deep learning based on the selected target, the spatial conditions at the time of shooting, and manual annotation.
9. The method of claim 2, wherein, The process of fusing the region of interest with the corresponding region to obtain a third image further includes: fusing the region of interest with the corresponding region only when the region of interest includes a specific object.
10. The method of claim 1 or 9, further comprising outputting a third image after fusion, or outputting an updated complete image after updating the complete image with the third image.
11. An apparatus for imaging, comprising: The image capturing component is configured as follows: The target field of view is captured and quantized at a first resolution to obtain a first image with a first width. as well as The target field of view is captured at a second resolution and differentially processed to obtain a second image with a second bit width. The differential processing includes: quantizing the difference between a pixel captured at the second resolution and its adjacent or nearby pixels to obtain the quantized difference as the value of the corresponding pixel in the second image; and a data processing component coupled to the image capturing component and configured to fuse the first image and the second image to obtain a third image. The first resolution is lower than the second resolution, and the first bit width is higher than the second bit width.
12. The apparatus of claim 11, wherein, The process of fusing the first image with the second image to obtain the third image also includes: Image recognition methods are used to determine the region of interest of the target field of view in the second image; Obtain the corresponding region in the first image that corresponds to the region of interest; and The corresponding region of the first image is fused with the region of interest of the second image to obtain the third image.
13. The apparatus of claim 11, further comprising an encoding component and a transmission component. The encoding component is coupled to the image capturing component and is configured to encode the second image prior to fusion. The transmission component is coupled to the encoding component and the data processing component, and is configured to transmit the encoded second image to the data processing component. The data processing component is also configured to decode the encoded second image after receiving it for fusion.
14. The apparatus of claim 13, wherein, The encoding component is further configured to encode the second image using run-length encoding, wherein the bit sequence of the second image is encoded into a count sequence L that records the number of repetitions of the repeating data and a data sequence D that records the repeating data itself.
15. The apparatus of claim 14, wherein, The encoding component is also configured to use Huffman coding to record the counting sequence L and fixed-length coding to record the data sequence D.
16. The apparatus of claim 11, wherein, Fusing the first image and the second image includes using a convolutional neural network to fuse the first image and the second image.
17. The apparatus of claim 12, wherein, Determining the region of interest (ROI) of the target field of view in the second image using an image recognition method includes: combining images taken within a specific time range for the target field of view or previously stored images taken for the target field of view to determine the ROI.
18. The apparatus of claim 12, wherein, The data processing component is also configured to train the image recognition method using deep learning based on the selected target, the spatial conditions at the time of shooting, and manual annotation.
19. A non-transitory computer-readable medium having program code recorded thereon, the program code performing the method as described in any one of claims 1-10 when executed by a computer.
Citation Information
Patent Citations
Composite dielectric grating metal-oxide-semiconductor field effect transistor (MOSFET) based dual-transistor light-sensitive detector and signal reading method thereof
CN102938409A
Analog-to-digital converter circuit based on composite dielectric gate dual transistor photodetector
CN111147078B
Multilevel shift circuit based on composite dielectric gate dual transistor photodetector
CN111541444B
Current subtraction circuit
CN111969983A
Image processing method based on HDR
CN106851138A
Cited By
Digital hardware implementation architecture based on image fusion algorithm and method thereof
CN117522709A
A digital hardware implementation device based on an image fusion algorithm and a method thereof
CN117522709B