Multi-spectral TDI sensor panchromatic sharpening method based on on-chip coding integration

By designing the encoding mask plate on the multispectral TDI sensor and performing on-chip integration, combined with convolutional neural network for full-color image fusion, the problem of difficult to reduce the cell size is solved, and the full-color sharpening of high-resolution multispectral images is achieved, which improves the stability and accuracy of image reconstruction.

CN120451026APending Publication Date: 2025-08-08NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510529075.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing multispectral TDI sensors are difficult to reduce the cell size during imaging, making it difficult to obtain full-color images and high-resolution multispectral data simultaneously, and deep learning methods do not perform well in terms of stability and accuracy.

Method used

The coded mask plate is designed and integrated on chip, and the multi-spectral full-color image fusion is performed through a convolutional neural network, and the whole-color sharpened image is recombined with the coded mask plate imaging model.

Benefits of technology

The pixel size reduction of multispectral TDI sensors at the current process level is achieved, which improves the high resolution and stability of image reconstruction, and has better generalization performance than deep learning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451026A_ABST
    Figure CN120451026A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multispectral TDI sensor panchromatic sharpening, in particular to a multispectral TDI sensor panchromatic sharpening method based on on-chip coding integration, which comprises the following steps: designing a coding mask plate, and carrying out on-chip integration of the coding mask plate on a multispectral TDI sensor B spectrum; obtaining a modulated B-spectrum two-dimensional measurement value image, and obtaining a panchromatic P-spectrum two-dimensional measurement value image through a P spectrum; establishing a coding mask plate imaging model; a B-spectrum two-dimensional measurement value image and a panchromatic P-spectrum two-dimensional measurement value image are preprocessed and then serve as network input, multispectral panchromatic image fusion is conducted through a convolutional neural network, network output is obtained, and the network output is recombined in combination with a coding mask imaging model to obtain a panchromatic sharpened image. According to the method, the problem that the physical pixel size of the multispectral TDI sensor is difficult to reduce by adopting the existing method is solved, and panchromatic sharpening of the multispectral image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of panchromatic sharpening of a multispectral TDI sensor, and in particular to a panchromatic sharpening method of a multispectral TDI sensor based on on-chip coding integration. Background Art

[0002] With the rapid development of remote sensing technology and national defense and military technology, the demand for imaging with high spatial resolution and fine spectral information is increasing. When TDI (Time Delay Integration) multispectral sensors perform imaging tasks, due to certain limitations of current semiconductor process levels, the pixel size is difficult to further reduce, making it quite difficult to simultaneously obtain full-color images and high-resolution multispectral data. To solve this problem, deep learning-based full-color sharpening imaging technology is now commonly used. It can generate high-resolution full-color sharpened images. However, due to the lack of a rigorous mathematical and logical foundation, deep learning methods perform poorly in generalization performance and are prone to artifacts and color distortion when fusing images. Therefore, deep learning methods have great limitations in the aerospace and defense fields, which have extremely high requirements for stability and accuracy. Summary of the Invention

[0003] In order to solve the technical problems that existing deep learning methods have poor model generalization performance for panchromatic sharpening technology, cannot simultaneously obtain panchromatic images and high-resolution multispectral data, and have low stability and accuracy, the purpose of the present invention is to provide a panchromatic sharpening method for multispectral TDI sensors based on on-chip coding integration. The technical solutions adopted are as follows:

[0004] Design a coding mask based on the multispectral TDI sensor, and integrate the coding mask on-chip for the B spectrum of the multispectral TDI sensor;

[0005] The modulated B spectrum two-dimensional measurement value image is obtained by the B spectrum of the multi-spectral TDI sensor integrated by on-chip coding, and the full-color P spectrum two-dimensional measurement value image is obtained by the P spectrum;

[0006] A coded mask imaging model is established based on the B spectrum two-dimensional measurement value image and the coded mask;

[0007] The B-spectrum two-dimensional measurement value image and the panchromatic P-spectrum two-dimensional measurement value image are preprocessed and used as the network input. A convolutional neural network is designed based on the coded mask imaging model. Multispectral panchromatic image fusion is performed through the convolutional neural network to obtain the network output. The network output is recombined with the coded mask imaging model to obtain the panchromatic sharpened image.

[0008] Preferably, a coding mask is designed based on a multispectral TDI sensor, and the coding mask is integrated on-chip for the B spectrum of the multispectral TDI sensor, including:

[0009] The sub-pixel encoding array scale and the single sub-pixel encoding unit size of the encoding mask in the B spectrum are determined according to the pixel array scale and the single physical pixel size of the physical pixel of the P spectrum in the multispectral TDI sensor;

[0010] Sequentially encode the sub-pixel encoding units of the sub-pixel encoding array of the B spectrum segment, and output the encoding mask plate of each B spectrum segment;

[0011] Each B spectrum segment is integrated on-chip with the corresponding coding mask.

[0012] Preferably, the pixel array size of the physical pixels of the P spectrum is the same as the sub-pixel coding array size of the coding mask; the size of a single physical pixel of the P spectrum is the same as the size of a single sub-pixel coding unit.

[0013] Preferably, obtaining a modulated B spectrum two-dimensional measurement value image through the B spectrum of the multi-spectral TDI sensor integrated with on-chip coding, and obtaining a full-color P spectrum two-dimensional measurement value image through the P spectrum, comprises:

[0014] After the coded mask modulates the incident light, the on-chip coded integrated multispectral TDI sensor performs push-sweep scanning to obtain electrical signal data. The on-chip coded integrated multispectral TDI sensor then performs time-delay integration on the electrical signal data to output a one-dimensional measurement value image.

[0015] Acquiring a preset condition, caching the one-dimensional measurement value image, and when the cached one-dimensional measurement value image meets the preset condition, combining the cached one-dimensional measurement value image to output a B-spectrum two-dimensional measurement value image;

[0016] Based on the time delay integration, a full-color P spectrum two-dimensional measurement value image is obtained in the same way.

[0017] Preferably, establishing a coding mask imaging model based on the B-spectrum two-dimensional measurement value image and the coding mask includes:

[0018] Determining a two-dimensional coded image according to the coded mask and the B-spectrum two-dimensional measurement value image;

[0019] Each pixel in the B-spectrum two-dimensional measurement value image is matched with a pixel in the two-dimensional coded image, and the measurement value of the B-spectrum two-dimensional measurement value image is calculated. The corresponding calculation formula is:

[0020]

[0021] in, L(x,y) represents the measurement value of the (x,y)th pixel; Y b represents the sum of the measured values of the B-spectrum two-dimensional measurement value image corresponding to the b-th B-spectrum segment; M1×P1 represents the pixel scale of the B-spectrum two-dimensional measurement value image corresponding to the b-th B-spectrum segment, M represents the column, and P represents the row; z1 and z2 represent that one pixel of the two-dimensional measurement value image corresponds to the z1×z2 pixels of the two-dimensional coded image; F(x i ,y j ) represents the (x,y)th 2D measurement image pixel corresponding to the (x,y)th 2D coded image pixel. i ,y j ) pixel values, which can be 0 or 1; S(x i ,y j ) represents the (x,y)th pixel in the expected high-resolution B-spectrum image corresponding to the (x,y)th pixel in the two-dimensional measurement image. i ,y j ) pixels.

[0022] Preferably, the B spectrum two-dimensional measurement value image and the panchromatic P spectrum two-dimensional measurement value image are preprocessed and used as network input, a convolutional neural network is designed based on the coded mask imaging model, multispectral panchromatic image fusion is performed through the convolutional neural network, network output is obtained, and the network output is recombined in combination with the coded mask imaging model to obtain a panchromatic sharpened image, including:

[0023] Obtain network input through preprocessing;

[0024] Design a convolutional neural network, which includes a feature extraction module, a feature enhancement module, and a feature reconstruction module. The feature extraction module, the feature enhancement module, and the feature reconstruction module respectively perform feature extraction, deep feature generation, and feature reconstruction. The preprocessed network input is passed through the convolutional neural network to obtain a network output;

[0025] The network output is divided into several groups of output results, and each group of output results is recombined to obtain a full-color sharpened image.

[0026] Preferably, the network input is obtained by preprocessing, including:

[0027] Based on the pixel scale of each B spectrum segment in the B spectrum two-dimensional measurement value image, the coded pixels of the corresponding two-dimensional coded image are selected from the first pixel to the last pixel. From the selected coded pixels, the coded pixels at the same relative position are selected to split and recombine into several coded images to construct the network input corresponding to each B spectrum segment of the B spectrum two-dimensional measurement value image. The corresponding calculation formula is:

[0028]

[0029] in, represents the middle value, M1×P1 represents the B spectrum segment currently being analyzed, i.e., the pixel scale of the B spectrum two-dimensional measurement value image corresponding to the b-th B spectrum segment, M represents the column, and P represents the row; Y represents the total measurement value of the B spectrum two-dimensional measurement value image corresponding to the b-th B spectrum segment; z1 and z2 represent that one pixel of the two-dimensional measurement value image corresponds to the z1×z2 pixels of the two-dimensional coded image; M t represents the t-th coded image after the two-dimensional coded image is split and reassembled; X b represents the network input corresponding to the bth B-spectrum segment after preprocessing, b∈{1,…,n}, n represents the total number of B-spectrum segments; represents Hadamard division, ⊙ represents Hadamard multiplication;

[0030] According to the B spectrum two-dimensional measurement value image, the network input corresponding to each B spectrum segment is merged to obtain the network input of the two-dimensional coded image after preprocessing:

[0031] Similarly, the full-color P spectrum two-dimensional measurement value image is split and reassembled to obtain the corresponding P spectrum preprocessed network input:

[0032] The pre-processed network inputs corresponding to the B spectrum and P spectrum are merged to determine the final network input as

[0033] Preferably, a convolutional neural network is designed, and feature extraction, deep feature generation and feature reconstruction are performed respectively through a feature extraction module, a feature enhancement module and a feature reconstruction module, and the preprocessed network input is passed through the convolutional neural network to obtain a network output, including:

[0034] The feature extraction module extracts original features from the preprocessed network input to generate a feature map;

[0035] Generate deep high-dimensional features through feature enhancement modules;

[0036] The deep high-dimensional features are reconstructed through the feature reconstruction module to obtain the network output.

[0037] Preferably, the network output is divided into several groups of output results, and each group of output results is recombined to obtain a full-color sharpened image, including:

[0038] The network output is Divide each z1z2 channel into a group in sequence, and divide it into n groups in total, that is,

[0039] Each group is split and reassembled according to the two-dimensional coded image The inverse process of the method is reorganized to obtain n groups of full-color sharpened images Where M2×P2 represents the pixel size of the high-resolution B-spectrum image, M represents the column, and P represents the row.

[0040] The present invention has the following beneficial effects:

[0041] By integrating the coded mask on-chip on the multispectral TDI sensor to obtain the corresponding measurement value image, the measurement value is input into the designed convolutional neural network model to reconstruct the high-resolution multispectral remote sensing image, thereby achieving full-color sharpening of the image. This solves the problem that the pixel size of the multispectral TDI sensor is difficult to further reduce under the existing process level. The proposed full-color sharpening method for the multispectral TDI sensor based on on-chip coding integration has better generalization performance than other deep learning methods, and can improve the quality of reconstructed high-resolution images. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 A flowchart of a method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration provided by one embodiment of the present invention;

[0044] Figure 2 A schematic diagram of an on-chip coded integrated multispectral TDI sensor according to a method for full color sharpening of a multispectral TDI sensor based on on-chip coded integration provided by one embodiment of the present invention;

[0045] Figure 3 A schematic structural diagram of a B-spectrum sensor for a multi-spectral TDI sensor full-color sharpening method based on on-chip coding integration provided by one embodiment of the present invention;

[0046] Figure 4 A schematic diagram of the structure of a convolutional neural network for a pan-sharpening method for a multispectral TDI sensor based on on-chip coding integration provided by one embodiment of the present invention;

[0047] Figure 5 Comparison of a pan-sharpened image obtained by a multispectral TDI sensor based on on-chip coding integration provided by an embodiment of the present invention and a traditional depth method Figure 1 ;

[0048] Figure 6 Comparison of a pan-sharpened image obtained by a multispectral TDI sensor based on on-chip coding integration provided by an embodiment of the present invention and a traditional depth method Figure 2 ;

[0049] Figure 7 Comparison of a pan-sharpened image obtained by a multispectral TDI sensor based on on-chip coding integration provided by an embodiment of the present invention and a traditional depth method Figure 3 . DETAILED DESCRIPTION

[0050] To further illustrate the technical means and effectiveness of the present invention in achieving its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effectiveness of a pan-sharpening method for a multispectral TDI sensor based on on-chip coding integration, according to the present invention. In the following description, references to different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0051] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0052] The following describes in detail a method for full-color sharpening of a multi-spectral TDI sensor based on on-chip coding integration provided by the present invention with reference to the accompanying drawings.

[0053] See also Figure 1 , which shows a flowchart of a method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration provided by one embodiment of the present invention, the method comprising:

[0054] Step S1: designing a coding mask based on the multispectral TDI sensor, and integrating the coding mask on-chip for the B spectrum of the multispectral TDI sensor;

[0055] Step S2: obtaining a modulated B spectrum two-dimensional measurement value image through the B spectrum of the multi-spectral TDI sensor integrated with on-chip coding, and obtaining a full-color P spectrum two-dimensional measurement value image through the P spectrum;

[0056] Step S3: establishing a coding mask imaging model based on the B spectrum two-dimensional measurement value image and the coding mask;

[0057] Step S4: The B-spectrum two-dimensional measurement value image and the panchromatic P-spectrum two-dimensional measurement value image are preprocessed and used as network input. A convolutional neural network is designed based on the coded mask imaging model. Multispectral panchromatic image fusion is performed through the convolutional neural network to obtain the network output. The network output is recombined with the coded mask imaging model to obtain a panchromatic sharpened image.

[0058] For better explanation, a multispectral TDI sensor with on-chip coding integration refers to a multispectral TDI sensor that uses an on-chip coding mask to distinguish it from a traditional multispectral TDI sensor without a coding mask. The multispectral TDI sensor with on-chip coding integration performs specific encoding on the acquired image and adds specific constraints to the image reconstruction algorithm, so that the reconstruction algorithm has better anti-interference ability in the face of noise and interference, and can handle more complex scenes.

[0059] Specifically, in the application of on-chip coded integrated multispectral TDI sensors, the B spectrum and P spectrum typically refer to different spectral ranges or bands collected by the sensor. The B spectrum is a multispectral band, typically consisting of several to a dozen bands, each relatively wide. For example, common land satellite multispectral imagery may include several bands, such as blue, green, red, and near-infrared. For multispectral TDI sensors, the B spectrum can be customized to include its own spectral bands and number. The P spectrum (Pan Band) is a relatively broad band, covering electromagnetic waves from visible light to the near-infrared. It generally does not have a specific spectral resolution, but can provide high spatial resolution and is often used to provide a basis for image sharpening or fusion. That is, by combining the high spatial resolution of the panchromatic band with information from other multispectral bands, the spatial detail and clarity of the image can be enhanced.

[0060] Furthermore, step S1 includes:

[0061] Step S11: determining the sub-pixel coding array size and the single sub-pixel coding unit size of the coding mask in the B spectrum according to the pixel array size and the single physical pixel size of the physical pixels in the P spectrum of the multispectral TDI sensor.

[0062] It is explained that the coding mask includes sub-pixel coding units, and the B spectrum and P spectrum include physical pixels; the sub-pixel coding array scale of the coding mask in the B spectrum of the multispectral TDI sensor is determined by the pixel array scale of the physical pixels of the P spectrum in the multispectral TDI sensor, where the physical pixel array scale of the P spectrum refers to the physical number or distribution of each physical pixel in the sensor, which can reflect the spatial resolution of the sensor, and high spatial resolution helps to improve the details and clarity of the image; the sub-pixel coding array of the coding mask is used in the multispectral sensor to perform spatial encoding on the B spectrum band in order to perform specific spectral encoding or modulation on the imaging area; the sub-pixel coding array refers to a subdivided area used to precisely control the spatial sampling of the image, which is usually smaller than a single pixel and can enhance the details and resolution of the image, especially for low light or complex environments; the scale of the sub-pixel coding array of the B spectrum needs to be consistent with the scale of the physical pixel array of the P spectrum to ensure that the information between different bands can be effectively integrated while optimizing image quality.

[0063] Furthermore, the pixel array scale of the physical pixels of the P spectrum is the same as the sub-pixel coding array scale of the coding mask; the size of a single physical pixel of the P spectrum is the same as the size of a single sub-pixel coding unit.

[0064] As an optional implementation, in this embodiment, the pixel array size of the physical pixel of the P spectrum in the multispectral TDI sensor is defined as M2×N2, where M represents columns and N represents rows; then the sub-pixel coding array size of the coding mask is correspondingly M2×N2, that is, the size of a single sub-pixel coding unit of the coding mask is determined by the size of a single physical pixel of the P spectrum; at this time, based on the pixel array size of the physical pixel of the P spectrum, the size of a single physical pixel is determined to be m2×n2, and the size of a single sub-pixel coding unit of the coding mask is also m2×n2; and then the B spectrum of the multispectral TDI sensor is B b ,b∈{1,…,n}; n represents the total number of spectral segments of B spectrum; P spectrum is P p ,p∈{1}; 1 represents the number of spectrum segments of the P spectrum.

[0065] Step S12: sequentially encode the sub-pixel coding units of the sub-pixel coding array of the B spectrum segment, and output the coding mask of each B spectrum segment.

[0066] Specifically, a single physical pixel of the multispectral TDI sensor's B spectrum corresponds to the z1×z2 sub-pixel coding units on the coding mask, where z1 and z2 represent the scale multiples of the corresponding columns and rows. Assuming that the pixel array size of a single physical pixel is M1×N1, the calculation formulas for z1 and z2 are:

[0067] z1=M2 / M1

[0068] z2=N2 / N1

[0069] Among them, M2 represents the number of sub-pixel coding units contained in each level of the sub-pixel coding array scale of the coding mask plate, that is, the number of columns; M1 represents the number of physical pixels contained in each level of the pixel array scale of the physical pixels of the B spectrum; N2 represents the number of levels of the sub-pixel coding array of the coding mask plate, that is, the number of rows; N1 represents the number of pixel array levels of the physical pixels of the B spectrum; and z1 and z2 are natural numbers that must be greater than or equal to 2.

[0070] At the same time, assuming that the size of a single physical pixel of the B spectrum of the multispectral TDI sensor is m1×n1, the calculation formulas corresponding to z1 and z2 are:

[0071] z1=m1 / m2

[0072] z2=n1 / n2

[0073] Among them, m2 represents the row width of the sub-pixel coding unit; m1 represents the row width of the physical pixel of the B spectrum; n2 represents the column height of the sub-pixel coding unit; n1 represents the column height of the physical pixel of the B spectrum.

[0074] See also Figure 2 , which shows a schematic diagram of an on-chip coding integrated multispectral TDI sensor of a multispectral TDI sensor full color sharpening method based on on-chip coding integration provided by an embodiment of the present invention; wherein, when z1=z2=2, the B spectrum of the multispectral TDI sensor is encoded, and the P spectrum and B spectrum in the multispectral TDI sensor are encoded as Figure 2 The forms are arranged in the horizontal direction along the x-axis, and the lengths of the two in the vertical direction of the y-axis are consistent, and the multispectral TDI sensor is pushed and scanned in the horizontal direction along the x-axis when working.

[0075] Specifically, the sub-pixel coding units of the sub-pixel coding array scale of each B spectrum segment are encoded, and all the sub-pixel coding units of each z2 level in the sub-pixel coding array scale are taken as a coding period, then the N2-level sub-pixel coding array can be divided into N1 coding periods, that is, the encoding of the sub-pixel coding units of one coding period is determined by using an imaging model that satisfies snapshot compression imaging, and the same encoding method is used to repeatedly encode each coding period of the sub-pixel coding array scale, so that the encoding of the sub-pixel coding units at the same position in each coding period is the same, and the encoding of all sub-pixel coding units in the sub-pixel coding array is obtained; wherein, the encoding values of the sub-pixel coding units include "0" and "1", the encoding pixel with the encoding value of "0" is a closed element, and the encoding pixel with the encoding value of "1" is a non-closed element; then based on the imaging model, the distribution of the coding units of the sub-pixel coding array satisfies the 0-1 random distribution, then the probability of each coding unit in the sub-pixel coding array taking the value of 0 is p0, and correspondingly, the probability of taking the value of 1 is 1-p0, and 0 <p0<1。

[0076] See also Figure 3 , which shows a schematic diagram of the structure of a B spectrum sensor of a multi-spectral TDI sensor full color sharpening method based on on-chip coding integration provided by one embodiment of the present invention; wherein, when z1=z2=2, a schematic diagram of the coding mask obtained after encoding the B spectrum in the multi-spectral TDI sensor is used to define B i As an example, i represents the i-th B spectrum segment. Since z1=2, it means that the magnification in the horizontal direction of the x-axis is 2, that is, 2 columns of sub-pixel coding units are taken as the first-level integral pixels in the horizontal direction of the x-axis, corresponding to Figure 3 For the i-th level integral pixel, it can be seen that the black square represents the code with a code value of "0", and the white square represents the code with a code value of "1". For the coding mask of each level of integral pixels, the coding distribution of the sub-coding pixel units is consistent; since z2=2, two rows in the y-axis vertical direction of each level of integral pixels are a physical pixel, and one physical pixel corresponds to four sub-pixel coding units. For each level of integral pixels, the coding mask of each physical unit in the y-axis vertical direction has different codes.

[0077] As an optional implementation, each B spectrum sensor in the multispectral TDI sensor has the same size, assuming that they are all 153.6mm×76.8mm, and the size of a single physical pixel of the sensor is 150um×150um. Then the resolution of each B spectrum sensor is 1024×512. The size of the P spectrum sensor is also 153.6mm×76.8mm, and the size of a single physical pixel of the sensor is 75um×75um. Then the resolution of the P spectrum sensor is 2046×1024. At this time, the corresponding designed coding mask plate size is 153.6mm×76.8mm, and the size of a single sub-pixel coding unit is 75um×75um. The resolution of the high-resolution spectral image of each spectral segment is finally obtained with a resolution of 2046×1024.

[0078] Step S13: Integrate each B spectrum segment with the corresponding coding mask on-chip.

[0079] It is explained that the B spectrum segment and the coding mask plate are integrated on-chip, that is, the B spectrum segment and the coding mask plate are combined so that each spectrum segment can be processed by the coding mask plate; specifically, the coding mask plate is tightly attached to the surface of the B spectrum photoelectric image sensor in an on-chip integration manner, and the distance between the coding mask plate and the photoelectric image sensor is 0. The "0" in the coding mask plate is the part that photons cannot pass through, and the "1" is the part that photons can pass through. The coding mask plate is a thin aluminum plate or other opaque material engraved with a coding pattern and covers the photosensitive surface of the B spectrum sensor of the multi-spectral TDI sensor.

[0080] Furthermore, step S2 includes:

[0081] Step S21: After the coding mask modulates the incident light, the on-chip coded integrated multispectral TDI sensor performs push-scanning to obtain electrical signal data, and the on-chip coded integrated multispectral TDI sensor performs time-delay integration on the electrical signal data to output a one-dimensional measurement value image.

[0082] Specifically, after the coded mask modulates the incident light, it is pushed and scanned by the multispectral TDI sensor in a push-scan direction parallel to the horizontal direction of the x-axis. The electrical signal data obtained by the scan is time-delayed and integrated by the multispectral TDI sensor, and the electrical signal data obtained after the time delay integration is output as a one-dimensional measurement value image with a pixel scale of M1×1.

[0083] Step S22: obtaining preset conditions, caching the one-dimensional measurement value images, and when the cached one-dimensional measurement value images meet the preset conditions, combining the cached one-dimensional measurement value images to output a B-spectrum two-dimensional measurement value image.

[0084] Specifically, the preset condition refers to a pre-set number of rows P1. The one-dimensional measurement value image is cached through a cache circuit to achieve fast access and storage, avoiding repeated calculation or reading. When the cached one-dimensional measurement value image meets the preset condition, the corresponding P1 rows of one-dimensional measurement value images are combined into a B-spectrum two-dimensional measurement value image with a pixel size of M1×P1 for output.

[0085] Step S23: similarly obtaining a full-color P spectrum two-dimensional measurement value image based on the time delay integration; that is, the pixel scale of the full-color P spectrum two-dimensional measurement value image is similarly M2×P2.

[0086] Furthermore, step S3 includes:

[0087] Step S31: determining a two-dimensional coded image according to the coded mask and the B-spectrum two-dimensional measurement value image;

[0088] It can be understood that it has the same mathematical expression as the aforementioned imaging model that satisfies snapshot compression imaging, and the imaging model has a complete mathematical convergence proof in mathematical form; specifically, the pixel scale of the two-dimensional coded image is determined according to the pixel scale of the B-spectrum two-dimensional measurement value image, and the encoding of the two-dimensional coded image is determined according to the encoding of the encoding pixel array of the encoding mask; the pixel scale of the two-dimensional coded image is M2×P2, and the encoding of the two-dimensional coded image is composed of continuously repeated encoding every z1 row, and the continuously repeated encoding is the same as the encoding pixel of the first-level sub-pixel encoding unit of the encoding mask.

[0089] Step S32: each pixel in the B-spectrum two-dimensional measurement value image is matched with a pixel in the two-dimensional coded image, and the measurement value of the B-spectrum two-dimensional measurement value image is calculated. The corresponding calculation formula is:

[0090]

[0091] in, L(x,y) represents the measurement value of the (x,y)th pixel; Y b represents the sum of the measured values of the B-spectrum two-dimensional measurement value image corresponding to the b-th B-spectrum segment; M1×P1 represents the pixel scale of the B-spectrum two-dimensional measurement value image corresponding to the b-th B-spectrum segment, M represents the column, and P represents the row; z1 and z2 represent that one pixel of the two-dimensional measurement value image corresponds to the z1×z2 pixels of the two-dimensional coded image; F(x i ,y j ) represents the (x,y)th 2D measurement image pixel corresponding to the (x,y)th 2D coded image pixel. i ,y j ) pixel values, which can be 0 or 1; S(x i ,y j) represents the (x,y)th pixel in the expected high-resolution B-spectrum image corresponding to the (x,y)th pixel in the two-dimensional measurement image. i ,y j ) pixels.

[0092] See also Figure 4 , which shows a structural schematic diagram of a convolutional neural network for a full-color sharpening method of a multi-spectral TDI sensor based on on-chip coding integration provided by an embodiment of the present invention.

[0093] Furthermore, step S4 includes:

[0094] Step S41: Obtain network input through preprocessing.

[0095] It is explained that the B spectrum two-dimensional measurement value image, the two-dimensional coding image and the full-color P spectrum two-dimensional measurement value image are preprocessed and used as network input, so that the network input is more suitable for fusion reconstruction. The preprocessed network input avoids the problem of unbalanced energy distribution caused by random coding mode.

[0096] Furthermore, step S41 includes:

[0097] Step S411: Based on the pixel scale of each B-spectrum segment in the B-spectrum two-dimensional measurement value image, the coded pixels corresponding to the two-dimensional coded image are selected from the first pixel to the last pixel. From the selected coded pixels, the coded pixels at the same relative position are selected to split and reassemble into several coded images, and the network input corresponding to each B-spectrum segment of the B-spectrum two-dimensional measurement value image is constructed. The corresponding calculation formula is:

[0098]

[0099] in, represents the middle value, M1×P1 represents the B spectrum segment currently being analyzed, i.e., the pixel scale of the B spectrum two-dimensional measurement value image corresponding to the b-th B spectrum segment, M represents the column, and P represents the row; Y represents the total measurement value of the B spectrum two-dimensional measurement value image corresponding to the b-th B spectrum segment; z1 and z2 represent that one pixel of the two-dimensional measurement value image corresponds to the z1×z2 pixels of the two-dimensional coded image; M t represents the t-th coded image after the two-dimensional coded image is split and reassembled; X b represents the network input corresponding to the bth B-spectrum segment after preprocessing, b∈{1,…,n}, n represents the total number of B-spectrum segments; represents Hadamard division, and ⊙ represents Hadamard multiplication.

[0100] Step S412: Merge the network input corresponding to each B spectrum segment according to the B spectrum two-dimensional measurement value image to obtain the network input after preprocessing of the two-dimensional coded image:

[0101] Step S413: Similarly, the full-color P spectrum two-dimensional measurement value image is split and reassembled to obtain the pre-processed network input corresponding to the P spectrum:

[0102] Step S414: Combine the pre-processed network inputs corresponding to the B spectrum and the P spectrum to determine the final network input as Specifically, the network inputs obtained in step S412 and step S413 are merged to obtain the final network input X.

[0103] Step S42: Design a convolutional neural network, which includes a feature extraction module, a feature enhancement module, and a feature reconstruction module. Feature extraction, deep feature generation, and feature reconstruction are performed through the feature extraction module, the feature enhancement module, and the feature reconstruction module respectively. The preprocessed network input is passed through the convolutional neural network to obtain the network output.

[0104] The feature extraction module consists of a 3×3 convolutional layer. The feature enhancement module uses a U-shaped network structure to generate deep features. The U-shaped network includes an encoder, a feature mixing module, and a decoder. The encoder consists of two convolutional layers and two downsampling modules. The downsampling module uses 3×3 convolutional layers to reduce the size of the feature map and expand the number of channels. The feature mixing module uses channel splitting and channel shuffling to efficiently mix features, enhancing features and reducing computational cost. The decoder consists of two convolutional layers and two upsampling modules. The upsampling module uses 3×3 convolutional layers to expand the size of the feature map and reduce the number of channels. The image reconstruction module consists of 3×3 2D convolutional layers. The feature reconstruction module reconstructs the deep features output by the feature enhancement module.

[0105] Furthermore, step S42 includes:

[0106] Step S421: extract original features from the pre-processed network input through the feature extraction module to generate a feature map; that is, extract original features from the pre-processed network input through the feature extraction module to generate a feature map. Perform original feature extraction.

[0107] Step S422: Generate deep high-dimensional features through the feature enhancement module.

[0108] Specifically, the encoder in the feature enhancement module encodes the feature map to generate a high-dimensional feature map, and the decoder decodes the feature map to reduce the feature dimension. The feature mixing module includes channel separation and channel shuffling mechanisms. Through channel separation, the high-dimensional feature map is split into X1 and X2 along the channel dimension. Among them, X1 passes through a 1×1 convolution layer, then passes through an activation function, and then passes through a 1×1 convolution layer to obtain the output feature Xc1 ; Then X c1 The X2 of the identity mapping is connected in the channel dimension, and the channel shuffling and reorganization allow the deep high-dimensional features to be fully integrated.

[0109] Step S423: reconstructing the deep high-dimensional features through the feature reconstruction module to obtain network output.

[0110] Specifically, the deep high-dimensional features output by the feature mixing module are reconstructed into an image through the feature reconstruction module, and the network output is a two-dimensional image of M1×P1 with a channel number of nz1z2. At this time, the convolution modules on the convolutional neural network exchange information through connections, and the multi-level cross-layer connections are used to improve the convergence speed of model training and the reconstruction quality of feature images.

[0111] Preferably, artificial constraints are added to the network input so that the convolutional neural network has better generalization performance between different data sets, and thus better image reconstruction effect between cross-data sets.

[0112] Step S43: Divide the network output into several groups of output results, and reorganize each group of output results to obtain a full-color sharpened image.

[0113] To clarify, pan-sharpened images generally refer to sharpening an image to enhance its details, especially in terms of color and texture, that is, a pan-sharpened image is a high-resolution multispectral fusion image; sharpening is generally achieved by enhancing the edges and details in the image, making the image appear clearer and more layered.

[0114] Furthermore, step S43 includes:

[0115] Step S431: The network output is Divide each z1z2 channel into a group in sequence, and divide it into n groups in total, that is,

[0116] Step S432: Split and reassemble each group into The inverse process of the method is reorganized to obtain n groups of full-color sharpened images Where M2×P2 represents the pixel size of the high-resolution B-spectrum image, M represents the column, and P represents the row.

[0117] Preferably, in this embodiment, the convolutional neural network is trained and inferred through the pytorch deep learning platform. The data set used is panchromatic multispectral remote sensing data, including multiple remote sensing scenes such as cities, water bodies, farmlands, forests, and wastelands. The loss function used in training is the mean square error loss MSE, and the corresponding loss function calculation formula is:

[0118]

[0119] Where n represents the number of pixels; y i Represents the image corresponding to the reconstruction, that is, the pixel value output by the network; Represents the pixel value of the real original image.

[0120] It is explained that during the training process, a simulation method is used to simulate the coding mask plate to obtain the measurement value of the multispectral TDI sensor. The real original image is dot-multiplied with the designed coding mask plate image and then down-sampled to a low resolution corresponding to the measurement value image. The downsampling method is to merge the pixel values of every z1z2 pixels of the image after the dot multiplication, and divide the merged pixel value by z1z2 to obtain the pixel value of the measurement value image.

[0121] It can be understood that by performing on-chip integration of the coding mask plate on the multispectral TDI sensor to obtain the corresponding measurement value image, the measurement value is input into the designed convolutional neural network model to reconstruct the high-resolution multispectral remote sensing image, thereby achieving full-color sharpening of the image. This solves the problem that the pixel size of the multispectral TDI sensor is difficult to further reduce under the existing process level. The proposed full-color sharpening method of the multispectral TDI sensor based on on-chip coding integration has better generalization performance than other deep learning methods, and can improve the quality of reconstructed high-resolution images.

[0122] For better explanation and to verify the feasibility of the full-color sharpening method of a multispectral TDI sensor based on on-chip coding integration proposed in this application, scientific demonstration is carried out through economic benefit calculation and simulation experiments, and a comparative experiment is conducted with the existing traditional deep learning method.

[0123] Please combine Figure 5-Figure 7 , which respectively show the comparison of a pan-sharpened image obtained by a multispectral TDI sensor based on on-chip coding integration provided by an embodiment of the present invention and a traditional deep method, that is, a deep learning method without coding. Figure 1 ,contrast Figure 2 and contrast Figure 3; Specifically, seven existing classic pan-sharpening methods based on deep learning are compared, where GT is the real original image and SCI is the reconstruction result of the pan-sharpening method of this application. In the figure, from left to right are PGCU (Probability-based Global Cross-modal Upsampling for Pan-sharpening, that is, global cross-modal upsampling pan-sharpening network), GPPNN (Deep Gradient Projection Networks for Pan-sharpening, that is, deep gradient projection network for pan-sharpening), Pannet (PanNet: A deep network architecture for pan-sharpening, that is, deep network architecture for pan-sharpening), PNN (Pansharpening by Convolutional Neural Networks, that is, convolutional neural network for pan-sharpening), MSDCNN (A Multi-Scale and Multi-Depth Convolutional Neural Network for Remote Sensing Imagery Pan-Sharpening, that is, multi-scale and multi-depth pan-sharpening convolutional neural network), Fusionnet (Detail Injection-Based Deep Convolutional Neural Network Seven comparative methods are presented, namely, Pansharpening via Detail Injection Based Convolutional Neural Networks (i.e., pansharpening network based on detail fusion), Dicnn (Pansharpening via Detail Injection Based Convolutional Neural Networks, i.e., pansharpening network based on detail injection). These methods upsample the multispectral image and input it into the network for pansharpening, which is different from the encoding-based pansharpening method in this paper that does not require upsampling.

[0124] It can be explained that the data set used in the experiment contains four spectral bands: R, G, B, and infrared. The result picture shown in the figure is a true color picture of the fusion of the three RGB bands in the network output; the two numbers below the picture represent the PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity) indicators between the true value and the reconstructed value of the multispectral data from left to right. The calculation method is to calculate the PSNR and SSIM for each spectral band, add them up, and then divide them by the average of the total number of spectral bands. The higher the values of the two, the higher the degree of restoration of the image.

[0125] Make an explanation, Figure 5 The scene in the game is a town, which contains a large number of buildings; Figure 5 It can be seen that SCI has excellent results and is superior to other comparison methods in terms of indicators. This shows that the full-color sharpening method proposed in this application has a very high degree of color and structure restoration, and can clearly distinguish buildings, which is better than other comparison methods. Figure 6 The scene in the image is farmland, containing a large number of crops. SCI better restores the dividing boundary lines between farmlands, and the boundaries are very clear, with fewer image artifacts. However, the farmland boundaries of other comparison methods are blurred, and it is impossible to distinguish between two farmlands. Some comparison methods even produce large artifacts. In addition, the indicators of the full-color sharpening method of this application are better than those of other comparison methods. Figure 7 The scene in the image is a forest mountain with a lot of green vegetation. SCI can better restore the vegetation zone and the non-vegetation area, with clearer boundaries and fewer image artifacts. Other contrast methods will produce more artifacts and the structure will be blurred.

[0126] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0127] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A multispectral TDI sensor full color sharpening method based on on-chip coding integration, characterized in that: The method comprises: Design a coding mask based on the multispectral TDI sensor, and integrate the coding mask on-chip for the B spectrum of the multispectral TDI sensor; The modulated B spectrum two-dimensional measurement value image is obtained by the B spectrum of the multi-spectral TDI sensor integrated by on-chip coding, and the full-color P spectrum two-dimensional measurement value image is obtained by the P spectrum; A coded mask imaging model is established based on the B spectrum two-dimensional measurement value image and the coded mask; The B-spectrum two-dimensional measurement value image and the panchromatic P-spectrum two-dimensional measurement value image are preprocessed and used as the network input. A convolutional neural network is designed based on the coded mask imaging model. Multispectral panchromatic image fusion is performed through the convolutional neural network to obtain the network output. The network output is recombined with the coded mask imaging model to obtain the panchromatic sharpened image.

2. The method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration according to claim 1, characterized in that: Design a coding mask based on the multispectral TDI sensor and integrate the coding mask on-chip for the B spectrum of the multispectral TDI sensor, including: The sub-pixel encoding array scale and the single sub-pixel encoding unit size of the encoding mask in the B spectrum are determined according to the pixel array scale and the single physical pixel size of the physical pixel of the P spectrum in the multispectral TDI sensor; Sequentially encode the sub-pixel encoding units of the sub-pixel encoding array of the B spectrum segment, and output the encoding mask plate of each B spectrum segment; Each B spectrum segment is integrated on-chip with the corresponding coding mask.

3. The method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration according to claim 2, characterized in that: The pixel array scale of the physical pixels of the P spectrum is the same as the sub-pixel coding array scale of the coding mask; the size of a single physical pixel of the P spectrum is the same as the size of a single sub-pixel coding unit.

4. The method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration according to claim 1, characterized in that: The modulated B-spectrum 2D measurement value image is obtained through the B-spectrum of the multispectral TDI sensor integrated with on-chip coding, and the full-color P-spectrum 2D measurement value image is obtained through the P-spectrum, including: After the coded mask modulates the incident light, the on-chip coded integrated multispectral TDI sensor performs push-sweep scanning to obtain electrical signal data. The on-chip coded integrated multispectral TDI sensor then performs time-delay integration on the electrical signal data to output a one-dimensional measurement value image. Acquiring a preset condition, caching the one-dimensional measurement value image, and when the cached one-dimensional measurement value image meets the preset condition, combining the cached one-dimensional measurement value image to output a B-spectrum two-dimensional measurement value image; Based on the time delay integration, a full-color P spectrum two-dimensional measurement value image is obtained in the same way.

5. The method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration according to claim 1, characterized in that: A coded mask imaging model is established based on the B-spectrum two-dimensional measurement value image and the coded mask, including: Determining a two-dimensional coded image according to the coded mask and the B-spectrum two-dimensional measurement value image; Each pixel in the B-spectrum two-dimensional measurement value image is matched with a pixel in the two-dimensional coded image, and the measurement value of the B-spectrum two-dimensional measurement value image is calculated. The corresponding calculation formula is: in, L(x,y) represents the measurement value of the (x,y)th pixel; Y b represents the sum of the measured values of the B-spectrum two-dimensional measurement value image corresponding to the b-th B-spectrum segment; M1×P1 represents the pixel scale of the B-spectrum two-dimensional measurement value image corresponding to the b-th B-spectrum segment, M represents the column, and P represents the row; z1 and z2 represent that one pixel of the two-dimensional measurement value image corresponds to the z1×z2 pixels of the two-dimensional coded image; F(x i ,y j ) represents the (x,y)th 2D measurement image pixel corresponding to the (x,y)th 2D coded image pixel. i ,y j ) pixel values, which can be 0 or 1; S(x i ,y j ) represents the (x,y)th pixel in the expected high-resolution B-spectrum image corresponding to the (x,y)th pixel in the two-dimensional measurement image. i ,y j ) pixels.

6. The method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration according to claim 5, characterized in that: The B-spectrum 2D measurement value image and the panchromatic P-spectrum 2D measurement value image are preprocessed and used as network input. A convolutional neural network is designed based on the coded mask imaging model. Multispectral panchromatic image fusion is performed through the convolutional neural network to obtain the network output. The network output is recombined with the coded mask imaging model to obtain a panchromatic sharpened image, including: Obtain network input through preprocessing; Design a convolutional neural network, which includes a feature extraction module, a feature enhancement module, and a feature reconstruction module. The feature extraction module, the feature enhancement module, and the feature reconstruction module respectively perform feature extraction, deep feature generation, and feature reconstruction. The preprocessed network input is passed through the convolutional neural network to obtain a network output; The network output is divided into several groups of output results, and each group of output results is recombined to obtain a full-color sharpened image.

7. The method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration according to claim 6, characterized in that: The network input is obtained through preprocessing, including: Based on the pixel scale of each B spectrum segment in the B spectrum two-dimensional measurement value image, the coded pixels of the corresponding two-dimensional coded image are selected from the first pixel to the last pixel. From the selected coded pixels, the coded pixels at the same relative position are selected to split and recombine into several coded images to construct the network input corresponding to each B spectrum segment of the B spectrum two-dimensional measurement value image. The corresponding calculation formula is: in, represents the middle value, M1×P1 represents the B spectrum segment currently being analyzed, i.e., the pixel scale of the B spectrum two-dimensional measurement value image corresponding to the b-th B spectrum segment, M represents the column, and P represents the row; Y represents the total measurement value of the B spectrum two-dimensional measurement value image corresponding to the b-th B spectrum segment; z1 and z2 represent that one pixel of the two-dimensional measurement value image corresponds to the z1×z2 pixels of the two-dimensional coded image; M t represents the t-th coded image after the two-dimensional coded image is split and reassembled; X b represents the network input corresponding to the bth B-spectrum segment after preprocessing, b∈{1,…,n}, n represents the total number of B-spectrum segments; represents Hadamard division, ⊙ represents Hadamard multiplication; According to the B spectrum two-dimensional measurement value image, the network input corresponding to each B spectrum segment is merged to obtain the network input of the two-dimensional coded image after preprocessing: Similarly, the full-color P spectrum two-dimensional measurement value image is split and reassembled to obtain the corresponding P spectrum preprocessed network input: The pre-processed network inputs corresponding to the B spectrum and P spectrum are merged to determine the final network input as 8. The method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration according to claim 6, characterized in that: Design a convolutional neural network, perform feature extraction, deep feature generation, and feature reconstruction through the feature extraction module, feature enhancement module, and feature reconstruction module, respectively. Pass the preprocessed network input through the convolutional neural network to obtain the network output, including: The feature extraction module extracts original features from the preprocessed network input to generate a feature map; Generate deep high-dimensional features through feature enhancement modules; The feature reconstruction module is used to reconstruct deep high-dimensional features to obtain network output.

9. The method for pan-sharpening a multispectral TDI sensor based on on-chip coding integration according to claim 6, characterized in that: The network output is divided into several groups of output results, and each group of output results is recombined to obtain a full-color sharpened image, including: The network output is Divide each z1z2 channel into a group in sequence, and divide it into n groups in total, that is, Each group is split and reassembled according to the two-dimensional coded image The inverse process of the method is reorganized to obtain n groups of full-color sharpened images Where M2×P2 represents the pixel scale of the high-resolution B-spectrum image, M represents the column, and P represents the row.