Image processing method and system based on near-infrared imaging
Through multi-branch convolutional neural network and dynamic afterimage correction technology, the problem of dynamic afterimage on high-speed production lines is solved, and effective separation and correction of background, target structure and dynamic afterimage are achieved, improving the accuracy and real-timeness of image detection.
Patent Information
- Application Number
- CN202510820627.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
On high-speed production lines, the dynamic afterimage phenomenon that occurs in the near-infrared image sequence causes the target edge or structure to superimpose positions in multi-frame images, affecting the accuracy of defect detection. The prior art is difficult to effectively deal with the time domain accumulation effect and multi-frame dependence characteristics of dynamic afterimage, and has high computational complexity and poor real-time performance.
Multi-branch convolutional neural network is used to learn spatial features of near-infrared images, separate background, target structure and dynamic afterimage areas, and use spatial alignment and neighborhood repair centered on the target structure area, and build a regional distribution prediction model with historical sample data to achieve rapid disappearance of dynamic afterimages.
The spatial resolution capability of dynamic afterimage areas is improved, the accuracy and image stability of multi-frame afterimage superposition analysis are improved, and the accuracy and real-time performance of visual inspection in production lines are ensured.
Smart Images

Figure CN120339137A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and more specifically, to an image processing method and system based on near-infrared imaging. Background Art
[0002] Near-infrared image processing technology is widely used in industrial production lines for defect detection, dimension measurement, and surface quality analysis of continuously moving targets. However, in a high-speed production environment, due to the rapid movement of the target and the integration characteristics of the image acquisition system, dynamic ghosting often occurs in near-infrared image sequences. This phenomenon is manifested as the superposition of the positions of the target edges or structures in multiple frames of images, resulting in "trailing" or "ghosting" residual signals in the images, seriously affecting the accuracy of defect detection. Existing image denoising or background modeling techniques mainly focus on static images or uniform motion blur, and it is difficult to effectively cope with the time-domain cumulative effect and multi-frame dependence characteristics of dynamic ghosting. Especially in high-frame-rate, pipeline-type detection scenarios, traditional algorithms often have high computational complexity, poor real-time performance, and are prone to loss of target structure information, so it is difficult to apply to specific production scenarios.
[0003] Therefore, there is an urgent need for a fast fading process of dynamic ghosting for near-infrared images in high-speed production lines, which can not only accurately separate dynamic ghosting from the real target structure, avoid detection errors caused by ghosting interference, but also improve the stability of target detection and the real-time performance of the system. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an image processing method based on near-infrared imaging to solve the problems proposed in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions: An image processing method based on near-infrared imaging, comprising the following steps: S1: Continuously collect near-infrared images of the production area based on the running speed of the production line to generate a pixel set of near-infrared image frames; S2: Obtain the pixel gray values of the pixel sets corresponding to all image frames and establish a pixel gray distribution map; S3: Use a multi-branch convolutional neural network to perform spatial feature learning on the pixel gray distribution map to separate the background, target structure, and dynamic ghosting area distributions; S4: Centering on the target structure area, perform spatial alignment on consecutive image frames, and stack all image frames to generate a global area distribution; S5: Obtain historical near-infrared image frame samples, train a neural network model through a training set constructed by the samples, and establish a regional distribution prediction model; S6: Predict the background and afterimage area distribution of the real-time image based on the area distribution prediction model, perform neighborhood repair on the afterimage area through the background area, and output the image after the afterimage is corrected.
[0006] In a preferred embodiment, in S1, continuous near-infrared image acquisition is performed on the production area based on the running speed of the production line, and a pixel set of the near-infrared image frame is generated, which specifically includes: Obtain the real-time running speed of the production line and dynamically adjust the near-infrared image acquisition frame rate based on the preset fixed total number of frames; Obtain external light source parameters, continuously collect near-infrared image frames of the target production area on the production line, record the acquisition timestamp of each frame, and use interpolation method to perform time axis homogenization on the image frames; The image frames corresponding to each timing index position on the time axis are calibrated with pixel numbers and row and column positions to generate a pixel set for each image frame.
[0007] In a preferred embodiment, in S2, obtaining the pixel grayscale values of the pixel sets corresponding to all image frames and establishing the pixel grayscale distribution map specifically includes: In the corresponding pixel set of the image frame, the light signal collected at each pixel row and column position is electronically converted to generate a corresponding pixel grayscale signal, and the pixel grayscale signal is quantized into a digital grayscale value; The digital grayscale value is converted into a two-dimensional grayscale matrix based on the row and column position arrangement of the pixels, and the two-dimensional grayscale matrix is marked as a pixel grayscale distribution map; The continuous image frame sequence is read and the image frame number index is marked to establish the corresponding relationship between the pixel grayscale distribution map and the image frame number to which it belongs.
[0008] In a preferred embodiment, in S3, a multi-branch convolutional neural network is used to perform spatial feature learning on a pixel grayscale distribution map, and separation of background, target structure, and dynamic residual image area distribution specifically includes: Scan the pixel grayscale distribution map pixel by pixel, and arrange the pixel grayscale values in row and column order to form a multi-dimensional input tensor; Set up a multi-branch convolutional neural network structure, retain the single-channel input format in the multi-dimensional input tensor, and configure convolution kernels of corresponding sizes for different branches respectively; Wherein, the multiple branches include background, target structure and dynamic residual image branches; Input the multi-dimensional input tensor into the multi-branch convolutional neural network, and perform local convolution operations in each branch based on the configured convolution kernel to extract pixel spatial features to construct feature maps of different branches, where the spatial features include texture features, edge features, and continuous area features; Perform cross-branch feature fusion on the feature maps of different branches, and separate the background, target structure, and dynamic afterimage area distributions through feature clustering.
[0009] In a preferred embodiment, the performing cross-branch feature fusion on the feature maps of different branches and separating the background, target structure, and dynamic afterimage area distributions through feature clustering specifically includes: In the output of the multi-branch convolutional neural network, sequentially read the feature maps of the background branch, target structure branch, and dynamic afterimage branch; Overlay the three types of feature maps, and construct a joint feature based on the background feature value, target structure feature value, and dynamic afterimage feature value at each pixel position to form a pixel-level joint feature set; Preset a joint feature classification template for the background, target structure, and dynamic afterimage areas, and establish corresponding category indexes respectively; Perform unsupervised clustering on the joint feature set within the entire pixel gray-scale distribution map, extract the local joint feature set of each category cluster, calculate the average feature of the local joint feature set, and label the category index of each clustering cluster through the distance between the average feature and the joint feature classification template; Map the labeled category indexes to the row and column positions of the pixel gray-scale distribution map to generate the background area, target structure area, and dynamic afterimage area distributions corresponding to a single image frame.
[0010] In a preferred embodiment, in S4, with the target structure area as the center, perform spatial alignment on consecutive image frames, and overlay all image frames to generate the global area distribution specifically includes: Use the target structure area as the central position of the pixel distribution map, and perform spatial alignment on the pixel distribution maps corresponding to all image frames; Based on the corresponding relationship of the pixel row and column positions of the background area, target structure area, and dynamic afterimage area distributions after spatial alignment, perform cross-frame overlay on all image frames to generate the global background area, global target structure area, and global dynamic afterimage area distributions.
[0011] In a preferred embodiment, in S5, obtain historical near-infrared image frame samples, and train a neural network model through the training set constructed by the samples to establish a regional distribution prediction model specifically includes: Obtain historical near-infrared image frame sample data, and each group of sample data covers different external light source parameters and production line running speeds; In each group of sample data, determine the global background area, global target structure area, and global dynamic afterimage area distributions for the pixel gray-scale distribution maps of all historical near-infrared image frames; Perform label marking processing on the sample data based on the determination results of the global background region, global target structure region, and global dynamic ghost region distribution, and construct the training set data of the region distribution prediction model; Select a neural network model, input the training set data to train the region distribution prediction model, and deploy the model to the production line end after the model training is completed.
[0012] In a preferred embodiment, in S6, predicting the background and ghost region distribution of the real-time image based on the region distribution prediction model, and performing neighborhood repair on the ghost region through the background region, and outputting the image with the ghost corrected specifically includes: Input the external light source parameters and real-time production line running speed at the production line end into the region distribution prediction model; Obtain the model output result, and perform background region and dynamic ghost region distribution annotation on the pixel gray value distribution map of the real-time collected near-infrared image frame; Extract the pixel position coordinates of the marked ghost region and background region in the pixel gray value distribution map; Perform neighborhood repair processing on the pixel gray values of the pixel gray value distribution map at the pixel position coordinates of the ghost region, perform pixel gray value correction on the ghost region according to the pixel gray values of the background region, and output the image frame with the ghost corrected.
[0013] On the other hand, the present invention provides an image processing method system based on near-infrared imaging, including a gray value construction module, a convolutional feature extraction module, a global distribution construction module, a prediction model training module, and a dynamic ghost correction module: Gray value construction module: Collect continuous near-infrared image frame samples during the operation of the production line, dynamically adjust the image frame rate in combination with the real-time production line speed and external light source parameters, and perform electronic conversion and quantization processing on the pixel optical signals of the image frames, generate digital gray values and mark the row and column positions, and construct a pixel gray value distribution map; Convolutional feature extraction module: Use a multi-branch convolutional neural network to perform spatial feature learning on the pixel gray value distribution map, configure convolutional kernels according to the background, target structure, and dynamic ghost branches and extract spatial features, generate branch feature maps, and output the pixel region distributions of the background, target structure, and dynamic ghost through feature fusion and unsupervised clustering to assign category indexes; Global distribution construction module: Use the target structure region as the alignment reference, perform spatial alignment on the continuous image frames, and based on the corresponding relationship of the pixel row and column positions, superimpose the background, target structure, and dynamic ghost region distributions across frames to generate the global background region, global target structure region, and global dynamic ghost region distributions; Prediction model training module: Obtain historical near-infrared image frame sample data and the corresponding global region distribution map, perform label marking on the sample data to construct a training set, select a neural network model to input the training set for training, and after training is completed, generate a region distribution prediction model and deploy it to the production line end; Dynamic afterimage correction module: Input the external light source parameters at the production line end and the real-time production line running speed into the region distribution prediction model, obtain the background and dynamic afterimage region distributions, perform neighborhood repair on the afterimage region, and use the pixel gray values of the background region to correct the afterimage region, and output the corrected image frame.
[0014] Technical effects and advantages of an image processing method and system based on near-infrared imaging according to the present invention: For the real-time near-infrared image acquisition data on the production line, through the multi-branch convolutional neural network to perform spatial feature learning on the pixel gray distribution map, it can effectively separate the background, target structure and dynamic afterimage region distributions, enhancing the spatial resolution ability of the dynamic afterimage region. Adopting the spatial alignment strategy centered on the target structure region solves the problem of regional distribution misalignment caused by target drift in consecutive image frames, improving the accuracy of multi-frame afterimage superposition analysis. Through the training set and region distribution prediction model constructed by historical sample data, combined with real-time light source parameters and production line running speed, the prediction ability of dynamic afterimage distribution is realized, significantly improving the adaptability of the model to dynamic afterimage interference under multiple working conditions. The neighborhood repair mechanism of the afterimage region effectively reduces the influence of dynamic afterimage interference on image quality, ensuring the image stability and accuracy of production line visual inspection and subsequent processes. Brief description of the drawings
[0015] Figure 1 It is a schematic diagram of an image processing method based on near-infrared imaging according to the present invention; Figure 2 It is a schematic structural diagram of an image processing system based on near-infrared imaging according to the present invention. Detailed implementation manners
[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0017] Embodiment 1, Figure 1 An image processing method based on near-infrared imaging according to the present invention is given, which includes the following steps: S1: Continuously collect near-infrared images of the production area based on the running speed of the production line to generate a pixel set of near-infrared image frames; S2: Obtain the pixel gray values of all image frame corresponding pixel sets, and establish a pixel gray distribution map; S3: Use a multi-branch convolutional neural network to perform spatial feature learning on the pixel gray distribution map, and separate the background, target structure, and dynamic afterimage area distributions; S4: Centering on the target structure area, perform spatial alignment on consecutive image frames, and stack all image frames to generate a global area distribution; S5: Obtain historical near-infrared image frame samples, train a neural network model through the training set constructed by the samples, and establish a regional distribution prediction model; S6: Based on the regional distribution prediction model, predict the background and afterimage area distributions of the real-time image, perform neighborhood repair on the afterimage area through the background area, and output the image with the afterimage corrected.
[0018] In S1, continuous near-infrared images of the production area are collected based on the running speed of the production line, and a pixel set of near-infrared image frames is generated.
[0019] Obtain real-time running speed information on the production line. The running speed of the production line is described by the length passed by the target area (such as a fixed pallet, a fixed device, a specific station on the material conveyor line, or an assembly area, etc.) per unit time, and the unit is meters per second or millimeters per second. The preset fixed total number of frames is default set to 100 frames. To ensure complete image acquisition cycles covering the target area, this value is determined before the start of the image acquisition process and stored in the parameter configuration, and does not change with the running speed. The setting of the fixed total number of frames supports adjustment according to actual needs. For example, in special process scenarios, it can be adjusted to 50 frames, 150 frames, or 300 frames. According to the length of the target area and the real-time running speed, calculate the passing time of the target area, and perform a numerical division of this passing time by the fixed total number of frames to obtain the frame rate value collected per second. This frame rate value is used to dynamically set the time interval during the image acquisition process to achieve dynamic synchronous acquisition of image frames. The frame rate calculation method is: divide the length of the target area (unit: meter) by the real-time running speed (meter per second) to obtain the passing time of the target area (second), then divide the passing time of the target area by the fixed total number of frames to calculate the time interval (second) for each frame acquisition, and finally take its reciprocal to obtain the acquisition frame rate (frames per second). For example, when the length of the target area is 2 meters and the real-time running speed is 0.2 meters per second, the passing time of the target area is 10 seconds. If the fixed total number of frames is set to 100 frames, the acquisition frame rate is 10 frames per second. This calculation method ensures a uniform frame acquisition time interval within the target area, so that the subsequent pixel gray value extraction steps have a high degree of time-axis consistency.
[0020] Collect and input the external light source parameters of the target production area. The external light source parameters include three indicators: light source intensity, color temperature, and irradiation angle. The set value of the default light source intensity is 500 lx, the set value of the default color temperature is 5000 K, and the set value of the default irradiation angle is 45°. The above values are configured before the image acquisition process starts and can be specifically adjusted according to different natural light effects. The specific value range can be referred to: light source intensity from 300 lx to 1000 lx, color temperature from 3500 K to 6000 K, and irradiation angle from 30° to 60°. Each image acquisition is carried out with reference to the external light source parameters. For the target production area, continuous near-infrared image frame acquisition is performed. During the acquisition process, a unique timestamp is marked for each frame of the image. The timestamp adopts the unified format of year-month-day hour-minute-second millisecond and is directly appended to the head of each frame of the image file or in the data structure to ensure the integrity and traceability of the acquisition order.
[0021] After the continuous image frame acquisition is completed, image frame homogenization processing is performed for the possible problem of uneven inter-frame time intervals. In the time axis index, the inter-frame interval is detected. When the inter-frame acquisition interval exceeds the preset threshold, for example, exceeds 1000 milliseconds (this threshold can be set according to requirements during the production line debugging stage and noted in the parameter document), the time axis equalization process is executed. The time axis equalization process method uses an interpolation algorithm to correct the time index between frames to make the sampling time distribution of the image frames uniform. The specific interpolation method adopts the previous frame image copying method, directly copying the previous frame image to the vacant time index position to ensure the continuity of the time axis and data integrity. After the time axis equalization process is completed, the pixels of each frame of the image are numbered and calibrated according to the row and column positions in the two-dimensional space. Each pixel has a unique row and column index. The pixel numbers are arranged in an increasing order row by row, indexed from left to right and from top to bottom, and a pixel set is constructed in combination with the row and column positions.
[0022] In S2, obtain the pixel gray values of all the pixel sets corresponding to the image frames and establish a pixel gray value distribution map.
[0023] After each frame of near-infrared image is acquired, the optical signals collected at each pixel row and column position within the image frame are directly subjected to electronic conversion. Using a standard optoelectronic conversion process, the received optical signals are converted into electrical signals, and then the corresponding pixel gray-scale signals are generated. Each pixel gray-scale signal is digitized using a quantizer, quantifying the analog electrical signal into a digital gray-scale value. The digital gray-scale value is in the format of an unsigned integer and can be represented by 8-bit, 10-bit, or 12-bit numbers. Usually, the 8-bit format is selected for near-infrared image processing, with a range of 0 - 255. After the quantization of the pixel gray-scale values, based on the row and column positions of the pixels, the digital gray-scale values are arranged in row and column order to construct a two-dimensional gray-scale matrix. The number of rows of this two-dimensional gray-scale matrix is consistent with the vertical resolution of the image frame, and the number of columns is consistent with the horizontal resolution, ensuring the integrity of the spatial correspondence relationship of the pixels. Each element in the two-dimensional gray-scale matrix contains its corresponding row and column indices and the digital gray-scale value. This two-dimensional gray-scale matrix is marked as a pixel gray-scale distribution map in the image processing flow, facilitating subsequent operations such as spatial feature extraction and convolution analysis. When the number of pixels is large (such as 1024×1024 or higher resolution), resulting in tight computing resources or storage resources, a matrix simplification strategy is adopted to simplify the pixel gray-scale matrix to reduce the computational burden. The matrix simplification process of this solution includes two methods: row and column downsampling and regional mean substitution. The specific steps are as follows: (1) Perform row and column downsampling according to a preset simplification factor. For example, the simplification factor can be set to 2, which means keeping one pixel every other pixel and removing adjacent pixels to form a downsampled pixel gray-scale matrix. If the simplification factor is set to 4, then one pixel is kept every 4 pixels, and the simplification ratio follows accordingly.
[0024] (2) In some scenarios, to retain the average gray-scale characteristics of the pixels, the regional mean substitution method can also be used. The adjacent pixels are divided into several small blocks (such as 4×4 or 8×8), and the pixel gray-scale mean of each small block is calculated. This mean is used to replace the original pixel values, simplifying the matrix size while ensuring the integrity of the overall gray-scale distribution characteristics. The above simplification factor and block size can be set during the process design stage. For example, the default value of the simplification factor is 2, and the default block size for regional mean substitution is 4×4, which can be adjusted according to the actual application.
[0025] After the construction of the pixel gray-scale distribution map is completed, a sequence of consecutive image frames is read, and each frame of the image is labeled with a frame number index. The numbering rule is to increment the number according to the image acquisition order. For example, the number of the first frame is 1, the number of the second frame is 2, and so on. Each pixel gray-scale distribution map is bound to the corresponding image frame number.
[0026] In S3, a multi-branch convolutional neural network is used to perform spatial feature learning on the pixel gray-scale distribution map, separating the background, target structure, and dynamic afterimage area distributions.
[0027] Perform a pixel-by-pixel scan operation on the pixel gray-scale distribution map formed after photoelectric conversion and digital processing. Pixel-by-pixel scanning means that, in the row-column order of the image frame, starting from the top-left pixel of the image, scan sequentially column by column horizontally to the right. After completing the scan of one row of pixels, continue to scan the next row of pixels downward until the entire image is scanned row by row and column by column. During the pixel-by-pixel scan process, the digital gray-scale value of each pixel is sequentially read and temporarily stored in an ordered array, keeping its row-column position in the two-dimensional plane unchanged. Then, the pixel gray-scale values in the ordered array are recombined in row-column order to construct a multi-dimensional input tensor. In this embodiment, the constructed multi-dimensional input tensor mainly includes three dimensions: the row dimension, the column dimension, and the channel dimension. Among them, the row dimension is consistent with the vertical resolution of the pixel gray-scale distribution map, the column dimension is consistent with the horizontal resolution, and the channel dimension is set to a single channel, that is, each pixel point corresponds to a digital gray-scale value.
[0028] The convolutional neural network structure adopts a multi-branch strategy, and transmits the input multi-dimensional input tensor to multiple branches simultaneously. Each branch corresponds to a target component, including a background branch, a target structure branch, and a dynamic ghost branch. The input format of each branch is retained as a single-channel format, that is, each input tensor only contains one layer of digital gray-scale value data to ensure the simplicity of the data structure and the efficiency of calculation. To ensure the focus of each branch on the target features, convolutional kernels of corresponding sizes are configured for each branch. The size of the convolutional kernel of the background branch is set to 5×5 pixels, the size of the convolutional kernel of the target structure branch is set to 3×3 pixels, and the size of the convolutional kernel of the dynamic ghost branch is set to 7×7 pixels. The selection of the convolutional kernel size is based on the spatial distribution characteristics of each target component: the background area usually has a large spatial distribution, so it is suitable to use a larger convolutional kernel; the target structure area is relatively compactly distributed, so a medium-sized convolutional kernel is used; the dynamic ghost area often has large-area disturbances, so using a larger convolutional kernel for feature extraction has more advantages. The above convolutional kernel sizes can be fine-tuned according to actual needs during the engineering debugging stage. For example, the background convolutional kernel size can be set between 3×3 and 9×9.
[0029] The constructed multi-dimensional input tensor is directly input into the multi-branch convolutional neural network. The input layer of the network maintains a single-channel format, and the grayscale value of each pixel corresponds to a separate channel dimension in the tensor, ensuring the correspondence with the image grayscale distribution. Each branch receives the input tensor and performs independent convolution operations. The convolution operation uses the configured convolution kernel for pixel-by-pixel calculation. Among them, the weights of each convolution kernel are determined through a large amount of data analysis and engineering optimization during the development and verification stages of the network. Specifically, in combination with historical image data, repeated convolution calculations and feature analyses are performed on the convolution kernels of each branch, and the weight distribution of the convolution kernels is determined through the statistical laws of grayscale distribution, enabling it to stably extract the features of their respective target components in different branches. In the background branch, target structure branch, and dynamic afterimage branch, the convolution operation performs sliding window convolution with their respective configured convolution kernel sizes, and the convolution step size is default set to 1 to ensure the complete extraction of pixel information. After the convolution operation is completed, the output result is used as the feature map of this branch. In the local convolution operation, the extracted pixel spatial features mainly include texture features, edge features, and continuous region features. The texture features are reflected by the local change patterns of pixel grayscale values, the edge features are expressed by the changes in pixel grayscale gradients, and the continuous region features are reflected by the similarity of grayscale values of adjacent pixels. Finally, through the local convolution operation of the multi-branch convolutional neural network, the feature maps of the background, target structure, and dynamic afterimage branches are constructed.
[0030] Perform cross-branch feature fusion on the feature maps of different branches, and separate the background, target structure, and dynamic afterimage area distributions through feature clustering. Sequentially read the feature maps of the background branch, target structure branch, and dynamic afterimage branch from the output results of the multi-branch convolutional neural network. The feature map of each branch is in the format of a two-dimensional matrix, maintaining the same row and column dimensions as the input multi-dimensional input tensor, and each pixel position contains the feature value output by the convolution kernel. The reading order of the feature maps strictly follows the defined order of the branch structure to avoid feature correspondence offsets caused by incorrect order. When reading, the pixel positions of each feature map correspond to the row and column positions of the original image frame to ensure the consistency of spatial positions. After the reading of the feature maps is completed, the feature values of the background branch, target structure branch, and dynamic afterimage branch at the same pixel position are numerically superimposed to form a joint feature in the form of a three-dimensional vector. The first dimension of this joint feature vector represents the background feature value, the second dimension represents the target structure feature value, and the third dimension represents the dynamic afterimage feature value. Through the construction of this joint feature, the complete fusion of features from different branches at the same pixel position is ensured. The joint features of all pixels are stored in the order of row and column positions, forming a pixel-level joint feature set for the entire image frame.
[0031] According to the characteristic distribution characteristics of the target area, a joint feature classification template for the background, target structure, and dynamic afterimage area is preset. This classification template is completed by manual sampling and annotation, that is, typical background areas, target structure areas, and dynamic afterimage areas are selected from historical samples, and the joint feature vectors of these areas are collected in sequence, and their numerical features are recorded. The joint feature template of each area is formed by the average value of the joint feature vectors of multiple samples, ensuring the representativeness and engineering repeatability of the template. On the basis of establishing the joint feature classification template, corresponding category indexes are established according to the area categories (background, target structure, dynamic afterimage). The category index is represented by an integer value. For example, the category index of the background area is 0, the category index of the target structure area is 1, and the category index of the dynamic afterimage area is 2. The category index corresponds to the joint feature template. The constructed pixel-level joint feature set performs unsupervised clustering analysis within the entire pixel gray level distribution map. The unsupervised clustering process uses the K-means clustering method to divide the pixel-level joint feature set into several category clusters. The joint feature vectors in each category cluster have high similarity, ensuring the local consistency of pixel features. The number of clusters in the clustering is preset to 3 in engineering, corresponding to the three target components of the background, target structure, and dynamic afterimage.
[0032] In each category cluster, a local joint feature set is extracted, and the average feature of all joint feature vectors in this set is calculated. The average feature is calculated by means of normalized averaging, that is, after each eigenvalue is converted to the range of 0-1 for expression, the pixels within the category cluster are summed and divided by the number of pixels to obtain each average feature value, and then an average feature vector is formed. The average feature vector of each category cluster is used to measure the distance from the preset joint feature classification template. The Euclidean distance or other distance measurement algorithms are used to calculate the numerical distance between the average feature vector and each template, and the template category with the smallest distance is selected as the category index of the category cluster. The category index obtained for each category cluster is mapped to the row and column positions of the pixel gray level distribution map. The mapping process strictly follows the pixel row and column order, and directly assigns the category index of each pixel to the corresponding pixel row and column position, ensuring the integrity and accuracy of the mapping relationship. The pixel position corresponding to each category index represents the area category to which the pixel belongs. The background area corresponds to the category index 0, the target structure area corresponds to the category index 1, and the dynamic afterimage area corresponds to the category index 2. Through this mapping, the pixel-level automatic annotation of the distribution of the background area, target structure area, and dynamic afterimage area of a single image frame is completed.
[0033] In S4, centered on the target structure area, spatial alignment is performed on consecutive image frames, and all image frames are superimposed to generate a global area distribution.
[0034] According to the distribution of the extracted target structure regions, determine the spatial positions of the target structure regions in each image frame. The spatial positions of the target structure regions are described by pixel row and column coordinates. Specifically, the center pixel point of the target structure region is used as the reference point. The row and column coordinates of the center pixel point can be obtained by traversing the pixel set of the target structure region and calculating the average values of its upper and lower boundaries and left and right boundaries, ensuring the repeatability and consistency of the spatial position of the center pixel point. After determining the center pixel point of the target structure region, use it as the alignment reference center of the pixel distribution map. For each frame of the image, calculate the offset of the center pixel point of its target structure region relative to the preset reference center (i.e., the row offset and column offset of the center point of the target structure region), and then perform an overall translation operation on the pixel distribution map of the entire image frame to align the center pixel point of the target structure region with the preset reference center. The translation operation is performed by adjusting the row and column index values of each pixel. The row offset and column offset are added to the original row coordinate and column coordinate of each pixel respectively to complete the pixel-level spatial alignment. For pixels that exceed the image boundary, in engineering, a boundary clipping strategy is adopted, that is, directly discard the pixel values that exceed the image boundary range, ensuring that the row and column dimensions of the pixel distribution map after spatial alignment are consistent with the initial input and avoiding data dimension mismatch. In addition, to prevent image area loss or black edges caused by spatial translation, a boundary compensation mechanism can be set for the pixel distribution map in engineering, such as filling with the gray values of adjacent pixels, thus ensuring the consistency and integrity of the image frame.
[0035] For each region category, according to the corresponding relationship of pixel row and column positions, statistically summarize the pixel category indices at the same row and column positions in all image frames to generate the region category probability distribution at this row and column position under cross-frame superposition. Specifically, for the background region, the target structure region, and the dynamic afterimage region, establish superposition counters for pixel row and column positions respectively. Add one to the count of the pixel positions of the corresponding category in each frame of the image. After accumulating and counting all frames, obtain the region category superposition distribution at each pixel row and column position. Through this statistical summarization process, cross-frame pixel region distribution analysis is realized. After completing the superposition statistics of pixel region categories, for the pixel category distribution at the same row and column position, select the category index with the largest cumulative count as the final region category at this row and column position, avoiding the problem of category conflicts in cross-frame region distribution. Finally, based on the superposition results of all frames, output the distribution maps of the global background region, the global target structure region, and the global dynamic afterimage region. Each distribution map is represented in the form of a two-dimensional matrix. The row and column dimensions of the matrix are consistent with the row and column dimensions of the input image frame. Each matrix element contains the final category index information, ensuring the integrity of the spatial correspondence of the region distribution and strict matching with the row and column positions of each image frame.
[0036] In S5, historical near-infrared image frame samples are obtained, and a neural network model is trained using the training set constructed from the samples to establish a regional distribution prediction model.
[0037] Collect historical near-infrared image frame sample data. The sample data should cover different external light source parameters and production line running speeds to comprehensively reflect the image acquisition characteristics under the actual operating conditions of the production line. The external light source parameters include light source intensity, color temperature, and irradiation angle. The specific value ranges can be referred to as follows: light source intensity from 300 lx to 1000 lx, color temperature from 3500 K to 6000 K, and irradiation angle from 30° to 60°. The production line running speed range covers from 0.2 m per second to 1.5 m per second, ensuring the diversity and representativeness of the sample data. Each group of sample data is managed by numbering. Each group of data contains several consecutive image frame samples and their corresponding external light source parameters and production line running speed information, facilitating subsequent data grouping and label annotation. In each group of sample data, the pixel gray-scale distribution maps of all historical near-infrared image frames are read and processed frame by frame. Each frame of the image has pixel gray-scale distribution information and frame number indexes. For the image frames in each group of sample data, according to the extracted background region, target structure region, and dynamic afterimage region distributions, cross-frame alignment of pixel row and column positions is performed, and combined with the spatial superposition statistical method, the region category indexes at the same row and column positions are summarized to determine the spatial distribution of the global background region, global target structure region, and global dynamic afterimage region in each group of sample data.
[0038] For the generated global region distribution results, label marking processing is performed on the sample data. The label marking process takes pixel-level labels as units, and each pixel point in the sample data is matched with the category index in the global region distribution according to its row and column coordinates. After the label marking processing is completed, the marked data is bound to the original pixel gray-scale distribution data to construct the training set data of the regional distribution prediction model. The training set data includes pixel gray-scale distribution features and corresponding category index labels, with high engineering operability and repeatability.
[0039] The neural network model adopts a convolutional neural network structure and consists of multiple convolutional layers, activation function layers, and an output layer. The convolutional layers are responsible for extracting the spatial features of the input data. The activation function layers use the ReLU function for the ability to extract non-linear features. The output layer uses the Softmax function to convert the model output into the probability distribution of each category. During the model training process, the cross-entropy loss function is used as the objective function of the model to calculate the difference between the model prediction result and the actual label. During the training process, the mini-batch stochastic gradient descent optimization algorithm is used to iteratively update the model weights. The initial value of the learning rate is set to 0.001, which can be dynamically adjusted according to the loss change during the training process to ensure the convergence speed and training stability of the model. In each round of training, the training set data is randomly shuffled and input into the model in batches for forward propagation and backward propagation. During the forward propagation process, the input data passes through the convolutional layer, activation function layer, and output layer in sequence, and outputs the probability distribution of the regional categories. During the backward propagation process, the gradient information is calculated according to the cross-entropy loss function, and the convolutional kernel weights and the bias parameters of each layer of the neural network are updated to make the model prediction result gradually approach the true label distribution. The entire training process sets multiple training epochs. Each epoch traverses the entire training set data once. Usually, the number of training epochs is set to 100 epochs or more, and the specific number of epochs can be dynamically adjusted according to the convergence curve of the model. When the validation set loss value of the model tends to be stable and reaches the expected convergence accuracy, the training process ends, and the trained model parameters are saved to form a regional distribution prediction model.
[0040] After the training is completed, the model is deployed to the production line side and integrated into the real-time image processing process to ensure that during the actual acquisition process, the real-time prediction of the regional categories of the pixel grayscale distribution map of each frame of image can be performed.
[0041] In S6, based on the regional distribution prediction model, the background and afterimage region distributions of the real-time image are predicted, and the afterimage region is patched in the neighborhood through the background region, and the image with the afterimage corrected is output.
[0042] The obtained external light source parameters and the real-time production line running speed are used as input features and input into the regional distribution prediction model. The regional distribution prediction model adopts the trained neural network model and has the functions of input feature processing and regional category output. In the model inference stage, the external light source parameters and the production line running speed are passed into the model through the input interface, and at the same time, the pixel grayscale distribution map of the current image frame is input. The category index matrix output by the model is bound to the input pixel grayscale distribution map to form a pixel position annotation map of the background region and the dynamic afterimage region.
[0043] At the pixel position coordinates in the afterimage area, neighborhood repair processing is performed on the pixel gray values of the pixel gray value distribution map. The repair processing uses the pixel gray values in the background area as a reference. By establishing a pixel gray value statistical model (local weighted average) in the background area, the average gray value and standard deviation of the background area are calculated, and combined with the pixel positions in the afterimage area, the weighted interpolation method is used to correct the pixel gray values in the afterimage area. The weighted interpolation method uses the neighboring pixels in the background area as a reference, and calculates the corrected gray value of the pixels in the afterimage area according to the distance weight function. Or directly replace the pixel gray values in the afterimage area with the average gray value of the background area. This method is enabled as a simple correction method when the system operation amount is too large. After the repair is completed, the corrected pixel gray values are used to replace the pixel values in the original afterimage area, and the pixel values in other areas remain unchanged, ensuring the integrity of the gray structure of the entire image. The output image frame is saved in the form of a two-dimensional gray matrix, containing the corrected gray values at all pixel positions, and the row and column indexes are consistent with the original input image, avoiding image alignment errors and regional offsets.
[0044] Through the neighborhood repair processing, the dynamic correction of the afterimage area is completed, ensuring the consistency and integrity of the corrected image frame in visual and engineering applications, and providing reliable data support for subsequent image analysis and product detection. The finally output image frame after afterimage correction has high image quality and engineering applicability, meeting the engineering requirements of dynamic afterimage correction.
[0045] Embodiment 2. The difference between Embodiment 2 and Embodiment 1 of the present invention is that this embodiment introduces an image processing system based on near-infrared imaging.
[0046] Figure 2 The structural schematic diagram of an image processing system based on near-infrared imaging according to the present invention is given. An image processing system based on near-infrared imaging includes a gray value construction module, a convolutional feature extraction module, a global distribution construction module, a prediction model training module, and a dynamic afterimage correction module: Gray value construction module: Collect continuous near-infrared image frame samples during the operation of the production line, dynamically adjust the image frame rate in combination with the real-time production line speed and external light source parameters, perform electronic conversion and quantization processing on the pixel optical signals of the image frame, generate digital gray values and mark the row and column positions, and construct a pixel gray value distribution map; Convolutional feature extraction module: Use a multi-branch convolutional neural network to perform spatial feature learning on the pixel gray value distribution map, configure convolutional kernels according to the background, target structure, and dynamic afterimage branches to extract spatial features, generate branch feature maps, and output the pixel area distributions of the background, target structure, and dynamic afterimage through feature fusion and unsupervised clustering to assign category indexes; Global distribution construction module: Using the target structure area as the alignment reference, spatially aligning consecutive image frames, and based on the corresponding relationship of pixel row and column positions, overlaying the background, target structure, and dynamic afterimage area distributions across frames to generate the global background area, global target structure area, and global dynamic afterimage area distributions; Prediction model training module: Obtaining historical near-infrared image frame sample data and the corresponding global area distribution maps, performing label marking on the sample data to construct a training set, selecting a neural network model to input the training set for training, and after training is completed, generating an area distribution prediction model and deploying it to the production line end; Dynamic afterimage correction module: Inputting the external light source parameters at the production line end and the real-time production line running speed into the area distribution prediction model, obtaining the background and dynamic afterimage area distributions, performing neighborhood repair on the afterimage area, and using the pixel gray values of the background area to correct the afterimage area, and outputting the corrected image frame.
[0047] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0048] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0049] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0050] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0051] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in an electrical, mechanical, or other form.
[0052] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0053] In addition, the functional modules in each embodiment of this application can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0054] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0055] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0056] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An image processing method based on near-infrared imaging, characterized in that, It includes the following steps: S1: Continuously collect near-infrared images of the production area based on the running speed of the production line to generate a pixel set of near-infrared image frames; S2: Obtain the pixel gray values of the pixel sets corresponding to all image frames and establish a pixel gray distribution map; S3: Use a multi-branch convolutional neural network to perform spatial feature learning on the pixel gray distribution map to separate the background, target structure, and dynamic afterimage area distributions; S4: Centering on the target structure area, perform spatial alignment on consecutive image frames, and stack all image frames to generate a global area distribution; S5: Obtain historical near-infrared image frame samples, train a neural network model through the training set constructed by the samples, and establish a regional distribution prediction model; S6: Predict the background and afterimage area distributions of real-time images based on the regional distribution prediction model, perform neighborhood repair on the afterimage area through the background area, and output the image with corrected afterimages.
2. The image processing method based on near-infrared imaging according to claim 1, wherein, In S1, continuously collecting near-infrared images of the production area based on the running speed of the production line to generate a pixel set of near-infrared image frames specifically includes: Obtain the real-time running speed of the production line and dynamically adjust the near-infrared image acquisition frame rate based on a preset fixed total number of frames; Obtain external light source parameters, continuously collect near-infrared image frames of the target production area on the production line, record the acquisition timestamp of each frame of image, and at the same time perform time-axis uniform-axis processing on the image frames using the interpolation method; Number the pixels and calibrate the row and column positions of each image frame corresponding to each time sequence index position on the time axis to generate a pixel set for each image frame.
3. The image processing method based on near-infrared imaging according to claim 1, wherein, In S2, obtaining the pixel gray values of the pixel sets corresponding to all image frames and establishing a pixel gray distribution map specifically includes: In the pixel set corresponding to the image frame, perform electronic conversion on the optical signals collected at each pixel row and column position to generate corresponding pixel gray signals, and quantize the pixel gray signals into digital gray values; Convert the digital gray values into a two-dimensional gray matrix based on the arrangement of the pixel row and column positions, and mark the two-dimensional gray matrix as the pixel gray distribution map; Read the continuous image frame sequence and perform image frame number indexing annotation to establish the corresponding relationship between the pixel gray distribution map and its corresponding image frame number.
4. The image processing method based on near-infrared imaging according to claim 1, characterized in that In S3, using a multi-branch convolutional neural network to perform spatial feature learning on the pixel gray distribution map to separate the background, target structure, and dynamic afterimage area distributions specifically includes: Perform pixel-by-pixel scanning on the pixel gray distribution map, and arrange the pixel gray values in row and column order to form a multi-dimensional input tensor; Set the multi-branch convolutional neural network structure, retain the single-channel input format in the multi-dimensional input tensor, and configure convolutional kernels of corresponding sizes for different branches respectively; Among them, the multi-branch includes a background, target structure, and dynamic afterimage branch; Input the multi-dimensional input tensor into the multi-branch convolutional neural network, and perform local convolution operations within each branch based on the configured convolutional kernels to extract pixel spatial features and construct feature maps of different branches, where the spatial features include texture features, edge features, and continuous region features; Perform cross-branch feature fusion on the feature maps of different branches, and separate the background, target structure, and dynamic afterimage area distributions through feature clustering.
5. A method for processing images based on near-infrared imaging according to claim 4, characterized in that, The cross-branch feature fusion of the feature maps of different branches, and the separation of the background, target structure, and dynamic afterimage region distributions through feature clustering specifically includes: In the output of the multi-branch convolutional neural network, sequentially read the feature maps of the background branch, target structure branch, and dynamic afterimage branch; Overlay the three types of feature maps, and construct a joint feature based on the background feature value, target structure feature value, and dynamic afterimage feature value at each pixel position to form a pixel-level joint feature set; Preset the joint feature classification templates for the background, target structure, and dynamic afterimage regions, and establish corresponding category indexes respectively; Perform unsupervised clustering on the joint feature set within the entire pixel grayscale distribution map, extract the local joint feature set of each category cluster, calculate the average feature of the local joint feature set, and label the category index of each clustering cluster through the distance between the average feature and the joint feature classification template; Map the labeled category indexes to the row and column positions of the pixel grayscale distribution map to generate the background region, target structure region, and dynamic afterimage region distributions corresponding to a single image frame.
6. The image processing method based on near-infrared imaging according to claim 1, wherein In S4, with the target structure region as the center, perform spatial alignment on consecutive image frames, and overlay all image frames to generate the global region distribution, which specifically includes: Use the target structure region as the central position of the pixel distribution map to perform spatial alignment on the pixel distribution maps corresponding to all image frames; Based on the corresponding relationship of the pixel row and column positions of the background region, target structure region, and dynamic afterimage region distributions after spatial alignment, perform cross-frame overlay on all image frames to generate the global background region, global target structure region, and global dynamic afterimage region distributions.
7. A method for processing an image based on near-infrared imaging according to claim 1, wherein In S5, obtain historical near-infrared image frame samples, and train the neural network model through the training set constructed by the samples to establish a region distribution prediction model, which specifically includes: Obtain the historical near-infrared image frame sample data, and each group of sample data covers different external light source parameters and production line running speeds; In each group of sample data, determine the global background region, global target structure region, and global dynamic afterimage region distributions for the pixel grayscale distribution maps of all historical near-infrared image frames; Perform label marking processing on the sample data based on the determination results of the global background region, global target structure region, and global dynamic afterimage region distributions to construct the training set data of the region distribution prediction model; Select a neural network model, input the training set data to train the region distribution prediction model, and deploy the model to the production line end after the model training is completed.
8. The image processing method based on near-infrared imaging according to claim 1, characterized in that In S6, based on the region distribution prediction model, predict the background and afterimage region distributions of the real-time image, and perform neighborhood repair on the afterimage region through the background region, and output the image with the afterimage corrected, which specifically includes: Input the external light source parameters and real-time production line running speed at the production line end into the region distribution prediction model; Obtain the model output result, and label the background region and dynamic afterimage region distributions for the pixel grayscale distribution map of the real-time acquired near-infrared image frame; Extract the pixel position coordinates of the labeled afterimage region and background region in the pixel grayscale distribution map; At the pixel position coordinates in the afterimage area, perform neighborhood patching on the pixel gray values of the pixel gray value distribution map, perform pixel gray correction on the afterimage area according to the pixel gray values in the background area, and output the image frame after afterimage correction.
9. An image processing system based on near-infrared imaging, which is used to implement the image processing method based on near-infrared imaging according to any one of claims 1-8, characterized in that, It includes a gray value construction module, a convolutional feature extraction module, a global distribution construction module, a prediction model training module, and a dynamic afterimage correction module: Gray value construction module: Collect continuous near-infrared image frame samples during the operation of the production line, dynamically adjust the image frame rate in combination with the real-time production line speed and external light source parameters, and perform electronic conversion and quantization processing on the pixel optical signals of the image frames to generate digital gray values and mark the row and column positions, and construct a pixel gray value distribution map; Convolutional feature extraction module: Use a multi-branch convolutional neural network to perform spatial feature learning on the pixel gray value distribution map, configure convolutional kernels according to the background, target structure, and dynamic afterimage branches to extract spatial features, generate branch feature maps, and assign class indexes through feature fusion and unsupervised clustering, and output the pixel area distributions of the background, target structure, and dynamic afterimage; Global distribution construction module: Use the target structure area as the alignment reference to perform spatial alignment on consecutive image frames, and based on the corresponding relationship of pixel row and column positions, stack the background, target structure, and dynamic afterimage area distributions across frames to generate the global background area, global target structure area, and global dynamic afterimage area distributions; Prediction model training module: Obtain the historical near-infrared image frame sample data and the corresponding global area distribution map, perform label marking on the sample data to construct a training set, select a neural network model to input the training set for training, and generate a regional distribution prediction model after training and deploy it to the production line end; Dynamic afterimage correction module: Input the external light source parameters and real-time production line operation speed at the production line end into the regional distribution prediction model, obtain the background and dynamic afterimage area distributions, perform neighborhood patching on the afterimage area, and use the pixel gray values in the background area to correct the afterimage area, and output the corrected image frame.
Citation Information
Patent Citations
Forest smoke and fire detection method based on video image analysis
CN106650600A
Semi-supervised image deblurring method based on fusion attention mechanism
CN113592736A
Infrared image deblurring algorithm based on attention mechanism residual network model
CN115345791A
Infrared fuzzy target identification method based on hybrid neural network model
CN116503745A
Infrared imaging non-uniform noise extraction and optimization method based on deep learning
CN117528274A
Cited By
Complex micro-nano structure parallel processing method based on image segmentation and area planning
CN120828201A