A method and system for online detection of non-smoking substances based on area array spectral imaging

Through the combination of surface array spectral imaging technology and deep learning model, the problems of low efficiency and poor accuracy of non-smoking substance detection in the prior art are solved, and efficient and accurate online detection and monitoring of non-smoking substances in tobacco leaves are achieved.

CN118657749BActive Publication Date: 2025-05-16CHANGSHA LUSONG INTELLIGENT INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410903714.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2025-05-16
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency, serious errors and poor adaptability in the detection of non-toxin substances in tobacco leaves. Especially when facing complex backgrounds and diverse non-toxin substances, it is difficult to achieve efficient and accurate detection.

Method used

The non-smoking substance online detection method based on surface matrix spectral imaging is adopted. Image features are extracted in multiple levels through encoder-decoder network MSFSNet and point-by-point convolution feature extraction network branches and SIPNet multi-layer perceptron network branches to achieve more accurate semantic information representation.

Benefits of technology

It significantly improves the accuracy and reliability of non-smoking substance detection, realizes real-time monitoring and early warning, reduces manual inspection costs, and ensures the quality and safety of tobacco leaf production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118657749B_ABST
    Figure CN118657749B_ABST
Patent Text Reader

Abstract

The present invention discloses an online detection method and system for non-smoking substances based on array spectral imaging, which relates to image segmentation technology for industrial detection in industrial artificial intelligence technology. The method comprises the following steps: hardware construction of the imaging system, selecting a 7-channel spectral camera and a 3-channel visible light camera to construct the imaging hardware; camera calibration, geometric calibration of the camera to achieve image alignment; data acquisition and annotation, image acquisition and image annotation of the tobacco leaf impurity removal pipeline to construct a data set; construction of a fusion feature extraction network model, design of the encoder-decoder network MSFSNet, use of 10-channel image data as network input, and addition of point-by-point convolution feature extraction branches and SIPNet multi-layer perceptrons. The present invention is based on array spectral imaging and image semantic segmentation technology, combined with spectral data and visible light data for analysis and processing, making full use of the target spectral characteristics, and significantly improving the detection accuracy while ensuring real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial-grade image segmentation based on deep learning, and in particular to an online detection method and system for non-smoking substances based on planar spectral imaging. Background Art

[0002] In the current international tobacco industry, although tobacco leaf threshing and redrying technology and equipment have been very mature and widely used, in-depth research and development of these technologies has not stopped. Research in recent years has achieved remarkable new results, promoting the further improvement of leaf threshing and redrying technology and equipment, and reaching a new technical level.

[0003] As global consumers pursue a higher and higher standard of healthy living, the quality standards of tobacco leaves have become increasingly stringent. Therefore, in the tobacco industry, the adoption of advanced tobacco production management methods and the continuous improvement of tobacco leaf quality and safety have become important issues that the tobacco industry needs to solve urgently. In particular, the control of non-tobacco substances in tobacco leaves is an important part of GAP management, and is an important task to ensure the quality of tobacco leaves, improve the use value of tobacco leaves and the credibility of goods. It is reported that in the world tobacco trade, more than 100 million US dollars worth of tobacco leaves are rejected by cigarette manufacturers every year because they contain non-tobacco substances. The international tobacco market requires zero non-tobacco substances in tobacco leaves.

[0004] The existing manual impurity removal methods have many problems such as low efficiency, high labor costs, and serious subjective judgment errors, which often lead to non-tobacco substances remaining in tobacco leaves, affecting the overall quality of tobacco leaves. In addition, the manual impurity removal process is also easily affected by factors such as operator fatigue and distraction, further aggravating the omission of non-tobacco substances.

[0005] Although some image segmentation methods based on visible light have been applied in actual industrial production, these methods often cause false detection or missed detection due to unclear surface characteristics of objects or insufficient segmentation feature information. In particular, when faced with non-smoking substances with complex backgrounds, different types and shapes, existing technologies often find it difficult to achieve efficient and accurate detection. In addition, traditional image segmentation technology also has problems such as sensitivity to lighting conditions and poor adaptability. When the lighting conditions in the production environment change, the detection accuracy of the system is often greatly affected.

[0006] Therefore, there is an urgent need to develop an efficient and high-quality online detection method and system for non-tobacco substances, which can monitor non-tobacco substances at key nodes of the production line in real time, provide real-time early warning and statistics on important impurities, and guide tobacco production to further improve quality and efficiency. Summary of the invention

[0007] The purpose of the present invention is to provide a non-smoking substance online detection method and system based on array spectral imaging, which extracts image features at multiple levels through the encoder-decoder network MSFSNet and the training point-by-point convolution feature extraction network branch and the SIPNet multi-layer perceptron network branch, and more accurately represents the semantic information of the image. The entire process is based on the image segmentation network framework, and the use of multispectral images for semantic segmentation still ensures the accuracy and speed of image segmentation, reduces the cost of manual detection, and helps the tobacco industry concept to flourish.

[0008] To achieve the above object, the technical solution adopted by the present invention is:

[0009] A method for online detection of non-smoking substances based on area array spectral imaging, the method comprising the following steps:

[0010] Step 1: Imaging system hardware construction: In order to make full use of different material compositions and imaging features to achieve identification and classification, a 7-channel spectral camera and a 3-channel visible light camera are used to build the imaging hardware;

[0011] Step 2: Camera calibration: Since the eight lenses are distributed differently in space, the center coordinates of the obtained imaging information cannot be mapped to the same world coordinates. Therefore, geometric calibration of the camera is required to ensure image position alignment.

[0012] Step 3: Data collection and annotation: Continuous image capture is performed on the tobacco leaf impurity removal production line in the real industrial scene to collect data, and the collected data is annotated at the pixel level to form a data set.

[0013] Step 4: Build a fusion feature extraction network model: Design the encoder-decoder network MSFSNet, use 10-channel image data as the network input, and add point-by-point convolutional feature extraction branches and SIPNet multi-layer perceptron.

[0014] As a preferred improvement of the present invention, the specific steps of step 2 are:

[0015] Collect image data: Use a multispectral camera to capture a checkerboard image and define the camera’s intrinsic parameter matrix;

[0016] Corner point detection and coordinate extraction: Corner point detection is achieved through image processing algorithms to determine the pixel coordinates of each checkerboard corner point. The optimal external parameter matrix is ​​calculated through the least squares method to minimize the difference between the actual pixel coordinates and the pixel coordinates calculated by the external parameter matrix;

[0017] Calculate the homography matrix and camera parameters: construct the homography matrix through the obtained intrinsic and extrinsic parameters, and perform SVD decomposition on the homography matrix to further obtain the accurate rotation matrix R and translation vector t;

[0018] Image alignment and system verification: Use the calculated homography matrix to align all images taken by the camera, and perform final evaluation and optimization of the system by comparing the alignment of overlapping areas in different camera views.

[0019] As a preferred improvement of the present invention, in step 3, the data set includes: a training set, a validation set, and a test set; the training set is used to update the model parameters during the model training process; the validation set is used to test the model performance after each round of training; the test set is used to test the parameters of the model that obtains the best results on the validation set and obtain the model evaluation results

[0020] As a preferred improvement of the present invention, in step three, the training set, the test set and the validation set are divided according to a quantity ratio of 6:2:2.

[0021] As a preferred improvement of the present invention, in step three, a non-tobacco substance dataset is constructed, and pixel-level image annotation is performed on the non-tobacco substance dataset.

[0022] As a preferred improvement of the present invention, the specific steps of step three are:

[0023] Image acquisition: Use a multispectral camera to photograph the material on the conveyor belt and ensure that the images obtained after camera calibration are aligned in the world coordinates, that is, the same pixel coordinate position in each image represents the same point in the world coordinates;

[0024] Fully supervised annotation: perform full pixel-level annotation on the collected image data, and annotate each pixel in the image. By annotating the image intensity, if the substance at the (x, y) pixel is labeled n in the sample library, then the pixel intensity of the pixel is labeled n;

[0025] Export annotation results: export the annotated image data for subsequent processing;

[0026] Dataset division: Divide the labeled dataset into training set, test set, and validation set to support the model training, testing, and validation process.

[0027] As a preferred improvement of the present invention, step 4 specifically includes the following sub-steps:

[0028] Step 4.1: Modify the number of input channels of the image segmentation network to 10, including 7 spectral channels and 3 visible light channels;

[0029] Step 4.2: Random cropping: During model training, the 10-channel data is randomly cropped to 640 pixels × 640 pixels.

[0030] Step 4.3: Train the encoder-decoder network MSFSNet, use cross entropy as the loss function, and obtain 11 category segmentation results;

[0031] Step 4.4: Train the point-by-point convolution feature extraction branch, use cross entropy as the loss function, and obtain 11 category segmentation results;

[0032] Step 4.5: Train the SIPNet multi-layer perceptron. After the input image is expanded into a one-dimensional vector after 6 convolution blocks, the light intensity feature is extracted by the multi-layer perceptron. Then, 6 convolution blocks are performed to calculate the output feature map, and 11 category segmentation results are obtained.

[0033] Step 4.6: Feature fusion and fusion training, load the pre-trained weights of 4.3, 4.4, and 4.5 models, perform feature fusion and feature extraction, use cross entropy as the loss function, obtain 11 category segmentation results, and train to obtain the final model.

[0034] As a preferred improvement of the present invention, the 10 channels described in step 4.1 are spectral channels of 450nm, 550nm, 650nm, 720nm, 750nm, 800nm, 850nm and three visible light channels.

[0035] As a preferred improvement of the present invention, the feature fusion and fusion training described in step 4.6 includes the following steps:

[0036] Load the pre-trained weights of the encoder-decoder network MSFSNet, the point-by-point convolutional feature extraction network, and the SIPNet multi-layer perceptron network;

[0037] The fused feature map is convolved to finally obtain a feature map containing segmentation results in 11 categories.

[0038] More specifically, in step 4, if Figure 2 As shown in the figure, building the model specifically includes the following steps:

[0039] Step 4.1. Scale the input image if any value of its length and width is less than 256, and change the number of input channels to 10, which are 7 spectral channels (450nm, 550nm, 650nm, 720nm, 750nm, 800nm, 850nm) and three visible light channels.

[0040] Step 4.2, random cropping: For image scaling, since the aligned image size will change after each camera calibration, it cannot be guaranteed that the aspect ratio of each input network is the same. In order to retain the shape characteristics of the material, the 10-channel data is randomly cropped to 640 pixels × 640 pixels during model training.

[0041] Step 4.3, train the designed encoder-decoder network MSFSNet. In this training phase, the input image is continuously convolved and downsampled, and then the feature map is deconvolved and upsampled until the feature map size is restored to the same size as the input image. The feature map of the input size is then convolved to obtain the final feature map. The result of the final feature map is the segmentation result of the 11 categories in the image. The cross entropy is used as the loss function during the training process, and the weights obtained during the training process are saved for loading the pre-trained model in step 4.6.

[0042] Step 4.4, train the point-by-point convolution feature extraction branch. In this stage, 10 layers of 1x1 convolution feature extraction are performed on the input image. The process of convolution feature extraction is also the process of increasing the feature map channels. The ReLU activation function is used in the multi-layer 1x1 convolution feature extraction process. The result after training is the feature map, and the result of the feature map is the segmentation result of 11 categories. The cross entropy is used as the loss function during the training process, and the weights obtained during the training process are saved for loading the pre-trained model in step 4.6.

[0043] Step 4.5, train the SIPNet multi-layer perceptron. In this stage, the input image is expanded into a one-dimensional vector after 6 convolution blocks, and then the features are processed by building a multi-layer perceptron, which can make full use of the feature information of spectral imaging. Perform another 6 convolution blocks to restore the output of the last layer of the multi-layer perceptron to a feature map, and continue to perform convolution calculations to obtain the segmentation results of 11 categories. The cross entropy is used as the loss function during the training process, and the weights obtained during the training process are saved for loading and use by the pre-trained model in step 4.6.

[0044] Step 4.6, feature fusion and fusion training. According to 4.3, 4.4, and 4.5, three models for multispectral image segmentation have been obtained, namely, the design of the encoder-decoder network MSFSNet, the training of the point-by-point convolution feature extraction network, and the SIPNet multi-layer perceptron network. In this step, the pre-trained weights of the three networks are first loaded before training, and then the features of the three networks are fused, and the fused feature map is convoluted, and finally a feature map containing segmentation results in 11 categories is obtained. During the training process, cross entropy is used as the loss function.

[0045] The present invention also provides an online detection system for non-tobacco substances in leaf threshing and redrying based on area array spectral imaging, which is applied to any of the above detection methods, and comprises:

[0046] Imaging system hardware: It consists of a 7-channel spectral camera and a 3-channel visible light camera, which are used to obtain spectral images and visible light images of substances on the tobacco leaf impurity removal line;

[0047] Processor: configured with at least one program. When the program is executed by the processor, the processor controls and coordinates the functions of the following modules:

[0048] Image processing unit: connected to the imaging system hardware, used to perform data acquisition, camera calibration, image annotation and image preprocessing tasks;

[0049] Feature extraction network module: connected to the image processing unit, including the fusion feature extraction network model, used to perform feature extraction and image segmentation;

[0050] Training and testing unit: connected to the feature extraction network module, used to train the encoder-decoder network MSFSNet, the point-by-point convolutional feature extraction branch and the SIPNet multi-layer perceptron, as well as perform model testing and verification;

[0051] Feature fusion unit: connected to the training and testing unit, used to perform fusion training and testing of feature maps to output the final image segmentation result;

[0052] Control unit: connected to all the above units, used to coordinate the operation of the entire system, including data flow management, model training supervision, and output of detection results.

[0053] The present invention has the following beneficial effects:

[0054] 1. Improve detection accuracy and reliability: This invention combines deep learning and spectral imaging technology to significantly improve the detection accuracy and reliability of non-tobacco substances. The deep learning model can learn complex features from spectral data, accurately distinguish different material components, reduce false detection and missed detection, and ensure the quality of tobacco production.

[0055] 2. Real-time monitoring and early warning: The present invention realizes real-time monitoring on the production line, instantly detects and warns of non-tobacco substances in tobacco leaves, speeds up production decision-making, improves the efficiency of handling abnormal situations, and ensures the quality and safety of tobacco leaves.

[0056] 3. Non-destructive, non-contact, fast and efficient detection: The present invention uses visible light and spectral information for image acquisition, and realizes non-destructive, non-contact, fast and efficient detection through the spectral characteristics of different parts or components of the target object. It improves the detection accuracy of industrial scene image segmentation, improves the efficiency of target segmentation, and effectively overcomes the problems of missed detection and false detection when the target is occluded and the feature information is not significant. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a basic flow chart of the method for online detection of non-smoking substances based on area array spectral imaging of the present invention.

[0058] Figure 2It is a schematic diagram of a model of the online detection method of non-smoking substances based on area array spectral imaging of the present invention.

[0059] Figure 3 This is an application effect diagram of the non-smoking substance online detection method based on area array spectral imaging of the present invention. DETAILED DESCRIPTION

[0060] The technical solutions in the embodiments of the present invention will be described clearly and completely below in combination with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] See also Figure 1 As shown, the present invention provides a method for online detection of non-smoking substances based on planar spectral imaging, the method comprising the following steps:

[0062] Step 1: Build the imaging system hardware: In order to make full use of different material compositions and imaging features to achieve identification and classification, a 7-channel spectral camera and a 3-channel visible light camera are selected to build the imaging hardware.

[0063] Specifically, the seven spectral bands selected are 450nm, 550nm, 650nm, 720nm, 750nm, 800nm, and 850nm. The image sensor uses a 1.55 million pixel 1 / 2.9-inch CMOS global shutter sensor, the optical lens focal length f=5.0mm, the aperture F / 2.0, the imaging resolution of the visible light camera and the spectral camera are both 1440*1080, and the acquisition frame rate is 20 frames / second. All lenses are installed on the same plane, and pulse signals are sent during data acquisition to ensure that all collected images have the same timestamp.

[0064] Step 2: Camera calibration: In the present invention, since the eight lenses selected have different spatial distribution positions, the center coordinates of the obtained imaging information cannot be mapped to the same world coordinates. Therefore, geometric calibration of the camera is required to ensure image position alignment.

[0065] The calibration steps are as follows:

[0066] 1. Collect image data

[0067] First, use a multispectral camera with seven spectral lenses and one visible light lens to capture a checkerboard image and define the camera's intrinsic parameter matrix:

[0068]

[0069] Among them, α and β are the focal lengths on the image plane, γ is the non-orthogonality factor of the image plane, and u0 and v0 are the coordinates of the optical center. The calculation of the camera's intrinsic parameter matrix is ​​achieved through the following relationship:

[0070] m′=K -1 m

[0071] where m = [u, v, 1] T is the pixel coordinate in the image, and m' is the ideal pixel coordinate after correction. This step provides the basis for subsequent corner point detection and coordinate extraction.

[0072] 2. Corner point detection and coordinate extraction

[0073] Corner point detection is achieved through image processing algorithms. First, the original image is converted into a grayscale image and smoothed by Gaussian filtering; then the gradient of the image is calculated by applying an edge detection operator. The gradient information is used to construct the structure tensor of each pixel, reflecting the directional change of the image at that point, and then the corner points in the image are identified through processing; after determining the candidate positions of the corner points, these positions are finely adjusted through non-maximum suppression technology to ensure that only the true local maximum points are marked as corner points.

[0074] After that, determine the pixel coordinates of each chessboard corner point m = [u, v, 1] T The mapping relationship between the pixel coordinates of the corner points and the world coordinates is described by the following model:

[0075]

[0076] Here, [r1r2t] is the camera's external parameter, which maps the world coordinates (X, Y, 0) of the chessboard to the camera's image coordinate system. The optimal external parameter matrix is ​​calculated by the least squares method to minimize the difference between the actual pixel coordinates and the pixel coordinates calculated by the external parameter matrix. The objective function to be minimized is:

[0077]

[0078] Where X i is the world coordinate of the i-th corner point.

[0079] 3. Calculate the homography matrix and camera parameters

[0080] Through the internal and external parameters obtained above, the homography matrix H is constructed, and its matrix form is:

[0081] H=λK[r1 r2 t]

[0082] Here λ is a normalization factor that can be solved by additional constraints, such as ensuring that the determinant of the H matrix is ​​1. This step of the calculation also includes RANSAC or other robust algorithms on the H matrix to remove outliers. SVD decomposition of H can further obtain the exact rotation matrix R and translation vector t:

[0083] 4. Image alignment and system verification

[0084] The images taken by all cameras are aligned using the calculated homography matrix H. The images are transformed into a common reference frame by the H matrix. The validity of the alignment is verified by calculating the reprojection error:

[0085]

[0086] Among them, m i are the actual observed image coordinates, are the image coordinates estimated by the homography matrix H. Finally, the final evaluation and optimization of the system is performed by comparing the alignment of the overlapping areas in different camera views.

[0087] If the camera image acquisition position changes subsequently, one-click calibration is required in the interactive interface.

[0088] Step 3: Data collection and annotation: Continuously capture images of the tobacco leaf impurity removal line in the real industrial scene to collect data, and perform pixel-level image annotation on the collected data.

[0089] Specifically, the shooting scene is to use a multispectral camera to shoot vertically the materials on the conveyor belt at a height of 1m under the lighting condition of a 100W halogen lamp. After the camera calibration in step 2, the image obtained is an image aligned with the world coordinates, that is, the same pixel coordinate position in each image represents the same point on the world coordinates.

[0090] Afterwards, the collected image data is annotated, and full-supervision annotation is performed for each pixel. For each image, its annotated image intensity is assigned according to the sample library. For example, if the substance at the (x, y) pixel point is labeled n in the sample library, then the pixel intensity of the point is labeled n, and finally the annotation result is exported in jpg format. After the data annotation is completed, the annotated data set is divided into training set, validation set, and test set in a ratio of 6:2:2.

[0091] The formats of the training set, validation set, and test set are exactly the same. The training set is used to update the model parameters during the model training process, the validation set is used to test the model performance after each round of training, and the test set is used to test the parameters of the model that obtains the best results on the validation set and obtain the model evaluation results.

[0092] The sample numbers are marked as shown below:

[0093]

[0094]

[0095] Step 4: Build a fusion feature extraction network model: Design the encoder-decoder network MSFSNet, use 10-channel image data as the network input, and add point-by-point convolutional feature extraction branches and SIPNet multi-layer perceptron.

[0096] Add a point-by-point convolution feature extraction branch, input the original image into the network, and extract the features of the pixels through 1x1 convolution, which can make full use of the spectral information; add a SIPNet multi-layer perceptron, expand the image matrix into a one-dimensional vector, and strengthen the full use of the spectral information through multiple layers of fully connected layers. Finally, the features extracted by the encoder-decoder network MSFSNet, the features extracted by the point-by-point convolution feature extraction branch, and the features extracted by the SIPNet multi-layer perceptron are spliced, and the final segmentation result is obtained through convolution calculation.

[0097] Specifically, Figure 2 As shown in Figure 1, building a fusion feature extraction network model specifically includes the following steps:

[0098] Step 4.1. Scale the input image if any value of its length and width is less than 256, and change the number of input channels to 10, which are 7 spectral channels (450nm, 550nm, 650nm, 720nm, 750nm, 800nm, 850nm) and three visible light channels.

[0099] Step 4.2, random cropping: For image scaling, since the aligned image size will change after each camera calibration, it cannot be guaranteed that the aspect ratio of each input network is the same. In order to retain the shape characteristics of the material, the 10-channel data is randomly cropped to 640 pixels × 640 pixels during model training.

[0100] Step 4.3, train the designed encoder-decoder network MSFSNet. In this training phase, the input image is continuously convolved and downsampled, and then the feature map is deconvolved and upsampled until the feature map size is restored to the same size as the input image. The feature map of the input size is then convolved to obtain the final feature map. The result of the final feature map is the segmentation result of the 11 categories in the image. The cross entropy is used as the loss function during the training process, and the weights obtained during the training process are saved for loading the pre-trained model in step 4.6.

[0101] Step 4.4, train the point-by-point convolution feature extraction branch. In this stage, 10 layers of 1x1 convolution feature extraction are performed on the input image. The process of convolution feature extraction is also the process of increasing the feature map channels. The ReLU activation function is used in the multi-layer 1x1 convolution feature extraction process. The result after training is the feature map, and the result of the feature map is the segmentation result of 11 categories. The cross entropy is used as the loss function during the training process, and the weights obtained during the training process are saved for loading the pre-trained model in step 4.6.

[0102] Step 4.5, train the SIPNet multi-layer perceptron. In this stage, the input image is expanded into a one-dimensional vector after 6 convolution blocks, and then the features are processed by building a multi-layer perceptron, which can make full use of the feature information of spectral imaging. Perform another 6 convolution blocks to restore the output of the last layer of the multi-layer perceptron to a feature map, and continue to perform convolution calculations to obtain the segmentation results of 11 categories. The cross entropy is used as the loss function during the training process, and the weights obtained during the training process are saved for loading and use by the pre-trained model in step 4.6.

[0103] Step 4.6, feature fusion and fusion training. According to 4.3, 4.4, and 4.5, three models for multispectral image segmentation have been obtained, namely, the design of the encoder-decoder network MSFSNet, the training of the point-by-point convolution feature extraction network, and the SIPNet multi-layer perceptron network. In this step, the pre-trained weights of the three networks are first loaded before training, and then the features of the three networks are fused, and the fused feature map is convoluted, and finally a feature map containing segmentation results in 11 categories is obtained. During the training process, cross entropy is used as the loss function.

[0104] The trained model was used to perform semantic segmentation tests on tobacco leaves and non-tobacco substances on the self-built test set images. The test accuracy is shown in Table 1. The average detection time for each image is 18ms.

[0105] Table 1 Detection effect of the method in this embodiment on non-smoking substances on the self-built data set

[0106] Target IoU Background (conveyor belt, tobacco leaves) 0.991 film 0.990 Cable Ties 0.991 feather 0.986 electric wire 0.988 Metal 0.995 fiber 0.986 hemp rope 0.985 Confetti 0.988 Eggs 0.966 Weeds 0.951

[0107] As can be seen from Table 1, after training on the self-built data set, this embodiment has a good effect on non-smoking substance detection, and the performance is significantly improved compared to the method of using visible light for non-smoking substance detection. In the actual deployment application in the factory environment, multi-channel images of the original image size are used for image segmentation, and a high accuracy rate is also achieved. The actual segmentation results are shown in Figure 1. Figure 3 shown.

[0108] In summary, in the actual application of non-smoking substance detection in industrial scenes, traditional visible light-based image segmentation technology is often prone to false detection or missed detection in actual industrial production scenes due to unclear surface characteristics of the object or insufficient segmentation feature information. Spectral cameras can combine the spatial image information and spectral information of the target object, and use the spectral characteristics of different parts or components of the target object for non-destructive, non-contact, fast and efficient identification, classification and analysis. Therefore, this implementation applies multi-spectral fusion detection to image segmentation, and obtains an online detection method for non-smoking substances based on array spectral imaging, which provides an effective target segmentation solution for the rapid detection of non-smoking substances in industrial production scenes and the industrial detection field behind it.

[0109] The present invention also provides an online detection system for non-tobacco substances in leaf threshing and redrying based on area array spectral imaging, which is applied to the above detection method and comprises:

[0110] Imaging system hardware: It consists of a 7-channel spectral camera and a 3-channel visible light camera, which are used to obtain spectral images and visible light images of substances on the tobacco leaf impurity removal line;

[0111] Processor: configured with at least one program. When the program is executed by the processor, the processor controls and coordinates the functions of the following modules:

[0112] Image processing unit: connected to the imaging system hardware, used to perform data acquisition, camera calibration, image annotation and image preprocessing tasks;

[0113] Feature extraction network module: connected to the image processing unit, including the fusion feature extraction network model, used to perform feature extraction and image segmentation;

[0114] Training and testing unit: connected to the feature extraction network module, used to train the encoder-decoder network MSFSNet, the point-by-point convolutional feature extraction branch and the SIPNet multi-layer perceptron, as well as perform model testing and verification;

[0115] Feature fusion unit: connected to the training and testing unit, used to perform fusion training and testing of feature maps to output the final image segmentation result;

[0116] Control unit: connected to all the above units, used to coordinate the operation of the entire system, including data flow management, model training supervision, and output of detection results.

[0117] The embodiment of the present invention further provides a storage medium storing instructions executable by a processor, and when the processor executes the instructions executable by the processor, the method for online detection of non-smoking substances based on planar array spectral imaging is executed.

[0118] It can also be seen that the contents in the above method embodiments are all applicable to the present storage medium embodiments, and the functions and beneficial effects achieved are the same as those in the method embodiments.

[0119] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0120] The logic and / or steps represented in the embodiments or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses.

[0121] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0122] In the description of this specification, the description with reference to the terms "one embodiment", "this embodiment", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0123] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A method for online detection of non-smoking substances based on array spectral imaging, characterized in that: The method comprises the following steps: Step 1: Build the imaging system hardware, using a 7-channel spectral camera and a 3-channel visible light camera to build the imaging hardware; Step 2: Camera calibration: perform geometric calibration on the camera to achieve image position alignment; Step 3: Data collection and annotation: collect images of the tobacco leaf impurity removal line and annotate the data at the pixel level to form a data set; Step 4: Build a fusion feature extraction network model, design the encoder-decoder network MSFSNet, use 10-channel image data as network input, and add point-by-point convolution feature extraction branches and SIPNet multi-layer perceptron; Step 4 specifically includes the following sub-steps: Step 4.1: Modify the number of input channels of the image segmentation network to 10, including 7 spectral channels and 3 visible light channels; Step 4.2: Random cropping: randomly crop the 10-channel data to 640 pixels × 640 pixels during model training; Step 4.3: Train the encoder-decoder network MSFSNet, use cross entropy as the loss function, and obtain 11 category segmentation results; Step 4.4: Train the point-by-point convolution feature extraction branch, use cross entropy as the loss function, and obtain 11 category segmentation results; Step 4.5: Train the SIPNet multi-layer perceptron. After the input image is expanded into a one-dimensional vector after 6 convolution blocks, the light intensity feature is extracted by the multi-layer perceptron. Then, 6 convolution blocks are performed to calculate the output feature map, and finally 11 categories of segmentation results are obtained. Step 4.6: Feature fusion and fusion training, load the pre-trained weights of 4.3, 4.4, and 4.5 models, perform feature fusion and feature extraction, use cross entropy as the loss function, obtain 11 category segmentation results, and train to obtain the final model.

2. The detection method according to claim 1, characterized in that: Step 2 The specific steps are: Collect image data: Use a multispectral camera to capture a checkerboard image and define the camera’s intrinsic parameter matrix; Corner point detection and coordinate extraction: Corner point detection is achieved through image processing algorithms to determine the pixel coordinates of each checkerboard corner point; the optimal extrinsic parameter matrix is ​​calculated through the least squares method to minimize the difference between the actual pixel coordinates and the pixel coordinates calculated by the extrinsic parameter matrix; Calculate the homography matrix and camera parameters: construct the homography matrix through the obtained intrinsic and extrinsic parameters, and perform SVD decomposition on the homography matrix to further obtain the accurate rotation matrix R and translation vector t; Image alignment and system verification: Use the calculated homography matrix to align all images taken by the camera, and perform final evaluation and optimization of the system by comparing the alignment of overlapping areas in different camera views.

3. The detection method according to claim 1, characterized in that: In step 3, the data set includes: a training set, a validation set, and a test set; the training set is used to update the model parameters during the model training process; the validation set is used to test the model performance after each round of training; The test set is used to test the parameters of the model that obtains the best results on the validation set and obtain the model evaluation results.

4. The detection method according to claim 3, characterized in that: The training set, test set, and validation set are divided in a ratio of 6:2:

2.

5. The detection method according to claim 1, characterized in that: In step three, a non-tobacco substance dataset is constructed, and pixel-level image annotation is performed on the non-tobacco substance dataset.

6. The detection method according to claim 1, characterized in that: Step 3 The specific steps are: Image acquisition: Use a multispectral camera to photograph the material on the conveyor belt and ensure that the images obtained after camera calibration are aligned in the world coordinates, that is, the same pixel coordinate position in each image represents the same point in the world coordinates; Fully supervised annotation: annotate the collected image data at the full pixel level, annotating each pixel in the image; By marking the image intensity, if the substance at the pixel point (x, y) is labeled n in the sample library, the pixel intensity of the pixel point is labeled as n; Export annotation results: export the annotated image data for subsequent processing; Dataset division: Divide the labeled dataset into training set, test set, and validation set to support the model training, testing, and validation process.

7. The detection method according to claim 1, characterized in that: The 10 channels described in step 1 are spectral channels of 450nm, 550nm, 650nm, 720nm, 750nm, 800nm, 850nm and three visible light channels.

8. The detection method according to claim 1, characterized in that: The feature fusion and fusion training described in step 6 include the following steps: Load the pre-trained weights of the encoder-decoder network MSFSNet, the point-by-point convolutional feature extraction network, and the SIPNet multi-layer perceptron network; The fused feature map is convolved to finally obtain a feature map containing segmentation results in 11 categories.

9. An online detection system for non-tobacco substances in leaf threshing and redrying based on array spectral imaging, applied to the detection method according to any one of claims 1 to 8, characterized in that: include: Imaging system hardware: It consists of a 7-channel spectral camera and a 3-channel visible light camera, which are used to obtain spectral images and visible light images of substances on the tobacco leaf impurity removal line; Processor: configured with at least one program. When the program is executed by the processor, the processor controls and coordinates the functions of the following modules: Image processing unit: connected to the imaging system hardware, used to perform data acquisition, camera calibration, image annotation and image preprocessing tasks; Feature extraction network module: connected to the image processing unit, including the fusion feature extraction network model, used to perform feature extraction and image segmentation; Training and testing unit: connected to the feature extraction network module, used to train the encoder-decoder network MSFSNet, the point-by-point convolutional feature extraction branch and the SIPNet multi-layer perceptron, as well as perform model testing and verification; Feature fusion unit: connected to the training and testing unit, used to perform fusion training and testing of feature maps to output the final image segmentation result; Control unit: connected to all the above units, used to coordinate the operation of the entire system, including data flow management, model training supervision, and output of detection results.

Citation Information

Patent Citations

  • Non-smoke debris detection method based on hyperspectral image segmentation technology

    CN117611828A

  • Spectrum degradation constrained multi-scale grouping feedback hyperspectral reconstruction method

    CN118212539A