Corn kernel quality detection model construction method and system based on multi-source data fusion
Through the construction of multi-source data fusion and lightweight model, the accuracy and deployment problems of multi-dimensional evaluation in corn kernel quality detection are solved, and efficient and flexible corn kernel quality detection is achieved, suitable for agricultural sites.
Patent Information
- Application Number
- CN202510534367.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-04
AI Technical Summary
Existing corn kernel quality detection methods are difficult to take into account the comprehensive evaluation of multidimensional indicators, especially in high-throughput and heterogeneous samples, and large-scale deep network architectures have problems of model redundancy and deployment difficulties.
Multi-source data is collected through multi-spectral imaging, near-infrared spectrometer and acoustic sensor, combined with three-dimensional point cloud reconstruction and spatial mapping, and dynamic weighted fusion and dual-channel neural network are used to build a lightweight corn kernel mass detection model, and model compression is used using knowledge distillation technology.
It realizes multi-source information collection of the apparent morphology, internal components and structural density of corn kernels, improves the accuracy and robustness of detection, and is suitable for embedded equipment with limited computing power, meeting the real-time and deployment flexibility needs of agricultural sites.
Smart Images

Figure CN120259826A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of corn kernel quality detection, and particularly to a method and system for constructing a corn kernel quality detection model based on multi-source data fusion. Background Art
[0002] As an important food and industrial raw material, the quality of corn is directly related to food safety, processing efficiency, and market value. Traditional methods for detecting the quality of corn kernels mostly rely on manual experience or single-sensor means, such as visible light image analysis, near-infrared spectroscopy detection, or ultrasonic testing. These methods have certain applicability in specific scenarios, but in actual applications, it is often difficult to comprehensively evaluate multi-dimensional indicators such as the appearance, composition, and structure of corn kernels. Especially when faced with high-throughput, heterogeneous samples, and complex defects, the detection accuracy and robustness are limited.
[0003] Existing multi-modal detection technologies usually adopt methods of multi-source data stitching or simple fusion. However, due to the lack of an effective spatial registration mechanism and a weight adjustment strategy between modalities, it is easy to cause feature redundancy, information imbalance, and loss of key features. In addition, some models adopt large-scale deep network architectures, which, although having strong expressive ability, have problems such as model redundancy, large consumption of computing resources, and difficulty in deployment, and are difficult to meet the actual requirements for real-time performance, embedded deployment, and energy efficiency ratio in the agricultural field. Therefore, how to construct an intelligent detection model that can not only fuse multi-modal data but also has the ability of lightweight deployment remains a technical difficulty and research hotspot in this field. Summary of the Invention
[0004] The present invention provides a method and system for constructing a corn kernel quality detection model based on multi-source data fusion to solve problems such as insufficient fusion, limited detection accuracy, and deployment difficulties in existing methods, and to improve the intelligent level and engineering practicability of the detection system.
[0005] A method for constructing a corn kernel quality detection model based on multi-source data fusion includes the following steps:
[0006] S1, multi-source data acquisition: Collect the apparent morphological data of corn kernels through a multi-spectral imaging device, simultaneously obtain the internal component spectral data of corn kernels using a near-infrared spectrometer, and combine an acoustic wave sensor to measure the structural density characteristic data of corn kernels;
[0007] S2, multi-modal feature construction: Perform three-dimensional point cloud reconstruction on the apparent morphological data to generate a spatial coordinate mapping model, and embed the spectral data and the structural density characteristic data into the spatial coordinate mapping model through a spatio-temporal registration algorithm to form a multi-modal feature matrix;
[0008] S3, dynamic weighted fusion: Perform adaptive weight allocation on the multi-modal feature matrix according to the feature confidence level to generate a mixed feature vector;
[0009] S4, Dual-channel neural network construction: Based on the generated hybrid feature vectors, a dual-channel deep neural network is constructed for high-level feature extraction and classification. In the visible light channel, the ResNeXt network is used to extract the apparent features, and in the near-infrared channel, a one-dimensional convolutional network is used to extract the spectral features. Feature interaction is achieved through a cross-modal attention mechanism;
[0010] S5, Model compression and deployment optimization: The knowledge distillation technique is used to compress the dual-channel neural network to generate a lightweight detection model suitable for embedded devices.
[0011] Optionally, the multi-source data acquisition in S1 includes:
[0012] S11, Apparent morphological data acquisition: Place the corn kernels on the imaging platform of the multispectral imaging device, and obtain the surface images of the corn kernels by irradiating with light sources of different bands , where, 、 are the spatial coordinates of the image, is the spectral wavelength, is the reflection intensity at wavelength . The Gaussian filtering algorithm is used to denoise and enhance the surface image, and the edge detection algorithm is used to extract the contour features of the corn kernels ;
[0013] S12, Internal component spectral data acquisition: Use a near-infrared spectrometer to scan the corn kernels and record their spectral data in the set band range, and calculate the normalized reflectance
[0014] S13, Structural density feature measurement: Use an acoustic wave sensor to emit acoustic waves to the corn kernels and record their propagation time , and combine the geometric thickness to estimate the sound velocity , and combine the material elastic modulus to estimate its structural density feature data .
[0015] Optionally, the multi-modal feature construction in S2 includes:
[0016] S21, 3D point cloud reconstruction: Based on the apparent morphological data, obtain the 3D point cloud data of the corn kernels through structured light or multi-view stereo vision reconstruction methods , where, is the total number of point clouds obtained by reconstruction, is the 3D spatial coordinate of the th point in the point cloud. The multi-view stereo matching method is used to calculate the point cloud depth and estimate its core depth ;
[0017] S22, Spatial coordinate mapping model construction: Perform spatial mapping on the three-dimensional point cloud data and the two-dimensional multispectral image to establish the mapping relationship between pixel points and three-dimensional spatial coordinates : , where is the two-dimensional image coordinate, is the corresponding three-dimensional point cloud spatial coordinate, is the coordinate mapping function;
[0018] S23, Multimodal feature embedding and spatio-temporal registration: Map the spectral data and the structural density feature data to the corresponding three-dimensional spatial coordinates, and align and fuse different modal data through a spatio-temporal registration algorithm to form a multimodal feature matrix , where is the number of point clouds, is the feature dimension of each point, is the th multimodal feature vector of the point.
[0019] Optionally, the dynamic weighted fusion in S3 includes:
[0020] S31, Feature confidence evaluation: Calculate the confidence score for each type of modal feature in the multimodal feature matrix ;
[0021] S32, Weight normalization and assignment: Normalize the confidence scores of all modalities and calculate the fusion weights of each modality ;
[0022] S33, Generation of mixed feature vector: Weight and fuse each modal feature vector to generate the final mixed feature vector .
[0023] Optionally, the construction of the dual-channel neural network in S4 includes:
[0024] S41, Feature channel separation: Divide the appearance feature and the spectral feature in the mixed feature vector by modality and input them into the visible light channel and the near-infrared channel respectively;
[0025] S42, Channel feature extraction: The visible light channel uses the ResNeXt network to extract the appearance features of the image class, and the near-infrared channel uses a one-dimensional convolutional network to extract the spectral features. The two are processed in parallel to obtain the deep information representation of multiple modalities;
[0026] S43, Cross-modal fusion: Through the cross-modal attention mechanism, information interaction and fusion between the visible light channel features and the near-infrared channel features are achieved, enhancing the correlation between features and improving the final classification effect.
[0027] Optionally, the feature sub-channels in S41 include:
[0028] S411, Hybrid feature vector analysis: Analyze the hybrid feature vector , where is the apparent feature part, including the surface morphology and texture information of the corn kernels, is the spectral feature part, including the near-infrared spectral reflection data of the corn kernels;
[0029] S412, Feature channel division: Divide and into two independent channels respectively, denoted as:
[0030] ;
[0031] where is to receive the apparent feature information from , is to receive the spectral feature information from ;
[0032] S413, Input feature channels: Input the divided apparent features and spectral features into the corresponding neural network channels for feature extraction and processing respectively, and set the input layer of each channel as: , where is the apparent feature input to the visible light channel, is the spectral feature input to the near-infrared channel.
[0033] Optionally, the channel feature extraction in S42 includes:
[0034] S421, Visible light channel feature extraction: In the visible light channel, use the ResNeXt network to extract features from the apparent features. The ResNeXt network improves the computational efficiency through grouped convolution and enhances the depth expression ability through residual connections;
[0035] S422, Near-infrared channel feature extraction: In the near-infrared channel, use a one-dimensional convolutional neural network (1D-CNN) to extract features from the spectral features. The 1D-CNN network performs one-dimensional convolution operations on each band of the spectral data to capture the local change patterns of the spectral curve;
[0036] S423, Parallel feature extraction: The feature extraction operations of the visible light channel and the near-infrared channel are performed in parallel to obtain multi-modal feature expressions.
[0037] Optionally, the cross-modal fusion in S43 includes:
[0038] S431, Feature correlation calculation: Calculate the correlation between the visible light channel features and the near-infrared channel features to obtain the weight distribution between modalities through a cross-modal attention mechanism, and calculate the similarity between features;
[0039] S432, Cross-modal attention weight calculation: Apply the softmax function to the calculated similarity matrix for normalization to obtain the attention weights for each pair of features ;
[0040] S433, Feature interaction and fusion: Use the calculated attention weights to perform weighted fusion on the visible light channel features and the near-infrared channel features to generate fused features .
[0041] Optionally, the model compression and deployment optimization in S5 includes:
[0042] S51, Teacher model training: Use a dual-channel neural network as the teacher model and train it with the fused features to obtain a classification model with the ability to jointly discriminate features, denoted as , whose output is the class probability distribution ;
[0043] S52, Student model construction and training: Construct a student model, denoted as , and use the output of the teacher model as the supervision signal for distillation training to output the output probability of the student model ;
[0044] S53, Optimization of the distillation loss function: Minimize the KL divergence between the outputs of the student model and the teacher model as the loss function to guide the student model to learn the prediction behavior of the teacher model.
[0045] A system for constructing a corn kernel quality detection model based on multi-source data fusion, which is used to implement the above-mentioned method for constructing a corn kernel quality detection model based on multi-source data fusion, includes the following modules:
[0046] Data acquisition module: Collect the apparent morphological data of corn kernels through a multispectral imaging device, simultaneously obtain the internal component spectral data using a near-infrared spectrometer, and combine the acoustic wave sensor to measure the structural density characteristic data;
[0047] Feature Construction and Fusion Module: It reconstructs the three-dimensional point cloud of the apparent morphological data to generate a spatial coordinate mapping model, and embeds the spectral data and structural density feature data into the spatial coordinate mapping model through a spatio-temporal registration algorithm to form a multi-modal feature matrix, and then performs weighted fusion according to the feature confidence to generate a mixed feature vector;
[0048] Neural Network Modeling Module: It constructs a two-channel deep neural network based on the mixed feature vector. Among them, the visible light channel uses the ResNeXt network to extract apparent features, and the near-infrared channel uses a one-dimensional convolutional network to extract spectral features, and realizes feature interaction through a cross-modal attention mechanism;
[0049] Model Compression and Deployment Module: It uses knowledge distillation technology to compress the two-channel deep neural network to generate a lightweight quality detection model suitable for deployment on embedded devices.
[0050] Advantages of the present invention:
[0051] In the present invention, by fusing multi-spectral images, near-infrared spectra and acoustic sensing data, multi-source information acquisition of the apparent morphology, internal components and structural density of corn kernels is realized. Combining three-dimensional point cloud reconstruction and spatial mapping methods, an accurate multi-modal feature matrix is constructed. Through a spatio-temporal registration algorithm, different modal data are aligned and fused, effectively retaining the fine-grained information of corn kernels in multiple dimensions such as morphology, spectrum and structure, and improving the integrity and accuracy of the original feature expression.
[0052] In the present invention, by adopting a dynamic weighted fusion strategy and a two-channel deep neural network structure, adaptive processing and parallel modeling of multi-modal features are carried out. It can dynamically adjust the modal feature weights according to the information confidence, improve the discriminant ability of the features and the model's perception ability of complex defects. At the same time, through a cross-modal attention mechanism, deep interaction between the visible light and near-infrared channels is guided to strengthen the complementarity and correlation between modalities, further enhancing the expression effect of the fusion features, and significantly improving the classification accuracy and model generalization ability.
[0053] In the present invention, by introducing knowledge distillation technology, a lightweight student model is constructed. While retaining the discriminant ability of the high-performance two-channel teacher model, the model scale and computational complexity are significantly compressed. The generated lightweight detection model has higher operating efficiency and resource adaptability, is suitable for deployment on embedded devices or edge ends with limited computing power, and meets the multiple requirements of real-time performance, accuracy and deployment flexibility in the agricultural field, and has good application promotion value and industrial transformation potential. Description of the Drawings
[0054] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0055] Figure 1 Schematic flow chart of the construction method of the embodiment of the present invention;
[0056] Figure 2 Schematic diagram of the system function modules of the embodiment of the present invention. Detailed implementation manners
[0057] The following will describe the present invention in detail with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0058] It should be noted that in the specification, when referring to "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc., it indicates that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining embodiments to describe specific features, structures or characteristics, implementing such features, structures or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0059] Generally, terms can be understood at least in part from their use in the context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but instead, at least in part depending on the context, allowing for the existence of other factors that may not be explicitly described.
[0060] As Figure 1 shown, a method for constructing a corn kernel quality detection model based on multi-source data fusion includes the following steps:
[0061] S1, multi-source data acquisition: Collect the apparent morphological data of corn kernels through a multi-spectral imaging device, synchronously obtain the spectral data of the internal components of corn kernels using a near-infrared spectrometer, and combine with an acoustic wave sensor to measure the structural density characteristic data of corn kernels;
[0062] S2, Multimodal Feature Construction: Perform 3D point cloud reconstruction on the apparent morphological data to generate a spatial coordinate mapping model. Embed the spectral data and structural density feature data into the spatial coordinate mapping model through a spatio-temporal registration algorithm to form a multimodal feature matrix;
[0063] S3, Dynamic Weighted Fusion: Perform adaptive weight assignment on the multimodal feature matrix according to the feature confidence to generate a mixed feature vector;
[0064] S4, Dual-channel Neural Network Construction: Based on the generated mixed feature vector, construct a dual-channel deep neural network for high-level feature extraction and classification. Among them, the visible light channel uses the ResNeXt network to extract apparent features, and the near-infrared channel uses a one-dimensional convolutional network to extract spectral features, and feature interaction is realized through a cross-modal attention mechanism;
[0065] S5, Model Compression and Deployment Optimization: Use knowledge distillation technology to compress the dual-channel neural network to generate a lightweight detection model suitable for embedded devices.
[0066] The multi-source data acquisition in S1 includes:
[0067] S11, Apparent Morphological Data Acquisition: Place the corn kernels on the imaging platform of the multispectral imaging device, and obtain the surface images of the corn kernels by irradiating with light sources of different bands , where, , are the spatial coordinates of the image, is the spectral wavelength, is the reflection intensity at wavelength . Use the Gaussian filtering algorithm to denoise and enhance the surface image, and use the edge detection algorithm to extract the contour features of the corn kernels , expressed as:
[0068] ;
[0069] Among them, is the Gaussian kernel function, * represents the two-dimensional convolution operation, is the image data after Gaussian smoothing at wavelength ;
[0070] ;
[0071] S12, Acquisition of Internal Component Spectral Data: Scan the corn kernels with a near-infrared spectrometer and record their spectral data within a set wavelength range , and calculate the normalized reflectance , expressed as:
[0072] ;
[0073] Among them, is the reflected light intensity of corn kernels, is the reflected light intensity of the reference whiteboard;
[0074] S13, Measurement of structural density characteristics: Use an acoustic wave sensor to emit acoustic waves to the corn kernels and record their propagation time , and combine the geometric thickness to estimate the sound velocity , and combine the material elastic modulus to estimate its structural density characteristic data , expressed as:
[0075] ;
[0076] ;
[0077] Through the above content, key characteristic information such as the apparent morphology, internal composition, and structural density of corn kernels can be comprehensively obtained, which not only improves the accuracy and robustness of quality inspection but also enhances the ability to identify complex defects or hidden problems, realizing a full-dimensional and multi-level evaluation of the quality of corn kernels.
[0078] The multi-modal feature construction in S2 includes:
[0079] S21, 3D point cloud reconstruction: Based on the apparent morphology data, obtain the 3D point cloud data of corn kernels through structured light or multi-view stereo vision reconstruction methods , among which, is the total number of point clouds obtained by reconstruction, is the 3D spatial coordinate of the th point in the point cloud. Use the multi-view stereo matching method to calculate the depth of the point cloud and estimate its core depth , expressed as:
[0080] ;
[0081] Among them, is the camera focal length, is the baseline length between cameras, is the disparity of the th pixel point;
[0082] S22, Construction of spatial coordinate mapping model: Map the 3D point cloud data and the 2D multi-spectral image in space to establish the mapping relationship between pixel points and 3D spatial coordinates : , among which, is the 2D image coordinate, is the corresponding 3D point cloud spatial coordinate, is the coordinate mapping function;
[0083] S23, Multimodal Feature Embedding and Spatiotemporal Registration: Map the spectral data and the structural density feature data to the corresponding three-dimensional space coordinates, and align and fuse the different modal data through the spatiotemporal registration algorithm to form a multimodal feature matrix , where is the number of point clouds, is the feature dimension of each point, is the th multimodal feature vector of the
[0084] point. The registration process uses the nearest neighbor interpolation matching algorithm, expressed as:
[0085] where is the spatial point corresponding to the spectral or acoustic data, represents the Euclidean distance;
[0086] Through the above content, not only the precise alignment of different modal data in space and time is achieved, but also the fine-grained features of each corn kernel in terms of morphology, composition, and internal structure are retained, which can improve the model's perception ability of complex quality indicators and enhance the integrity and discriminability of feature expression.
[0087] The dynamic weighted fusion in S3 includes:
[0088] S31, Feature Confidence Evaluation: Calculate the confidence score for each type of modal feature in the multimodal feature matrix , expressed as:
[0089] ;
[0090] where is the modal index, corresponding to apparent morphology, spectral features, and structural density, is the number of samples, is the th sample's feature vector under the th modality, is the confidence evaluation function (feature variance), is the confidence score of the th modality;
[0091] S32, Weight Normalization and Assignment: Normalize the confidence scores of all modalities and calculate the fusion weights of each modality, expressed as:
[0092] ;
[0093] Among them, is the confidence score of the th modality;
[0094] S33, generation of mixed feature vectors: The modality feature vectors are weighted and fused according to the weights to generate the final mixed feature vector , expressed as:
[0095] ;
[0096] Among them, is the modality feature vector after alignment;
[0097] Through the above content, the fusion strategy can be dynamically adjusted according to the reliability and contribution degree of each modality information, avoiding the interference of low-quality or redundant features on the model judgment, significantly improving the expression effect and discriminant performance of the mixed features, and thus enhancing the robustness and accuracy of the model in complex environments.
[0098] The construction of the dual-channel neural network in S4 includes:
[0099] S41, feature channel separation: The appearance features and spectral features in the mixed feature vector are divided according to the modality and input into the visible light channel and the near-infrared channel respectively, retaining the independence and expression advantages of their respective features;
[0100] S42, channel feature extraction: The ResNeXt network is used in the visible light channel to extract the appearance features of the image class, and the one-dimensional convolutional network is used in the near-infrared channel to extract the spectral features. The two are processed in parallel to obtain the deep information representation of multiple modalities;
[0101] S43, cross-modal fusion: Through the cross-modal attention mechanism, information interaction and fusion between the features of the visible light channel and the near-infrared channel are realized, enhancing the correlation between the features and improving the final classification effect.
[0102] The feature channel separation in S41 includes:
[0103] S411, parsing of the mixed feature vector: Parse the mixed feature vector , among which, is the appearance feature part, including the surface morphology and texture information of the corn kernels, is the spectral feature part, including the near-infrared spectral reflection data of the corn kernels;
[0104] S412, feature channel division: Divide and into two independent channels respectively, expressed as:
[0105] ;
[0106] Among them, is to receive the apparent feature information from , is to receive the spectral feature information from ;
[0107] S413, Input Feature Channel: Input the divided apparent features and spectral features into the corresponding neural network channels for feature extraction and processing respectively. Set the input layer of each channel as: Among them, is the apparent feature input to the visible light channel, is the spectral feature input to the near-infrared channel;
[0108] Through the above content, the independent information of each modality can be fully retained, enabling each channel to focus on extracting deep features under a specific modality, avoiding interference between modalities, improving the accuracy and efficiency of feature extraction. At the same time, it enhances the expression ability of the neural network, enabling features of different modalities to be better independently modeled and optimized, thereby improving the overall performance and robustness after multi-modal data fusion.
[0109] The channel feature extraction in S42 includes:
[0110] S421, Visible Light Channel Feature Extraction: In the visible light channel, use the ResNeXt network to extract features from the apparent features. The ResNeXt network improves the calculation efficiency through grouped convolution and enhances the depth expression ability through residual connection, expressed as:
[0111] ;
[0112] Among them, is the input feature of the th layer, is the output feature of the th layer, represents the convolution operation in the residual module, represents the weight of the convolutional layer;
[0113] S422, Near-Infrared Channel Feature Extraction: In the near-infrared channel, use a one-dimensional convolutional neural network (1D-CNN) to extract features from the spectral features. The 1D-CNN network performs one-dimensional convolution operations on each band of the spectral data to capture the local change patterns of the spectral curve, expressed as:
[0114] ;
[0115] Among them, is the spectral feature of the th layer, is the The spectral characteristics of the layer is the convolution kernel weight is the bias term is the activation function
[0116] S423, Parallel feature extraction: The feature extraction operations of the visible light channel and the near-infrared channel are carried out in parallel to ensure that the two types of features can independently extract deep information representations from their respective channels, obtaining a multi-modal feature expression, expressed as:
[0117] ;
[0118] where is the deep apparent feature output from the visible light channel is the deep spectral feature output from the near-infrared channel
[0119] Through the above content, the apparent features and spectral features can be efficiently extracted respectively. The grouped convolution and residual structure of the ResNeXt network can effectively improve the expression ability of the apparent features, while the one-dimensional convolution network can accurately capture the subtle changes in the spectral data, not only improving the efficiency of feature extraction, avoiding interference between different modal data, but also making full use of the advantages of each channel to extract richer and more diverse deep information, ultimately improving the performance and accuracy of the model.
[0120] The cross-modal fusion in S43 includes:
[0121] S431, Feature correlation calculation: For the visible light channel feature and the near-infrared channel feature calculate their correlation, obtain the weight distribution between the modalities through the cross-modal attention mechanism, and calculate the similarity between the features, expressed as:
[0122] ;
[0123] where represents the similarity between the th visible light feature and the th near-infrared feature is the similarity calculation function (cosine similarity) and respectively represent the th visible light feature and the th near-infrared feature
[0124] S432, Cross-modal attention weight calculation: Through the calculated similarity matrix, apply the softmax function to normalize it to obtain the attention weight of each pair of features, expressed as:
[0125] ;
[0126] Among them, represents the attention weight between the th visible light feature and the th near-infrared feature, is the number of near-infrared channel features, represents the similarity between the th visible light feature and all near-infrared features;
[0127] S433, Feature Interaction and Fusion: Using the calculated attention weight to perform weighted fusion on the visible light channel features and the near-infrared channel features to generate fused features , expressed as:
[0128] ;
[0129] Among them, is the feature vector after cross-modal fusion, is the number of visible light channel features, is the number of near-infrared channel features, and respectively represent the th near-infrared feature and the th visible light feature;
[0130] Through the above content, deep interaction and collaborative modeling of different modal information are achieved, which can automatically assign weights according to the similarity and correlation between modalities, effectively strengthen complementary information, suppress redundant interference, thereby enhancing the expression ability and discriminative performance of the fused features, and significantly improving the final classification accuracy and the generalization ability of the model.
[0131] The model compression and deployment optimization in S5 include:
[0132] S51, Teacher Model Training: Using a dual-channel neural network as the teacher model, training with the fused features to obtain a classification model with joint feature discrimination ability, denoted as , whose output is the class probability distribution , expressed as:
[0133] ;
[0134] ;
[0135] Among them, is the temperature coefficient, used to smooth the probability distribution, , is the teacher output probability distribution after softmax normalization, is a multi-layer feedforward neural network;
[0136] S52, Student model construction and training: Construct a student model, denoted as , and use the teacher model output as the supervision signal for distillation training to output the output probability of the student model , expressed as:
[0137] ;
[0138] ;
[0139] where, is a neural network including several layers of lightweight fully connected structures;
[0140] S53, Distillation loss function optimization: Minimize the KL divergence between the student model and the teacher model output as the loss function to guide the student model to learn the prediction behavior of the teacher model. The distillation loss function is expressed as:
[0141] ;
[0142] where, is the knowledge distillation loss, is the Kullback-Leibler divergence function, are the teacher and student output probabilities of the th class respectively, is the square term of the temperature coefficient;
[0143] Through the above content, the knowledge distillation technology is introduced, and the high-performance dual-channel teacher model is used to guide the lightweight student model to learn, so as to significantly reduce the model parameter quantity and computational overhead while maintaining the detection accuracy. This method can effectively transfer the discriminative ability of the fused features to the student model with a smaller structure and faster inference, and is suitable for deployment on embedded or edge devices with limited computing power, improving the model practicability and application flexibility, and meeting the real-time and efficient corn kernel quality detection requirements.
[0144] A corn kernel quality detection model construction system based on multi-source data fusion is used to implement the above-mentioned corn kernel quality detection model construction method based on multi-source data fusion, including the following modules:
[0145] Data acquisition module: Collect the apparent morphological data of corn kernels through a multi-spectral imaging device, simultaneously obtain the internal component spectral data using a near-infrared spectrometer, and combine the acoustic wave sensor to measure the structural density characteristic data;
[0146] Feature construction and fusion module: Perform 3D point cloud reconstruction on the apparent morphological data to generate a spatial coordinate mapping model, and embed the spectral data and structural density feature data into the spatial coordinate mapping model through a spatio-temporal registration algorithm to form a multi-modal feature matrix, and then perform weighted fusion according to the feature confidence to generate a mixed feature vector;
[0147] Neural network modeling module: Construct a two-channel deep neural network based on the mixed feature vector, where the visible light channel uses a ResNeXt network to extract apparent features, the near-infrared channel uses a one-dimensional convolutional network to extract spectral features, and feature interaction is achieved through a cross-modal attention mechanism;
[0148] Model compression and deployment module: Use knowledge distillation technology to compress the two-channel deep neural network to generate a lightweight quality detection model suitable for deployment on embedded devices.
[0149] As Figure 2 shown, a corn kernel quality detection model construction system based on multi-source data fusion is used to implement the above-mentioned corn kernel quality detection model construction method based on multi-source data fusion, including the following modules:
[0150] Data acquisition module: Collect the apparent morphological data of corn kernels through a multi-spectral imaging device, synchronously obtain the internal component spectral data using a near-infrared spectrometer, and combine with an acoustic wave sensor to measure the structural density feature data;
[0151] Feature construction and fusion module: Perform 3D point cloud reconstruction on the apparent morphological data to generate a spatial coordinate mapping model, and embed the spectral data and structural density feature data into the spatial coordinate mapping model through a spatio-temporal registration algorithm to form a multi-modal feature matrix, and then perform weighted fusion according to the feature confidence to generate a mixed feature vector;
[0152] Neural network modeling module: Construct a two-channel deep neural network based on the mixed feature vector, where the visible light channel uses a ResNeXt network to extract apparent features, the near-infrared channel uses a one-dimensional convolutional network to extract spectral features, and feature interaction is achieved through a cross-modal attention mechanism;
[0153] Model compression and deployment module: Use knowledge distillation technology to compress the two-channel deep neural network to generate a lightweight quality detection model suitable for deployment on embedded devices.
[0154] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0155] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for constructing a corn kernel quality detection model based on multi-source data fusion, characterized in that, Including the following steps: S1, Multi-source data acquisition: Collect the apparent morphological data of corn kernels through a multi-spectral imaging device, simultaneously obtain the internal component spectral data of corn kernels using a near-infrared spectrometer, and combine with an acoustic wave sensor to measure the structural density characteristic data of corn kernels; S2, Multi-modal feature construction: Perform 3D point cloud reconstruction on the apparent morphological data to generate a spatial coordinate mapping model, and embed the spectral data and structural density characteristic data into the spatial coordinate mapping model through a spatio-temporal registration algorithm to form a multi-modal feature matrix; S3, Dynamic weighted fusion: Perform adaptive weight allocation on the multi-modal feature matrix according to the feature confidence to generate a mixed feature vector; S4, Dual-channel neural network construction: Based on the generated mixed feature vector, construct a dual-channel deep neural network for high-level feature extraction and classification. Among them, the visible light channel uses the ResNeXt network to extract apparent features, and the near-infrared channel uses a one-dimensional convolutional network to extract spectral features, and feature interaction is realized through a cross-modal attention mechanism; S5, Model compression and deployment optimization: Use knowledge distillation technology to compress the dual-channel neural network to generate a lightweight detection model suitable for embedded devices.
2. A method for constructing a corn kernel quality detection model based on multi-source data fusion according to claim 1, characterized in that, The multi-source data acquisition in S1 includes: S11, Apparent Morphology Data Acquisition: Place the corn kernels on the imaging platform of the multispectral imaging device, and obtain the surface images of the corn kernels by irradiating with light sources of different bands. , where , are the spatial coordinates of the image, is the spectral wavelength, is the reflection intensity at the wavelength . Apply the Gaussian filtering algorithm to denoise and enhance the surface image, and use the edge detection algorithm to extract the contour features of the corn kernels. ; S12, Acquisition of internal component spectral data: Use a near-infrared spectrometer to scan corn kernels and record their spectral data within the set wavelength range , and calculate the normalized reflectance ; S13, Measurement of structural density characteristics: Use an acoustic wave sensor to emit acoustic waves towards corn kernels and record their propagation time , and combine with the geometric thickness to estimate the sound velocity , combine with the elastic modulus of the material to estimate its structural density characteristic data .
3. The method for constructing a corn kernel quality detection model based on multi-source data fusion according to claim 2, wherein The multi-modal feature construction in S2 includes: S21, 3D point cloud reconstruction: Based on the apparent morphological data, the 3D point cloud data of corn kernels is obtained through structured light or multi-view stereo vision reconstruction methods , where is the total number of point clouds obtained by reconstruction, is the 3D spatial coordinates of the th point in the point cloud. The depth of the point cloud is calculated using the multi-view stereo matching method to estimate its core depth ; S22, Spatial coordinate mapping model construction: Perform spatial mapping on the three-dimensional point cloud data and the two-dimensional multispectral image to establish the mapping relationship between pixel points and three-dimensional spatial coordinates : , where is the two-dimensional image coordinate, is the corresponding three-dimensional point cloud spatial coordinate, is the coordinate mapping function; S23, Multimodal Feature Embedding and Spatiotemporal Registration: Map the spectral data and the structural density feature data to the corresponding three-dimensional space coordinates, and align and fuse different modal data through the spatiotemporal registration algorithm to form a multimodal feature matrix , where is the number of point clouds, is the feature dimension of each point, is the -th multimodal feature vector of the point.
4. A method for constructing a corn kernel quality detection model based on multi-source data fusion according to claim 3, characterized in that, The dynamic weighted fusion in S3 includes: S31, Feature confidence evaluation: Calculate the confidence score for each type of modal feature in the multi-modal feature matrix ; S32, Weight normalization and allocation: Normalize the confidence scores of all modalities and calculate the fusion weights of each modality ; S33, Generation of Hybrid Feature Vector: The feature vectors of each modality are weighted and fused according to the weights to generate the final hybrid feature vector .
5. A method for constructing a corn kernel quality detection model based on multi-source data fusion according to claim 4, characterized in that, The dual-channel neural network construction in S4 includes: S41, Feature channel separation: Divide the apparent features and spectral features in the mixed feature vector according to the modality and input them into the visible light channel and the near-infrared channel respectively; S42, Channel feature extraction: The visible light channel uses the ResNeXt network to extract image-like apparent features, and the near-infrared channel uses a one-dimensional convolutional network to extract spectral features. The two are processed in parallel to obtain a multi-modal deep information representation; S43, Cross-modal fusion: Through a cross-modal attention mechanism, realize information interaction and fusion between the visible light channel features and the near-infrared channel features, enhance the correlation between features, and improve the final classification effect.
6. A method for constructing a corn kernel quality detection model based on multi-source data fusion according to claim 5, characterized in that, The feature channel separation in S41 includes: S411, Mixed feature vector analysis: Analyze the mixed feature vector wherein is the apparent feature part, including the surface morphology and texture information of corn kernels, is the spectral feature part, including the near-infrared spectral reflection data of corn kernels; S412, Feature channel division: Divide and into two independent channels respectively, expressed as: ; Among them, is to receive the apparent feature information from , is to receive the spectral feature information from . S413, Input Feature Channels: The divided appearance features and spectral features are respectively input into the corresponding neural network channels for feature extraction and processing. The input layer of each channel is set as: , where is the appearance feature input into the visible light channel, is the spectral feature input into the near-infrared channel.
7. A method for constructing a corn kernel quality detection model based on multi-source data fusion according to claim 6, characterized in that, The channel feature extraction in S42 includes: S421, Visible light channel feature extraction: In the visible light channel, use the ResNeXt network to extract features from the apparent features. The ResNeXt network improves the calculation efficiency through grouped convolution, and at the same time uses residual connections to enhance the deep expression ability; S422, Near-infrared channel feature extraction: In the near-infrared channel, use a one-dimensional convolutional neural network to extract features from the spectral features. The 1D-CNN network performs one-dimensional convolution operations on each band of the spectral data to capture the local change patterns of the spectral curve; S423, Parallel feature extraction: The feature extraction operations of the visible light channel and the near-infrared channel are performed in parallel to obtain a multi-modal feature representation.
8. A method for constructing a corn kernel quality detection model based on multi-source data fusion according to claim 7, characterized in that The cross-modal fusion in S43 includes: S431, Feature correlation calculation: Calculate the correlation between the visible light channel features and the near-infrared channel features Calculate their correlation, obtain the weight distribution between modalities through the cross-modal attention mechanism, and calculate the similarity between features; S432, Cross-modal attention weight calculation: Through the calculated similarity matrix, apply the softmax function to normalize it to obtain the attention weights for each pair of features ; S433, Feature Interaction and Fusion: Using the Computed Attention Weights Weightedly fuse the visible light channel features and the near-infrared channel features to generate fused features .
9. A method for constructing a corn kernel quality detection model based on multi-source data fusion according to claim 8, characterized in that The model compression and deployment optimization in S5 includes: S51, Teacher model training: Using a dual-channel neural network as the teacher model, training is performed using fused features to obtain a classification model with the ability to discriminate joint features, denoted as , whose output is the class probability distribution ; S52, Student model construction and training: Construct a student model, denoted as , and use the output of the teacher model as a supervision signal for distillation training to output the output probability of the student model ; S53, Distillation loss function optimization: Minimize the KL divergence between the outputs of the student model and the teacher model as the loss function to guide the student model to learn the prediction behavior of the teacher model.
10. A corn kernel quality detection model construction system based on multi-source data fusion is used to implement a corn kernel quality detection model construction method based on multi-source data fusion according to any one of claims 1-9, characterized in that, Including the following modules: Data acquisition module: Collect the apparent morphological data of corn kernels through a multispectral imaging device, simultaneously obtain the internal component spectral data using a near-infrared spectrometer, and combine with an acoustic wave sensor to measure the structural density characteristic data; Feature construction and fusion module: Perform 3D point cloud reconstruction on the apparent morphological data to generate a spatial coordinate mapping model, and embed the spectral data and structural density characteristic data into the spatial coordinate mapping model through a spatio-temporal registration algorithm to form a multi-modal feature matrix. Furthermore, perform weighted fusion according to the feature confidence to generate a mixed feature vector; Neural network modeling module: Construct a two-channel deep neural network based on the mixed feature vector. Among them, the visible light channel uses the ResNeXt network to extract apparent features, the near-infrared channel uses a one-dimensional convolutional network to extract spectral features, and feature interaction is achieved through a cross-modal attention mechanism; Model compression and deployment module: Use knowledge distillation technology to compress the two-channel deep neural network to generate a lightweight quality detection model suitable for deployment on embedded devices.
Citation Information
Cited By
Agricultural product quality identification method and system based on multi-modal fusion and feature enhancement
CN120599546A
Mobile device control system and method
CN120802764A
Multi-mode sensing fusion non-contact etchant solution concentration measurement method
CN120820497A
Transform-based wet tissue surface defect real-time detection method
CN120976205A
White spirit knowledge and graph combined question-answering system and test method
CN121255993A