Pulmonary vein isolation using a novel balloon-based ablation system

By employing a hyperspectral dual-stream multi-scale CNN method, the problems of low detection efficiency and insufficient accuracy of effective components in Typha pollen are solved, achieving rapid, accurate, and non-destructive analysis, which is suitable for the quality control of traditional Chinese medicinal materials.

CN121527631BActive Publication Date: 2026-05-05SHANDONG INST FOR FOOD & DRUG CONTROL +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610049058.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-05-05
Estimated Expiration
2046-01-15

AI Technical Summary

Technical Problem

In existing technologies, the methods for determining the content of effective components in Typha pollen are cumbersome, time-consuming, environmentally unfriendly, and difficult to achieve rapid, online analysis. Traditional machine learning models cannot effectively utilize the differences in multi-band spectral characteristics, resulting in insufficient prediction accuracy.

Method used

A hyperspectral dual-stream multi-scale CNN-based approach is adopted, which extracts features through VNIR and NIR branches respectively, and performs adaptive weighted fusion using a cross-band attention mechanism. Combined with centroid ROI extraction algorithm and multi-scale feature extraction, accurate prediction of the effective components of Typha pollen is achieved.

Benefits of technology

It significantly improves detection accuracy and efficiency, meets pharmacopoeia standards, enables rapid, non-destructive, in-situ analysis, adapts to stable prediction of different samples, and has good universality and environmental friendliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121527631B_ABST
    Figure CN121527631B_ABST
Patent Text Reader

Abstract

This invention proposes a method and system for predicting the content of Typha pollen based on a hyperspectral dual-stream multi-scale CNN, belonging to the field of quality detection of traditional Chinese medicinal materials. The method includes: acquiring hyperspectral image data of Typha pollen samples; preprocessing the acquired hyperspectral image data and extracting the region of interest (ROI) spectral data of the Typha pollen samples; inputting the extracted ROI spectral data into a trained dual-stream multi-scale residual fusion CNN model, using VNIR and NIR multi-scale branches for multi-scale feature extraction and fusion, and outputting multi-scale fused features; inputting the multi-scale fused features into an attention fusion module, and outputting the final fused features; based on the final fused features, outputting the predicted content of the effective components of Typha pollen through a content prediction layer. This invention solves the problems of low efficiency, sample destruction, and insufficient accuracy caused by existing spectral models ignoring band characteristic differences in traditional detection methods, achieving rapid, non-destructive, and high-precision prediction of Typha pollen content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of quality testing technology of Chinese medicinal materials, and in particular relates to a method and system for predicting the content of Typha pollen based on hyperspectral dual-stream multi-scale CNN. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Typha pollen, the dried pollen of Typha orientalis or related species in the Typhaceae family, possesses hemostatic, blood-activating, and urinary-diuretic properties. Its pharmacological activity and its active ingredients (such as isorhamnetin-3-) are key components. O The content of neohesperidin and typhain is closely related. According to relevant regulations, the total amount of the above components in Typha pollen shall not be less than 0.50%. Therefore, establishing an accurate and efficient method for content determination is crucial to ensuring the quality of Typha pollen and its clinical efficacy.

[0004] Currently, the determination of the effective components in Typha pollen mainly relies on traditional chemical analysis methods such as high-performance liquid chromatography (HPLC). While these methods offer high accuracy, they generally suffer from the following drawbacks: cumbersome procedures requiring complex sample extraction and purification pretreatment steps; long detection cycles, with the analysis of a single sample typically taking several hours; the need for large amounts of organic solvents, which contradicts the concept of green testing; and difficulty in achieving rapid, online analysis of large batches of samples, failing to meet the urgent needs of the modern Chinese medicine industry for real-time, in-situ quality monitoring of raw materials and finished products during production.

[0005] Hyperspectral imaging, as an emerging non-destructive testing technology, can simultaneously acquire spatial image information and continuous spectral information of the tested object, forming a three-dimensional data cube that integrates image and spectrum, providing a new technical path for rapid and non-destructive analysis of traditional Chinese medicine. However, while hyperspectral data brings rich information, it also faces challenges such as high data dimensionality, a large amount of redundant information, and complex inter-band correlations. Especially in the field of traditional Chinese medicine, different spectral ranges have different interaction mechanisms with different chemical groups in medicinal materials, and the chemical information they carry has different focuses. In existing technologies, when using hyperspectral technology combined with traditional machine learning or single-stream deep convolutional neural networks for content prediction, all band data are often simply mixed and processed, ignoring the differences in physical characteristics and information value between different band intervals. This results in insufficient feature extraction, weak discriminative power, and limited prediction accuracy and model generalization ability, making it difficult to meet the high requirements of pharmacopoeia standards for quantitative analysis. Summary of the Invention

[0006] To overcome the shortcomings of the prior art, the present invention provides a method and system for predicting cattail pollen content based on hyperspectral dual-stream multi-scale CNN, so as to solve the problems of low detection efficiency, insufficient accuracy and inability to adapt to the differences in multi-band spectral characteristics in the prior art.

[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0008] The first aspect of this invention provides a method for predicting Typhae pollen content based on hyperspectral dual-stream multi-scale CNN;

[0009] A method for predicting Typha pollen content based on hyperspectral dual-stream multi-scale CNN includes:

[0010] Hyperspectral image data of Typha pollen samples were collected, including visible-near infrared data and short-wave infrared data.

[0011] The acquired hyperspectral image data were preprocessed, and the region of interest spectral data of the cattail pollen sample was extracted using a centroid-based region of interest extraction algorithm.

[0012] The extracted spectral data of the region of interest is input into a trained two-stream multi-scale residual fusion CNN model, which outputs the final fused features. The trained two-stream multi-scale residual fusion CNN model utilizes VNIR and NIR multi-scale branches to extract and fuse multi-scale features from the spectral data of the region of interest in the corresponding bands, outputting multi-scale fused features. These multi-scale fused features are then input into a cross-band attention fusion module for interactive and dynamically weighted fusion, outputting the final fused features.

[0013] Based on the final fusion features, the predicted content values ​​of the effective components of Typha pollen are output through the content prediction layer.

[0014] As a further technical solution, the acquired hyperspectral image data is preprocessed, including:

[0015] To eliminate noise interference, the hyperspectral image data of the collected cattail pollen samples were subjected to black and white correction. The correction formula is shown below:

[0016]

[0017] in, The corrected reflectivity. Original reflectivity Black reference image, White reference image.

[0018] As a further technical solution, a centroid-based region of interest extraction algorithm is used to extract the region of interest spectral data of the *Typha orientalis* sample, including:

[0019] Extract the three-dimensional data from the preprocessed hyperspectral image and calculate the centroid coordinates;

[0020] Draw a circle based on the centroid coordinates, determine whether the relevant pixels are inside the circle based on the Euclidean distance, and use the pixels inside the circle as ROI candidate pixels. Extract the spectral data of all ROI candidate pixels and sort them according to wavelength priority.

[0021] The wavelength points of each batch of samples are combined sequentially to form the region of interest spectral data of the cattail pollen sample.

[0022] As a further technical solution, both the VNIR multi-scale branch and the NIR multi-scale branch include multi-scale feature extraction pathways;

[0023] The multi-scale feature extraction pathway includes:

[0024] The first path is used to extract local detail features, including a 1×1 convolutional layer, a batch normalization layer, a ReLU activation function layer and a max pooling layer connected in sequence;

[0025] The second pathway is used to extract mesoscale correlation features, which includes a 1×1 convolutional layer, a one-dimensional dilated convolutional layer with a first dilation rate, a batch normalization layer, a ReLU activation function layer, and a max pooling layer connected in sequence.

[0026] The third path, used to extract global trend features, consists of a 1×1 convolutional layer, a one-dimensional dilated convolutional layer with a second dilation rate, a batch normalization layer, a ReLU activation function layer, and a max pooling layer connected in sequence.

[0027] As a further technical solution, the multi-scale feature extraction path also includes a multi-scale feature fusion unit, which is used to concatenate the output features of the first path, the second path and the third path, and perform dynamic weight fusion through a 1×1 convolutional layer to obtain multi-scale fused features.

[0028] As a further technical solution, the cross-band attention fusion module includes:

[0029] The location encoding unit is used to add a learnable location encoding to the input multi-scale fused features;

[0030] The multi-head self-attention unit includes multiple parallel attention heads, which are used to extract cross-band correlated features from the features after adding positional encoding and output attention features.

[0031] The dynamic weight fusion unit is used to calculate the dynamic weight between the attention features output by the VNIR multi-scale branch and the NIR multi-scale branch, and to perform a weighted summation of the attention features of the two branches based on the dynamic weight to obtain the final fused feature.

[0032] As a further technical solution, based on the final fusion features, the content prediction value of the effective components of Typha pollen is output through the content prediction layer, including:

[0033] The final fused features output by the cross-band attention fusion module are used as the input features of the content prediction layer. The first fully connected layer operation is performed on the input final fused features. The feature space is mapped by a linear transformation function, and then batch normalization is performed to reduce the feature distribution difference. The ReLU activation function is used to introduce non-linear feature expression to obtain the first refined features.

[0034] Using the features refined in the first step as input, a second fully connected layer operation is performed. The linear transformation function is used to further map the features to a low-dimensional feature space. Batch normalization and ReLU activation function are repeated to eliminate feature redundancy and strengthen the association of effective features, resulting in the features refined in the second step.

[0035] A final linear transformation is performed on the features after the second refinement. The dimensional features are mapped to one-dimensional values ​​through the linear transformation function, and the final output is the predicted content of the effective components of Typha pollen.

[0036] The second aspect of this invention provides a system for predicting the content of Typhae pollen based on hyperspectral dual-stream multi-scale CNN.

[0037] A system for predicting Typha pollen content based on hyperspectral dual-stream multi-scale CNN includes:

[0038] The hyperspectral data acquisition module is configured to acquire hyperspectral image data of the cattail pollen sample, wherein the hyperspectral image data includes visible-near infrared band data and shortwave infrared band data.

[0039] The data processing module is configured to: preprocess the acquired hyperspectral image data and extract the region of interest spectral data of the cattail pollen sample using a centroid-based region of interest extraction algorithm;

[0040] The final fusion feature output module is configured to: input the extracted spectral data of the region of interest into a trained dual-stream multi-scale residual fusion CNN model, and output the final fusion features; wherein, the trained dual-stream multi-scale residual fusion CNN model uses VNIR multi-scale branches and NIR multi-scale branches to perform multi-scale feature extraction and fusion on the spectral data of the region of interest in the corresponding band, and outputs multi-scale fusion features; the multi-scale fusion features are input into the cross-band attention fusion module for interactive and dynamic weighted fusion, and output the final fusion features;

[0041] The content prediction output module is configured to output the content prediction value of the effective components of Typha pollen through the content prediction layer based on the final fusion features.

[0042] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the method for predicting cattail pollen content based on hyperspectral dual-stream multi-scale CNN as described in the first aspect of the present invention.

[0043] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the method for predicting cattail pollen content based on hyperspectral dual-stream multi-scale CNN as described in the first aspect of the present invention.

[0044] The above one or more technical solutions have the following beneficial effects:

[0045] (1) This invention employs a deep learning architecture with dual-stream multi-scale residual fusion. This architecture extracts the optimal features of the visible-near-infrared and short-wave infrared bands through independent VNIR and NIR branches, respectively, and performs adaptive weighted fusion through a cross-band attention mechanism, fully exploring the deep physicochemical information of different band spectra. Experimental results show that the method achieves a high coefficient of determination (R²) of 0.9405 and a low root mean square error (RMSE) of 0.0359 on the validation set, with prediction accuracy significantly superior to traditional spectral analysis models and single-stream deep learning models, fully meeting the stringent requirements of the pharmacopoeia for the accuracy of content determination. Furthermore, the model in this invention enhances its comprehensive perception of local spectral details, mid-range correlations, and global trends through multi-scale feature extraction. The dynamic adaptive fusion mechanism enables the model to intelligently adjust the dependence weights on information from different bands based on the spectral characteristics of specific samples. After parameter optimization and training with a large number of samples, the model exhibits stable prediction performance for samples of Typha pollen from different origins and batches, effectively overcoming the performance fluctuation problem caused by sample differences in traditional models, and possessing good practical applicability and reliability.

[0046] (2) The detection process of this invention relies solely on hyperspectral imaging equipment to acquire spectral information, without the need for any chemical reagents, and does not damage the sample itself. This allows precious medicinal material samples to continue to be used for other purposes after detection, avoiding sample loss and environmental pollution caused by chemical methods, and fully conforming to the modern analytical concept of green, environmentally friendly, and sustainable development. This invention greatly reduces the detection cycle of traditional high-performance liquid chromatography (HPLC) and eliminates the need for complex sample pretreatment. It enables rapid, in-situ analysis of Typha pollen samples, greatly improving detection throughput and providing a practical and feasible technical means for raw material acceptance, intermediate product monitoring, and rapid release of finished products in the process of traditional Chinese medicine production, strongly supporting the intelligent and automated upgrading of the traditional Chinese medicine industry.

[0047] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0048] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0049] Figure 1 This is a flowchart of the method in the first embodiment.

[0050] Figure 2 This is an architecture diagram of the two-stream multi-scale residual fusion CNN model in the first embodiment.

[0051] Figure 3 This is a scatter plot of the actual and predicted values ​​of the Puhuang validation set data in the first embodiment.

[0052] Figure 4 This is a system structure diagram of the second embodiment. Detailed Implementation

[0053] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0054] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0055] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0056] Example 1

[0057] This embodiment discloses a method for predicting the content of Typhae pollen based on a hyperspectral dual-stream multi-scale CNN. Visible-near-infrared (400-1000nm) and short-wave infrared (900-1700nm) dual-band spectral data of Typhae pollen samples are acquired using a hyperspectral imaging system. After preprocessing, the data is input into a dual-stream adaptive feature fusion CNN model to achieve accurate prediction of the effective component content. By constructing a specialized, intelligently collaborative dual-stream feature extraction framework and a dynamic adaptive fusion mechanism, the method overcomes the feature extraction limitations of traditional single-stream models, significantly improving the detection efficiency, accuracy, and robustness.

[0058] Specifically, such as Figure 1 As shown, the method for predicting Typhae pollen content based on hyperspectral dual-stream multi-scale CNN includes:

[0059] Step S1: Collect hyperspectral image data of the cattail pollen sample. The hyperspectral image data includes visible-near infrared band data and shortwave infrared band data.

[0060] In this embodiment, a hyperspectral imaging system was constructed to acquire hyperspectral image data of the *Typha orientalis* sample. The constructed hyperspectral imaging system consisted of an MV.C VNIR camera, an NIR camera, eight halogen tungsten lamps, a mobile platform, a white calibration plate, and a computer. Data acquisition was completed in a sealed darkroom. The VNIR camera acquired 301 data points in the 400-1000 nm band with a sampling interval of 2.0 nm, an exposure time of 5.0 ms, and a camera height of 32 cm above the mobile platform. The NIR camera acquired 201 data points in the 900-1700 nm band with a sampling interval of 4.0 nm, an exposure time of 6.5 ms, and a camera height of 37 cm above the mobile platform.

[0061] After the hyperspectral imaging system is set up, dark current calibration and spectral response calibration are performed. Specifically, all light sources are turned off, the camera lens cap is placed on, and dark current images are acquired to eliminate camera noise. Then, a standard white calibration plate with 95% reflectivity is placed in the center of the platform, and a white reference image is acquired for black and white correction of subsequent spectral data.

[0062] Step S2: Preprocess the acquired hyperspectral image data and use a centroid-based region of interest extraction algorithm to extract the region of interest spectral data of the cattail pollen sample.

[0063] Due to interference from camera dark current, light source intensity fluctuations, and ambient stray light reflection during hyperspectral imaging, the acquired hyperspectral image data must first undergo black-and-white correction to convert the original grayscale values ​​into standardized reflectance. Specifically, after the hyperspectral imaging system is set up, black and white reference images are selected for dark current and spectral response calibration. The original reflectance of each pixel is then corrected point-by-point using the following formula:

[0064]

[0065] in, The corrected reflectivity. Original reflectivity Black reference image, White reference image.

[0066] Since the corrected spectral data may still contain random fluctuations caused by camera electronic noise, this embodiment also uses moving average filtering for smoothing. Using 5 wavelength points as a window, the continuous band reflectance curve of each pixel is calculated by moving average to initially eliminate high-frequency random noise. The 5 wavelength points correspond to 10nm in the VNIR band and 20nm in the NIR band, balancing noise reduction and feature preservation.

[0067] Subsequently, a centroid-based region of interest (ROI) extraction algorithm was used to obtain the spectral data of the ROI of the Typha pollen sample by spatial range filtering.

[0068] First, the three-dimensional data cube of the preprocessed hyperspectral image can be loaded using ENVI software. The resulting four-dimensional raw dataset contains spatial X coordinates, spatial Y coordinates, wavelength bands, and corresponding reflectance. The three-dimensional core data of spatial X, spatial Y, and reflectance used for ROI extraction can then be selected.

[0069] The 3D core data is divided into several data blocks according to the pixel size, and the centroid coordinates of each data block are calculated separately:

[0070]

[0071]

[0072] in, This represents the total number of pixels in the current data block. and These are the x and y coordinates of the k-th point within the current data block. To ensure integer operations, the centroid is an integer coordinate.

[0073] Subsequently, with A circle is drawn with the centroid as the reference point and a preset radius. In this embodiment, the radius is set to 4. The point is determined based on the Euclidean distance. Whether it is inside the circle. The formula is as follows:

[0074]

[0075] Pixels within the circle satisfying the above inequality are retained as candidate pixels for the Region of Interest (ROI). Complete reflectance data for all ROI candidate pixels in the 400-1000nm (VNIR band) and 900-1700nm (NIR band) wavelengths are extracted and sorted according to wavelength priority. First, reflectance data for all candidate pixels at 400nm wavelength is extracted and arranged by pixel number; then, reflectance data at 402nm wavelength is extracted and arranged by the same pixel number, and so on, until the reflectance data for all wavelengths in the VNIR band (301 wavelengths at 2nm intervals) and the NIR band (201 wavelengths at 4nm intervals) are sorted and sequentially combined to form the region of interest spectral data for the *Pyrrosia lingua* sample.

[0076] Step S3: Input the extracted spectral data of the region of interest into the trained two-stream multi-scale residual fusion CNN model, and output the final fused features. For example, Figure 2 As shown, the dual-stream multi-scale residual fusion CNN model consists of three parts: VNIR multi-scale branch, NIR multi-scale branch, and cross-band attention fusion module.

[0077] Step S31: The extracted region of interest spectral data of the cattail pollen sample is split into VNIR band data and NIR band data according to the band range. After data dimension transformation and normalization, the data is input into the corresponding VNIR multi-scale branch and NIR multi-scale branch respectively. Data dimension transformation and normalization are used to ensure that the distribution of input data is consistent with the training data and avoid model prediction bias.

[0078] Step S32: The VNIR multi-scale branch and the NIR multi-scale branch have the same structure, both using the MSRF-ACB architecture to complete local-mesoscale-global feature extraction and fusion. Specifically, the multi-scale feature extraction pathways of the VNIR multi-scale branch and the NIR multi-scale branch include:

[0079] The first path consists of a 1×1 convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer connected in sequence. It is primarily used to extract local detail features, capturing local peak features within a 5-7 wavelength range in the spectrum. Channel transformation is performed through 1×1 convolution, reducing computation while preserving key local information. Subsequently, batch normalization and ReLU activation function processing eliminate distribution differences between channels and enhance the model's ability to express local features by introducing nonlinearity. Finally, max pooling downsamples the activated features, retaining the maximum response value within the local region and outputting the local detail features, as shown in the following formula:

[0080]

[0081] in, For local detail features; This is a max pooling operation; For batch normalization layer; This corresponds to the preprocessed spectral data features; The kernel size indicates that the kernel length of a 1×1 convolution in this path is 1. This represents the number of output channels.

[0082] The second path comprises a 1×1 convolutional layer, a one-dimensional dilated convolutional layer with a first dilation rate, a batch normalization layer, a ReLU activation function layer, and a max pooling layer, connected sequentially. It is used for mesoscale correlation feature extraction, with its core function being to capture mid-band correlation features within a wavelength range of 10-12 wavelengths. Specifically, it first uses a 1×1 convolution, then a one-dimensional dilated convolution with a first dilation rate d=2 for feature expansion, followed by batch normalization, ReLU activation function, and max pooling, as shown in the following formula:

[0083]

[0084] in, The extracted mesoscale correlation features.

[0085] The third path comprises a 1×1 convolutional layer, a one-dimensional dilated convolutional layer with a second dilation rate, a batch normalization layer, a ReLU activation function layer, and a max pooling layer, connected sequentially. It is used to extract global baseline trend features within a 15-18 wavelength range in the spectrum. The technical workflow is as follows: after a 1×1 convolution, a one-dimensional dilated convolution with a second dilation rate d=4 is applied, followed by batch normalization, ReLU activation function, and max pooling, as shown in the following formula:

[0086]

[0087] in, This is for extracting global trend features.

[0088] Furthermore, the multi-scale feature extraction pathway also includes a multi-scale feature fusion unit, which integrates the channel features output from the above three pathways. , , Channel concatenation is performed, followed by learning dynamic weights through 1×1 convolution to achieve intelligent fusion and obtain multi-scale fused features. The specific fusion formula is as follows:

[0089]

[0090] in, Multi-scale fusion features. This represents the number of output channels after multi-scale feature fusion.

[0091] Step S33: Input the multi-scale fusion features into the cross-band attention fusion module for interactive and dynamic weighted fusion, and output the final fusion features.

[0092] Learnable positional encoding is added to the multi-scale fused features to highlight the physical meaning corresponding to the spectral band order. Specifically, the multi-scale fused features are added element-wise to the learnable positional encoding parameter P to obtain the feature (xpos) with added positional encoding. The positional encoding parameter P has a dimension of 1×Lpos×D, where Lpos is the number of spectral bands, D=64 is the feature embedding dimension, and its initial value is a random number within the interval [-0.1, 0.1].

[0093] Eight parallel attention heads are employed. First, cross-band correlation feature capture is performed on the features after adding positional encoding. Each attention head independently performs query, key, and value projection transformations on the features. The features after adding positional encoding are simultaneously used as query, key, and value matrices, and linear transformations are performed on the query projection matrix, key projection matrix, and value projection matrix corresponding to that attention head, respectively. Second, the product of the query projection result and the transpose of the key projection result in each head is scaled, and then the result is converted into attention weights in the 0-1 interval using the Softmax function. The attention weights are then multiplied by the value projection result in a weighted sum to obtain the output feature of a single attention head, as shown below:

[0094]

[0095] in, For query vector; The key vector; It is a value vector; This is a transpose.

[0096] The output features of the eight attention heads are concatenated and then linearly transformed using the output projection matrix to obtain the multi-head self-attention features of the VNIR and NIR branches. The query projection matrix, key projection matrix, and value projection matrix corresponding to each attention head are all 64×8 in dimension, with a single head dimension of 8; the output projection matrix has a dimension of 64×64.

[0097] By calculating the attention weights of the VNIR and NIR branches, adaptive weighted fusion of features from both branches is achieved. The specific formula is as follows:

[0098]

[0099]

[0100] In the formula, The dynamic weights of the VNIR branch features; The dynamic weights of the NIR branch features; The cosine similarity function; The multi-head self-attention features output by the VNIR branch; The multi-head self-attention features output by the NIR branch; These are preliminary interaction features between VNIR and NIR branch features, used as a reference benchmark for similarity calculation.

[0101] Furthermore, the dynamic weights of the VNIR branch features and the dynamic weights of the NIR branch features are weighted and fused, as shown in the following formula:

[0102]

[0103] in, The final fused feature is obtained by weighting the features of the two branches.

[0104] Step S4: Based on the final fusion features, output the predicted content values ​​of the effective components of Typha pollen through the content prediction layer.

[0105] The final fused features output by the cross-band attention fusion module are used as the input features of the content prediction layer. The first fully connected layer operation is performed on the input final fused features, and the feature space is mapped through a linear transformation function. Subsequently, batch normalization is performed to reduce the feature distribution difference. The ReLU activation function is used to introduce non-linear feature expression, resulting in the first refined features, as shown below:

[0106]

[0107] in, These are the features after the first feature refinement; This represents the number of neurons in the first fully connected layer. It is a linear mapping function.

[0108] Secondly, using the features refined in the first step as input, a second fully connected layer operation is performed. A linear transformation function is used to further map the features to a low-dimensional feature space. Batch normalization and ReLU activation are then repeated to eliminate feature redundancy and strengthen effective feature associations, resulting in the features refined in the second step, as shown below:

[0109]

[0110] in, These are the features refined in the second stage. This represents the number of neurons in the second fully connected layer.

[0111] Finally, a final linear transformation is performed on the features after the second refinement. The linear transformation function maps the dimensional features to one-dimensional values, resulting in the final output predicted value of the effective component content of Typha pollen.

[0112]

[0113] in, This is the predicted content of the effective components in Typha pollen in the final output.

[0114] Finally, to verify the effectiveness of the method of this invention, its performance was compared with that of traditional high-performance liquid chromatography (HPLC). During model training, the Adam optimizer was used to minimize the mean squared error (MSE), combined with the ReduceLROnPlateau learning rate scheduler to avoid overfitting and improve generalization ability. Experimental results show that the training set Rc² is 0.9448, RMSEc is 0.0331, and the validation set Rv² is 0.9405, with RMSEv = 0.0359, demonstrating excellent prediction accuracy. The scatter plot of the true and predicted values ​​on the validation set is shown below. Figure 3 As shown in the figure, the true value represents the content of cattail pollen measured by traditional high-performance liquid chromatography (HPLC). It can be seen that the method used in this invention has a high degree of fit with the data measured by traditional HPLC. Compared with the traditional method, this invention has a faster detection speed, is simple to operate, and can realize rapid online analysis of large batches of samples. It also has robustness and higher generalization ability.

[0115] Example 2

[0116] This embodiment discloses a Typhae pollen content prediction system based on hyperspectral dual-stream multi-scale CNN;

[0117] like Figure 4 As shown, the Typhae pollen content prediction system based on hyperspectral dual-stream multi-scale CNN includes:

[0118] The hyperspectral data acquisition module is configured to acquire hyperspectral image data of the cattail pollen sample, wherein the hyperspectral image data includes visible-near infrared band data and shortwave infrared band data.

[0119] The data processing module is configured to: preprocess the acquired hyperspectral image data and extract the region of interest spectral data of the cattail pollen sample using a centroid-based region of interest extraction algorithm;

[0120] The final fusion feature output module is configured to: input the extracted spectral data of the region of interest into a trained dual-stream multi-scale residual fusion CNN model, and output the final fusion features; wherein, the trained dual-stream multi-scale residual fusion CNN model uses VNIR multi-scale branches and NIR multi-scale branches to perform multi-scale feature extraction and fusion on the spectral data of the region of interest in the corresponding band, and outputs multi-scale fusion features; the multi-scale fusion features are input into the cross-band attention fusion module for interactive and dynamic weighted fusion, and output the final fusion features;

[0121] The content prediction output module is configured to output the content prediction value of the effective components of Typha pollen through the content prediction layer based on the final fusion features.

[0122] Example 3

[0123] The purpose of this embodiment is to provide a computer-readable storage medium.

[0124] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the method for predicting cattail pollen content based on hyperspectral dual-stream multi-scale CNN as described in Example 1.

[0125] Example 4

[0126] The purpose of this embodiment is to provide an electronic device.

[0127] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for predicting cattail pollen content based on hyperspectral dual-stream multi-scale CNN as described in Example 1.

[0128] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0129] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0130] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for predicting Typhae pollen content based on hyperspectral dual-stream multi-scale CNN, characterized in that, include: Hyperspectral image data of Typha pollen samples were collected, including visible-near infrared data and short-wave infrared data. The acquired hyperspectral image data were preprocessed, and the region of interest spectral data of the cattail pollen sample was extracted using a centroid-based region of interest extraction algorithm. The extracted spectral data of the region of interest is input into a trained two-stream multi-scale residual fusion CNN model, which outputs the final fused features. The trained two-stream multi-scale residual fusion CNN model utilizes VNIR and NIR multi-scale branches to extract and fuse multi-scale features from the spectral data of the region of interest in the corresponding bands, outputting multi-scale fused features. These multi-scale fused features are then input into a cross-band attention fusion module for interactive and dynamically weighted fusion, outputting the final fused features. The cross-band attention fusion module includes: The location encoding unit is used to add a learnable location encoding to the input multi-scale fused features; The multi-head self-attention unit includes multiple parallel attention heads, which are used to extract cross-band correlated features from the features after adding positional encoding and output attention features. The dynamic weight fusion unit is used to calculate the dynamic weight between the attention features output by the VNIR multi-scale branch and the NIR multi-scale branch, and to perform a weighted summation of the attention features of the two branches based on the dynamic weight to obtain the final fused feature. Based on the final fusion features, the predicted content values ​​of the effective components of Typha pollen are output through the content prediction layer, including: The final fused features output by the cross-band attention fusion module are used as the input features of the content prediction layer. The first fully connected layer operation is performed on the input final fused features. The feature space is mapped by a linear transformation function, and then batch normalization is performed to reduce the feature distribution difference. The ReLU activation function is used to introduce non-linear feature expression to obtain the first refined features. Using the features refined in the first step as input, a second fully connected layer operation is performed. The linear transformation function is used to further map the features to a low-dimensional feature space. Batch normalization and ReLU activation function are repeated to eliminate feature redundancy and strengthen the association of effective features, resulting in the features refined in the second step. A final linear transformation is performed on the features after the second refinement. The dimensional features are mapped to one-dimensional values ​​through the linear transformation function, and the final output is the predicted content of the effective components of Typha pollen.

2. The method for predicting Typhae pollen content based on hyperspectral dual-stream multi-scale CNN as described in claim 1, characterized in that, The acquired hyperspectral image data is preprocessed, including: To eliminate noise interference, the hyperspectral image data of the collected cattail pollen samples were subjected to black and white correction. The correction formula is shown below: in, The corrected reflectivity. Original reflectivity Black reference image, White reference image.

3. The method for predicting Typhae pollen content based on hyperspectral dual-stream multi-scale CNN as described in claim 1, characterized in that, A centroid-based region of interest (ROI) extraction algorithm was used to extract the spectral data of the region of interest (ROI) in the *Typha orientalis* sample, including: Extract the three-dimensional data from the preprocessed hyperspectral image and calculate the centroid coordinates; Draw a circle based on the centroid coordinates, determine whether the relevant pixels are inside the circle based on the Euclidean distance, and use the pixels inside the circle as ROI candidate pixels. Extract the spectral data of all ROI candidate pixels and sort them according to wavelength priority. The wavelength points of each batch of samples are combined sequentially to form the region of interest spectral data of the cattail pollen sample.

4. The method for predicting Typhae pollen content based on hyperspectral dual-stream multi-scale CNN as described in claim 1, characterized in that, Both the VNIR multi-scale branch and the NIR multi-scale branch include multi-scale feature extraction pathways; The multi-scale feature extraction pathway includes: The first path is used to extract local detail features, including a 1×1 convolutional layer, a batch normalization layer, a ReLU activation function layer and a max pooling layer connected in sequence; The second pathway is used to extract mesoscale correlation features, which includes a 1×1 convolutional layer, a one-dimensional dilated convolutional layer with a first dilation rate, a batch normalization layer, a ReLU activation function layer, and a max pooling layer connected in sequence. The third path, used to extract global trend features, consists of a 1×1 convolutional layer, a one-dimensional dilated convolutional layer with a second dilation rate, a batch normalization layer, a ReLU activation function layer, and a max pooling layer connected in sequence.

5. The method for predicting Typhae pollen content based on hyperspectral dual-stream multi-scale CNN as described in claim 4, characterized in that, The multi-scale feature extraction pathway also includes a multi-scale feature fusion unit, which is used to concatenate the output features of the first, second and third pathways and perform dynamic weight fusion through a 1×1 convolutional layer to obtain multi-scale fused features.

6. A Typhae Pollen Content Prediction System Based on Hyperspectral Dual-Stream Multi-Scale CNN, characterized in that: include: The hyperspectral data acquisition module is configured to acquire hyperspectral image data of the cattail pollen sample, wherein the hyperspectral image data includes visible-near infrared band data and shortwave infrared band data. The data processing module is configured to: preprocess the acquired hyperspectral image data and extract the region of interest spectral data of the cattail pollen sample using a centroid-based region of interest extraction algorithm; The final fusion feature output module is configured to: input the extracted spectral data of the region of interest into a trained dual-stream multi-scale residual fusion CNN model, and output the final fusion features; wherein, the trained dual-stream multi-scale residual fusion CNN model uses VNIR multi-scale branches and NIR multi-scale branches to perform multi-scale feature extraction and fusion on the spectral data of the region of interest in the corresponding band, and outputs multi-scale fusion features; the multi-scale fusion features are input into a cross-band attention fusion module for interactive and dynamically weighted fusion, and output the final fusion features; the cross-band attention fusion module includes: The location encoding unit is used to add a learnable location encoding to the input multi-scale fused features; The multi-head self-attention unit includes multiple parallel attention heads, which are used to extract cross-band correlated features from the features after adding positional encoding and output attention features. The dynamic weight fusion unit is used to calculate the dynamic weight between the attention features output by the VNIR multi-scale branch and the NIR multi-scale branch, and to perform a weighted summation of the attention features of the two branches based on the dynamic weight to obtain the final fused feature. The content prediction output module is configured to: based on the final fusion features, output the predicted content values ​​of the effective components of Typha pollen through the content prediction layer, including: The final fused features output by the cross-band attention fusion module are used as the input features of the content prediction layer. The first fully connected layer operation is performed on the input final fused features. The feature space is mapped by a linear transformation function, and then batch normalization is performed to reduce the feature distribution difference. The ReLU activation function is used to introduce non-linear feature expression to obtain the first refined features. Using the features refined in the first step as input, a second fully connected layer operation is performed. The linear transformation function is used to further map the features to a low-dimensional feature space. Batch normalization and ReLU activation function are repeated to eliminate feature redundancy and strengthen the association of effective features, resulting in the features refined in the second step. A final linear transformation is performed on the features after the second refinement. The dimensional features are mapped to one-dimensional values ​​through the linear transformation function, and the final output is the predicted content of the effective components of Typha pollen.

7. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the method for predicting cattail pollen content based on hyperspectral dual-stream multi-scale CNN as described in any one of claims 1-5.

8. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for predicting cattail pollen content based on hyperspectral dual-stream multi-scale CNN as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Rheum officinale age limit identification method based on multi-modal map fusion

    CN116343037A

  • Grading method and system for quality detection of radix paeoniae rubra decoction pieces, terminal and medium

    CN120913705A