Method for detecting functional components of multiple medicinal and edible substances based on image analysis

Through multispectral image acquisition and analysis technology, combined with multispectral image acquisition equipment and deep learning models, the problem of non-destructive and high-throughput detection of functional components of medicinal and edible substances has been solved, and high-precision and low-cost detection of samples of different types and sources has been achieved.

CN120672675AInactive Publication Date: 2025-09-19JINGFUDA BIOTECHNOLOGY (SHANGHAI) CO LTD

Patent Information

Application Number
CN202510713665.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies make it difficult to conduct high-throughput, multi-index detection of the functional components of edible and medicinal substances quickly, at low cost, and non-destructively, especially for samples of different types, sources, and processing methods. Traditional methods have the problems of complex sample pretreatment, reliance on expensive instruments, and destructive testing.

Method used

Multispectral image acquisition and analysis technology is used to obtain multispectral images of samples through multispectral image acquisition equipment. Gray-level co-occurrence matrix, local binary pattern, principal component analysis and convolutional neural network are combined to extract texture and spectral features. Support vector machine and multi-layer perceptron are used for feature fusion and classification to achieve non-destructive testing.

Benefits of technology

It realizes non-destructive, high-throughput, multi-dimensional detection of medicine-food homologous substances, improves the accuracy and robustness of detection, reduces costs, is applicable to samples of different types and sources, and has an automated detection process, making it suitable for large-scale applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672675A_ABST
    Figure CN120672675A_ABST
Patent Text Reader

Abstract

The invention discloses a method for detecting functional components of multiple medicinal and edible substances based on image analysis, which comprises the following steps: acquiring n multispectral images of a target medicinal and edible substance sample, and preprocessing the acquired multispectral images to obtain a standardized multispectral image set; for an image Ii in the standardized multispectral image set, outputting texture features used for representing surface structure distribution and particle arrangement, extracting spectral features reflecting sample color change and chemical component reaction capability, and outputting structural features used for identifying micro-textures, edge contours and color spot distribution; generating a fused composite feature based on a multi-image channel fusion strategy; and classifying the composite features based on a support vector machine and / or a multi-layer perceptron, and outputting a component detection result of the target medicinal and edible substance sample. Through a multispectral image acquisition and analysis technology, a detection method suitable for samples with homology of medicine and food in different types, sources and processing modes is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for detecting functional components of multiple medicine and food homologous substances based on image analysis. Background Art

[0002] Food-drug substances are widely used in food, health supplements, and traditional Chinese medicine. Their nutritional and therapeutic potential makes them crucial in the modern health industry. These substances come from a wide variety of sources, and the types, content, and activity of their functional components are significantly influenced by factors such as the variety, origin, storage, and processing. Accurate and comprehensive testing of these functional components is crucial for product quality evaluation and traceability.

[0003] Chinese invention patent CN119246702A discloses a "UPLC-MS / MS method for detecting active ingredients in a traditional Chinese medicine composition." This method quantitatively analyzes 14 active ingredients (such as saikosaponin A, paeoniflorin, and salvianolic acid B) potentially present in the specific compound Chinese medicine Ganshuang Granules. This method utilizes gradient elution and electrospray ionization mass spectrometry to detect these functional components. However, this method is specific to the specific compound Chinese medicine Ganshuang Granules and is difficult to directly apply to other medicinal and edible substances.

[0004] Therefore, there is an urgent need for a non-destructive, rapid, low-cost, and scalable method for detecting functional ingredients of food and medicine to meet the actual needs of high-throughput, multi-index evaluation of diverse samples in modern industry. Summary of the Invention

[0005] The present invention provides a method for detecting the functional components of multiple medicinal and edible substances based on image analysis. By using multispectral image acquisition and analysis technology, a detection method suitable for medicinal and edible samples of different types, sources and processing methods is realized.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] The present invention provides a method for detecting functional components of multiple medicinal and edible substances based on image analysis, comprising:

[0008] Collect n multispectral images of target medicine-food homologous substance samples, and preprocess the collected multispectral images to obtain a standardized multispectral image set I = {I1, I2, ..., I n The preprocessing includes: geometric correction, illumination normalization, structure cleaning, and / or atlas enhancement;

[0009] For the image I in the standardized multispectral image set i, based on the gray-level co-occurrence matrix GLCM and / or local binary pattern LBP, texture pattern modeling is performed to output texture features used to characterize surface structure distribution and particle arrangement; based on principal component analysis PCA and / or linear discriminant analysis LDA, multi-channel spectral information is reduced and fused to extract spectral features reflecting the color change and chemical composition reactivity of the sample; based on the convolutional neural network CNN, deep spatial feature extraction is performed on the local area of ​​the image, and structural features used to identify micro-texture, edge contour and color spot distribution are output;

[0010] Based on a multi-image channel fusion strategy, the texture features, the spectral features and the structural features are jointly modeled to generate fused composite features; based on a support vector machine and / or a multi-layer perceptron, the composite features are classified to output the component detection results of the target medicine-food homologous substance sample.

[0011] In an optional embodiment, the image acquisition device for acquiring the multispectral image includes: an acquisition environment control device; a high-resolution imaging sensor with a resolution of not less than 5 million pixels, and a photosensitive chip with a CMOS or InGaAs structure; and a filter component including a rotating filter wheel structure and / or a liquid crystal tunable filter LCTF structure, with a central wavelength distributed in the range of 400-1000 nm and a band interval of 10-50 nm;

[0012] The method of collecting n multispectral images of the target medicine-food homologous substance sample includes:

[0013] The image acquisition environment is controlled by the acquisition environment control device to have a background grayscale value of [0, 20], a continuous spectrum LED array for illumination, a color temperature of 5300–5800K, a color rendering index of not less than 90, an illumination stability fluctuation of less than ±2%, and an exposure time of t e is [1ms,100ms];

[0014] The filter component acquires multi-channel images in a single acquisition, and acquires n multispectral images of the target medicine-food homologous substance sample.

[0015] In an optional embodiment, the multi-image channel fusion strategy includes: channel splicing fusion strategy, feature-level splicing fusion strategy, attention mechanism fusion strategy, tensor and low-rank decomposition fusion strategy, and / or transformer's cross-image channel feature encoding network fusion strategy.

[0016] In an optional embodiment, a fusion strategy based on an attention mechanism is used to jointly model the texture feature, the spectral feature, and the structural feature to generate a fused composite feature, including:

[0017] Constructing the texture features The spectral characteristics and the structural features The three-channel feature matrix Where d is the dimension after unified mapping;

[0018] The three-channel feature matrix is ​​input into the multi-head self-attention module, and the attention weight α used to represent the relative contribution between different features is calculated based on the following formula i ;

[0019]

[0020] Among them, q is a trainable query vector used to judge the importance of each type of feature, f j Indicates the traversal of all features;

[0021] Based on the calculated attention weight α i , the fused composite feature f is calculated by the following formula fused :

[0022]

[0023] In an optional embodiment, the texture feature, the spectral feature, and the structural feature are jointly modeled based on a fusion strategy of tensor and low-rank decomposition to generate a fused composite feature, including:

[0024] The texture features The spectral characteristics and the structural features Mapping to a 3D tensor Among them, d is the feature dimension;

[0025] The three-dimensional tensor is extracted using the singular value decomposition method The principal component information of is obtained to obtain a low-rank core tensor as the composite feature after fusion.

[0026] In an optional embodiment, the texture feature, the spectral feature, and the structural feature are jointly modeled based on a tensor and low-rank decomposition fusion strategy and an attention mechanism fusion strategy to generate a fused composite feature, including:

[0027] The texture features The spectral characteristics and the structural features Mapping to a 2D tensor Among them, d is the feature dimension;

[0028] The two-dimensional tensor is decomposed by singular value decomposition method Perform low-rank decomposition:

[0029]

[0030] in, r is the compression rank;

[0031] Select the two-dimensional tensor after low-rank decomposition Select the first k principal component eigenvectors in The principal component feature set {h1, h2, ..., h k}, calculate the attention weights of the k principal component eigenvectors based on the following formula:

[0032]

[0033] Among them, q is a trainable query vector used to judge the importance of each type of feature, h j Indicates the traversal of all features;

[0034] Based on the calculated attention weight α i , the fused composite feature f is calculated by the following formula fused :

[0035]

[0036] In an optional embodiment, the texture features, the spectral features, and the structural features are jointly modeled based on a tensor and low-rank decomposition fusion strategy and a transformer cross-image channel feature encoding network fusion strategy to generate a fused composite feature, including:

[0037] Constructing the texture features The spectral characteristics and the structural features The input sequence matrix Where d is the dimension after unified mapping, and each row of the input sequence matrix X corresponds to a feature type;

[0038] Add a learnable type embedding to each feature vector Get the encoded sequence input matrix X′:

[0039]

[0040] Input the encoded sequence input matrix X′ into a Transformer encoder containing at least one layer of multi-head self-attention module to calculate the global dependency between each feature vector;

[0041] Extract the average pooling vector output by the Transformer As a composite feature after fusion.

[0042] In an optional embodiment, the method further includes:

[0043] Before jointly modeling the texture features, the spectral features, and the structural features based on the multi-image channel fusion strategy, each type of feature f in the texture features, the spectral features, and the structural features is fused based on the following formula: i To perform standardization:

[0044]

[0045] Among them, μ i is the feature f i The mean value in the training sample, σ i is the feature f i Standard deviation in the training sample;

[0046] The normalized feature f i 'Mapping into a unified feature space through a unified projection matrix W so that the dimensions of the texture feature, the spectral feature and the structural feature are aligned.

[0047] In an optional embodiment, the method further includes:

[0048] Before jointly modeling the texture features, the spectral features, and the structural features based on a multi-image channel fusion strategy, the texture features, the spectral features, and the structural features are respectively input into three independent feature embedding networks, and the structure of the embedding network is a residual neural network with a shared layer or a convolution block plus a fully connected structure, wherein each embedding network includes: an input normalization layer; at least two convolution layers or residual modules for local feature modeling; and a fully connected projection layer for outputting an embedding vector of uniform dimension.

[0049] In an optional embodiment, the method further includes:

[0050] Before classifying the composite features based on the support vector machine SVM and / or the multi-layer perceptron MLP, the composite features are reconstructed using the residual denoising autoencoder structure, and the reconstruction error loss function is:

[0051]

[0052] in, is the reconstructed output vector.

[0053] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0054] 1. The present invention adopts multispectral image acquisition and analysis technology, which avoids the destructive treatment of samples such as crushing and extraction in traditional detection methods, and realizes non-contact and non-destructive detection of sample characterization, which can perform non-destructive detection and preserve the integrity of the sample.

[0055] 2. The present invention is based on the multimodal feature extraction and fusion strategy of images, which can simultaneously extract texture, spectrum and structural information. It is applicable to medicinal and edible samples of different types, sources and processing methods, and realizes high-throughput, multi-dimensional feature extraction and analysis, with high-throughput and broad-spectrum adaptability.

[0056] 3. This invention realizes adaptive modeling of complex image features by introducing advanced deep learning models such as convolutional neural network (CNN), attention mechanism, and multi-head Transformer encoder, which significantly improves the accuracy and robustness of functional component recognition and possesses high-precision discrimination capability driven by deep learning.

[0057] 4. The strategies of channel splicing, low-rank decomposition and attention fusion proposed in this invention are flexible and efficient information fusion strategies. They can make full use of the complementarity between various features, improve the interpretation ability of image features in different dimensions, and enhance the model's adaptability to the variability and heterogeneity of actual samples.

[0058] 5. The detection process of the present invention is highly automated, and the image acquisition device structure can be integrated into a standard imaging platform. No high-end chemical reagents and expensive equipment are required. The detection process is fast, efficient, and low-cost, achieving high automation and low cost, which is convenient for large-scale promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0060] Figure 1 1 is a flow chart of a method for detecting functional components of multiple medicinal and edible substances based on image analysis provided by an embodiment of the present invention;

[0061] Figure 2 It is a structural schematic diagram of a detection system for functional components of multiple medicinal and edible substances based on image analysis provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0062] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operating processes, but the protection scope of the present invention is not limited to the following embodiments.

[0063] In the present invention, words such as "in one possible embodiment," "exemplary," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in the present invention as "in one possible embodiment," "exemplary," or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "in one possible embodiment," "exemplary," or "for example" is intended to present the relevant concepts in a concrete manner.

[0064] Food-drug substances are widely used in food, health supplements, and traditional Chinese medicine. Their nutritional and therapeutic potential makes them crucial in the modern health industry. These substances come from a wide variety of sources, and the types, content, and activity of their functional components are significantly influenced by factors such as the variety, origin, storage, and processing. Accurate and comprehensive testing of these functional components is crucial for product quality evaluation and traceability.

[0065] Currently, the detection of functional components in food-drug homologs primarily relies on physical and chemical analysis methods, such as high-performance liquid chromatography (HPLC), gas chromatography (GC), and liquid chromatography-mass spectrometry (LC-MS). Ultra-high-performance liquid chromatography-tandem mass spectrometry (UPLC-MS / MS) has been widely used for the quantitative detection of components in complex natural products due to its high sensitivity, high selectivity, and low detection limit. However, this technology has the following shortcomings:

[0066] 1. The sample pretreatment steps are complicated and often require crushing, extraction, centrifugation, and filtration, which is time-consuming and prone to human error;

[0067] 2. The detection process relies on expensive precision instruments, which are costly and require high operation, limiting its popularity in grassroots laboratories and the industry.

[0068] 3. This type of method is mostly destructive testing and cannot achieve rapid screening of the entire product or full-process monitoring.

[0069] Therefore, a non-destructive, rapid, low-cost, and scalable method for detecting functional ingredients in food and medicine is urgently needed to meet the practical needs of high-throughput, multi-index evaluation of diverse samples in modern industry. In particular, by leveraging image analysis and artificial intelligence modeling to establish a correlation between sample appearance characteristics and functional ingredients, it is expected that intelligent detection and classification can be achieved without relying on complex instrumentation, which has important application value and research significance.

[0070] Figure 1 The present invention provides a flow chart of a method for detecting functional components of multiple medicinal and edible substances based on image analysis. Figure 1 As shown, a method for detecting functional components of multiple medicinal and edible substances based on image analysis comprises:

[0071] S101, collecting n multispectral images of target medicine-food homologous substance samples, and preprocessing the collected multispectral images to obtain a standardized multispectral image set I = {I1, I2, ..., I n The preprocessing includes: geometric correction, illumination normalization, structure cleaning, and / or atlas enhancement;

[0072] S102, for image I in the standardized multispectral image set, i , based on the gray-level co-occurrence matrix GLCM and / or local binary pattern LBP, texture pattern modeling is performed to output texture features used to characterize surface structure distribution and particle arrangement; based on principal component analysis PCA and / or linear discriminant analysis LDA, multi-channel spectral information is reduced and fused to extract spectral features reflecting the color change and chemical composition reactivity of the sample; based on the convolutional neural network CNN, deep spatial feature extraction is performed on the local area of ​​the image, and structural features used to identify micro-texture, edge contour and color spot distribution are output;

[0073] S103, jointly modeling the texture features, the spectral features, and the structural features based on a multi-image channel fusion strategy to generate a fused composite feature; classifying the composite feature based on a support vector machine and / or a multi-layer perceptron, and outputting a component detection result of the target medicine-food homologous substance sample.

[0074] In this embodiment, the target medicine-food homologous substance sample can be photographed in a standardized manner to collect visible light images. For example, the image acquisition device may include at least one high-resolution industrial camera (resolution not less than 5 million pixels), and the image acquisition is performed under a uniform light source (color temperature 5500K, light intensity 5000lux±5%) and a fixed background.

[0075] In an optional embodiment, the image acquisition device for acquiring the multispectral image includes: an acquisition environment control device; a high-resolution imaging sensor with a resolution of not less than 5 million pixels, and a photosensitive chip with a CMOS or InGaAs structure; and a filter component including a rotating filter wheel structure and / or a liquid crystal tunable filter LCTF structure, with a central wavelength distributed in the range of 400-1000 nm and a band interval of 10-50 nm;

[0076] The method of collecting n multispectral images of the target medicine-food homologous substance sample includes:

[0077] The image acquisition environment is controlled by the acquisition environment control device to have a background grayscale value of [0, 20], a continuous spectrum LED array for illumination, a color temperature of 5300–5800K, a color rendering index of not less than 90, an illumination stability fluctuation of less than ±2%, and an exposure time of t e is [1ms,100ms];

[0078] The filter component acquires multi-channel images in a single acquisition, and collects n multispectral images of the target medicine-food homologous substance sample.

[0079] Image set I refers to a collection of multiple images collected for the same sample. In actual detection, to improve detection accuracy and feature extraction capabilities, each image Ii is collected from different angles, different spectral channels, different viewing angles, or different time points, and can complement each other.

[0080] For example, multispectral imaging can be performed, and imaging can be performed separately at different wavelengths (such as red light, green light, blue light, near-infrared light, and ultraviolet light), such as visible light RGB combined with NIR, with a total of 4 channels, and one image is collected for each channel; multi-angle shooting can be performed, and samples can be photographed from multiple directions such as the front, side, and 45-degree angle to prevent missing some texture or morphology information; multi-time point sampling can be performed, and images of easily changeable samples (such as drying process, fermentation process) can be collected at different time nodes; images can be collected in multiple transformation spaces, and the same image can be transformed into multiple color spaces (such as RGB to HSV to Lab), and each space corresponds to a group of image channels (which can be regarded as expanded into multiple images); image blocks can be collected in different regions, and a large image can be cut into multiple small blocks, and each block is analyzed separately.

[0081] For example, you can use a multispectral camera to take the following 5 pictures:

[0082] I1: red light band (620nm), I2: green light band (550nm), I3: blue light band (470nm), I4: near-infrared band (850nm), I5: ultraviolet band (370nm).

[0083] Then the image set of this sample can be expressed as:

[0084]

[0085] Exemplarily, preprocessing may include: geometric correction, illumination normalization, structure cleaning, and / or atlas enhancement.

[0086] For example, geometric correction includes geometric distortion correction, image registration (image alignment), and size normalization. Geometric distortion correction, such as backprojection based on the camera's intrinsic parameter matrix, can be used to correct barrel and pincushion distortion introduced by the camera. Image registration (image alignment) can be used to align multi-channel images or multi-frame images. Size normalization, such as center cropping followed by scaling, can be used to adjust the input image to a unified input format.

[0087] Lighting standardization includes white balance correction, gamma correction, and color space transformation. White balance correction methods, such as the gray world method, total reflection method, and statistical white point estimation, can be used to eliminate the effects of different light sources on color; gamma correction can be used to adjust image contrast to make it more similar to human vision; and color space transformations, such as chromaticity separation and independent brightness processing, can optimize subsequent extraction of color or texture information.

[0088] Structural cleaning includes background removal, noise suppression, and artifact removal. Background removal, such as image segmentation models based on threshold segmentation, can be used to remove non-target areas such as background plates or stray objects. Noise suppression, such as mean filtering, Gaussian filtering, median filtering, and bilateral filtering, can be used to remove high-frequency noise in images. Artifact removal, such as deblocking filters, wavelet denoising, and CNN denoising models (such as DnCNN), can be used to eliminate compression noise or light transmission artifacts.

[0089] Image enhancement includes contrast enhancement, edge enhancement, and multispectral fusion. Among them, contrast enhancement, such as histogram equalization and adaptive histogram, can be used to improve dark details or the dynamic range of the entire image; edge enhancement, such as Laplacian filtering, Sobel / Scharr operators, and Canny edge detection, can be used to enhance texture edge information; multispectral fusion, such as PCA fusion, weighted average fusion, CNN automatic fusion, and atlas stitching, can be used to fuse information from different bands to enhance feature expression.

[0090] In addition, preprocessing can also be combined with data enhancement, image segmentation, principal component compression, pseudo-color mapping, etc. Among them, data enhancement, such as image rotation, flipping, cropping, brightness perturbation, etc., can be used to improve the robustness of the model and adapt to actual environmental changes; image segmentation, such as cutting the image into multiple patches for local feature extraction; principal component compression, can be used to reduce dimensions and retain the main spectral / color change information; pseudo-color mapping, such as converting a single-channel image into a color image to enhance features, can use bands such as thermal imaging to enhance visibility.

[0091] For example, for the characteristics of multiple medicinal and edible substances with diverse morphologies (such as slices, powders, blocks, strips) and complex morphological structures, structural feature extraction combined with shape description can be used; for the characteristics of difficult color recognition (existence of similar colors) and indistinguishable RGB, Lab color space conversion combined with ΔE color difference analysis can be used; for the characteristics of large differences in surface texture, texture differentiation can be extracted, such as using GLCM, Gabor filter or deep CNN texture channel; for the characteristics of large environmental interference factors, uneven lighting and complex background, data enhancement and background removal models can be used to improve adaptability.

[0092] In summary, for example, the target medicine-food homologous substance sample can be placed in an environment with uniform illumination and uniform background, and its visible light-near infrared image in the 400-1000nm band can be collected using a multispectral image acquisition system. The image distortion correction, background removal, size normalization, gamma correction and spectrum fusion are then performed to obtain a standardized multispectral image set I = {I1, I2, ..., I n For example, gamma correction γ = 2.2 can be used, and the resolution can be uniformly adjusted to 224×224 in the RGB space.

[0093] Another exemplary method is to preprocess the collected multispectral images, which may include color space conversion (such as RGB to HSV or Lab), image denoising (using Gaussian filtering or bilateral filtering), image enhancement (including histogram equalization or adaptive contrast enhancement), and ROI area extraction. After preprocessing, the image size is normalized to W×H pixels (W≥224, H≥224).

[0094] From the above, we can know that the standardized multispectral image set I={I1,I2,...,I n Each image I in i Corresponding to an imaging condition, for example: image versions in different spectral bands (such as RGB, near-infrared, ultraviolet); different angles (front, side); or different pre-processing spaces (such as RGB space, Lab space, etc.).

[0095] For each image I i , multi-dimensional image feature extraction can be performed: texture feature extraction, color and spectrum feature extraction, and deep structure feature extraction.

[0096] For example, texture feature extraction can calculate the gray-level co-occurrence matrix (GLCM) in the image grayscale space to extract statistics such as contrast, entropy, and correlation; or apply local binary pattern (LBP) encoding to texture patterns. When extracting color and spectral features, if the image is multi-channel (such as RGB or Lab), color histograms can be extracted for each channel separately. For multispectral images (>3 channels), principal component analysis (PCA) can be used to compress the high-dimensional spectrum into several principal component channels, and then spectral shape features can be extracted from these principal components. Deep structural feature extraction can use pre-trained convolutional neural networks (such as ResNet and EfficientNet) to extract image embedding vectors. The embeddings of multiple images are then concatenated or averaged to obtain the composite structural features of the final sample.

[0097] Exemplarily, image features can be extracted based on the preprocessed image using a convolutional neural network (CNN) model, where the CNN model is a pre-trained model (such as ResNet, EfficientNet, or MobileNet) and can be fine-tuned through transfer learning, and the extracted image feature vector dimension is N×D (N is the number of image blocks, and D is the feature dimension of each image block).

[0098] In an optional embodiment, the method further includes:

[0099] Before jointly modeling the texture features, the spectral features, and the structural features based on the multi-image channel fusion strategy, each type of feature f in the texture features, the spectral features, and the structural features is fused based on the following formula: i To perform standardization:

[0100]

[0101] Among them, μ i is the feature f i The mean value in the training sample, σ i is the feature f i Standard deviation in the training sample;

[0102] The normalized feature f i 'Mapping into a unified feature space through a unified projection matrix W so that the dimensions of the texture feature, the spectral feature and the structural feature are aligned.

[0103] For example, satisfying the dimension alignment d′∈[64,512] facilitates subsequent fusion processing.

[0104] In an optional embodiment, the method further includes:

[0105] Before jointly modeling the texture features, the spectral features, and the structural features based on a multi-image channel fusion strategy, the texture features, the spectral features, and the structural features are respectively input into three independent feature embedding networks, and the structure of the embedding network is a residual neural network with a shared layer or a convolution block plus a fully connected structure, wherein each embedding network includes: an input normalization layer; at least two convolution layers or residual modules for local feature modeling; and a fully connected projection layer for outputting an embedding vector of uniform dimension.

[0106] Through the above embedding structure, modality adaptive embedding between different types of image features is achieved.

[0107] In an optional embodiment, the multi-image channel fusion strategy includes: channel splicing fusion strategy, feature-level splicing fusion strategy, attention mechanism fusion strategy, tensor and low-rank decomposition fusion strategy, and / or transformer's cross-image channel feature encoding network fusion strategy.

[0108] For example, the extracted image features can be input into a functional component recognition model to establish a mapping relationship between the image features and the functional components. The model may include, for example, a support vector machine (SVM), a random forest (RF), a multi-layer perceptron (MLP), or an ensemble learning method. The model training uses a labeled sample set, and the loss function may be:

[0109]

[0110] Among them, y i is the functional component concentration label of the i-th sample, x i is the image feature vector, θ is the model parameter, and λ is the regularization coefficient.

[0111] In an optional embodiment, a fusion strategy based on an attention mechanism is used to jointly model the texture feature, the spectral feature, and the structural feature to generate a fused composite feature, including:

[0112] Constructing the texture features The spectral characteristics and the structural features The three-channel feature matrix Where d is the dimension after unified mapping;

[0113] The three-channel feature matrix is ​​input into the multi-head self-attention module, and the attention weight α used to represent the relative contribution between different features is calculated based on the following formula i ;

[0114]

[0115] Among them, q is a trainable query vector used to judge the importance of each type of feature, f j Indicates the traversal of all features;

[0116] Based on the calculated attention weight α i , the fused composite feature f is calculated by the following formula fused :

[0117]

[0118] In an optional embodiment, the texture feature, the spectral feature, and the structural feature are jointly modeled based on a fusion strategy of tensor and low-rank decomposition to generate a fused composite feature, including:

[0119] The texture features The spectral characteristics and the structural features Mapping to a 3D tensor Among them, d is the feature dimension;

[0120] The three-dimensional tensor is extracted using the singular value decomposition method The principal component information of is obtained to obtain a low-rank core tensor as the composite feature after fusion.

[0121] In an optional embodiment, the texture feature, the spectral feature, and the structural feature are jointly modeled based on a tensor and low-rank decomposition fusion strategy and an attention mechanism fusion strategy to generate a fused composite feature, including:

[0122] The texture features The spectral characteristics and the structural features Mapping to a 2D tensor Among them, d is the feature dimension;

[0123] The two-dimensional tensor is decomposed by singular value decomposition method Perform low-rank decomposition:

[0124]

[0125] in, r is the compression rank;

[0126] Select the two-dimensional tensor after low-rank decomposition Select the first k principal component eigenvectors from The principal component feature set {h1, h2, ..., h k}, calculate the attention weights of the k principal component eigenvectors based on the following formula:

[0127]

[0128] Among them, q is a trainable query vector used to judge the importance of each type of feature, h j Indicates the traversal of all features;

[0129] Based on the calculated attention weight α i , the fused composite feature f is calculated by the following formula fused :

[0130]

[0131] In an optional embodiment, the texture features, the spectral features, and the structural features are jointly modeled based on a tensor and low-rank decomposition fusion strategy and a transformer cross-image channel feature encoding network fusion strategy to generate a fused composite feature, including:

[0132] Constructing the texture features The spectral characteristics and the structural features The input sequence matrix Where d is the dimension after unified mapping, and each row of the input sequence matrix X corresponds to a feature type;

[0133] Add a learnable type embedding to each feature vector Get the encoded sequence input matrix X′:

[0134]

[0135] Input the encoded sequence input matrix X′ into a Transformer encoder containing at least one layer of multi-head self-attention module to calculate the global dependency between each feature vector;

[0136] Extract the average pooling vector output by the Transformer As a composite feature after fusion.

[0137] In an optional embodiment, the method further includes:

[0138] Before classifying the composite features based on the support vector machine SVM and / or the multi-layer perceptron MLP, the composite features are reconstructed using the residual denoising autoencoder structure, and the reconstruction error loss function is:

[0139]

[0140] in, is the reconstructed output vector.

[0141] In addition, a dropout layer can be added to randomly mask a certain proportion of the fused features. The dropout ratio is p∈[0.1,0.5] to prevent model overfitting. The composite features after regularization and noise reduction are output as the final input of the classification model.

[0142] Exemplarily, the trained model is used to predict the image of the sample to be tested, and the quantitative or semi-quantitative detection results of each target functional component are output, such as polyphenols, alkaloids, flavonoids, saponins, polysaccharides or volatile components; the detection results are output in the form of concentration values ​​(mg / g or %) or scoring levels (such as high / medium / low).

[0143] For example, after extracting features from the image of the sample to be tested and inputting them into the trained prediction model, the quantitative estimation results of each functional component are obtained, and the key components (such as total flavonoids, polyphenols, antioxidant factors, etc.) are ranked by concentration and classified by labels; the output results can be in the form of: concentration prediction value (unit: mg / g or %); classification level (high / medium / low); ingredient radar chart (visual display).

[0144] For example, the model detection results can also be compared with the detection results of traditional physical and chemical methods such as high performance liquid chromatography (HPLC) or gas chromatography-mass spectrometry (GC-MS) to verify the accuracy of the model. The difference index is the root mean square error (RMSE) or correlation coefficient (R 2 ) for evaluation.

[0145] For example, the prediction results are paired with the physical and chemical test results of some samples (such as UPLC-MS / MS, HPLC, etc.), and the transfer learning mechanism is introduced to fine-tune the model parameters to adapt to the image differences under different varieties / origins / processing methods, so as to realize the model's generalization ability for cross-batch samples.

[0146] In one embodiment, three common samples of medicinal and edible substances, such as wolfberry, Chinese yam, and lotus seeds, were imaged under LED lighting with a background grayscale value of 15 and a color temperature of 5500K. Using an imaging device containing an InGaAs photosensitive chip and a liquid crystal tunable filter (LCTF), 61 multispectral images were collected in the 400–1000 nm range with a wavelength interval of 10 nm.

[0147] Gray-level co-occurrence matrix (GLCM) and local binary pattern (LBP) were used to extract surface texture features of the samples after preprocessing, quantifying their roughness and particle distribution. Principal component analysis (PCA) was also performed on the multispectral channel data to extract the spectral features of the first five principal components, which were used to describe color variation trends and chemical responsiveness. A CNN model (consisting of three convolutional and pooling layers and two fully connected layers) was constructed to extract structural features from local image regions, identifying detailed contours, texture edges, and areas of color spots.

[0148] The extracted texture, spectral, and structural features are fed into a tensor fusion module. Low-rank decomposition techniques are used to extract principal component information. A multi-head self-attention mechanism is used to assign feature weights, generating a unified dimensional fused composite feature. This fused composite feature is then fed into a trained SVM classifier to identify the active ingredient content range within each sample type and output the functional component detection results.

[0149] Experimental results show that this method achieved a detection accuracy of over 93% on 50 samples from different sources, and the average detection time for a single sample was less than 2 seconds, demonstrating excellent accuracy and real-time performance.

[0150] Geometric correction is used to eliminate shooting angle deviation, histogram matching is used for illumination standardization, and image filtering and false color enhancement are used to clean and enhance the structure of the atlas.

[0151] The beneficial effects of the technical solution provided by the present invention include at least the following: the present invention adopts multispectral image acquisition and analysis technology, which avoids the destructive treatment of samples such as crushing and extraction in traditional detection methods, realizes non-contact and non-destructive detection of sample characterization, can perform non-destructive detection, and preserves the integrity of the sample. The image-based multimodal feature extraction and fusion strategy can simultaneously extract texture, spectrum and structural information, and is suitable for medicinal and edible samples of different types, sources and processing methods, realizing high-throughput, multi-dimensional feature extraction and analysis, with high-throughput and broad-spectrum adaptability. By introducing advanced deep learning models such as convolutional neural networks (CNNs), attention mechanisms, and multi-headed Transformer encoders, adaptive modeling of complex image features is achieved, significantly improving the accuracy and robustness of functional component identification, and possessing high-precision discrimination capabilities driven by deep learning. The proposed strategies such as channel splicing, low-rank decomposition and attention fusion are flexible and efficient information fusion strategies that can fully utilize the complementarity between various features, improve the interpretation ability of image features in different dimensions, and enhance the model's adaptability to the variability and heterogeneity of actual samples. The detection process is highly automated, and the image acquisition device structure can be integrated into a standard imaging platform. No high-end chemical reagents or expensive equipment are required. The detection process is fast, efficient, and low-cost, achieving high automation and low cost, which is convenient for large-scale promotion and application.

[0152] Figure 2 Schematic diagram of a detection system for functional components of multiple medicinal and edible substances based on image analysis provided by an embodiment of the present invention. Figure 2 The above-mentioned system 10 for detecting functional components of multiple medicinal and edible substances based on image analysis is characterized by comprising: a data processing module 101, a feature processing module 102, and a component detection module 103;

[0153] The data processing module 101 is used to collect n multispectral images of target medicine and food homologous substance samples, and pre-process the collected multispectral images to obtain a standardized multispectral image set I = {I1, I2, ..., I n The preprocessing includes: geometric correction, illumination normalization, structure cleaning, and / or atlas enhancement;

[0154] The feature processing module 102 is used to process the image I in the standardized multispectral image set. i , based on the gray-level co-occurrence matrix GLCM and / or local binary pattern LBP, texture pattern modeling is performed to output texture features used to characterize surface structure distribution and particle arrangement; based on principal component analysis PCA and / or linear discriminant analysis LDA, multi-channel spectral information is reduced and fused to extract spectral features reflecting the color change and chemical composition reactivity of the sample; based on the convolutional neural network CNN, deep spatial feature extraction is performed on the local area of ​​the image, and structural features used to identify micro-texture, edge contour and color spot distribution are output;

[0155] The component detection module 103 is used to jointly model the texture features, the spectral features and the structural features based on a multi-image channel fusion strategy to generate a fused composite feature; classify the composite feature based on a support vector machine and / or a multi-layer perceptron, and output the component detection result of the target medicine-food homologous substance sample.

[0156] Furthermore, it should be noted that the present invention may be provided as a method, apparatus, or computer program product. Thus, embodiments of the present invention may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention may take the form of a computer program product embodied on one or more computer-usable storage media containing computer-usable program code.

[0157] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0158] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0159] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0160] Finally, it should be noted that the above is a preferred embodiment of the present invention. It should be noted that although the preferred embodiment of the present invention has been described, it is clear that those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered as within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. A method for detecting functional components of multiple medicinal and edible substances based on image analysis, characterized in that: include: Collect n multispectral images of target medicine-food homologous substance samples, and preprocess the collected multispectral images to obtain a standardized multispectral image set I = {I1, I2, ..., I n The preprocessing includes: geometric correction, illumination normalization, structure cleaning, and / or atlas enhancement; For the image I in the standardized multispectral image set i , based on the gray-level co-occurrence matrix GLCM and / or local binary pattern LBP, texture pattern modeling is performed to output texture features used to characterize surface structure distribution and particle arrangement; based on principal component analysis PCA and / or linear discriminant analysis LDA, multi-channel spectral information is reduced and fused to extract spectral features reflecting the color change and chemical composition reactivity of the sample; based on the convolutional neural network CNN, deep spatial feature extraction is performed on the local area of ​​the image, and structural features used to identify micro-texture, edge contour and color spot distribution are output; Based on a multi-image channel fusion strategy, the texture features, the spectral features and the structural features are jointly modeled to generate fused composite features; based on a support vector machine and / or a multi-layer perceptron, the composite features are classified to output the component detection results of the target medicine-food homologous substance sample.

2. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 1, wherein: The image acquisition equipment for acquiring the multispectral image includes: an acquisition environment control device; a high-resolution imaging sensor with a resolution of not less than 5 million pixels, and a photosensitive chip with a CMOS or InGaAs structure; and a filter component, including a rotating filter wheel structure and / or a liquid crystal tunable filter LCTF structure, with a central wavelength distributed in the range of 400-1000nm and a band interval of 10-50nm; The method of collecting n multispectral images of the target medicine-food homologous substance sample includes: The image acquisition environment is controlled by the acquisition environment control device to have a background grayscale value of [0, 20], a continuous spectrum LED array for illumination, a color temperature of 5300–5800K, a color rendering index of not less than 90, an illumination stability fluctuation of less than ±2%, and an exposure time of t e is [1ms,100ms]; The filter component acquires multi-channel images in a single acquisition, and acquires n multispectral images of the target medicine-food homologous substance sample.

3. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 1, wherein: The multi-image channel fusion strategy includes: channel splicing fusion strategy, feature-level splicing fusion strategy, attention mechanism fusion strategy, tensor and low-rank decomposition fusion strategy, and / or transformer cross-image channel feature encoding network fusion strategy.

4. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 3, wherein: The texture feature, the spectral feature, and the structural feature are jointly modeled using a fusion strategy based on an attention mechanism to generate fused composite features, including: Constructing the texture features The spectral characteristics and the structural features The three-channel feature matrix Where d is the dimension after unified mapping; The three-channel feature matrix is ​​input into the multi-head self-attention module, and the attention weight α used to represent the relative contribution between different features is calculated based on the following formula i ; Among them, q is a trainable query vector used to judge the importance of each type of feature, f j Indicates the traversal of all features; Based on the calculated attention weight α i , the fused composite feature f is calculated by the following formula fused :

5. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 3, wherein: The texture feature, the spectral feature, and the structural feature are jointly modeled based on a fusion strategy of tensor and low-rank decomposition to generate a fused composite feature, including: The texture features The spectral characteristics and the structural features Mapping to a 3D tensor Among them, d is the feature dimension; The three-dimensional tensor is extracted using the singular value decomposition method The principal component information of is obtained to obtain a low-rank core tensor as the composite feature after fusion.

6. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 1, wherein: The texture features, the spectral features, and the structural features are jointly modeled based on the tensor and low-rank decomposition fusion strategy and the attention mechanism fusion strategy to generate fused composite features, including: The texture features The spectral characteristics and the structural features Mapping to a 2D tensor Among them, d is the feature dimension; The two-dimensional tensor is decomposed by singular value decomposition method Perform low-rank decomposition: in, r is the compression rank; Select the two-dimensional tensor after low-rank decomposition Select the first k principal component eigenvectors from The principal component feature set {h1,h2,...,h k }, calculate the attention weights of the k principal component eigenvectors based on the following formula: Among them, q is a trainable query vector used to judge the importance of each type of feature, h j Indicates the traversal of all features; Based on the calculated attention weight α i , the fused composite feature f is calculated by the following formula fused :

7. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 3, wherein: The texture features, the spectral features, and the structural features are jointly modeled based on the tensor and low-rank decomposition fusion strategy and the transformer's cross-image channel feature encoding network fusion strategy to generate fused composite features, including: Constructing the texture features The spectral characteristics and the structural features The input sequence matrix Where d is the dimension after unified mapping, and each row of the input sequence matrix X corresponds to a feature type; Add a learnable type embedding to each feature vector Get the encoded sequence input matrix X′: Input the encoded sequence input matrix X′ into a Transformer encoder containing at least one layer of multi-head self-attention module to calculate the global dependency between each feature vector; Extract the average pooling vector output by the Transformer As a composite feature after fusion.

8. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 1, wherein: The method further comprises: Before jointly modeling the texture features, the spectral features, and the structural features based on the multi-image channel fusion strategy, each type of feature f in the texture features, the spectral features, and the structural features is fused based on the following formula: i To perform standardization: Among them, μ i is the feature f i The mean value in the training sample, σ i is the feature f i Standard deviation in the training sample; The normalized feature f i 'Mapping into a unified feature space through a unified projection matrix W so that the dimensions of the texture feature, the spectral feature and the structural feature are aligned.

9. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 1, wherein: The method further comprises: Before jointly modeling the texture features, the spectral features, and the structural features based on a multi-image channel fusion strategy, the texture features, the spectral features, and the structural features are respectively input into three independent feature embedding networks, and the structure of the embedding network is a residual neural network with a shared layer or a convolution block plus a fully connected structure, wherein each embedding network includes: an input normalization layer; at least two convolution layers or residual modules for local feature modeling; and a fully connected projection layer for outputting an embedding vector of uniform dimension.

10. The method for detecting functional components of multiple medicinal and edible substances based on image analysis according to claim 1, wherein: The method further comprises: Before classifying the composite features based on the support vector machine SVM and / or the multi-layer perceptron MLP, the composite features are reconstructed using the residual denoising autoencoder structure, and the reconstruction error loss function is: in, is the reconstructed output vector.

Citation Information

Patent Citations

  • Medicinal component detection method of traditional Chinese medicine composition UPLC-MS / MS

    CN119246702A

Cited By

  • Bimodal non-contact optical microscopic imaging device for detecting structures and components of traditional Chinese medicinal materials

    CN121298615A