Intelligent identification and classification method and system for traditional Chinese medicine decoction pieces based on deep learning

By improving the RT-DETR network and feature extraction module, the problems of low efficiency and insufficient accuracy in the identification of Chinese herbal medicine pieces are solved, and efficient and accurate identification and classification of Chinese herbal medicine pieces are achieved.

CN120997827AActive Publication Date: 2025-11-21CHANGCHUN VOCATIONAL INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511512204.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-21
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Traditional identification of Chinese herbal medicine pieces relies on human experience, which is inefficient and easily affected by human factors. Existing deep learning methods lack accuracy and robustness in multi-scale and multi-angle identification of Chinese herbal medicine pieces.

Method used

An improved RT-DETR network is used, which extracts features through the LGLB module and performs upsampling in combination with the C2D module to generate high-resolution enhanced feature maps. The detection results, including bounding boxes with spatial location information and category prediction, are generated through decoding.

Benefits of technology

It achieves efficient and accurate recognition and classification of Chinese herbal medicine slices images of different scales and angles, improving the accuracy and robustness of the intelligent recognition system for Chinese herbal medicine slices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997827A_ABST
    Figure CN120997827A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine decoction piece intelligent identification and classification method and system based on deep learning, and relates to the technical field of intelligent classification, and the method comprises the steps: collecting a traditional Chinese medicine RGB image, carrying out the preprocessing of the traditional Chinese medicine RGB image, and obtaining a traditional Chinese medicine RGB image sequence; inputting the traditional Chinese medicine RGB image sequence into an improved RT-DETR network, and performing feature extraction on the traditional Chinese medicine RGB image sequence through an LGLB module to obtain an enhanced feature map; performing up-sampling on the enhanced feature map through a C2D module to generate a high-resolution enhanced feature map; and decoding the high-resolution enhanced feature map to generate a detection result. According to the invention, by combining the LGLB module and the C2D module, the detail features of the traditional Chinese medicine decoction pieces are automatically extracted and classified.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent classification, in particular to a traditional Chinese medicine decoction piece intelligent recognition and classification method and system based on deep learning. BACKGROUND

[0002] In recent years, with the rapid development of artificial intelligence technology, especially the progress in the field of deep learning and computer vision, more and more intelligent technologies have begun to be applied in drug identification, quality detection and other fields. In the field of intelligent recognition of traditional Chinese medicine decoction pieces, deep learning, especially convolutional neural networks (CNN), has been widely used in image classification, object detection and other tasks, and has achieved remarkable results. Traditional Chinese medicine decoction pieces are the forms of medicinal materials obtained after cutting and processing of traditional Chinese medicinal materials during preparation, and are an important basis for traditional Chinese medicine treatment. The quality of traditional Chinese medicine decoction pieces directly affects the curative effect of traditional Chinese medicine.

[0003] However, traditional identification of traditional Chinese medicine decoction pieces mainly relies on manual experience, which is time-consuming and laborious, and is easily affected by human factors, resulting in unstable and large error identification results, which cannot meet the needs of modern and intelligent drug management. Due to the diversity of the appearance of traditional Chinese medicine decoction pieces and the high similarity of the appearance, and the influence of factors such as shooting angle and light, the existing traditional Chinese medicine decoction piece recognition system based on deep learning still has some challenges in accuracy and robustness, especially when dealing with multi-scale and multi-angle traditional Chinese medicine decoction piece images, the performance of the model is often affected. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a traditional Chinese medicine decoction piece intelligent recognition and classification method based on deep learning to solve the problems of low efficiency of manual identification of traditional Chinese medicine decoction pieces and insufficient accuracy and robustness of existing deep learning methods in multi-scale and multi-angle traditional Chinese medicine decoction piece identification.

[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a traditional Chinese medicine decoction piece intelligent recognition and classification method based on deep learning, which comprises: collecting traditional Chinese medicine RGB images, pre-processing the traditional Chinese medicine RGB images, and obtaining a sequence of traditional Chinese medicine RGB images; inputting the sequence of traditional Chinese medicine RGB images into an improved RT-DETR network, extracting features of the sequence of traditional Chinese medicine RGB images through an LGLB module, and obtaining an enhanced feature map; up-sampling the enhanced feature map through a C2D module to generate a high-resolution enhanced feature map; decoding the high-resolution enhanced feature map to generate a detection result.

[0007] As a preferred scheme of the intelligent identification and classification method of traditional Chinese medicine decoction pieces based on deep learning, the preprocessing includes denoising processing, normalization processing, image uniform size processing, and expansion of the training data set.

[0008] As a preferred scheme of the intelligent identification and classification method of traditional Chinese medicine decoction pieces based on deep learning, the traditional Chinese medicine RGB image is preprocessed to obtain a traditional Chinese medicine RGB image sequence, and the steps are as follows. The traditional Chinese medicine RGB image is collected by a high-resolution image acquisition device. The traditional Chinese medicine RGB image is denoised by a denoising algorithm. The denoised traditional Chinese medicine RGB image is normalized in pixel value, and the normalized traditional Chinese medicine RGB image is adjusted to a uniform size. The traditional Chinese medicine RGB image with uniform size is subjected to image enhancement processing to obtain a preprocessed traditional Chinese medicine RGB image sequence.

[0009] As a preferred scheme of the intelligent identification and classification method of traditional Chinese medicine decoction pieces based on deep learning, the improved RT-DETR network includes a backbone network, an LGLB module, a C2D module, and an RTDETRDecoder.

[0010] As a preferred scheme of the intelligent identification and classification method of traditional Chinese medicine decoction pieces based on deep learning, the training process of the improved RT-DETR network includes the following steps. Set the training experiment environment, the input size of the traditional Chinese medicine RGB image sequence, and the training hyperparameters. Use mAP50 and mAP50:95 as the main performance indicators, and combine Precision and Recall to comprehensively evaluate the performance of the improved RT-DETR network on the validation set. Use the joint loss of classification cross-entropy and feature regularization to iteratively train the parameters of the improved RT-DETR network end-to-end, and obtain the trained improved RT-DETR network.

[0011] As a preferred scheme of the intelligent identification and classification method of traditional Chinese medicine decoction pieces based on deep learning, the feature extraction of the traditional Chinese medicine RGB image sequence by the LGLB module includes the following steps. Perform local feature aggregation operation on the preprocessed traditional Chinese medicine RGB image sequence to obtain a feature map after local feature aggregation. Perform multi-head self-attention calculation on the feature map after local feature aggregation by low-rank projection to generate a feature map after global attention refinement. The feature map after global attention refinement is input into an MLP for dimension increasing and dimension decreasing operation to generate an MLP processed feature map; The feature map after local feature aggregation, the feature map after global attention refinement and the MLP processed feature map are subjected to a gated tensor operation, and the enhanced feature map is generated through the decomposable constraint of channel attention and spatial attention for gated fusion.

[0012] As a preferred scheme of the intelligent recognition and classification method of traditional Chinese medicine decoction pieces based on deep learning, wherein: the enhanced feature map is up-sampled through a C2D module to generate a high-resolution enhanced feature map, and the steps are as follows, The enhanced feature map is input into the C2D module to expand through a zero interpolation up-sampling operator to generate a sparse up-sampled enhanced feature map; The sparse up-sampled enhanced feature map is mapped to the frequency domain through Fourier transform, and the frequency domain filter is subjected to frequency domain filtering through an optical transfer function to generate a frequency domain filtered enhanced feature map; The frequency domain filtered enhanced feature map is input into a PadShift operation to perform centering circular shift and zero padding on the point spread function to generate a PadShift processed enhanced feature map; The PadShift processed enhanced feature map is subjected to frequency domain restoration, and the high frequency amplification and noise suppression are performed through the frequency band weight and the channel weight to generate a frequency domain weight restored enhanced feature map; The frequency domain weight restored enhanced feature map is input into a block rearrangement operator to perform local frequency domain aggregation, and the high-resolution enhanced feature map is generated through inverse rearrangement operation.

[0013] As a preferred scheme of the intelligent recognition and classification method of traditional Chinese medicine decoction pieces based on deep learning, wherein: the detection result includes a bounding box containing spatial position information and a class prediction.

[0014] As a preferred scheme of the intelligent recognition and classification method of traditional Chinese medicine decoction pieces based on deep learning, wherein: the high-resolution enhanced feature map is decoded to generate a detection result, and the steps are as follows, The high-resolution enhanced feature map of different scales is subjected to bounding box regression operation to obtain a bounding box containing spatial position information; The high-resolution enhanced feature map of different scales is input into a classification head to generate a class probability.

[0015] In a second aspect, the present application provides an intelligent recognition and classification system of traditional Chinese medicine decoction pieces based on deep learning, which comprises a preprocessing module, a traditional Chinese medicine RGB image is collected, and denoising, normalization and uniform size processing are performed to generate a traditional Chinese medicine RGB image sequence; The network training module trains the improved RT-DETR network through the labeled traditional Chinese medicine RGB image sequence, continuously optimizes the network parameters, and obtains the trained improved RT-DETR network. The feature extraction module calls the LGLB module in the trained improved RT-DETR network, extracts deep features of the traditional Chinese medicine RGB image sequence, and generates an enhanced feature map. The up-sampling module receives the enhanced feature map, realizes up-sampling and detail recovery through C2D operation, and generates a high-resolution enhanced feature map. The classification and detection module inputs the high-resolution enhanced feature map into the decoder, and generates a detection result of traditional Chinese medicine decoction pieces.

[0016] The LGLB module and the C2D module effectively process the images of traditional Chinese medicine decoction pieces, realize automatic extraction and classification of the detail features of the traditional Chinese medicine decoction pieces, and effectively and accurately identify and classify the traditional Chinese medicine decoction piece images of different scales and angles through the combination of multi-level feature extraction, composite scaling deep neural network architecture and frequency domain operation technology, thereby improving the accuracy and robustness of the traditional Chinese medicine decoction piece intelligent identification system. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0018] Figure 1 The flowchart of the intelligent identification and classification method of traditional Chinese medicine decoction pieces based on deep learning.

[0019] Figure 2 The schematic diagram of the intelligent identification and classification system of traditional Chinese medicine decoction pieces based on deep learning.

[0020] Figure 3 The structure diagram of the improved RT-DETR network.

[0021] Figure 4 The mAP diagram of the original RT-DETR.

[0022] Figure 5 The mAP diagram of the improved RT-DETR network.

[0023] Figure 6 The loss trend diagram of the original RT-DETR.

[0024] Figure 7 The loss trend diagram of the improved RT-DETR network.

[0025] Figure 8 A detection effect diagram of the improved RT-DETR network. DETAILED DESCRIPTION

[0026] In order to make the above objectives, characteristics and advantages of the present application more apparent, concrete embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0027] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details given herein, that the present application can be practiced with other than the described implementations, and that variations from the particular examples given can be made and practiced within the scope of the present application.

[0028] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an independent or alternative embodiment that excludes other embodiments.

[0029] Reference Figures 1-8 For one embodiment of the present application, the embodiment provides a deep learning-based intelligent identification and classification method for traditional Chinese medicine decoction pieces, including the following steps: Collecting traditional Chinese medicine RGB images, pre-processing the traditional Chinese medicine RGB images, and obtaining a sequence of traditional Chinese medicine RGB images.

[0030] Further, a high-resolution image acquisition device is used to collect traditional Chinese medicine RGB images of traditional Chinese medicine decoction pieces.

[0031] The traditional Chinese medicine RGB images are denoised by a denoising algorithm.

[0032] Further, the background noise in the traditional Chinese medicine RGB images is removed by the denoising algorithm, and the pixel values of the traditional Chinese medicine RGB images are normalized to the range of [0, 1].

[0033] The denoised traditional Chinese medicine RGB images are normalized in pixel value, and the normalized traditional Chinese medicine RGB images are adjusted to a uniform size.

[0034] The traditional Chinese medicine RGB images of uniform size are subjected to image enhancement processing by an image enhancement technique, and a sequence of pre-processed traditional Chinese medicine RGB images is obtained.

[0035] The sequence of traditional Chinese medicine RGB images is input into an improved RT-DETR network, the traditional Chinese medicine RGB image sequence is subjected to feature extraction by an LGLB module, and an enhanced feature map is obtained.

[0036] Further, the improved RT-DETR network comprises a backbone, an LGLB module, a C2D module, and an RTDETRDecoder (decoding and multi-scale feature fusion). The structure diagram of the improved RT-DETR network is shown in Figure 3 , wherein, Figure 3 The serial numbers 0-26 in the above table represent the numbers of the corresponding layers of the improved RT-DETR network, and P1, P2, P3, P4, and P5 represent the feature maps output by the improved RT-DETR network at different levels.

[0037] The training process of the improved RT-DETR network is as follows. Specifically, the training experimental environment, the input size of the training image sequence, and the training hyperparameters are set. The mAP50 and mAP50:95 are used as the main performance indicators, and the Precision and Recall are combined to comprehensively evaluate the performance of the improved RT-DETR network on the validation set. The classification cross-entropy and feature regularization joint loss are used to perform end-to-end iterative training on the parameters of the improved RT-DETR network in the iteration, and the trained improved RT-DETR network is obtained.

[0038] Further, the training hyperparameters include cache, imgsz, epochs, batch, workers, and optimizer. The training experimental environment is set to Windows 11, PyTorch 1.13.1, torchvision 0.13.1, Python 3.9, and CUDA 11.6. The input size of the training image is set to 640x640, and the training hyperparameters are set as follows: ; ; ; ; ; ; The initial learning rate is set to 0.01, and the weight decay coefficient is set to 0.0005. The mAP50 and mAP50:95 are used as the main performance indicators, and the Precision and Recall are combined to comprehensively evaluate the performance of the improved RT-DETR network on the validation set, to ensure the detection accuracy and generalization ability of the network in the target detection task. The classification cross-entropy and feature regularization joint loss are used to perform end-to-end iterative training on the parameters of the improved RT-DETR network in the iteration, which is represented as: ; wherein, denotes the joint loss of the classification cross-entropy and the feature regularization, denotes the cross-entropy loss, denotes the probability distribution predicted by the model, denotes the actual label, denotes the trade-off coefficient, denotes the F1 norm, denotes the gradient operation, denotes the enhanced feature map generated by the gated fusion, denotes the feature map after upsampling, denotes the L2 norm, i.e., the square of the Euclidean distance, for measuring the difference between the enhanced feature map generated by the gated fusion and the feature map after upsampling.

[0039] It should be noted that, The network training and validation performance are set to meet the reasonable balance between the cross-entropy term and the feature regularization term during the training process.

[0040] After the improved RT-DETR network parameters in the iteration are iteratively trained end-to-end, the trained improved RT-DETR network is obtained.

[0041] The pre-processed traditional Chinese medicine RGB image sequence is subjected to local feature aggregation operation to obtain a feature map after local feature aggregation.

[0042] Further, the improved RT-DETR network is used to extract features from the pre-processed traditional Chinese medicine RGB image sequence, learn important structural features in the pre-processed traditional Chinese medicine RGB image sequence, and perform multi-scale local feature aggregation on the pre-processed traditional Chinese medicine RGB image sequence through group convolution or depth separable convolution to obtain a feature map after local feature aggregation. .

[0043] It should be noted that the important structural features in the pre-processed traditional Chinese medicine RGB image sequence include the morphology, boundary, and texture information of the traditional Chinese medicine decoction pieces.

[0044] It should be noted that local feature aggregation is used to process the local area of the input feature, and through the processing of the local area, it helps the improved RT-DETR network to capture fine-grained local structures. When sr_ratio>1, the LGLB module enables local feature aggregation, which affects the number of input channels without changing the spatial size. The local feature aggregation of the improved RT-DETR network helps to enhance the local information in the image, especially when the resolution of the pre-processed traditional Chinese medicine RGB image sequence is low, which can improve the ability to capture details.

[0045] It should be noted that the pre-processed traditional Chinese medicine RGB image sequence is aggregated by multi-scale local features through group convolution or depth separable convolution, denoted as: ; wherein, denotes the index of the convolution kernel of different scales, denotes the total number of the scale of the convolution kernel, denotes the depth convolution, denotes the size of the convolution kernel, denotes the point-by-point convolution weight.

[0046] The multi-head self-attention calculation of low-rank projection is performed on the feature map after local feature aggregation to generate the feature map after global attention refinement.

[0047] Further, the SelfAttention is adopted to model the relationship between each region in the feature map after local feature aggregation by weighted summation of each position of the feature map after local feature aggregation, to generate the feature map after global attention refinement.

[0048] It should be noted that the SelfAttention is used to capture long-distance dependencies, which can effectively model the relationship between each region in the feature map after local feature aggregation. In the LGLB module, the SelfAttention further improves the global expression ability of the improved RT-DETR network, so that the improved RT-DETR network effectively focuses on the long-range dependencies in the feature map after local feature aggregation.

[0049] It should be noted that the multi-head self-attention calculation of low-rank projection is used to obtain the feature map after global attention refinement , denoted as: ; ; ; ; ; wherein, denotes the query, denotes the weight matrix of the query, denotes the key, denotes the weight matrix of the key, denotes the value, denotes the weight matrix of the value, denotes the th attention head, denotes the output of splicing multiple attention heads, denote the learnable parameters that map the multi-head output after concatenation to the final feature space.

[0050] It should be noted that, , and are the core of the self-attention mechanism in the Transformer, which are used to calculate the attention score respectively.

[0051] It should be noted that, , , and are learnable parameters, which are automatically learned by backpropagation and optimizer on the labeled data during the training of the improved RT-DETR. The dimensions of the learnable parameters are determined according to the network structure requirements of the input feature channel number, the number of attention heads, and the output mapping dimension.

[0052] The feature map after refining the global attention is input into the MLP for dimension expansion and dimension reduction operation to generate the feature map after MLP processing.

[0053] Further, the feature map after refining the global attention is input into the MLP, the channel number of the feature map after refining the global attention is expanded through dimension expansion operation, the features in the feature map after refining the global attention are transformed through the full connection layer, and the expanded feature map channel number is restored to the original dimension through dimension reduction operation to generate the feature map after MLP processing.

[0054] It should be noted that the expansion of the channel dimension is determined by mlp_ratio.

[0055] It should be noted that the MLP is used in the LGLB module to enhance the nonlinear mapping ability of the improved RT-DETR network, and to enhance the exploration ability of the improved RT-DETR network to the feature space.

[0056] The feature map after local feature aggregation, the feature map after refining the global attention, and the feature map after MLP processing are subjected to gated tensor operation, and are fused through the decomposable constraint of channel attention and spatial attention to generate an enhanced feature map.

[0057] It should be noted that through the gated tensor operation, the gated fusion and residual stability are realized, and the gated tensor satisfies the decomposable constraint of channel attention and spatial attention.

[0058] The gated tensor operation is represented as: ; wherein, denotes batch normalization, denotes the weight parameter for adjusting the channel dimension feature, Indicates global average pooling. This represents the convolution weight parameters.

[0059] It should be noted that, During the training process of the improved RT-DETR network, the backpropagation algorithm and optimizer learning settings on labeled data are adjusted according to the response relationship of channel dimension features during gated fusion. By learning settings for the backpropagation algorithm and optimizer on labeled data during the training process of the improved RT-DETR network, the spatial feature distribution in gated fusion is adjusted.

[0060] It should be noted that the LGLB module handles intermediate feature maps. After performing local feature aggregation, global modeling, and gated fusion, the changes in the intermediate feature map are represented as follows: ; ; ; ; ; in, A multi-scale kernel set representing the aggregation of local features. Weight coefficients representing local features Indicates a specific scale, This represents the intermediate feature map. Represents the attention matrix. Indicates the transpose of the key. This indicates the dimensions of the query and the key. This represents the Sigmoid activation function. Represents the gate weight matrix. This represents the final output feature. Represents convolution. This indicates element-wise multiplication.

[0061] It should be noted that, The training process is set through multi-scale local feature aggregation. The value of is determined by the response intensity of each scale feature to image details; The training process is set up through gating tensor operations. The value of is determined based on the decomposable constraints of channel attention and spatial attention.

[0062] It should be noted that, , , Depend on Obtained through linear mapping.

[0063] It should be noted that the LGLB module is a deep learning module combining local feature aggregation and global self-attention mechanism, and the LGLB module is mainly used to extract deep local and global features from the input pre-processed traditional Chinese medicine RGB image sequence. Through local feature aggregation and self-attention mechanism, the LGLB module can effectively capture the detailed information of different regions in the image and enhance the representation of global context information.

[0064] The enhanced feature map is upsampled by the C2D module to generate a high-resolution enhanced feature map.

[0065] The enhanced feature map is input into the C2D module and expanded by the zero interpolation upsampling operator to generate a sparse upsampled enhanced feature map.

[0066] Further, the enhanced feature map is input into the C2D module, the resolution is expanded by the zero interpolation upsampling operator to generate a sparse upsampled enhanced feature map.

[0067] It should be noted that the resolution is expanded by the zero interpolation upsampling operator Us to generate a sparse upsampled enhanced feature map, which is represented as: ; wherein, the upsampling operator is used to expand the input feature .

[0068] It should be noted that the upsampling operator satisfies: ; ; wherein, the new position of the original coordinate (x, y) after upsampling is represented as (x', y'), and (x', y') represents the pixel coordinate in the new map.

[0069] The sparse upsampled enhanced feature map is mapped to the frequency domain by Fourier transform, and the frequency domain filter is filtered by the optical transfer function to generate the frequency domain filtered enhanced feature map.

[0070] Further, the sparse upsampled enhanced feature map is mapped to the frequency domain by Fourier transform, the optical transfer function OTF is constructed, the frequency domain filter is filtered by the optical transfer function, and the frequency domain enhanced feature is obtained by the Wiener type recovery with a stabilizing term ε>0 to generate the frequency domain filtered enhanced feature map.

[0071] ​​​It should be noted that the Fourier transform for mapping the enhanced feature map after sparse up-sampling to the frequency domain adopts a multi-dimensional fast Fourier transform; the optical transfer function OTF is derived from a point spread function PSF.

[0072] It should be noted that the multi-dimensional fast Fourier transform is denoted as: ; wherein, denotes the up-sampled feature after fast Fourier operation, denotes the fast Fourier operation.

[0073] It should be noted that the optical transfer function OTF is denoted as: ; wherein, denotes a function of zero padding and spectrum shift, denotes the padding size in the vertical direction, denotes the padding size in the horizontal direction.

[0074] It should be noted that, for ensuring the correctness of the fast Fourier transform (FFT), zero padding is performed on the PSF, and a shift operation is performed as needed to meet the appropriate matching of the actual optical transfer function, and for controlling the number of pixels filled in the operation.

[0075] The frequency domain enhanced feature is obtained by using the Wiener type recovery with a stabilizing term ε>0, and the result is as follows: ; ; wherein, denotes the feature map after frequency domain enhancement, denotes the optical transfer function, denotes the Fourier transform of the OTF, denotes a stabilizing term for avoiding division by zero error, ensuring that the denominator is not zero, and avoiding numerical instability, denotes the recovered spatial domain image, denotes the real part of the Fourier transform result, denotes the inverse Fourier transform.

[0076] The enhanced feature map after frequency domain filtering is input into the PadShift operation, the point spread function is centrally circularly shifted and zero-padded, and the enhanced feature map after PadShift processing is generated.

[0077] Further, the enhanced feature map after frequency domain filtering is input into a PadShift operation, the point spread function is centrally circularly shifted and zero-padded to (Hs, Ws) to generate an enhanced feature map after PadShift processing.

[0078] It should be noted that the point spread function is centrally circularly shifted and zero-padded to (Hs, Ws) to generate an enhanced feature map after PadShift processing, denoted as: ; wherein, denotes the enhanced feature map after PadShift processing, denotes the circular shift, denotes the zero padding operation.

[0079] It should be noted that by ensuring that the OTF and F(X↑) are phase-aligned under the circular convolution model and reducing boundary artifacts.

[0080] The enhanced feature map after PadShift processing is subjected to frequency domain restoration, high frequency amplification and noise suppression are performed through frequency band weight and channel weight, and a frequency domain weight restored enhanced feature map is generated.

[0081] It should be noted that the enhanced feature map after PadShift processing is subjected to frequency domain restoration, denoted as: ; wherein, denotes the enhanced feature map after frequency domain weight restoration, denotes a constant factor for adjusting the amplitude of the restoration result, denotes a regularization term for preventing overfitting or instability in the restoration process, and smoothing the optimization process.

[0082] It should be noted that, can be learned or set according to the prior spectrum to balance high frequency amplification and noise suppression.

[0083] The enhanced feature map after frequency domain weight restoration is input into a block rearrangement operator to perform local frequency domain aggregation, and is restored through an inverse rearrangement operation to generate a high-resolution enhanced feature map.

[0084] It should be noted that the enhanced feature map after frequency domain weight restoration is input into a block rearrangement operator to perform local frequency domain aggregation, denoted as: ; ; wherein, denotes the enhanced feature map, denotes the rearrangement operation, denotes the feature map after the frequency domain enhancement processing and is finally restored to the spatial domain, denotes the frequency domain, denotes the spatial adaptive weight.

[0085] It should be noted that, The training process is set by the frequency domain recovery, The value of is determined according to the difference between the spatial position and the frequency band response.

[0086] It should be noted that, by inverse rearrangement operation, a high-resolution enhanced feature map is generated, denoted as: ; wherein, denotes the inverse rearrangement operation.

[0087] The high-resolution enhanced feature map is decoded to generate a detection result.

[0088] The detection result includes a bounding box containing spatial position information and class prediction.

[0089] The bounding box regression operation is performed on the high-resolution enhanced feature maps of different scales to obtain the bounding box containing spatial position information.

[0090] Further, the RTDETRDecoder receives the high-resolution enhanced feature maps (P3, P4, P5, etc.) from different scales and decodes them, and uses a regression algorithm to predict the bounding box (bounding box) of each target, each bounding box is represented by four parameters (x, y, w, h), which are the center coordinates, width and height of the bounding box.

[0091] The high-resolution enhanced feature maps of different scales are input into the classification head to generate class probabilities.

[0092] Further, the classification head classifies the high-resolution enhanced feature maps (P3, P4, P5, etc.) of different scales, and outputs the category to which the traditional Chinese medicine decoction piece belongs, It should be noted that, by the classification head output, denoted as: ; wherein, denotes the probability of each category, denotes the combination of global pooling and fully connected mapping.

[0093] Further, according to the detection result, the high-resolution enhanced feature maps of different scales are drawn with bounding boxes and classification labels, and the detection result is fed back to the database for storage and quality control.

[0094] The embodiment also provides a traditional Chinese medicine decoction piece intelligent recognition and classification system based on deep learning, comprising: a preprocessing module that collects traditional Chinese medicine RGB images, performs denoising, normalization and uniform size processing, and generates a traditional Chinese medicine RGB image sequence; A network training module trains the improved RT-DETR network through the labeled traditional Chinese medicine RGB image sequence, continuously optimizes network parameters, and obtains the trained improved RT-DETR network; A feature extraction module calls the LGLB module in the trained improved RT-DETR network, performs deep feature extraction on the traditional Chinese medicine RGB image sequence, and generates an enhanced feature map; An up-sampling module receives the enhanced feature map, performs up-sampling and detail recovery through C2D operation, and generates a high-resolution enhanced feature map; A classification and detection module inputs the high-resolution enhanced feature map into a decoder, and generates a detection result of traditional Chinese medicine decoction piece recognition.

[0095] To sum up, the present application effectively processes the images of traditional Chinese medicine decoction pieces by combining the LGLB module and the C2D module, automatically extracts and classifies the detail features of traditional Chinese medicine decoction pieces, efficiently and accurately recognizes and classifies traditional Chinese medicine decoction piece images of different scales and angles through the multi-level feature extraction, composite scaling deep neural network architecture and frequency domain operation technology, and improves the accuracy and robustness of the traditional Chinese medicine decoction piece intelligent recognition system.

[0096] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning, characterized by: include, Acquire RGB images of Chinese medicinal herbs, preprocess the RGB images of Chinese medicinal herbs, and obtain RGB image sequences of Chinese medicinal herbs; The RGB image sequence of Chinese herbal medicine is input into the improved RT-DETR network, and the LGLB module is used to extract features from the RGB image sequence of Chinese herbal medicine to obtain enhanced feature maps; The enhanced feature map is upsampled using the C2D module to generate a high-resolution enhanced feature map; The high-resolution enhanced feature map is decoded to generate the detection results.

2. The method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning as described in claim 1, characterized in that: The preprocessing includes denoising, normalization, image resizing, and expanding the training dataset.

3. The method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning as described in claim 2, characterized in that: The steps for acquiring RGB images of traditional Chinese medicine, preprocessing these RGB images, and obtaining RGB image sequences of traditional Chinese medicine are as follows. Acquire RGB images of traditional Chinese medicine using high-resolution image acquisition equipment; Denoising is performed on the RGB images of traditional Chinese medicine using a denoising algorithm; The pixel values ​​of the denoised RGB images of Chinese medicine are normalized, and the normalized RGB images of Chinese medicine are adjusted to a uniform size. Image enhancement processing is performed on the RGB images of Chinese herbal medicines after they have been standardized in size to obtain a preprocessed RGB image sequence of Chinese herbal medicines.

4. The method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning as described in claim 3, characterized in that: The improved RT-DETR network includes a backbone network, an LGLB module, a C2D module, and an RTDETRDecoder.

5. The method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning as described in claim 4, characterized in that: The training process of the improved RT-DETR network includes the following steps: Set up the training experimental environment, the input size of the training RGB image sequence of Chinese herbal medicine, and the training hyperparameters; mAP50 and mAP50:95 are used as the main performance indicators for the improved RT-DETR network; The improved RT-DETR network was comprehensively evaluated on the validation set using Precision and Recall. By employing a joint loss of classification cross-entropy and feature regularization, the parameters of the improved RT-DETR network are trained end-to-end to obtain the trained improved RT-DETR network.

6. The method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning as described in claim 5, characterized in that: The steps for extracting features from the RGB image sequence of traditional Chinese medicine using the LGLB module to obtain enhanced feature maps are as follows. Local feature aggregation operation is performed on the preprocessed RGB image sequence of Chinese herbal medicine to obtain the feature map after local feature aggregation. Multi-head self-attention calculation with low-rank projection is performed on the feature map after local feature aggregation to generate a feature map after global attention refinement. The feature map refined by global attention is input into the MLP for dimensionality upscaling and dimensionality downscaling operations to generate the feature map after MLP processing. The feature map after local feature aggregation, the feature map after global attention refinement, and the feature map after MLP processing are subjected to gating tensor operations. Gating fusion is performed through the decomposable constraints of channel attention and spatial attention to generate an enhanced feature map.

7. The method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning as described in claim 6, characterized in that: The steps for upsampling the enhanced feature map using the C2D module to generate a high-resolution enhanced feature map are as follows: The enhanced feature map is input into the C2D module and expanded by the zero-interpolation upsampling operator to generate a sparsely upsampled enhanced feature map. By using Fourier transform, the sparsely upsampled enhanced feature map is mapped to the frequency domain, and frequency domain filtering is performed using the optical transfer function to generate the frequency domain filtered enhanced feature map. The enhanced feature map after frequency domain filtering is input into the PadShift operation, and the point spread function is centered cyclically shifted and zero-filled to generate the enhanced feature map processed by PadShift. The enhanced feature map processed by PadShift is restored in the frequency domain. High-frequency amplification and noise suppression are performed by frequency band weights and channel weights to generate an enhanced feature map after frequency domain weight restoration. By inputting the enhanced feature map after frequency domain weight recovery into the block rearrangement operator for local frequency domain aggregation, and then restoring it through inverse rearrangement operation, a high-resolution enhanced feature map is generated.

8. The method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning as described in claim 7, characterized in that: The detection results include bounding boxes containing spatial location information and category predictions.

9. The method for intelligent identification and classification of traditional Chinese medicine decoction pieces based on deep learning as described in claim 8, characterized in that: The steps for decoding the high-resolution enhanced feature map and generating the detection result are as follows: Boundary box regression operations are performed on high-resolution enhanced feature maps at different scales to obtain bounding boxes containing spatial location information; High-resolution enhanced feature maps at different scales are input into the classification head to generate class probabilities.

10. A deep learning-based intelligent identification and classification system for traditional Chinese medicine decoction pieces, based on the deep learning-based intelligent identification and classification method for traditional Chinese medicine decoction pieces according to any one of claims 1 to 9, characterized in that: include, The preprocessing module acquires RGB images of traditional Chinese medicine, performs denoising, normalization, and size unification processing, and generates a sequence of RGB images of traditional Chinese medicine. The network training module trains the improved RT-DETR network using labeled RGB image sequences of traditional Chinese medicine, continuously optimizes the network parameters, and obtains the trained improved RT-DETR network. The feature extraction module calls the LGLB module in the trained improved RT-DETR network to perform deep feature extraction on the RGB image sequence of traditional Chinese medicine and generate enhanced feature maps. The upsampling module receives the enhanced feature map, performs upsampling and detail restoration through C2D operations, and generates a high-resolution enhanced feature map. The classification and detection module inputs high-resolution enhanced feature maps into the decoder to generate detection results for the identification of Chinese herbal medicine pieces.

Citation Information

Patent Citations

  • Traditional Chinese medicine decoction piece identification method and system based on deep residual network

    CN113361564A

  • Traditional Chinese medicine decoction piece image recognition method and system based on bidirectional multi-scale module

    CN120411696A

  • Data processing method and device based on multistage contrast learning, equipment and medium

    CN120544204A

  • Multi-scale context enhancement small target detection method based on improved RT-DETR

    CN120599503A

  • Land change detection method combining matrix decomposition and adaptive propagation, and system

    WO2025111921A1