Traditional Chinese medicine image recognition method and system, electronic equipment and storable medium

By constructing and processing Chinese medicinal material slice data, extracting multiple features and optimizing the anchor frame generation mechanism, the problems of low recognition efficiency, high error recognition rate and poor recognition accuracy of Chinese medicinal material slice slices are solved, and efficient and accurate recognition effect is achieved.

CN120472438APending Publication Date: 2025-08-12JIANGXI FAR EAST PHARMA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510563431.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing image recognition technology has low efficiency, high error recognition rate and poor recognition accuracy in the recognition of Chinese medicinal materials slice recognition, and large model parameters, making the operation cumbersome.

Method used

The image collection data of Chinese medicinal material slices are constructed, pre-processed and enhanced, color, shape and texture features are extracted, and the image recognition model is built using a two-stage object detection algorithm. Through multi-feature fusion and adjustment of the anchor box generation mechanism, the candidate boxes are optimized using the improved object detection algorithm.

Benefits of technology

It improves the recognition accuracy and efficiency of Chinese medicinal material slices, reduces model calculations, reduces redundant anchor frames, avoids missed detection, and improves recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472438A_ABST
    Figure CN120472438A_ABST
Patent Text Reader

Abstract

The invention provides a traditional Chinese medicine image recognition method and system, electronic equipment and a storage medium, and belongs to the field of image data reading and recognition. The method comprises the following steps: constructing picture set data of traditional Chinese medicinal material slices; preprocessing the picture set data to obtain preprocessed data; enhancing the preprocessed data; extracting image features of the preprocessed data after enhancement processing; constructing an image recognition model by adopting a double-stage target detection algorithm; fusing the image features through a multi-feature fusion strategy to obtain multi-vector fusion features, and carrying out convolution on the feature extraction network; adjusting an anchor frame generation mechanism of the regional recommendation network; an improved target detection algorithm is adopted to iteratively optimize candidate boxes generated by the adjusted region recommendation network; and inputting a to-be-identified traditional Chinese medicinal material image into the optimized image identification model to output an identification result. According to the invention, the defects of low recognition efficiency, high error recognition rate, poor recognition precision and the like of the traditional Chinese medicinal material slices recognized by the existing machine vision can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image data reading and recognition, and specifically relates to a method, system, electronic device, and storage medium for recognizing images of traditional Chinese medicine. (G06V) Background Art

[0002] Traditional Chinese medicine (TCM) resources are plentiful and diverse. Naturally, they are mostly derived from herbal plants, with a few also derived from minerals, animals, chemicals, or biological products. Due to this rich variety, many herbs share similar appearances, making their identification and classification difficult. Traditional TCM identification methods rely heavily on manual judgment based on the clinical experience of pharmacists. This approach relies heavily on the subjective perceptions and experience of professionals, resulting in significant drawbacks such as high workload, low efficiency, and unsuitability for large-scale testing.

[0003] With the development of information technology, machine learning technology has made rapid progress, and image recognition technology has achieved significant success in the biomedical field. In recent years, image recognition technology has been widely used for the identification and classification of traditional Chinese medicines. However, currently used image recognition technology requires that the Chinese medicines being identified have distinct characteristics and regional characteristics, and does not cover a wide range of Chinese medicines. Even if a wide range of types are covered, the recognition accuracy is still unsatisfactory. Furthermore, Chinese medicinal material slices come in many different types and have high similarity. Using a single feature can only express a specific characteristic of the image and cannot take into account multiple features such as color, shape, and texture, resulting in a high recognition error rate. In addition, the detection model parameters of existing image recognition technologies are large, resulting in a large amount of model calculations, cumbersome operation, and low recognition efficiency.

[0004] Therefore, how to improve the shortcomings of existing machine vision in identifying Chinese medicinal material slices, such as low recognition efficiency, high misrecognition rate and poor recognition accuracy, is an urgent issue to be solved by technical personnel in this field. Summary of the Invention

[0005] In order to solve at least one of the above technical problems, the present invention provides a method, system, electronic device and storage medium for traditional Chinese medicine image recognition.

[0006] In a first aspect, the invention provides a method for identifying images of Chinese medicinal materials, comprising: Construct image collection data of Chinese herbal medicine slices; Preprocessing the picture set data to obtain preprocessed data; enhancing processing of the pre-processed data; Extracting image features of the pre-processed data after enhancement processing, wherein the image features include color features, shape features and texture features; A two-stage target detection algorithm is used to build an image recognition model, where the image recognition model includes a feature extraction network and a region recommendation network; fusing the image features through a multi-feature fusion strategy to obtain multi-vector fusion features, and convolving the feature extraction network; Adjusting the anchor box generation mechanism of the region recommendation network; Iteratively optimizing the candidate boxes generated by the adjusted region recommendation network using an improved target detection algorithm; The image of the Chinese medicinal material to be identified is input into the optimized image recognition model to output the recognition result.

[0007] Preferably, the step of preprocessing the picture set data to obtain preprocessed data specifically includes: Performing normalization processing on the picture set data using a linear function conversion algorithm; Gray-scaling the normalized image set data using a weighted average algorithm; The grayscaled picture set data is denoised using a median filtering algorithm.

[0008] Preferably, the step of enhancing the pre-processed data specifically includes: Randomly flipping and rotating the filtered pre-processed data; Selecting a scaling factor within a predetermined range to scale the pre-processed data after flipping and rotating; A color dithering factor is used to add Gaussian noise to the pre-processed data after the scaling process.

[0009] Preferably, the step of extracting image features of the pre-processed data after enhancement processing specifically includes: Extracting the color feature vector of the pre-processed data after enhancement processing by a spatial transformation algorithm, and performing normalization processing to obtain color features; Using a histogram of oriented gradients algorithm to extract shape features of the pre-processed data after enhancement processing; The texture features of the pre-processed data after enhancement processing are extracted based on a local binary pattern algorithm.

[0010] Preferably, the steps of fusing the image features to obtain multi-vector fusion features through a multi-feature fusion strategy and convolving the feature extraction network specifically include: A multi-feature fusion algorithm is used to fuse the image features to obtain a multi-vector fusion feature; Performing a convolution operation on the feature extraction network in a bottom-up manner to obtain a corresponding feature layer; The feature layer is respectively subjected to 1×1 and 3×3 convolution processing with the multi-vector fusion feature in sequence to output the corresponding feature map; The aliasing parts of adjacent feature maps are eliminated through 3×3 convolution.

[0011] Preferably, the step of adjusting the anchor box generation mechanism of the region recommendation network specifically includes: Utilizing the region recommendation network to perform a traversal operation on the multi-vector fusion feature to generate a series of anchor frames of different sizes; Setting the aspect ratio and size of the corresponding anchor frame according to the scale of each feature layer to generate the corresponding candidate frame; A prefabricated adjustment algorithm and the width and height of the candidate frame are combined to make the candidate frame correspond to the corresponding feature layer one by one.

[0012] Preferably, the step of using the improved target detection algorithm to iteratively optimize and adjust the candidate boxes generated by the region recommendation network specifically includes: Sort the generated candidate boxes in descending order according to their confidence scores, and find the candidate box with the highest score; Using an improved target detection algorithm to compare the score of the candidate box with the highest score and the degree of overlap with its neighboring candidate boxes; Whether to change the confidence score of the candidate box is determined based on whether the overlap threshold is exceeded. Repeat the above steps to find the optimal candidate box.

[0013] In a second aspect, a Chinese medicinal material image recognition system includes: Construction module, used to construct image set data of Chinese herbal medicine slices; A preprocessing module, configured to preprocess the image set data to obtain preprocessed data; An enhancement module, configured to enhance processing of the pre-processed data; An extraction module, configured to extract image features of the pre-processed data after enhancement processing, wherein the image features include color features, shape features, and texture features; A building module for building an image recognition model using a two-stage object detection algorithm, wherein the image recognition model includes a feature extraction network and a region recommendation network; A fusion module, configured to fuse the image features through a multi-feature fusion strategy to obtain multi-vector fusion features, and convolve the feature extraction network; An adjustment module, configured to adjust the anchor box generation mechanism of the region recommendation network; An iterative module, configured to iteratively optimize the candidate boxes generated by the adjusted region recommendation network using an improved object detection algorithm; The recognition module is used to input the image of the Chinese medicinal material to be recognized into the optimized image recognition model to output the recognition result.

[0014] Preferably, the pre-processing module specifically includes: A normalization unit, configured to perform normalization processing on the picture set data using a linear function conversion algorithm; A grayscale unit, configured to grayscale the normalized picture set data using a weighted average algorithm; The filtering unit is used to perform denoising on the grayscaled picture set data using a median filtering algorithm.

[0015] Preferably, the enhancement module specifically includes: A flipping unit, configured to randomly flip and rotate the pre-processed data after screening; a scaling unit, configured to select a scaling factor within a predetermined range to scale the pre-processed data after flipping and rotating; The noise adding unit is used to add Gaussian noise to the pre-processed data after the scaling process by using a color dithering factor.

[0016] Preferably, the extraction module specifically includes: A color extraction unit, configured to extract the color feature vector of the pre-processed data after enhancement processing by a spatial conversion algorithm, and perform normalization processing to obtain color features; a shape extracting unit, configured to extract shape features of the pre-processed data after enhancement processing by using a histogram of directional gradients algorithm; The texture extraction unit is used to extract texture features of the pre-processed data after enhancement based on a local binary pattern algorithm.

[0017] Preferably, the fusion module specifically includes: A fusion unit, configured to fuse the image features using a multi-feature fusion algorithm to obtain a multi-vector fusion feature; A convolution unit, configured to perform a convolution operation on the feature extraction network in a bottom-up manner to obtain a corresponding feature layer; A convolution unit is used to process the feature layer with the multi-vector fusion feature in sequence by 1×1 and 3×3 convolution to output the corresponding feature map; The elimination unit is used to eliminate the aliasing parts of adjacent feature maps through 3×3 convolution.

[0018] Preferably, the adjustment module specifically includes: A traversal unit, configured to perform a traversal operation on the multi-vector fusion feature using the region recommendation network to generate a series of anchor frames of different sizes; A setting unit, configured to set the aspect ratio and size of the corresponding anchor frame according to the scale of each feature layer to generate a corresponding candidate frame; An adjustment unit is used to combine a prefabricated adjustment algorithm and the width and height of the candidate frame to make the candidate frame correspond to the corresponding feature layer one by one.

[0019] Preferably, the iteration module specifically includes: The descending unit is used to sort the generated candidate boxes in descending order according to the confidence score and find the candidate box with the highest score; A comparison unit, configured to compare the score of the candidate box with the highest score with the degree of overlap of its neighboring candidate boxes using an improved object detection algorithm; A conditional unit, used to determine whether to change the confidence score of a candidate box based on whether the overlap threshold is exceeded; The optimization unit is used to repeatedly execute the above steps to find the optimal candidate box.

[0020] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method for recognizing images of Chinese medicinal materials as described in the first aspect is implemented.

[0021] In a fourth aspect, the present application provides a storable medium having a computer program stored thereon, which, when executed by a processor, implements the method for recognizing images of traditional Chinese medicines as described in the first aspect.

[0022] Compared with the prior art, the present invention has the following beneficial effects: 1. Preprocess the image data to remove artifacts, reduce the impact of minor information on the image, enhance important information, and retain key features, thereby improving image quality. Furthermore, image data augmentation is the most convenient way to expand the data size. The data before and after augmentation is highly correlated, and this can address overfitting issues during subsequent training.

[0023] 2. Based on the low accuracy of automatic recognition of Chinese medicinal material slices in complex backgrounds, in order to further improve the accuracy and efficiency of recognition, an extraction algorithm for Chinese medicinal material slice images suitable for complex backgrounds is analyzed and established to extract the color, texture and shape features of visual features, providing an important basis for subsequent image classification.

[0024] 3. Considering that when using convolutional neural networks to extract image features, the acquired feature information will continue to be lost as the number of network layers increases, and considering that Chinese medicinal materials are of many types and have high similarity, using a single feature can only express a certain characteristic of the image. Multi-feature fusion processing will provide a more comprehensive description of the image features, taking into account multiple features such as color, shape, and texture, thus avoiding the disadvantage of reduced feature information obtained by high-level networks.

[0025] 4. By adjusting the anchor frame generation mechanism of the region recommendation network, it can be made to better correspond to the input feature maps of corresponding scales, rather than using the same size and proportion to generate a large number of redundant anchor frames on all feature maps. The adjusted anchor frames can achieve efficient detection of Chinese medicinal materials of different scales, effectively improving detection results while reducing the number of network parameters.

[0026] 5. Since the region recommendation network generates a large number of candidate boxes on multiple input feature maps, but not all candidate boxes contain valid feature information, the use of an improved target detection algorithm can eliminate the redundant selection boxes generated in the local range and avoid the missed detection phenomenon that may occur in conventional target detection algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 Flowchart of the method for recognizing images of Chinese medicinal materials provided in Example 1 of the present invention; Figure 2 This is a structural block diagram of a Chinese medicinal material image recognition system corresponding to the method of Example 1 and provided in Example 2 of the present invention; Figure 3 3 is a schematic diagram of the hardware structure of the computer provided in Example 3 of the present invention.

[0029] Description of reference numerals: 10-Building modules.

[0030] 20-preprocessing module, 21-normalization unit, 22-grayscale unit, 23-filtering unit.

[0031] 30-enhancement module, 31-rotation unit, 32-scaling unit, 33-noise adding unit.

[0032] 40 - extraction module, 41 - color extraction unit, 42 - shape extraction unit, 43 - texture extraction unit.

[0033] 50-Build the module.

[0034] 60- fusion module, 61- fusion unit, 62- convolution unit, 63- mixed product unit, 64- elimination unit.

[0035] 70-adjustment module, 71-traversal unit, 72-setting unit, 73-adjustment unit.

[0036] 80-iteration module, 81-descending unit, 82-comparison unit, 83-condition unit, 84-optimization unit.

[0037] 90-Identification module.

[0038] 100 - bus, 101 - processor, 102 - memory, 103 - communication interface. DETAILED DESCRIPTION

[0039] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.

[0040] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0041] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0042] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0043] Traditional Chinese medicine (TCM) applications rely heavily on the judgment and identification of professional technicians based on their own experience. With the widespread use of TCM, identification has evolved from empirical to objective. With the development of information technology, image recognition technology has achieved significant success in the biomedical field. Applying image recognition technology to the identification and classification of TCMs, computer-generated data enhancement and feature extraction methods and techniques are used to create a complete and clear image description.

[0044] Example 1 Specifically, Figure 1 FIG. 1 is a flow chart of the method for recognizing images of Chinese medicinal materials provided in this embodiment.

[0045] like Figure 1 As shown, the Chinese medicinal material image recognition method of this embodiment includes the following steps: S101, constructing image set data of Chinese medicinal material slices.

[0046] Specifically, in the field of image recognition of slices of Chinese medicinal materials, there is currently no public standard data set. This embodiment uses shooting and crawler technology to obtain image data of slices of Chinese medicinal materials. Among them, the device used for shooting is the rear camera of a mobile phone, and the pictures are taken in places such as herbal medicine stores or museums, and the collected slice images are labeled and classified. At the same time, in order to facilitate storage and calculation, the image size is uniformly adjusted to 512×512 pixels using image processing software. Compared with manually downloading a single image, using a web crawler to capture images of slices of Chinese medicinal materials has many conveniences. The Selenium-based image crawling mode is used to download data from major Chinese medicinal materials websites and Bing pictures. The crawled pictures are screened, and pictures that are unclear or have incorrect information are removed. Then the image size is uniformly adjusted to 512×512 pixels using image processing software.

[0047] S102: Preprocess the picture set data to obtain preprocessed data.

[0048] Specifically, image quality is directly linked to the efficiency and accuracy of detection models. Preprocessing data before image recognition can improve image quality and maximize the effectiveness of detection models. The goal of image preprocessing is to remove artifacts from the image, reduce the impact of minor information on the image, enhance important information, retain key features, and further simplify the data.

[0049] Furthermore, the specific steps of step S102 include: S1021: Perform normalization processing on the picture set data using a linear function conversion algorithm.

[0050] Specifically, image normalization is a widely used image preprocessing technology in the fields of pattern recognition and computer vision. Image normalization is to perform multiple transformations on the original image. The format image after the transformation is a standard form, and the format image after the transformation is not affected by the geometric transformation of the image. This embodiment uses linear function transformation to normalize the image. This is because the linear function transformation method does not involve distance measurement and covariance calculation, is relatively simple, and is suitable for image normalization processing in image processing. Among them, the linear function transformation method used is as follows: y =(x-MinValue) / (MaxValue-MinValue); Where: x, y represent the values before and after conversion, respectively; MaxValue and MinValue are the maximum and minimum values of the samples, respectively.

[0051] S1022: grayscale the normalized picture set data using a weighted average algorithm.

[0052] Specifically, image grayscale processing refers to the process of converting a color image into a grayscale image. It is a dimensionality reduction process that converts a three-dimensional channel into a one-dimensional channel. The grayscale image retains the characteristics of the image to the greatest extent and requires less calculation. The grayscale image, like the color image, still reflects the overall and local characteristics of the entire image. This embodiment uses the weighted average method to grayscale the medicinal material image. According to the importance, the three components R, G, and B in the color image are weighted averaged with different weights. The weights are set according to the different sensitivities of the human eye to the three colors red, green, and blue. The conversion formula is as follows: Gray(i,j)=0.3R(i,j)+0.59G(i,j)+0.11B(i,j); Where R(i,j), G(i,j), and B(i,j) represent the R, G, and B components of the color image, respectively, and Gray(i,j) represents the final grayscale value after grayscale processing.

[0053] S1023, performing denoising on the grayscaled picture set data using a median filtering algorithm.

[0054] Specifically, image data may be interfered with and contaminated by various noises during the acquisition and transmission process, resulting in a decrease in image quality. The purpose of image denoising is to improve the quality of a specified image and solve the problem of image quality degradation caused by noise interference. The quality of the denoising effect is directly related to the effect of subsequent image processing. This embodiment uses median filtering to perform denoising on medicinal material images. Median filtering is also a template filter. Instead of replacing the central target pixel with the mean of the pixels in the template, all pixels in the template are sorted, and the median of the sorted template pixel sequence is taken as the value of the target pixel. Median filtering retains more details of the image, does not cause image blurring problems, and has a particularly good removal effect on salt and pepper noise.

[0055] S103: Enhance the pre-processed data.

[0056] Specifically, after manual screening of data crawled from the Internet, there may be only dozens of images of each Tibetan medicinal material slice, which is far from enough. In addition, the number of images that can be obtained through field photography is very limited and insufficient to support experimental research. Data augmentation can increase training samples and reduce the overfitting phenomenon of the network to a certain extent.

[0057] Furthermore, the specific steps of step S103 include: S1031, randomly flipping and rotating the filtered pre-processed data.

[0058] Specifically, image rotation involves rotating an image clockwise or counterclockwise by a certain angle. This rotation may alter the image's dimensions, potentially leading to image scaling issues. This embodiment uses 90°, 180°, and 270° image rotations, all performed clockwise. This ensures that the image's size after rotation is consistent with the original image.

[0059] S1032: Select a scaling factor within a predetermined range to scale the pre-processed data after flipping and rotating.

[0060] Specifically, image scaling refers to the process of reducing or enlarging the size of an image; enlarging an image usually requires cropping the enlarged image, and the image quality will be reduced, requiring interpolation processing to improve the clarity of the enlarged image. Reducing an image is to reduce the image size by deleting pixels and making assumptions about the boundary content of the image after reduction. When the image is reduced, the image will become clearer. This embodiment specifies a jitter factor d, randomly selects a value s as the brightness scaling factor in the range of max (0, 2-factor) to 1+factor, and multiplies the original image by the scaling factor s.

[0061] S1033: Use a color dithering factor to add Gaussian noise to the pre-processed data after the scaling process.

[0062] Specifically, the purpose of adding Gaussian noise is to make the synthesized image look more realistic, and adding noise in machine learning and deep learning can be used as a data enhancement method to improve the generalization ability of the model. The mathematical expression of Gaussian noise in this embodiment is as follows: N(z,y)=\mu+\sigma\cdot\mathcal{N}(0,1) Where \mu represents the mean, \sigma represents the standard deviation, and \mathcal{N}(0,1) represents the standard normal distribution.

[0063] S104: Extracting image features of the pre-processed data after enhancement.

[0064] The image features include color features, shape features and texture features.

[0065] Specifically, Chinese medicinal material slices are of many types and have high similarity. Using a single feature can only express a certain characteristic of the image. In order to make the feature description of the image more comprehensive, it is necessary to take into account multiple features such as color, shape and texture.

[0066] Furthermore, the specific steps of step S104 include: S1041 , extracting the color feature vector of the pre-processed data after enhancement processing by a spatial transformation algorithm, and performing normalization processing to obtain color features.

[0067] Specifically, color features are an important factor in identifying the type of Chinese herbal medicine slices. This embodiment first uses the RGB color feature method to extract the color features of the image. After encoding, the RGB color space is converted to the HSI color space for image feature vector extraction, and the obtained color feature vector is normalized. Among them, the HSI color space components correspond to the three element colors of human visual characteristics. Each channel is independent of each other, and the changes in each color component can be independently perceived.

[0068] S1042: Using a histogram of oriented gradients algorithm to extract shape features of the pre-processed data after enhancement.

[0069] Specifically, the directional gradient histogram algorithm forms features by calculating and statistically analyzing the directional gradient histogram of the local area of the image. The specific method for shape feature extraction in this embodiment is as follows: first, the image is normalized, and then the color image is gamma compressed. When performing gradient calculation, the gradient calculation is performed on the normalized color image to obtain the horizontal and vertical gradient components, and the gradient amplitude of the current pixel is calculated. After extracting the gradient value of the pixel point, the gradient histogram is calculated for each of the 9 small blocks after the block is divided and vector normalization is performed. The 9 normalized block feature vectors are combined into a directional gradient histogram feature vector.

[0070] S1043: Extracting texture features of the enhanced preprocessed data based on a local binary pattern algorithm.

[0071] Specifically, the Local Binary Patterns algorithm is ideal for extracting texture features from Tibetan medicinal material slice images against complex backgrounds, even when subject to occlusion caused by lighting, angle, or overlapping elements. This embodiment first resizes the Chinese medicinal material slice images to a uniform size before performing texture extraction, which helps fully capture the local features of the Chinese medicinal material slice images against complex backgrounds.

[0072] S105, build an image recognition model using a two-stage target detection algorithm.

[0073] Specifically, the two-stage target detection algorithm first uses a selective search or region recommendation network to generate candidate frames for the image, then uses a feature extraction network to extract features of the target within the candidate frame, and then uses a classifier to perform detection and classification to obtain the detection results. This embodiment adopts the Faster R-CNN algorithm, and uses a feature extraction network to replace the traditional SS selective search method to obtain candidate frames, effectively improving the extraction speed of candidate frames. Among them, the image recognition model includes a feature extraction network and a region recommendation network. The feature extraction network is Faster R-CNN using a convolutional neural network to perform a convolution operation on the input image to obtain image features, and then the generated feature map is used for subsequent classification and prediction. The region recommendation network also performs a convolution operation on the input feature map and generates a series of anchor frames of different sizes and proportions with confidence scores and positive and negative sample categories. Then, after preliminary screening of the generated candidate frames by a set threshold, the N candidate frames with higher scores are output for subsequent calculations.

[0074] S106, fusing the image features through a multi-feature fusion strategy to obtain multi-vector fusion features, and convolving the feature extraction network.

[0075] Specifically, in the process of extracting features from Chinese medicinal materials images, the underlying neural network extracts feature information such as the edges and textures of the medicinal materials. Then, the subsequent network layers extract feature information from other parts of the Chinese medicinal materials. Finally, the Chinese medicinal materials images are abstracted into high-dimensional features in the high-level neural network. This process is accompanied by the loss of feature information from the Chinese medicinal materials images, making it difficult to effectively improve the model recognition effect. In order to better express the image feature information, this embodiment uses a multi-feature fusion strategy to fuse the local feature map with the global feature map. After multiple convolution and pooling operations, the different feature information of the high and low layers can be effectively integrated.

[0076] Furthermore, the specific steps of step S106 include: S1061: Use a multi-feature fusion algorithm to fuse the image features to obtain multi-vector fusion features.

[0077] Specifically, the basic strategy of the multi-feature fusion algorithm is to fuse the feature information extracted by a certain layer in the previous network layer with the feature map of the current layer before extracting the feature information of the Chinese medicinal material image in the current network layer, so that the feature information extracted by the previous network layer can be partially retained, minimizing the loss of feature information. The multi-feature fusion algorithm of this embodiment adopts different weights for different features, and the overall weight of the fused features is 1. The multi-feature fusion algorithm is as follows: F=aF RGB +bF Hoc +cF LBP ; Among them, F represents the fusion feature, F Rcz Indicates color features, F Hoc Represents the shape feature, F Lsp Represents texture features, and a, b, and c represent the weight coefficients of each feature respectively.

[0078] S1062: Perform a convolution operation on the feature extraction network in a bottom-up manner to obtain a corresponding feature layer.

[0079] Specifically, the feature extraction network of this embodiment uses ResNet-50 as the new backbone feature extraction network, which has stronger feature extraction capabilities and fewer parameters. ResNet-50 has a total of five convolutional layers, but considering that the P1 layer contains less feature information, it is composed of the remaining four convolutional layers and three connection channels of the ResNet-50 network. The corresponding feature layer is obtained by performing convolution operations on the ResNet-50 network from bottom to top.

[0080] S1063: Process the feature layer and the multi-vector fusion feature with 1×1 and 3×3 convolutions in sequence to output corresponding feature maps.

[0081] Specifically, the ResNet-50 network convolution obtains the highest layer feature map, and then uses 1×1 convolution on the highest layer feature map to obtain convolution layer 5, and at the same time uses 3×3 convolution on convolution layer 5 to output feature map 5, and then adjusts the fourth layer to the second layer layer by layer to obtain convolution layer 4 to convolution layer 2, and then adds convolution layer 5 to convolution layer 2 in pairs and uses 3×3 convolution to eliminate the aliasing part generated in the feature fusion process to obtain feature map 4 to feature map 5. Figure 2 .

[0082] S1064: Eliminate aliasing parts of adjacent feature maps through 3×3 convolution.

[0083] Specifically, the intermediate feature map convolution layer 5 is downsampled by a factor of two and then a 3×3 convolution is used to eliminate aliasing.

[0084] S107: Adjust the anchor box generation mechanism of the region recommendation network.

[0085] Specifically, by adjusting the anchor frame generation strategy of the region recommendation network, it can be made to better correspond to the input feature maps of the corresponding scale, thereby achieving efficient detection of Chinese medicinal materials of different scales, instead of conventionally using the same size and proportion on all feature maps to generate a large number of redundant anchor frames.

[0086] Furthermore, the specific steps of step S107 include: S1071: Utilize the region recommendation network to perform a traversal operation on the multi-vector fusion feature to generate a series of anchor frames of different sizes.

[0087] Specifically, the region proposal network generates a size and aspect ratio of (128 2 , 256 2 , 512 2 ) and (1:1, 1:2, 2:1). After feature fusion, the output of the image recognition model changes from P5 to five new feature maps {Z2, Z3, Z4, Z5, Z6} with the same number of channels but different scales.

[0088] S1072: Setting the aspect ratio and size of the corresponding anchor frame according to the scale of each feature layer to generate a corresponding candidate frame.

[0089] Specifically, because the five new feature maps mentioned above all contain fused feature information, if the default anchor box generation mechanism of the region recommendation network is used to set multi-scale anchor boxes at each feature layer, it will lead to redundancy in the extracted regions of interest and increase the computational complexity of the model. This embodiment adjusts the anchor box generation mechanism of the region recommendation network and sets the corresponding anchor box aspect ratio and size according to the scale of each feature layer. The specific sizes of each layer {Z2, Z3, Z4, Z5, Z6} are {322, 642, 1282, 2562, 5122}, and the corresponding aspect ratios are (1:2, 1:1, 2:1).

[0090] S1073: Combine a prefabricated adjustment algorithm and the width and height of the candidate frame to make the candidate frame correspond to the corresponding feature layer one by one.

[0091] Specifically, to pool the candidate regions generated by anchor frames of different scales on the {Z2, Z3, Z4, Z5} feature maps onto the original image to obtain a fixed-size feature map, it is necessary to establish a one-to-one correspondence between the candidate region coordinates on the original image and the four feature maps. The adjusted anchor frames enable efficient detection of Chinese medicinal materials of varying scales, significantly improving detection performance while reducing network parameters. The prefabricated adjustment algorithm is as follows: ; Where k0 represents the default reference level number, w and h represent the width and height of the candidate region, input represents the size of the original image input matrix, and k represents the feature map corresponding to the candidate region.

[0092] S108, using an improved object detection algorithm to iteratively optimize the candidate boxes generated by the adjusted region recommendation network.

[0093] Specifically, the region proposal network generates a large number of candidate boxes on multiple input feature maps. Not all candidate boxes contain valid feature information. Therefore, an improved target detection algorithm is needed to eliminate redundant boxes generated in the local range.

[0094] Furthermore, the specific steps of step S108 include: S1081: Arrange the generated candidate boxes in descending order according to their confidence scores, and find the candidate box with the highest score.

[0095] Specifically, the candidate boxes generated by the region recommendation network will have corresponding confidence scores. The higher the score, the closer the candidate box is to the ideal value. This embodiment finds the candidate box with the highest score, the purpose of which is to provide judgment conditions for the subsequent object detection algorithm.

[0096] S1082: Using an improved target detection algorithm, compare the score of the candidate box with the highest score with the degree of overlap of its neighboring candidate boxes.

[0097] Specifically, the traditional target detection algorithm adopts the idea of "either 0 or 1", that is, the candidate frame is either retained or deleted. At this time, the setting of the threshold will have a great impact on the detection result. If the threshold is set too low, the redundant frames cannot be effectively eliminated, and if the threshold is set too high, it is easy to cause the adjacent Chinese medicinal materials to be missed. The improved target detection algorithm adopted in this embodiment does not directly delete the overlapping candidate frames, but multiplies a weight coefficient based on the candidate frame score. The score of the adjacent candidate frame that highly overlaps with the candidate frame with the highest score can be reduced by weighted calculation, and the score is inversely proportional to the degree of overlap. Among them, the improved target detection algorithm is as follows: ; Where s f Indicates the score of the candidate box, M represents the candidate box with the highest score, IoU(M,b i ) represents the overlap between the highest-scoring candidate box and its neighboring candidate boxes, N t Indicates the set IoU threshold.

[0098] S1083: Determine whether to change the confidence score of the candidate box based on whether the confidence score exceeds a set overlap threshold.

[0099] Specifically, the improved target detection algorithm calculates the neighboring candidate boxes with M. If the two candidate boxes only overlap slightly, their confidence scores will not be changed. If the overlap with M is high, their confidence scores will be linearly attenuated.

[0100] S1084: Repeat the above steps to find the optimal candidate box.

[0101] Specifically, through continuous iteration, the anchor frames of the region of interest are screened to reduce redundant anchor frames, increase positioning accuracy, and thus enhance the detection capability of the image recognition model.

[0102] S109: Input the image of the Chinese medicinal material to be identified into the optimized image recognition model to output a recognition result.

[0103] Specifically, by conducting relevant experiments and tests on the optimized image recognition model, the effectiveness of the improvement and optimization measures was proved, and the average recognition accuracy of the image recognition model was increased by 2.7%, which is a significant improvement compared with the existing image recognition model; and the optimized image recognition model has good detection capabilities and robustness in test environments under different conditions.

[0104] In summary, first, capturing and crawling images together to collect as much image data as possible. Preprocessing reduces the impact of secondary image information, and image data enhancement is performed to expand the data size and acquire more image data. Second, through analysis, a feature extraction algorithm suitable for Chinese herbal medicine slice images in complex backgrounds is established to extract visual features such as color, texture, and shape, providing an important basis for subsequent image classification. Third, multi-feature fusion processing provides a more comprehensive description of image features, taking into account multiple features such as color, shape, and texture, thereby avoiding the drawback of reduced feature information obtained by high-level networks. Third, by adjusting the anchor box generation mechanism of the region recommendation network, efficient detection of Chinese herbal medicines of different scales is achieved, effectively improving detection performance while reducing network parameters. Finally, an improved object detection algorithm eliminates redundant selection boxes generated within a local area and avoids the potential for missed detections that can occur with conventional object detection algorithms. These steps effectively address the drawbacks of existing machine vision methods in identifying Chinese herbal medicine slices, such as low recognition efficiency, high false positive rates, and poor recognition accuracy.

[0105] Example 2 This embodiment provides a structural block diagram of a system corresponding to the method described in Embodiment 1. Figure 2 is a structural block diagram of the Chinese medicinal material image recognition system according to this embodiment, such as Figure 2 As shown, the system includes: The construction module 10 is used to construct the image set data of the slices of Chinese medicinal materials.

[0106] The preprocessing module 20 is configured to preprocess the picture set data to obtain preprocessed data.

[0107] The enhancement module 30 is configured to enhance the pre-processed data.

[0108] The extraction module 40 is used to extract image features of the pre-processed data after enhancement processing.

[0109] The image features include color features, shape features and texture features.

[0110] The building module 50 is used to build an image recognition model using a two-stage target detection algorithm.

[0111] The image recognition model includes a feature extraction network and a region recommendation network.

[0112] A fusion module 60 is configured to fuse the image features using a multi-feature fusion strategy to obtain multi-vector fusion features and convolve the feature extraction network; An adjustment module 70, configured to adjust the anchor box generation mechanism of the region recommendation network; Iterative module 80, configured to iteratively optimize the candidate boxes generated by the adjusted region recommendation network using an improved object detection algorithm; The recognition module 90 is used to input the image of the Chinese medicinal material to be recognized into the optimized image recognition model to output a recognition result.

[0113] Furthermore, the pre-processing module 20 specifically includes: A normalization unit 21, configured to perform normalization processing on the picture set data using a linear function conversion algorithm; A grayscale unit 22 is configured to grayscale the normalized picture set data using a weighted average algorithm; The filtering unit 23 is configured to perform denoising on the grayscaled picture set data using a median filtering algorithm.

[0114] Furthermore, the enhancement module 30 specifically includes: A flipping unit 31 is used to randomly flip and rotate the pre-processed data after screening; A scaling unit 32 is configured to select a scaling factor within a predetermined range to scale the pre-processed data after flipping and rotating; The noise adding unit 33 is configured to add Gaussian noise to the pre-processed data after the scaling process by using a color dithering factor.

[0115] Furthermore, the extraction module 40 specifically includes: A color extraction unit 41 is used to extract the color feature vector of the pre-processed data after the enhancement process by using a spatial conversion algorithm, and perform normalization processing to obtain color features; a shape extracting unit 42, configured to extract shape features of the pre-processed data after enhancement processing by using a histogram of directional gradients algorithm; The texture extraction unit 43 is configured to extract texture features of the enhanced pre-processed data based on a local binary pattern algorithm.

[0116] Preferably, the fusion module 60 specifically includes: A fusion unit 61 is configured to fuse the image features using a multi-feature fusion algorithm to obtain a multi-vector fusion feature; A convolution unit 62 is used to perform a convolution operation on the feature extraction network in a bottom-up manner to obtain a corresponding feature layer; A convolution unit 63 is configured to process the feature layer and the multi-vector fusion feature with 1×1 and 3×3 convolutions in sequence to output corresponding feature maps; The elimination unit 64 is configured to eliminate aliasing portions of adjacent feature maps through a 3×3 convolution.

[0117] Furthermore, the adjustment module 70 specifically includes: A traversal unit 71 is configured to perform a traversal operation on the multi-vector fusion feature using the region recommendation network to generate a series of anchor frames of different sizes; A setting unit 72 is configured to set the aspect ratio and size of the corresponding anchor frame according to the scale of each feature layer to generate a corresponding candidate frame; The adjustment unit 73 is configured to combine a prefabricated adjustment algorithm and the width and height of the candidate frame to ensure a one-to-one correspondence between the candidate frame and the corresponding feature layer.

[0118] Furthermore, the iteration module 80 specifically includes: A descending order unit 81 is used to sort the generated candidate boxes in descending order according to the confidence scores and find the candidate box with the highest score; A comparison unit 82 is configured to compare the score of the candidate box with the highest score with the degree of overlap of its neighboring candidate boxes using an improved object detection algorithm; A conditional unit 83 is used to determine whether to change the confidence score of the candidate box based on the condition that the confidence score exceeds a set overlap threshold; The optimization unit 84 is used to repeatedly execute the above steps to find the optimal candidate frame.

[0119] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0120] Example 3 Combine Figure 1 The described method for recognizing images of traditional Chinese medicines can be implemented by a computer. Figure 3 FIG. 4 is a schematic diagram of the hardware structure of a computer according to this embodiment.

[0121] The computer may include a processor 101 and a memory 102 storing computer program instructions.

[0122] Specifically, the processor 101 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the present application.

[0123] Memory 102 may include a large-capacity memory for data or instructions. By way of example, and not limitation, memory 102 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 102 may include removable or non-removable (or fixed) media. Where appropriate, memory 102 may be internal or external to the data processing device. In certain embodiments, memory 102 is non-volatile memory. In certain embodiments, memory 102 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0124] The memory 102 may be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 101 .

[0125] The processor 101 implements the Chinese medicinal material image recognition method of the above-mentioned embodiment 1 by reading and executing the computer program instructions stored in the memory 102 .

[0126] In some embodiments, the computer may further include a communication interface 103 and a bus 100. Figure 3 As shown, the processor 101 , the memory 102 , and the communication interface 103 are connected via a bus 100 and communicate with each other.

[0127] The communication interface 103 is used to implement communication between the various modules, devices, units and / or equipment in this application. The communication interface 103 can also implement data communication with other components such as: external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.

[0128] The bus 100 includes hardware, software, or both, and couples the components of the computer to each other. The bus 100 includes, but is not limited to, at least one of the following: a data bus (DataBus), an address bus (AddressBus), a control bus (ControlBus), an expansion bus (ExpansionBus), and a local bus (LocalBus). By way of example and not limitation, the bus 100 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, or a 100-bit GPIO bus. The bus 100 may include a bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 100 may include one or more buses. Although this application describes and illustrates a specific bus, this application contemplates any suitable bus or interconnect.

[0129] The computer can obtain the Chinese medicinal material image recognition system and execute the Chinese medicinal material image recognition method of Example 1.

[0130] In addition, in conjunction with the method for recognizing Chinese medicinal materials images in the above-mentioned embodiment 1, the present application may provide a storage medium for implementation. The storage medium stores computer program instructions; when the computer program instructions are executed by a processor, the method for recognizing Chinese medicinal materials images in the above-mentioned embodiment 1 is implemented.

[0131] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0132] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for recognizing images of Chinese medicinal materials, characterized in that: include: Construct image collection data of Chinese herbal medicine slices; Preprocessing the picture set data to obtain preprocessed data; enhancing processing of the pre-processed data; Extracting image features of the pre-processed data after enhancement processing, wherein the image features include color features, shape features and texture features; A two-stage target detection algorithm is used to build an image recognition model, where the image recognition model includes a feature extraction network and a region recommendation network; fusing the image features through a multi-feature fusion strategy to obtain multi-vector fusion features, and convolving the feature extraction network; Adjusting the anchor box generation mechanism of the region recommendation network; Iteratively optimizing the candidate boxes generated by the adjusted region recommendation network using an improved target detection algorithm; The image of the Chinese medicinal material to be identified is input into the optimized image recognition model to output the recognition result.

2. The method for recognizing images of Chinese medicinal materials according to claim 1, wherein: The step of preprocessing the picture set data to obtain preprocessed data specifically includes: Performing normalization processing on the picture set data using a linear function conversion algorithm; Gray-scaling the normalized image set data using a weighted average algorithm; The grayscaled picture set data is denoised using a median filtering algorithm.

3. The method for recognizing images of Chinese medicinal materials according to claim 1, wherein: The step of enhancing the processing of the pre-processed data specifically includes: Randomly flipping and rotating the filtered pre-processed data; Selecting a scaling factor within a predetermined range to scale the pre-processed data after flipping and rotating; A color dithering factor is used to add Gaussian noise to the pre-processed data after the scaling process.

4. The method for recognizing images of Chinese medicinal materials according to claim 1, wherein: The step of extracting the image features of the pre-processed data after enhancement processing specifically includes: Extracting the color feature vector of the pre-processed data after enhancement processing by a spatial transformation algorithm, and performing normalization processing to obtain color features; Using a histogram of oriented gradients algorithm to extract shape features of the pre-processed data after enhancement processing; The texture features of the pre-processed data after enhancement processing are extracted based on a local binary pattern algorithm.

5. The method for recognizing images of Chinese medicinal materials according to claim 1, wherein: The steps of fusing the image features through a multi-feature fusion strategy to obtain multi-vector fusion features and convolving the feature extraction network specifically include: A multi-feature fusion algorithm is used to fuse the image features to obtain a multi-vector fusion feature; Performing a convolution operation on the feature extraction network in a bottom-up manner to obtain a corresponding feature layer; The feature layer is respectively subjected to 1×1 and 3×3 convolution processing with the multi-vector fusion feature in sequence to output the corresponding feature map; The aliasing parts of adjacent feature maps are eliminated through 3×3 convolution.

6. The method for recognizing images of Chinese medicinal materials according to claim 1, characterized in that: The step of adjusting the anchor box generation mechanism of the region recommendation network specifically includes: Utilizing the region recommendation network to perform a traversal operation on the multi-vector fusion feature to generate a series of anchor frames of different sizes; Setting the aspect ratio and size of the corresponding anchor frame according to the scale of each feature layer to generate the corresponding candidate frame; A prefabricated adjustment algorithm and the width and height of the candidate frame are combined to make the candidate frame correspond to the corresponding feature layer one by one.

7. The method for recognizing images of Chinese medicinal materials according to claim 1, characterized in that: The step of using the improved target detection algorithm to iteratively optimize and adjust the region recommendation network to generate the candidate box specifically includes: Sort the generated candidate boxes in descending order according to their confidence scores, and find the candidate box with the highest score; Using an improved target detection algorithm to compare the score of the candidate box with the highest score and the degree of overlap with its neighboring candidate boxes; Whether to change the confidence score of the candidate box is determined based on whether the overlap threshold is exceeded. Repeat the above steps to find the optimal candidate box.

8. A Chinese herbal medicine image recognition system, characterized in that: include: Construction module, used to construct image set data of Chinese herbal medicine slices; A preprocessing module, configured to preprocess the image set data to obtain preprocessed data; An enhancement module, configured to enhance processing of the pre-processed data; An extraction module, configured to extract image features of the pre-processed data after enhancement processing, wherein the image features include color features, shape features, and texture features; A building module for building an image recognition model using a two-stage object detection algorithm, wherein the image recognition model includes a feature extraction network and a region recommendation network; A fusion module, configured to fuse the image features through a multi-feature fusion strategy to obtain multi-vector fusion features, and convolve the feature extraction network; An adjustment module, configured to adjust the anchor box generation mechanism of the region recommendation network; An iterative module, configured to iteratively optimize the candidate boxes generated by the adjusted region recommendation network using an improved object detection algorithm; The recognition module is used to input the image of the Chinese medicinal material to be recognized into the optimized image recognition model to output the recognition result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for recognizing images of traditional Chinese medicines according to any one of claims 1 to 7 is implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for recognizing images of traditional Chinese medicines according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Medicinal material identification method based on visual system

    CN121121274A