Method for quickly retrieving traditional Chinese medicinal material components based on image matching

By acquiring images from multiple angles and constructing feature parameters, the problem of inaccurate feature matching under surface reflection and texture noise of Chinese medicinal materials was solved, and the identification of Chinese medicinal material components with high stability and high accuracy was achieved.

CN121010784APending Publication Date: 2025-11-25CANGNAN COUNTY QIUSHI TRADITIONAL CHINESE MEDICINE INNOVATION RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511138822.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately capture discriminative local information under conditions of surface reflection, high texture noise, or local variations in Chinese medicinal materials. This leads to distorted recognition results, especially in batch recognition or database matching where accuracy declines.

Method used

Under preset lighting conditions, multi-angle image acquisition is performed around the Chinese medicinal material sample. Geometric size adjustment and color space calibration are performed to generate a standardized multi-view image set. Highlight area detection is performed on the image of the predetermined reference view. The area of ​​the highlight area and the overall gray standard deviation are calculated to construct representative surface feature parameters. Combined with supplementary view feature parameters, comprehensive surface characteristic data is generated to establish initial medicinal material feature pairs for matching.

Benefits of technology

By acquiring images from multiple angles and constructing feature parameters, the stability and accuracy of identifying components of Chinese medicinal materials are improved, the risk of interference from individual perspective anomalies on the overall judgment is reduced, and the stability of feature expression and the reliability of recognition judgment are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010784A_ABST
    Figure CN121010784A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image matching, in particular to a traditional Chinese medicinal material component quick retrieval method based on image matching, which comprises the following steps of: performing multi-angle image information acquisition around a target traditional Chinese medicinal material sample under a preset illumination condition, and performing geometric dimension adjustment and color space calibration on each image, and generating a standardized multi-view image set. According to the invention, multi-angle image acquisition is carried out around the traditional Chinese medicinal material sample under the preset illumination condition, and geometric dimension adjustment and color space calibration are carried out on each image, so that an image source has a unified standard, data deviation caused by angle, size and illumination is controlled from the source, and the stability of subsequent image processing and the accuracy of comparison are improved. After a predetermined reference view angle image is selected, through highlight area detection and boundary pixel extraction, the area of a highlight area and the overall gray scale standard deviation are further analyzed, representative surface feature parameters are constructed, and specific physical attributes and optical responses of a sample are reflected.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image matching, in particular to a Chinese medicinal material component rapid retrieval method based on image matching. BACKGROUND

[0002] In the prior art, under the condition that there are complex samples such as Chinese medicinal materials, reflection, high texture noise or local variation on the surface, it is difficult to accurately capture the local information with discriminative power only by relying on single view or overall feature point matching mode. For example, when there are multiple highlight areas on the surface of the sample, the feature points will be incorrectly focused on the non-key area, resulting in distorted recognition results, especially in batch recognition or database matching, the accuracy is reduced. Therefore, improvement is needed. SUMMARY

[0003] The purpose of the present application is to solve the shortcomings in the prior art, and a Chinese medicinal material component rapid retrieval method based on image matching is proposed.

[0004] In order to achieve the above-mentioned purpose, the present application adopts the following technical scheme, a Chinese medicinal material component rapid retrieval method based on image matching, comprising the following steps: Under the preset lighting condition, multi-angle image information of the target Chinese medicinal material sample is collected, and the geometric size adjustment and color space calibration are performed on each image to generate a standardized multi-view image set; From the standardized multi-view image set, a predetermined reference view image is selected, the highlight area of the selected image is detected and the boundary pixel coordinates of the area are obtained, the reference image highlight area boundary data is established, and the highlight area and the overall image gray standard deviation are calculated based on the reference image highlight area boundary data. The representative surface feature parameters are jointly constituted; Each image in the standardized multi-view image set is extracted, the highlight area and the overall image gray standard deviation of each image are calculated, the calculation results are summarized, the supplementary view feature parameter set is established, and the average highlight area and the average gray standard deviation are calculated based on the supplementary view feature parameter set. The comprehensive surface characteristic data is generated by combining the representative surface feature parameters; The average highlight area and the average gray standard deviation in the comprehensive surface characteristic data are taken and combined into an initial numerical pair to establish an initial medicinal material feature pair, and based on the initial medicinal material feature pair, the medicinal material surface characteristic identification result is matched and generated.

[0005] Preferably, the acquisition step of the standardized multi-view image set is: Under preset lighting conditions, surround the target Chinese medicinal material sample and use a camera to acquire multi-angle images of the Chinese medicinal material sample at preset angle intervals. Each acquisition records the camera angle parameters and the real-time brightness of the lighting source. All acquired raw images are numbered one by one with the corresponding angle parameters and brightness values ​​to generate a multi-angle raw image set with angle parameter and brightness value labels. Based on the multi-angle original image set with angle parameters and brightness values, each original image is geometrically scaled and cropped using bilinear interpolation according to the set target pixel aspect ratio and target resolution parameters. All size-adjusted images are archived using the angle parameters and brightness values ​​of each image as unique identifiers to obtain a standard-size multi-angle image set. Based on the standard-sized multi-angle image set, the color deviation coefficient of each channel is obtained by comparing the RGB components of the pixels in the standard color card area with the standard values ​​of each standard-sized image. The color deviation coefficient is then used to correct the color components of all pixels in the entire image. The corrected color space parameters of each image are marked, and all images with completed color space calibration are organized to form a standardized multi-view image set.

[0006] Preferably, the step of obtaining the highlight region definition data of the reference image is as follows: Based on the angle parameters corresponding to each image in the standardized multi-view image set, the image closest to the target angle is extracted by image number indexing, and this image is used as the predetermined reference view image. A reference image index entry is generated with the image number and shooting angle as unique identifiers. Based on the reference image index entries, the RGB channel values ​​of all pixels in the image are extracted and converted into corresponding brightness values. At the same time, pixel regions with brightness values ​​higher than the preset highlight recognition threshold are identified, the boundary pixel coordinates of the region are extracted and a boundary contour set is constructed, and the boundary discrepancy factor of the region is calculated. Based on the boundary discrete factor, each boundary discrete factor is judged one by one, and the set of boundary contours with all boundary discrete factors below the compactness threshold is retained. The pixel coordinates are extracted as the highlight boundary pixel data to establish the highlight area definition data of the reference image.

[0007] Preferably, the steps for obtaining the representative surface feature parameters are as follows: Based on the highlight area definition data of the reference image, the coordinates of the closed boundary pixels of all highlight areas are extracted, all connected pixels within the boundary pixels are marked in sequence, and each point is judged to be within the closed boundary area and marked as a pixel inside the highlight. The total number of pixels marked as pixels inside the highlight is counted to form the pixel count value of the highlight area of ​​the reference image. Based on the highlight area definition data of the reference image, all pixels in the highlight area of ​​the reference image are masked, and only the pixel values ​​in the non-highlight area are retained. The RGB channel values ​​of each pixel in the remaining area are converted into gray values ​​to construct a gray value vector sequence. Based on the gray value vector sequence, the sum of the mean of all elements and the sum of their respective squared deviations is calculated. The square root is taken and divided by the number of effective pixels to form the overall gray value standard deviation of the image. Based on the number of pixels in the highlight area of ​​the reference image and the standard deviation of the overall grayscale of the image, data pairs are constructed and bound to the image number identifier to generate representative surface feature parameters.

[0008] Preferably, the steps for obtaining the supplementary viewpoint feature parameter set are as follows: The predetermined reference view images are excluded from the standardized multi-view image set, all remaining non-reference view images are extracted, the angle number and image path information of each image are recorded in sequence, and an identification code is generated for each image to generate a non-reference view image index list. Based on the non-reference viewpoint image index list, the image content is read one by one, the pixel enclosure range of the highlight area is calculated, the area of ​​the highlight area is obtained, and at the same time, all pixels in the image are converted into gray values ​​and a gray value vector is constructed. Mean and variance analysis is performed to obtain the overall gray value standard deviation of the image. Two feature values ​​are bound according to the image identification code to generate non-reference viewpoint image feature parameter pairs. Based on the non-reference viewpoint image feature parameter pairs, all image feature parameter pairs are batch-classified to generate a supplementary viewpoint feature parameter set.

[0009] Preferably, the steps for obtaining the comprehensive surface characteristic data are as follows: Based on the supplementary viewpoint feature parameter set, the image number, angle identifier, highlight area value and grayscale standard deviation value of each non-reference viewpoint image are read sequentially. The highlight area values ​​of all images are written into a unified sequence array A, and the grayscale standard deviation values ​​are written into a unified sequence array B. Each array element is bound to the corresponding image number to form a supplementary viewpoint highlight area sequence and a supplementary viewpoint grayscale standard deviation sequence. Based on the supplementary viewpoint highlight area sequence and the supplementary viewpoint grayscale standard deviation sequence, the highlight area value of the reference image extracted from the representative surface feature parameters is appended to the end of the unified sequence array A, and the overall grayscale standard deviation value of the image is appended to the end of the unified sequence array B to generate the merged total highlight area sequence and the total grayscale standard deviation sequence. Based on the merged total sequence of highlight area and the total sequence of grayscale standard deviation, the average of all values ​​in the sequence is calculated to construct the average highlight area value and the average grayscale standard deviation value, thereby generating comprehensive surface characteristic data.

[0010] Preferably, the step of obtaining the initial medicinal material characteristic pair is as follows: Based on the comprehensive surface characteristic data, the average highlight area value and the average gray standard deviation value are extracted. The corresponding fields are read and the consistency of the value type, value range and unit is verified to generate single-value indexes of highlight area and gray standard deviation. Based on the single-value index of highlight area and the single-value index of grayscale standard deviation, a two-element data structure is constructed. The data is written into the structure in the order of "area-grayscale" and field names and identifiers are assigned to the structure to form the initial medicinal material feature pairs.

[0011] Preferably, the steps for obtaining the identification results of the surface characteristics of the medicinal material are as follows: Based on the average highlight area and average gray standard deviation extracted from the initial medicinal material feature pair, the mean edge gradient parameter corresponding to the same image is selected from the standardized multi-view image set, and data type consistency verification and normalization conversion are performed on the three parameters to obtain the highlight area value, gray standard deviation value and mean edge gradient value. The characteristic response value of the medicinal material is calculated based on the area of ​​the highlight region, the standard deviation of gray level, and the mean value of the edge gradient. Based on the medicinal material characteristic response value, a zone-by-zone matching is performed with the established standard medicinal material characteristic response interval. When the medicinal material characteristic response value falls into any interval, the medicinal material type label of that interval is extracted and bound to the current image number to generate the medicinal material surface characteristic identification result.

[0012] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention acquires images of Chinese medicinal herbs from multiple angles under preset lighting conditions, and performs geometric size adjustment and color space calibration on each image to ensure a unified standard at the source. This controls data deviations caused by angle, size, and lighting from the source, improving the stability of subsequent image processing and the accuracy of comparison. After selecting a predetermined reference viewpoint image, highlight area detection and boundary pixel extraction are used to further analyze the highlight area and overall grayscale standard deviation, constructing representative surface feature parameters that reflect the specific physical properties and optical response of the sample. Features from non-reference viewpoint images are extracted and aggregated using the same indicators to form a supplementary feature parameter set. After being combined with representative features, this forms an average highlight area and average grayscale standard deviation with overall expressive power, reducing the risk of interference from individual viewpoint anomalies to the overall judgment. Based on the initial medicinal herb feature pair composed of dual indicators, a comprehensive evaluation method of multi-dimensional image performance factors is introduced during the matching process to improve the matching accuracy between samples and targets, enhance the stability of feature expression, and improve the reliability of recognition judgment. It overcomes the impact of single-dimensional, localized, and uncalibrated data on recognition accuracy and stability, and improves the accuracy and anti-interference ability of medicinal material image feature matching in actual retrieval scenarios. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the steps of the present invention. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0015] Please see Figure 1 This invention provides a technical solution: a method for rapid retrieval of components in traditional Chinese medicine based on image matching, comprising the following steps: Under preset lighting conditions, multi-angle image information is collected around the target Chinese medicinal material sample. Geometric size adjustment and color space calibration are performed on each image to generate a standardized multi-view image set. Select a predetermined reference view image from a standardized multi-view image set, perform highlight region detection on the selected image and obtain the boundary pixel coordinates of the region, establish reference image highlight region definition data, and calculate the highlight region area and the overall grayscale standard deviation of the image based on the reference image highlight region definition data, which together constitute representative surface feature parameters. Extract each image from the non-reference viewpoint in the standardized multi-view image set, calculate the area of ​​the highlight region and the standard deviation of the overall gray level of each image, summarize the calculation results, establish a supplementary viewpoint feature parameter set, and calculate the average area of ​​the highlight region and the average standard deviation of the gray level based on the supplementary viewpoint feature parameter set and representative surface feature parameters to generate comprehensive surface characteristic data. The average area of ​​the highlight region and the average standard deviation of gray level in the comprehensive surface characteristic data are taken and combined into an initial value pair to establish an initial medicinal material feature pair. Based on the initial medicinal material feature pair, the surface characteristic identification results of the medicinal material are generated.

[0016] The steps for obtaining a standardized multi-view image set are as follows: Under preset lighting conditions, surround the target Chinese medicinal material sample and use a camera to acquire multi-angle images of the Chinese medicinal material sample at preset angle intervals. Each acquisition records the camera angle parameters and the real-time brightness of the lighting source. All acquired raw images are numbered one by one with the corresponding angle parameters and brightness values ​​to generate a multi-angle raw image set with angle parameter and brightness value labels. Based on a multi-angle original image set with angle parameters and brightness values, each original image is geometrically scaled and cropped using bilinear interpolation according to the set target pixel aspect ratio and target resolution parameters. All size-adjusted images are archived using the angle parameters and brightness values ​​of each image as unique identifiers to obtain a standard-sized multi-angle image set. Based on a standard-sized multi-angle image set, the color deviation coefficient of each channel is obtained by comparing the RGB components of the pixels in the standard color card area with the standard values ​​of each standard-sized image. The color deviation coefficient is then used to correct the color components of all pixels in the entire image. The corrected color space parameters of each image are marked, and all images with completed color space calibration are organized to form a standardized multi-angle image set.

[0017] Specifically, under defined preset lighting conditions and targeting the medicinal herb sample, the subsequent image acquisition process is executed. The preset lighting conditions refer to a set of parameters pre-configured and calibrated to ensure the stability and consistency of the lighting environment during image acquisition. These parameters include a high-frequency flicker-free LED flat panel light selected as the light source type, with a color temperature set to 5500K (simulating sunlight color temperature to capture the true color of the object), and a total illumination intensity set to 1000 lux. This value was determined through preliminary experiments on various representative medicinal herb samples, observing the effects of different illumination intensities (e.g., 500 lux, 800 lux, 1000 lux, 1200 lux). To assess image sharpness, texture detail, and the balance between highlights and shadows, 1000 lux was ultimately selected as the intensity that best suits a broad range and maximizes detail. A symmetrical four-corner lighting layout was used, with each lamp incident at a 45-degree angle to the sample's horizontal plane to achieve uniform illumination and minimize shadows. Subsequently, a 24-megapixel industrial camera equipped with a 50mm fixed-focus macro lens was used to automatically acquire multi-angle images of the target medicinal herb sample, fixed at the center of a rotating stage, at preset 10-degree intervals. This interval was determined to satisfy a 360-degree... Under the premise of comprehensive coverage, by comparing the redundancy and key feature capture integrity of image sets acquired at different intervals (such as 5 degrees, 10 degrees, 15 degrees, and 20 degrees), it was found that a 10-degree interval can achieve a balance between ensuring the capture of subtle feature changes of medicinal materials from various angles and controlling the total amount of data. For example, too small a 5-degree interval will produce a large number of highly similar images, increasing the processing burden, while too large a 20-degree interval may miss locally rapidly changing features. Therefore, 10 degrees was selected. After each camera capture, the system automatically records the angle parameters fed back by the rotating platform, for example, starting from 0 degrees, recording 10 degrees, 20 degrees, up to 350 degrees. A calibrated brightness sensor monitors and records the actual brightness value of the illumination source at the sample location in real time, such as 1005 lux, 998 lux, etc., to ensure that even small fluctuations in lighting conditions are captured. Then, each acquired raw image is strictly associated with its corresponding angle parameter and real-time brightness value and assigned a unique number. The numbering rule can be set as "sample ID_acquisition date and time_angle value_brightness value", for example, "DangShen_202505211030_010_1002lux". In this way, a multi-angle raw image set with angle parameter and brightness value identifiers is generated.

[0018] Based on the multi-angle raw image set with angle parameters and brightness values ​​obtained in the previous step, fine-grained size and content adjustments are performed on each raw image in the set. First, the images are standardized according to the set target pixel aspect ratio and target resolution parameters. The target pixel aspect ratio is set to 1:1, i.e., a square image. This setting takes into account the preference of subsequent feature analysis algorithms for a uniform input size, and the fact that the effective information area of ​​most Chinese medicinal material samples can be roughly surrounded by a square frame during rotation. The target resolution parameter is set to 2048x2048 pixels. This resolution selection is... The optimal resolution (2048x2048 pixels) was determined after evaluating the degree of detail preservation and sensitivity to subsequent processing algorithms (such as texture analysis and highlight region calculation) at different resolutions (e.g., 1024x1024, 2048x2048, 3072x3072 pixels). It was found that 2048x2048 pixels could fully display the microstructure of the medicinal herb surface while avoiding the data storage and computational burden caused by excessively high resolution. In practice, for each original image, image content analysis methods, such as algorithms based on color contrast or edge detection, were first used to initially locate the main area of ​​the medicinal herb sample in the image and obtain its most prominent features. A small bounding rectangle is used as the reference. Then, to ensure the image content is centered and adapts to the target pixel aspect ratio of 1:1, the cropping area is determined based on the center of this smallest bounding rectangle, according to the original image's width and height and the target aspect ratio of 2048x2048. If the original image's aspect ratio does not conform to 1:1, symmetrical cropping is performed on the longer side, while ensuring the integrity of the main content. Alternatively, pixels of a preset background color (e.g., pure black) are added to the shorter side to make it a square. Then, bilinear interpolation is used uniformly to perform geometric scaling, precisely scaling the cropped or adjusted square image to 2048x2048. 2048 pixels. Bilinear interpolation is chosen for its good balance between image quality and computational efficiency. It determines the new pixel value by calculating the weighted average of the pixel values ​​of the four nearest pixels around the sampling point, effectively reducing jagged edges and mosaic effects during scaling. After scaling and cropping, the system retains the angle parameters and brightness values ​​in the original image set as unique identifiers for the image after size adjustment. However, these are not directly used as part of the file name, but are stored as metadata in a structured way with the image data. All images processed through these steps are aggregated to obtain a standard-sized multi-angle image set.

[0019] Based on the standard-sized multi-angle image set generated in the previous stage, color calibration is performed on each standard-sized image in the set to eliminate color deviations introduced by camera sensor characteristics, lighting conditions, and environmental factors. Specifically, before acquiring each multi-angle image, a standard color chart, such as the 24-color patch standard chart of the X-Rite ColorChecker Classic Mini, is placed within the same field of view of the target medicinal herb sample. This color chart contains various colors and grayscale patches with known spectral characteristics. When processing each standard-sized image, the system first automatically identifies and extracts the standard color chart area based on the preset fixed position coordinates of the color chart in the image (e.g., a specific rectangular area in the upper left corner of the image, whose coordinates are set through initial calibration, such as from pixel positions (x1, y1) to (x2, y2)). Subsequently, for each standard color patch within the standard color chart area, such as selecting six standard grayscale patches (from white to black) and the main color patches (red, green, blue, yellow, cyan, magenta, etc.), the average value of the RGB components of the pixels within these color patches is calculated and denoted as . , , Simultaneously, the system pre-stores the theoretical standard RGB values ​​of each color patch of this model's standard color chart under a standard D65 light source (close to the 5500K color temperature of the system's preset lighting conditions), denoted as... , , These standard values ​​are provided by the color chart manufacturer and obtained through table lookup. For example, the standard value for a certain medium gray patch might be (R: 128, G: 128, B: 128), while the average value measured in the image is (R: 120, G: 135, B: 115). By comparing the measured values ​​with the standard values, the color deviation coefficient for each color channel (R, G, B) is calculated. For each channel The calculation method involves selecting multiple color patches (especially grayscale color patches). The ratios are averaged to improve robustness. For example, for the red channel, if the three grayscale blocks... The values ​​are 1.07, 1.06, and 1.08 respectively, then the color deviation coefficient of the red channel is... Set it to their average value (1.07 + 1.06 + 1.08) / 3 Similarly, the color deviation coefficient of the green channel is calculated as 1.07. Color deviation coefficient of the blue channel After obtaining these three color deviation coefficients, each pixel in the standard-sized image... Original color components Make corrections, and the corrected color components The calculation is as follows: ; ; ; The calculation results will be truncated to the effective pixel value range (e.g., 0-255). After the color correction of the entire image is completed, the system will mark the image as having completed color calibration and will specify the color deviation coefficient used. , , The color card model and standard light source information used are recorded as the corrected color space parameters as the image's metadata, but the original angle and brightness labels are not changed. After performing the above color space calibration process on all images in the standard-sized multi-angle image set, these color space calibrated images are organized to form a standardized multi-view image set.

[0020] The steps for obtaining the highlight region delineation data of the reference image are as follows: Based on the angle parameters corresponding to each image in the standardized multi-view image set, the image closest to the target angle is extracted by image number indexing, and this image is used as the predetermined reference view image. Reference image index entries are generated with image number and shooting angle as unique identifiers. Based on the baseline image index entries, the RGB channel values ​​of all pixels in the image are extracted and converted into corresponding brightness values. Simultaneously, pixel regions with brightness values ​​exceeding a preset highlight recognition threshold are identified. The boundary pixel coordinates of these regions are extracted, and a boundary contour set is constructed. The boundary discrepancy factor of these regions is calculated using the following formula: ; in, For the first The boundary discrepancy factor of each highlight region This represents the total number of pixels in the highlight area. , For the first The horizontal and vertical coordinates of each pixel in the image , These are the mean horizontal and vertical coordinates of all pixels in the highlight region, respectively. , The image width and image height are the reference image. For the first The brightness value of each pixel. This is the average brightness of all pixels in the highlight area; Based on the boundary discrepancy factor, each boundary discrepancy factor is judged one by one, and the set of boundary contours with all boundary discrepancy factors below the compactness threshold is retained. The pixel coordinates are extracted as the highlight boundary pixel data to establish the highlight area definition data of the reference image.

[0021] Specifically, based on the angle parameters recorded in each image of the standardized multi-view image set obtained in the previous steps, a target angle needs to be determined first. This target angle is set based on a perspective that is generally considered the most representative or easiest to observe for key features in traditional Chinese medicine identification. For example, for most medicinal materials, the target angle can be set to 0 degrees, i.e., the perspective facing the camera directly, or 45 degrees, to take into account both frontal and partial side information. Here, we will use 0 degrees as the target angle for explanation. We will iterate through the angle parameters recorded in each image of the standardized multi-view image set, calculate the absolute difference between the angle parameters of each image and the set target angle of 0 degrees, and select the image with the smallest difference. If multiple images have the same minimum difference between their angle and the target angle (e.g., the target angle is 0 degrees, and the image set contains both 0-degree and 360-degree images, all considered equivalent), the image with the smaller number is selected by default, or one is selected according to a preset rule (e.g., prioritizing images with brightness values ​​closer to the standard brightness). This selected image is then used as the predetermined reference viewpoint image. Subsequently, the system extracts the unique image number of the predetermined reference viewpoint image, such as "sample ID_collection date and time_000_brightness value", and its confirmed shooting angle, i.e., 0 degrees, together forming an information record. This information record is assigned a specific identifier, such as "REF_IMG_001", thereby generating a reference image index entry.

[0022] Based on the reference image index entry generated in the previous step, the system first loads the predetermined reference viewpoint image pointed to by the entry, which has a known image width. and image height Next, for each pixel in the image, its RGB channel value is extracted, and the RGB value is converted into a single luminance value according to the standard luminance conversion formula. The luminance value range is usually from 0 to 255. Then, the luminance value of each pixel is compared with a preset highlight recognition threshold to identify highlight areas. The preset highlight recognition threshold is set to 235. This value was determined by analyzing 1000 test images of various typical Chinese medicinal herb samples taken under standard lighting. Experts marked the areas they considered to be highlights, and the distribution of pixel luminance values ​​in these areas was statistically analyzed. It was found that 95% of the highlight areas marked by the experts had pixel luminance values ​​higher than 235. To reduce interference from bright background specks other than the medicinal materials themselves and to more accurately capture the strong reflections on the surface of the medicinal materials due to their material properties, further testing was conducted on the impact of different thresholds (220, 225, 230, 235, 240, 245) on the accuracy of highlight region extraction (compared with expert labeling) and the stability of subsequent feature parameters. 235 was selected as the threshold that best balances the integrity of highlight detection with noise suppression. All pixels with brightness values ​​higher than 235 were initially marked as highlight pixels. Then, through connected component analysis, spatially adjacent highlight pixels were organized into several independent highlight regions. For each identified highlight region (denoted as the ), (A set of highlight regions) is identified, and the pixel coordinates constituting the boundary of these regions are extracted to construct the boundary contour set of the regions. Then, the boundary discrepancy factor of the regions is calculated. .

[0023] formula: The advantage of the formula lies in the boundary discrepancy factor. This method combines the spatial geometric dispersion of highlight regions with the uniformity of their internal brightness. It normalizes the deviation of pixel coordinates from the region's geometric center point (divided by the overall width and height of the image) and the deviation of pixel brightness from the region's average brightness (divided by the region's average brightness), thus... As a dimensionless parameter, it can be used for relative comparison across highlight areas of different sizes and even different images. The smaller the value, the more compact, regular in shape, and uniform in internal brightness of the highlight area. This helps to distinguish between specular highlights caused by dense structure on the surface of medicinal materials and diffuse bright spots caused by loose or multi-planar structure, providing a more refined basis for subsequent surface characteristic analysis. In terms of formula design, the square root of the sum of squares of various deviations is used, similar to the calculation idea of ​​standard deviation, which can effectively amplify the influence of outliers and make the measurement of regional irregularity more sensitive. At the same time, the square term ensures that the direction of the deviation does not affect its contribution.

[0024] For the first The total number of pixels in a specific highlight region is obtained through the following steps: After initially identifying highlight pixels using a preset highlight recognition threshold (e.g., 235), these highlight pixels are grouped using an eight-neighbor connected component labeling algorithm to form several independent highlight region blocks. For a specific [number]th [highlight region], the total number of pixels in the highlight region is determined by the following steps: A highlight area block This refers to the number of highlight pixels contained within the specified region. For example, if connected component analysis of a highlight region reveals that it contains 150 pixels, then the highlight region... The value is 150.

[0025] , The first in the highlight area The steps to obtain the horizontal and vertical coordinates of a pixel in the image are as follows: After determining a highlight region and its constituent pixels... After identifying each pixel, iterate through every pixel within that region, reading its horizontal (column number) and vertical (row number) coordinates in the entire reference image coordinate system. The image coordinate system typically has its origin at the top left corner (0, 0), with the positive X-axis pointing horizontally to the right and the positive Y-axis pointing vertically downwards. For example, if a pixel within the highlight region is located at row 100 and column 150 of the image, its coordinates would be... , (If counting starts from 0) or , (If counting starts from 1), here we uniformly use counting starting from 0, so the example value is... , .

[0026] , All of the highlights The average of the horizontal and vertical coordinates of the nth pixel is obtained by the following steps: For the nth pixel... All within the highlight area 1 pixel Sum of coordinates and divide by get Similarly, for all Sum of coordinates and divide by get For example, if a highlight region contains 3 pixels with coordinates (10, 20), (11, 21), and (12, 20), then... , .

[0027] , These are the image width and image height of the current reference image being processed. The steps to obtain these parameters are as follows: These two parameters are read directly from the metadata of the loaded, predetermined reference viewpoint image. Based on the previous step (size adjustment in the generation of the standardized multi-view image set), the image has been uniformly processed to a standard size. For example, if the target resolution was set to 2048x2048 pixels in a previous step, then here... Pixels Pixel.

[0028] The first in the highlight area The steps for obtaining the brightness value of each pixel are as follows: After preprocessing the reference image and converting the RGB values ​​of each pixel into brightness values, for each pixel in the highlight area... Its corresponding brightness value is This value is calculated and used when identifying highlight areas. The brightness value ranges from 0 to 255. For example, if a pixel in a highlight area has a converted brightness value of 240, then... .

[0029] For all of this highlight area The average brightness value of the nth pixel is obtained by the following steps: For the nth pixel... All within the highlight area Brightness value of each pixel Summing and dividing by get For example, if a highlight area contains 3 pixels with brightness values ​​of 240, 245, and 250, then... .

[0030] Calculation process: Define a highlight area containing 1 pixel, baseline image width pixels, height Pixel.

[0031] Pixel 1: ,brightness ; Pixel 2: ,brightness ; Pixel 3: ,brightness .

[0032] First, calculate the mean: ; ; ; Next, we calculate the contribution value for each pixel: For pixel 1 : ; ; ; The contribution of pixel 1: ; For pixel 2 : ; ; ; The contribution of pixel 2: ; For pixel 3 : ; ; ; The contribution of pixel 3:

[0033] Total contribution of all pixels: ; calculate : ; This result indicates that the boundary discrepancy factor of the highlight region in this example is... Approximately 0.006556, this is a relatively small value, indicating that the highlight area is spatially concentrated and has little internal brightness difference, tending to be compact and uniform. This value will serve as a quantitative description of the morphology and brightness characteristics of this specific highlight area, and will be used for subsequent comparison with the compactness threshold value.

[0034] The boundary discretization factor is calculated based on each identified highlight region in the predetermined reference viewpoint image in the previous step. Next, these highlight areas will be filtered based on a pre-set compactness threshold, for example, 0.05. This threshold is determined by analyzing highlight areas in a large number of known medicinal herb images (e.g., 10 samples of each of 500 different herbs from a reference viewpoint). Values ​​were calculated, and combined with expert annotations on whether these highlight areas belonged to "effective characteristic highlights" (usually referring to those with relatively regular and compact shapes that reflect the specific texture of the medicinal material's surface, rather than large areas of overexposure or fine noise), statistical analysis was conducted (e.g., constructing...). A histogram of values ​​was used to observe the effective characteristic highlights and ineffective highlights. The threshold is determined by the distribution difference in values ​​(or by training a simple classifier using machine learning methods to find the optimal segmentation threshold). The goal is to select a threshold that retains the most effective characteristic highlights while removing most of the highlights caused by diffusion, irregular shapes, or drastic changes in internal brightness. For highlight regions with larger values, selecting 0.05 as the threshold means that if the boundary dispersion factor of a highlight region is large... If the value is less than 0.05, the highlight area is considered sufficiently compact and uniform, possessing the potential to serve as a surface feature; conversely, if... If the value is greater than or equal to 0.05, the area is considered too diffuse or irregular and will be discarded. The system will check each highlight area one by one. Values, only those that are retained. The system extracts a set of boundary contours corresponding to highlight regions with values ​​less than 0.05. For each retained set of boundary contours, the system extracts a list of precise pixel coordinates that constitute these contours. These coordinate lists are then combined to form highlight boundary pixel data. Finally, these filtered highlight boundary pixel data are integrated to establish the reference image highlight region definition data for the current predetermined reference viewpoint image.

[0035] The steps for obtaining representative surface feature parameters are as follows: Based on the highlight region definition data of the reference image, the pixel coordinates of the closed boundary of all highlight regions are extracted, and all connected pixels within the boundary pixels are marked in turn. Each point is judged to see if it falls within the closed boundary region and is marked as a pixel inside the highlight. The total number of pixels marked as pixels inside the highlight is counted to form the pixel count value of the highlight region area of ​​the reference image. Based on the highlight area definition data of the reference image, all pixels in the highlight area of ​​the reference image are masked, and only the pixel values ​​of the non-highlight area are retained. The RGB channel values ​​of each pixel in the remaining area are converted into gray values ​​and a gray vector sequence is constructed. Based on the gray vector sequence, the sum of the mean of all elements and the sum of the squared deviations of each element is calculated. The square root is taken and divided by the number of effective pixels to form the overall gray standard deviation value of the image. Based on the number of pixels in the highlight area of ​​the benchmark image and the standard deviation of the overall grayscale value, data pairs are constructed and bound to the image number identifier to generate representative surface feature parameters.

[0036] Specifically, based on the reference image highlight region definition data obtained in the previous steps, this data contains the closed boundary pixel coordinates of each highlight region that has been filtered and considered as valid features. First, the sequence of closed boundary pixel coordinates for each highlight region is extracted from this data. For each extracted highlight region boundary, a scanline filling algorithm is used to mark and count the pixels within it. Specifically, for a given highlight region boundary... Given the described closed polygon boundary, the algorithm determines the minimum bounding rectangle of the polygon, and then within the scan line range of that rectangle (i.e., from the minimum bounding rectangle)... Value to maximum (value), for each horizontal scan line (e.g., ), calculate all intersections between the scan line and the polygon boundary, and sort these intersections according to their The coordinate values ​​are sorted, and then these intersections are processed in pairs from left to right, for each pair of intersections (e.g., and All pixels on the scan line segment between () and (i.e., coordinates are) and All pixels are determined to be located inside the closed region of the boundary and are marked as pixels inside the highlight. After traversing all the highlight regions defined by the highlight region boundary data of the reference image and completing the internal marking and counting of pixels in each region, the internal pixel count values ​​obtained from the statistics of all these highlight regions are summed to obtain a total number of pixels. This total number forms the pixel count value of the highlight region area of ​​the reference image.

[0037] Based on the specific pixel locations of all valid highlight regions determined in the highlight region definition data of the reference image obtained in the previous steps, and retrieving the corresponding predetermined reference viewpoint image, a masking operation is first performed. This involves generating a binary mask on the predetermined reference viewpoint image, where all pixels belonging to valid highlight regions are marked as to be excluded (e.g., pixel values ​​are set to 0), while all pixels not belonging to any valid highlight region are marked as to be retained (e.g., pixel values ​​are set to 1). This mask is applied to the predetermined reference viewpoint image, thereby logically or practically removing pixel information within the highlight regions, retaining and focusing only on the pixel values ​​of non-highlight regions in the image. Then, for each pixel in these retained non-highlight regions, its RGB color values ​​are extracted, and a standard luminance conversion formula is used to convert the RGB value of each pixel into a single grayscale value. The grayscale values ​​typically range from 0 to 255. All these grayscale values ​​calculated from pixels in the non-highlight regions are collected to form a one-dimensional grayscale vector sequence. Subsequently, the overall grayscale standard deviation of the non-highlight regions of the image is calculated based on this grayscale vector sequence. The calculation process includes: First, calculating the arithmetic mean of all grayscale values ​​in the grayscale vector sequence to obtain the average grayscale value. Second, for each grayscale value in the sequence, calculating the difference (i.e., deviation) between it and the aforementioned average grayscale value, and then squaring each difference to obtain the squared value of the deviation. Third, summing up all the squared values ​​of the deviations to obtain the total sum of squared deviations. Finally, dividing this total sum of squared deviations by the total number of elements in the grayscale vector sequence (i.e., the number of effective pixels in the non-highlight regions) to obtain the variance, and taking the square root of this variance, the result is the desired overall grayscale standard deviation of the image.

[0038] Based on the calculated and obtained pixel count of the highlight area in the reference image (representing the total number of pixels in all valid highlight areas of the image at the predetermined reference viewpoint) and the overall grayscale standard deviation of the image (reflecting the grayscale variation or texture complexity of non-highlight areas in the image at the predetermined reference viewpoint), the system then performs a data integration operation. This integrates these two values—the pixel count of the highlight area in the reference image and the overall grayscale standard deviation of the image—in a predefined order (e.g., area first, standard deviation second) into a binary data pair. For example, if the calculated area is 580 pixels and the grayscale standard deviation is 25.3, the resulting data pair would be (580, 25.3). To ensure this... The feature data is explicitly associated with its source image. This data pair is bound to a unique image number identifier of a predetermined reference view image, which is generated or recorded earlier when the predetermined reference view image is selected (e.g., derived from the image number and shooting angle information in the "Reference Image Index Entry"). For example, if the number of the predetermined reference view image is "SampleX_Angle0_Brightness1000lux", then the feature data pair (580, 25.3) will be associated with this number. In this way, two key quantitative indicators describing the highlight characteristics and background texture characteristics of the image surface are combined and traced back to generate representative surface feature parameters that can represent the surface characteristics of the predetermined reference view image.

[0039] The steps for obtaining the supplementary viewpoint feature parameter set are as follows: The predetermined reference view images are excluded from the standardized multi-view image set, all remaining non-reference view images are extracted, the angle number and image path information of each image are recorded in sequence, and an identification code is generated for each image to generate a non-reference view image index list. Based on the non-reference viewpoint image index list, the image content is read one by one, the pixel enclosure range of the highlight area is calculated, the area of ​​the highlight area is obtained, and all pixels in the image are converted into gray values ​​to construct a gray value vector. Mean and variance analysis is performed to obtain the overall gray value standard deviation of the image. Two feature values ​​are bound according to the image identification code to generate non-reference viewpoint image feature parameter pairs. Based on the non-reference viewpoint image feature parameter pairs, all image feature parameter pairs are batch-classified to generate a supplementary viewpoint feature parameter set.

[0040] Specifically, from the standardized multi-view image set generated in the previous steps, it is first necessary to identify and exclude the specific image that has been selected as the predetermined reference view image in the previous process. This exclusion operation is completed by comparing the unique image number or its corresponding angle parameter of each image in the image set with the identification information of the predetermined reference view image. After the exclusion is completed, all the remaining images in the image set constitute the non-reference view image set. Subsequently, the system will traverse each image in this non-reference view image set. For each non-reference view image, the system will record its existing angle number in the standardized multi-view image set. This angle number is the angle parameter recorded during acquisition, such as 10 degrees, 20 degrees, 30 degrees, etc. (the reference view such as 0 degrees has been excluded). At the same time, the system will also record the complete image path information of the image file in the storage system, such as " / data / tcm_samples / sample001 / Next, to facilitate rapid retrieval and management, the system generates a unique identifier for each non-reference view image. The generation rule for this identifier can be designed to combine the sample identifier, angle information, and sequence number. For example, if a sample is identified as "SJ001", the angle of the currently processed non-reference view image is 10 degrees, and it is the first non-reference view image of that sample, then its identifier can be set to "SJ001_NonRef_010_001". This identifier ensures that each non-reference view image has a unique reference name, even in different samples or different processing batches. All of these, including the angle number, image path information, and newly generated identifier recorded for each non-reference view image, will together constitute an index record. After summarizing all the index records of non-reference view images, a structured non-reference view image index list is formed.

[0041] Based on the non-reference viewpoint image index list generated in the previous step, the system will process each non-reference viewpoint image one by one in the list order. For each entry in the list, the corresponding image content is first read according to the image path information it records. After loading the image data, two core features are extracted. The first is to calculate the area of ​​the highlight region. Here, the "highlight region" is defined as the area in the image where the pixel brightness value exceeds a preset highlight recognition threshold. This threshold is consistent with the threshold used when processing the predetermined reference viewpoint image, for example, it is set to 235 (within the brightness range of 0-255). The setting is based on statistical analysis and expert visual evaluation of a large number of medicinal sample images to select a value that can effectively distinguish between real highlights and ordinary bright areas. In specific calculation, the brightness value of all pixels in the image is calculated first, and then all pixels with a brightness higher than 235 are identified. Through connected component analysis, these highlight pixels are divided into several independent highlight region blocks. Then, the number of pixels contained in each highlight region block is calculated, that is, the area within its pixel enclosed range. Finally, all these independent... The total area of ​​the highlight regions is obtained by summing the areas of the highlight regions in the non-reference view image. The second feature is to obtain the overall grayscale standard deviation of the image. To do this, the RGB color values ​​of all pixels (regardless of whether they belong to the highlight region) in the current non-reference view image need to be converted into grayscale values ​​to form a grayscale vector containing the corresponding grayscale values ​​of all pixels in the image. Then, statistical analysis is performed on this grayscale vector: first, the arithmetic mean of all these grayscale values ​​is calculated; then, the square of the difference between each grayscale value and its mean is calculated; then, all these squared differences are summed and divided by the total number of pixels to obtain the variance; finally, the square root of the variance is taken, which is the overall grayscale standard deviation of the image. After completing the calculation of these two feature values ​​(highlight region area and overall grayscale standard deviation of the image), they are bound to the identifier of the non-reference view image being processed to form a record, for example ("SJ001_NonRef_010_001", (area value, standard deviation value)). After all non-reference view images are processed in this way, a series of non-reference view image feature parameter pairs are generated.

[0042] Based on the non-reference viewpoint image feature parameter pairs generated in the previous step for each non-reference viewpoint image, where each pair contains the image's unique identifier and the corresponding calculated highlight area and overall image grayscale standard deviation, the next step is to integrate and manage these scattered feature parameter pairs. The system will perform a batch classification operation. Here, "batch classification" mainly refers to collecting and organizing all non-reference viewpoint image feature parameter pairs into a unified data structure. The design of this data structure should facilitate subsequent querying, retrieval, and further statistical analysis. For example, a list or form can be created where each row represents a non-reference viewpoint image, and the columns store the image's identifier, highlight area, and grayscale standard deviation, respectively. The system classifies data into regional area values ​​and overall grayscale standard deviations, and can include the angle number or image path information of the original records as auxiliary index fields. This classification operation does not involve complex classification or clustering based on the feature values ​​themselves, but focuses on the systematic collection and structured storage of data. It ensures that all surface characteristic quantification indicators extracted from non-reference viewpoint images have a clear and orderly set. During the classification process, the system verifies the data integrity and format consistency of each feature parameter pair. For example, it checks whether the area and standard deviation are valid numerical types. Through such batch classification processing, all independent non-reference viewpoint image feature parameter pairs are integrated into a whole data resource, generating a supplementary viewpoint feature parameter set.

[0043] The steps for obtaining comprehensive surface property data are as follows: Based on the supplementary viewpoint feature parameter set, the image number, angle identifier, highlight area value and gray standard deviation value of each non-reference viewpoint image are read sequentially. The highlight area values ​​of all images are written into a unified sequence array A, and the gray standard deviation values ​​are written into a unified sequence array B. Each array element is bound to the corresponding image number to form a supplementary viewpoint highlight area sequence and a supplementary viewpoint gray standard deviation sequence. Based on the supplementary viewpoint highlight area sequence and the supplementary viewpoint grayscale standard deviation sequence, the highlight area value of the reference image extracted from the representative surface feature parameters is appended to the end of the unified sequence array A, and the overall grayscale standard deviation value of the image is appended to the end of the unified sequence array B, generating the merged total highlight area sequence and the total grayscale standard deviation sequence. Based on the merged total sequence of highlight area and the total sequence of grayscale standard deviation, the average of all values ​​in the sequence is calculated to construct the average highlight area value and the average grayscale standard deviation value, thereby generating comprehensive surface characteristic data.

[0044] Specifically, based on the supplementary viewpoint feature parameter set generated in the previous steps, this parameter set has systematically collected the identifier code, corresponding angle number (i.e., angle parameter), calculated highlight area value, and overall grayscale standard deviation value of each non-reference viewpoint image. Now, this data needs to be reorganized for subsequent merging calculations. The system will traverse each record in the supplementary viewpoint feature parameter set and accurately extract four key pieces of information from each record: the "image number" of the non-reference viewpoint image (i.e., the previously generated unique "identifier code"), the "angle identifier" of the image (i.e., the angle parameter at the time of shooting), the "highlight area value" of the image, and the "grayscale standard deviation value" of the image. During the extraction process, the system will initialize two dynamic arrays or sequences, named unified sequence array A and unified sequence array B, respectively. The unified sequence array A is specifically used to store the highlight area values ​​of all images, while the unified sequence array B... Sequence array B is specifically used to store the grayscale standard deviation of all images. When a highlight area value is read from each record of the supplementary viewpoint feature parameter set, the area value itself is added to the unified sequence array A. At the same time, in order to maintain data traceability, the "image number" of the non-reference viewpoint image corresponding to the area value is also bound or associated with the corresponding entry in array A. For example, each element of array A can be a structure or tuple containing (image number, area value). Similarly, when a grayscale standard deviation value is read, the standard deviation value and its corresponding non-reference viewpoint image "image number" are also added to the unified sequence array B and associated in the same way. After completing the reading and writing operations of all non-reference viewpoint image data in the supplementary viewpoint feature parameter set, a supplementary viewpoint highlight area sequence and a supplementary viewpoint grayscale standard deviation sequence are finally formed, with the data content corresponding one-to-one with the image source.

[0045] Based on the supplementary viewpoint highlight area sequence (i.e., a unified sequence array A containing area data of all non-reference viewpoint images and their corresponding image numbers) and the supplementary viewpoint grayscale standard deviation sequence (i.e., a unified sequence array B containing standard deviation data of all non-reference viewpoint images and their corresponding image numbers) generated in the previous step, the next step is to integrate the feature parameters previously calculated for the predetermined reference viewpoint image. The system will retrieve the representative surface feature parameters generated in an earlier step. These parameters encapsulate the image number identifier of the predetermined reference viewpoint image, the number of pixels in the corresponding reference image highlight area (hereinafter referred to as the reference image highlight area value), and the overall grayscale standard deviation value of the image. From this representative surface feature parameter, the system accurately extracts the reference image highlight area value and combines this area value with its... The image ID of the corresponding predetermined reference view image is appended as a new element to the end of the supplementary view highlight area sequence (unified sequence array A). Similarly, the system also extracts the overall grayscale standard deviation of the predetermined reference view image from the representative surface feature parameters, and appends this standard deviation along with its image ID as a new element to the end of the supplementary view grayscale standard deviation sequence (unified sequence array B). Through these two append operations, the two sequences that originally only contained non-reference view image data have now been expanded to include highlight area information and grayscale standard deviation information from images from all viewpoints (including the reference view and all non-reference viewpoints), as well as the image ID corresponding to each information item. This generates a more comprehensive merged total highlight area sequence and total grayscale standard deviation sequence.

[0046] Based on the merged total highlight area sequence and total grayscale standard deviation sequence formed in the previous step, these two sequences now contain the highlight area values ​​from the predetermined reference view image and all non-reference view images, as well as the overall grayscale standard deviation of the image. Each value is associated with the ID of its source image. Next, the system will perform final statistical calculations on these two complete sequences. First, for the merged total highlight area sequence, the system will iterate through each element in the sequence, extracting the recorded highlight area values ​​(ignoring the associated image IDs and using only the values ​​themselves for calculation). All these area values ​​will be summed, and then this sum will be divided by the total number of elements in the sequence (i.e., the total number of images, including one reference view image). (Images and all non-reference viewpoint images) are calculated using this standard arithmetic mean method to obtain a single value, which is the average highlight area value. In the same way, the system processes the total sequence of grayscale standard deviations, that is, iterates through the sequence, extracts the overall grayscale standard deviation values ​​of all images, sums them up, and then divides by the total number of images to calculate the average grayscale standard deviation value. The average highlight area value and the average grayscale standard deviation value obtained by these two calculations together characterize the overall average highlight characteristics and average grayscale texture characteristics of the Chinese herbal medicine sample images acquired from multiple perspectives. Combining these two average values ​​(for example, forming a simple data structure or record containing two floating-point numbers) generates the final comprehensive surface characteristic data.

[0047] The steps for obtaining the initial medicinal material characteristic pairs are as follows: Based on comprehensive surface characteristic data, the average highlight area value and the average gray standard deviation value are extracted. The corresponding fields are read and the consistency of the value type, value range and unit is verified to generate single-value indexes of highlight area and gray standard deviation. Based on the single-value index of highlight area and the single-value index of grayscale standard deviation, a two-element data structure is constructed. The data is written into the structure in the order of "area-grayscale" and the structure is assigned field names and identifier numbers to form the initial medicinal material feature pairs.

[0048] Specifically, based on the comprehensive surface characteristic data generated in the preceding steps, which includes the calculated average highlight area and average grayscale standard deviation, these two core values ​​need to be extracted from the comprehensive surface characteristic data structure. The system will access the designated fields storing the average highlight area and the average grayscale standard deviation, respectively, and read their respective values. After successful reading, a series of strict verification procedures are immediately performed on these two values. The first step is value type verification, ensuring that the average highlight area (e.g., expected to be floating-point or integer) and the average grayscale standard deviation (e.g., expected to be floating-point) both conform to the predefined value data type. If the types do not match, an error is recorded or a safe conversion is attempted. The second step is value range verification. For the average highlight area, its theoretical value range should be non-negative and not exceed the total number of pixels in the image (e.g., if the standard image size is 2048x2048 pixels, then the upper limit is 4,194,304 pixels). In actual operation, a more accurate range will be set based on historical data statistics. A reasonable upper limit for verification, such as 500,000 pixels, is set to exclude anomalies, i.e., to verify whether it falls within the [0, 500,000] pixel range. This upper limit of 500,000 pixels is derived by analyzing the average highlight area value of a large number of different types of Chinese medicinal herb sample images, covering 99.9% of the normal sample range. The average grayscale standard deviation value represents the dispersion of the grayscale distribution. Theoretically, for grayscale levels of 0-255, its value should be within the range of [0, 127.5]. This is also based on empirical data, such as setting actual... The verification range is [0, 80] to ensure that it falls within the range of typical medicinal material surface texture. This upper limit of 80 is based on the statistical analysis of the average grayscale standard deviation of more than 10,000 different medicinal material sample images. It was found that the vast majority (99.5%) of the samples had this index below 80. The third item is unit consistency verification to ensure that the unit of the average highlight area value is "pixel". Only when both values ​​successfully pass all the above verifications are they confirmed as valid, and the single value index of highlight area and the single value index of grayscale standard deviation are generated respectively.

[0049] Based on the valid highlight area and grayscale standard deviation metrics generated and confirmed in the previous step, the system will now integrate these two independent numerical metrics into a single structured data entity. Specifically, a new two-element data structure instance will be created in memory. This structure is specifically designed to store this pair of feature values. During data writing, the preset "area-grayscale" order will be strictly followed: the highlight area metric will be stored in the first data slot of the structure, and the grayscale standard deviation metric will be stored in the second data slot. To improve data readability and ease of subsequent program calls, the system will assign appropriate field names to this newly created structure and its two internal data slots. For example, the structure itself could be named "MedicinalMaterialFeatures," and its first field (corresponding to the highlight area metric) could be named "AverageHighlightPixelArea." The second field (corresponding to the single-value index of grayscale standard deviation) is named "AverageGrayscaleStandardDeviation". After data filling and field naming are completed, in order to ensure the uniqueness and traceability of this feature pair in the entire data system, the system will also assign it a globally unique identifier. The generation rule of this identifier can combine the batch number, sample number and timestamp of the currently processed Chinese medicinal material sample, for example, "MMF_Batch02_Sample105_20250521183000", where "MMF" represents the medicinal material feature pair, "Batch02_Sample105" refers to the specific sample information, and "20250521183000" is the generation time accurate to the second. This identifier will serve as the primary key or unique index of the feature pair. Through the above construction, assignment, naming and numbering process, an initial medicinal material feature pair with a clear structure and complete information is finally formed.

[0050] The steps for obtaining the surface characteristic identification results of medicinal materials are as follows: Based on the average highlight area and average gray standard deviation extracted from the initial medicinal material feature pairs, the mean edge gradient parameter corresponding to the same image is selected from the standardized multi-view image set, and data type consistency verification and normalization conversion are performed on the three parameters to obtain the highlight area value, gray standard deviation value and mean edge gradient value. The characteristic response value of the medicinal material is calculated based on the area of ​​the highlight region, the standard deviation of gray level, and the mean edge gradient. The calculation formula is as follows: ; in, This represents the response value of the medicinal material's characteristics. This represents the area of ​​the highlight region (in pixels). This represents the standard deviation of gray levels (unit: dimensionless gray level). This represents the average edge gradient value (unit: grayscale difference / pixel). It represents the absolute value of the relative difference between brightness structure and texture change, and the denominator is the nonlinear control term of the coupling relationship between brightness and texture; Based on the response value of medicinal material characteristics, a zone-by-zone matching is performed with the established standard medicinal material characteristic response intervals. When the response value of medicinal material characteristics falls into any interval, the medicinal material type label of that interval is extracted and bound to the current image number to generate the identification result of the surface characteristics of the medicinal material.

[0051] Specifically, based on the initial medicinal material feature pair generated in the previous steps, this feature pair encapsulates two summary indicators calculated for the current medicinal material sample: the average highlight area and the average grayscale standard deviation. First, these two values, namely the average highlight area and the average grayscale standard deviation, are extracted from this initial medicinal material feature pair. Simultaneously, to introduce a more direct measurement of surface texture, the system also selects a specific image from the previously established standardized multi-view image set to calculate the mean edge gradient parameter. Typically, the image previously determined as the predetermined reference viewpoint is selected because it is considered to best represent the typical viewpoint of the sample. For this selected reference image, it is first converted to grayscale. The image is processed by first analyzing the gradient of each pixel in the image, then using the Sobel operator to calculate the gradient in the horizontal and vertical directions, thus obtaining the gradient magnitude of each pixel. Finally, the arithmetic mean of the gradient magnitudes of all pixels in the entire image is calculated to obtain the mean edge gradient parameter corresponding to the image. After obtaining the three parameters—average highlight area, average grayscale standard deviation, and mean edge gradient—the system performs a data type consistency check on them to ensure that all three are floating-point numbers or numerical types that can be safely converted to floating-point numbers. Subsequently, these three parameters are normalized to eliminate dimensional differences and map them to a unified [0, 1] interval. The conversion method uses min-max normalization, and its calculation formula is as follows: ,in These are the original parameter values. and These are the minimum and maximum values ​​of this parameter observed in a large number of representative sample data. and The value was obtained in advance through statistical analysis of a database containing, for example, one thousand different Chinese medicinal herbs, with one hundred samples of each herb. Specifically, it represents the average area of ​​the highlight region. Set to 0 pixels. Set to 30,000 pixels; average grayscale standard deviation Set to 5 (grayscale level). Set to 75 (grayscale level); mean edge gradient Set to 2 (grayscale difference / pixel). The value is set to 60 (grayscale difference / pixel). All raw values ​​outside this range will be truncated to the boundary of this range before normalization. After this normalization transformation, the highlight area value, grayscale standard deviation value, and edge gradient mean value required for subsequent calculations are obtained.

[0052] formula: The advantage of the formula is that it... The aim is to generate a single response value sensitive to the comprehensive surface characteristics of Chinese medicinal materials by nonlinearly combining three core image features: highlight area (reflecting the macrostructure related to surface gloss and smoothness), grayscale standard deviation (reflecting the overall texture roughness or uniformity of the surface), and mean edge gradient (reflecting the clarity and richness of fine surface textures). Captured surface micro-texture characteristics With macroscopic gloss structure (composed of high gloss area) Logarithmic transformation The relative differences or imbalances between (represented by) the terms, where the relative differences or imbalances between (represented by) the terms, are the ... Taking the logarithm and adding 1 is to compress its value range and process it. In this case, the absolute value ensures that the degree of difference is positive, while the square root adjusts the sensitivity of that difference; the denominator term This constitutes a texture roughness based on the overall texture. and close-up texture The composite regulatory factor, Similar to the combined strength of these two texture features, adding 1 is to avoid the denominator being too small or zero, while the exponent of 1.5 enhances the non-linear effect of the adjustment, making the response value more significantly suppressed when the combined texture feature is strong.

[0053] This is the area of ​​the highlight region, which is the average highlight region area after normalization transformation. The unit is dimensionless (the original unit is pixels). The steps to obtain it are as follows: First, extract the average highlight region area from the initial medicinal material feature pair (for example, the original value is 6000 pixels). Then, based on the minimum possible value of this preset parameter... (e.g., 0 pixels) and the maximum possible value (For example, 30,000 pixels, this range is derived through statistical analysis of a large number of samples), calculated using the min-max normalization formula. For example, if the original average highlight area is 6,000 pixels, then the normalized area will be... .

[0054] This is the grayscale standard deviation value, which is the average grayscale standard deviation after normalization transformation, and the unit is dimensionless (the original unit is grayscale level). The steps to obtain it are as follows: First, extract the average grayscale standard deviation from the initial medicinal material feature pairs (for example, the original value is 26), and then calculate the minimum possible value of this parameter based on the preset parameters. (e.g., 5 gray levels) and the maximum possible value (For example, 75 gray levels, this range is derived through statistical analysis of a large number of samples), calculated using the min-max normalization formula. For example, if the original average gray level standard deviation is 26, then the normalized gray level... .

[0055] This is the mean edge gradient value, calculated for a specific image (e.g., a predetermined reference viewpoint image) and normalized. The unit is dimensionless (the original unit is grayscale difference / pixel). The steps to obtain it are: first, calculate the original mean edge gradient of the selected image (e.g., using the Sobel operator to obtain an original value of 17 grayscale difference / pixel); then, based on the minimum possible value of this preset parameter... (e.g., 2 grayscale difference / pixel) and the maximum possible value (For example, 60 grayscale difference / pixel, this range is derived from statistical analysis of a large number of samples), calculated using the max-min normalization formula. For example, if the original edge gradient mean is 17, then the normalized value is... .

[0056] Calculation process: Substituting the normalized parameter values ​​obtained in the above example: , , .

[0057] Calculate the logarithmic term in the numerator: ; Calculate the absolute value term in the numerator: ; Calculate the square root of the numerator: ; Calculate the denominator item: ; Calculate the denominator item: ; Calculate the denominator item: ; Calculate the 1.5th power term in the denominator: ; Calculate the response value of medicinal material characteristics : ; The results show that for the current Chinese medicinal material sample, the characteristic response value of the medicinal material is 0.1674, calculated by a specific formula based on the features of the highlighted area, gray standard deviation and mean edge gradient extracted and normalized from the image. This response value is a comprehensive quantitative indicator, which will be used to match the feature response intervals in the standard medicinal material database to determine the possible type of the medicinal material.

[0058] Based on the medicinal material characteristic response value calculated in the previous step This is a numerical value representing the overall surface characteristics of the current Chinese medicinal herb sample (e.g., 0.1674). The system then performs a type matching operation, which relies on a pre-established database of standard medicinal herb characteristic response intervals. This database is created by performing the same image acquisition and feature extraction process on a large number (e.g., thousands) of authoritatively identified and known categories of standard Chinese medicinal herbs (including calculating their respective...). After (value), for each standard medicinal material The database is constructed by performing statistical analysis on the value distribution (e.g., calculating the 5th and 95th quantiles, or the mean plus or minus a certain number of standard deviations). Each entry in the database contains a type label for a medicinal material (such as "ginseng" or "angelica") and its corresponding... One or more valid ranges of values ​​(e.g., the range corresponding to "ginseng") The interval may be [0.15, 0.25]), and the system will calculate the response value of the medicinal material characteristics for the current sample. The characteristic response range of each standard medicinal material stored in this database is compared one by one to determine its characteristics. Does the value fall within one or more intervals? If a value (e.g., 0.1674) falls within the characteristic response range of a certain standard medicinal material (e.g., "Astragalus membranaceus", whose standard response range is set to [0.16, 0.19]), the system extracts the medicinal material type label "Astragalus membranaceus" for that range. If multiple ranges contain this value, the system will extract the medicinal material type label "Astragalus membranaceus". The value can be processed according to preset rules, such as selecting the label corresponding to the narrowest matching interval, or marking it as multiple possible types. Then, this extracted medicinal material type label is bound to the unique identifier of the sample being processed (i.e., the "identifier number" generated from the initial medicinal material feature pair, such as "MMF_Batch02_Sample105_20250521183000", which is associated with the original sample information) to form a record containing the sample identifier and the presumed medicinal material type. Summarizing these records generates the identification result of the surface characteristics of the medicinal material in this analysis.

[0059] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A rapid retrieval method for components of traditional Chinese medicine based on image matching, characterized in that, Includes the following steps: Under preset lighting conditions, multi-angle image information is collected around the target Chinese medicinal material sample. Geometric size adjustment and color space calibration are performed on each image to generate a standardized multi-view image set. Select a predetermined reference view image from the standardized multi-view image set, perform highlight region detection on the selected image and obtain the boundary pixel coordinates of the region, establish reference image highlight region definition data, and calculate the highlight region area and the overall grayscale standard deviation of the image based on the reference image highlight region definition data, which together constitute representative surface feature parameters. Extract each image from the non-reference viewpoint in the standardized multi-view image set, calculate the area of ​​the highlight region and the standard deviation of the overall gray level of each image, summarize the calculation results, establish a supplementary viewpoint feature parameter set, and calculate the average highlight region area and average gray level standard deviation based on the supplementary viewpoint feature parameter set and the representative surface feature parameters to generate comprehensive surface characteristic data. The average area of ​​the highlight region and the average standard deviation of gray level in the comprehensive surface characteristic data are taken and combined into an initial value pair to establish an initial medicinal material feature pair. Based on the initial medicinal material feature pair, the surface characteristic identification result of the medicinal material is generated.

2. The method for rapid retrieval of Chinese medicinal herb components based on image matching according to claim 1, characterized in that, The steps for obtaining the standardized multi-view image set are as follows: Under preset lighting conditions, surround the target Chinese medicinal material sample and use a camera to acquire multi-angle images of the Chinese medicinal material sample at preset angle intervals. Each acquisition records the camera angle parameters and the real-time brightness of the lighting source. All acquired raw images are numbered one by one with the corresponding angle parameters and brightness values ​​to generate a multi-angle raw image set with angle parameter and brightness value labels. Based on the multi-angle original image set with angle parameters and brightness values, each original image is geometrically scaled and cropped using bilinear interpolation according to the set target pixel aspect ratio and target resolution parameters. All size-adjusted images are archived using the angle parameters and brightness values ​​of each image as unique identifiers to obtain a standard-size multi-angle image set. Based on the standard-sized multi-angle image set, the color deviation coefficient of each channel is obtained by comparing the RGB components of the pixels in the standard color card area with the standard values ​​of each standard-sized image. The color deviation coefficient is then used to correct the color components of all pixels in the entire image. The corrected color space parameters of each image are marked, and all images with completed color space calibration are organized to form a standardized multi-angle image set.

3. The method for rapid retrieval of Chinese medicinal herb components based on image matching according to claim 1, characterized in that, The steps for obtaining the highlight region definition data of the reference image are as follows: Based on the angle parameters corresponding to each image in the standardized multi-view image set, the image closest to the target angle is extracted by image number indexing, and this image is used as the predetermined reference view image. A reference image index entry is generated with the image number and shooting angle as unique identifiers. Based on the reference image index entries, the RGB channel values ​​of all pixels in the image are extracted and converted into corresponding brightness values. At the same time, pixel regions with brightness values ​​higher than the preset highlight recognition threshold are identified, the boundary pixel coordinates of the region are extracted and a boundary contour set is constructed, and the boundary discrepancy factor of the region is calculated. Based on the boundary discrete factor, each boundary discrete factor is judged one by one, and the set of boundary contours with all boundary discrete factors below the compactness threshold is retained. The pixel coordinates are extracted as the highlight boundary pixel data to establish the highlight area definition data of the reference image.

4. The method for rapid retrieval of Chinese medicinal herb components based on image matching according to claim 1, characterized in that, The steps for obtaining the representative surface feature parameters are as follows: Based on the highlight area definition data of the reference image, the coordinates of the closed boundary pixels of all highlight areas are extracted, all connected pixels within the boundary pixels are marked in sequence, and each point is judged to be within the closed boundary area and marked as a pixel inside the highlight. The total number of pixels marked as pixels inside the highlight is counted to form the pixel count value of the highlight area of ​​the reference image. Based on the highlight area definition data of the reference image, all pixels in the highlight area of ​​the reference image are masked, and only the pixel values ​​in the non-highlight area are retained. The RGB channel values ​​of each pixel in the remaining area are converted into gray values ​​to construct a gray value vector sequence. Based on the gray value vector sequence, the sum of the mean of all elements and the sum of their respective squared deviations is calculated. The square root is taken and divided by the number of effective pixels to form the overall gray value standard deviation of the image. Based on the number of pixels in the highlight area of ​​the reference image and the standard deviation of the overall grayscale of the image, data pairs are constructed and bound to the image number identifier to generate representative surface feature parameters.

5. The method for rapid retrieval of Chinese medicinal herb components based on image matching according to claim 1, characterized in that, The steps for obtaining the supplementary viewpoint feature parameter set are as follows: The predetermined reference view images are excluded from the standardized multi-view image set, all remaining non-reference view images are extracted, the angle number and image path information of each image are recorded in sequence, and an identification code is generated for each image to generate a non-reference view image index list. Based on the non-reference viewpoint image index list, the image content is read one by one, the pixel enclosure range of the highlight area is calculated, the area of ​​the highlight area is obtained, and at the same time, all pixels in the image are converted into gray values ​​and a gray value vector is constructed. Mean and variance analysis is performed to obtain the overall gray value standard deviation of the image. Two feature values ​​are bound according to the image identification code to generate non-reference viewpoint image feature parameter pairs. Based on the non-reference viewpoint image feature parameter pairs, all image feature parameter pairs are batch-classified to generate a supplementary viewpoint feature parameter set.

6. The method for rapid retrieval of Chinese medicinal herb components based on image matching according to claim 1, characterized in that, The steps for obtaining the comprehensive surface characteristic data are as follows: Based on the supplementary viewpoint feature parameter set, the image number, angle identifier, highlight area value and grayscale standard deviation value of each non-reference viewpoint image are read sequentially. The highlight area values ​​of all images are written into a unified sequence array A, and the grayscale standard deviation values ​​are written into a unified sequence array B. Each array element is bound to the corresponding image number to form a supplementary viewpoint highlight area sequence and a supplementary viewpoint grayscale standard deviation sequence. Based on the supplementary viewpoint highlight area sequence and the supplementary viewpoint grayscale standard deviation sequence, the highlight area value of the reference image extracted from the representative surface feature parameters is appended to the end of the unified sequence array A, and the overall grayscale standard deviation value of the image is appended to the end of the unified sequence array B to generate the merged total highlight area sequence and the total grayscale standard deviation sequence. Based on the merged total sequence of highlight area and the total sequence of grayscale standard deviation, the average of all values ​​in the sequence is calculated to construct the average highlight area value and the average grayscale standard deviation value, thereby generating comprehensive surface characteristic data.

7. The method for rapid retrieval of Chinese medicinal herb components based on image matching according to claim 1, characterized in that, The steps for obtaining the initial medicinal material feature pairs are as follows: Based on the comprehensive surface characteristic data, the average highlight area value and the average gray standard deviation value are extracted. The corresponding fields are read and the consistency of the value type, value range and unit is verified to generate single-value indexes of highlight area and gray standard deviation. Based on the single-value index of highlight area and the single-value index of grayscale standard deviation, a two-element data structure is constructed. The data is written into the structure in the order of "area-grayscale" and the structure is assigned field names and identifier numbers to form the initial medicinal material feature pairs.

8. The method for rapid retrieval of Chinese medicinal herb components based on image matching according to claim 1, characterized in that, The steps for obtaining the surface characteristic identification results of the medicinal materials are as follows: Based on the average highlight area and average gray standard deviation extracted from the initial medicinal material feature pair, the mean edge gradient parameter corresponding to the same image is selected from the standardized multi-view image set, and data type consistency verification and normalization conversion are performed on the three parameters to obtain the highlight area value, gray standard deviation value and mean edge gradient value. The characteristic response value of the medicinal material is calculated based on the area of ​​the highlight region, the standard deviation of gray level, and the mean value of the edge gradient. Based on the medicinal material characteristic response value, a zone-by-zone matching is performed with the established standard medicinal material characteristic response interval. When the medicinal material characteristic response value falls into any interval, the medicinal material type label of that interval is extracted and bound to the current image number to generate the medicinal material surface characteristic identification result.