Vision-based defect detection method for traditional Chinese medicine

By performing pixel alignment and three-dimensional reconstruction on the time-series images and multi-view images of Chinese medicinal material samples, combined with dynamic change feature analysis, the problems of insufficient three-dimensional model accuracy and texture difference recognition in existing technologies are solved, and high-precision detection of surface defects of Chinese medicinal materials is achieved.

CN120374870BActive Publication Date: 2025-09-12CANGNAN COUNTY QIUSHI TRADITIONAL CHINESE MEDICINE INNOVATION RES INST

Patent Information

Application Number
CN202510856241.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-12
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In the existing technology, the vision-based defect detection method for traditional Chinese medicine relies on the manual selection of feature points. The accuracy of the reconstructed three-dimensional model is affected by the operator's experience, and it is difficult to accurately reflect the tiny depressions or protrusions on the surface. In addition, feature extraction focuses on color and texture dimensions, ignoring geometric deformation parameters, making it difficult to distinguish between similar textures such as insect holes and natural wrinkles.

Method used

By regularly collecting time-series images of Chinese medicinal material samples and performing pixel alignment, a registered time-series image sequence and a multi-view image set are generated. The color value difference between adjacent frame images is calculated, and a frame-by-frame change data set is obtained. A sparse three-dimensional point cloud and camera posture are established. Combined with the dynamic change feature map and the three-dimensional model, the surface defect characteristics of Chinese medicinal materials are analyzed and the defect category is output.

Benefits of technology

It improves the ability to identify dynamic changes on the surface of Chinese medicinal materials, enhances the ability to quantitatively characterize the depth of insect holes and the concave and convex features of oil-filled areas, reduces the risk of missed detection due to misjudgment of a single indicator, and improves the accuracy of identifying complex defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374870B_ABST
    Figure CN120374870B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of defect detection of Chinese medicinal materials, specifically a vision-based method for defect detection of Chinese medicinal materials, comprising the following steps: regularly collecting time-series images of Chinese medicinal material samples, collecting multi-perspective images, pixel-aligning the time-series images, structurally organizing and associating camera parameters of the multi-perspective images according to the shooting orientation, and generating a registered time-series image sequence and a multi-perspective image set. The present invention eliminates image offsets caused by ambient light fluctuations and device jitter by regularly collecting time-series images of Chinese medicinal material samples and performing pixel alignment, thereby ensuring the spatiotemporal consistency of dynamic change calculations. The multi-perspective images are structurally organized and associated with camera parameters according to the shooting orientation, and a geometric constraint relationship between perspectives is established to solve the problem of insufficient three-dimensional reconstruction accuracy caused by perspective isolation in traditional methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Chinese medicinal material defect detection, and in particular to a vision-based Chinese medicinal material defect detection method. Background Art

[0002] The vision-based defect detection method for traditional Chinese medicines is used to automatically identify physical defects (such as moisture absorption deformation, oily patches, and insect holes) and abnormal color areas on the surface of traditional Chinese medicines through computer vision technology, so as to classify the defect types and assess their severity.

[0003] Existing techniques rely on manually selected feature points for matching between viewpoints. The accuracy of the reconstructed 3D model is affected by operator experience, making it difficult to accurately represent subtle surface depressions and protrusions. Feature extraction focuses on color and texture dimensions, neglecting quantitative analysis of geometric deformation parameters. This makes it difficult to distinguish between similar textures such as insect holes and natural wrinkles. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a vision-based defect detection method for Chinese medicinal materials.

[0005] In order to achieve the above objectives, the present invention adopts the following technical solution: a vision-based method for detecting defects in traditional Chinese medicine, comprising the following steps:

[0006] Collect time-series images of Chinese medicinal material samples at regular intervals, collect multi-view images, align the pixels of the time-series images, and structurally organize and associate camera parameters of the multi-view images according to the shooting orientation to generate a registered time-series image sequence and a multi-view image set;

[0007] Based on the registered time-series image sequence, the color value difference between the corresponding pixel positions of adjacent frame images in the sequence is calculated to obtain a frame-by-frame variation data set. Based on the frame-by-frame variation data set, the total variation value of each pixel position is accumulated along the time dimension to obtain a dynamic change feature map of the surface of the Chinese medicinal material;

[0008] Based on the multi-view image set, a sparse three-dimensional point cloud and a camera pose are obtained by matching feature points of the same name between images from different viewpoints and combining them with camera parameters. Based on the sparse three-dimensional point cloud and the camera pose, adjacent points in the point cloud are connected to construct triangular mesh patches to establish a three-dimensional model of the sample to be tested;

[0009] Based on the dynamic change characteristic map of the Chinese medicinal material surface and the three-dimensional model of the sample to be tested, the pattern of the change area in the dynamic change characteristic map of the Chinese medicinal material surface is analyzed to see whether it conforms to the dynamic defect characteristics of moisture absorption or oiliness, and a multi-dimensional defect indicator set is obtained. Based on the multi-dimensional defect indicator set, the qualification or defect category of the Chinese medicinal material sample is output, and a Chinese medicinal material defect category determination is established.

[0010] Preferably, the steps of obtaining the registration time-series image sequence and the multi-view image set are:

[0011] Time-series images and multi-view images of the same batch of Chinese medicinal materials are collected regularly. Pixel alignment is performed on the time-series images. SIFT feature points are extracted from adjacent frames and bidirectionally matched. The spatial Euclidean distance of matching point pairs is calculated. Outlier matching point pairs are removed using the RANSAC algorithm. The affine transformation matrix parameters are calculated and applied to the entire image to generate a sequence of registered time-series images.

[0012] Based on the registered time-series image sequence and the multi-view images, the multi-view images are structurally organized, the focal length, principal point coordinates, and radial distortion coefficient in the camera parameter file of each image are parsed, a mapping relationship between the azimuth angle, the pitch angle, and the camera parameters is established, and a multi-view image set is generated;

[0013] Based on the registered temporal image sequence and the multi-view image set, the validity of the temporal alignment is verified by calculating the SSIM structural similarity of adjacent frames in the temporal sequence, and the registered temporal image sequence and the multi-view image set are generated.

[0014] Preferably, the steps of acquiring the frame-by-frame variation data set are:

[0015] Based on the registered time-series image sequence, extracting the RGB three-channel color value of each pixel position in the adjacent frame image, establishing the current frame color value matrix and the previous frame color value matrix, and generating an adjacent frame color value matrix group;

[0016] Calculating the color value difference of corresponding pixel positions based on the adjacent frame color value matrix group;

[0017] Based on the color value difference, all pixel positions and frame sequence indexes are traversed, and the color value difference is superimposed and arranged in frame sequence to generate a frame-by-frame variation data set.

[0018] Preferably, the steps for obtaining the characteristic map of dynamic changes on the surface of the Chinese medicinal material are:

[0019] Based on the frame-by-frame variation data set, extract the color value difference sequence of each pixel position in all frames, establish a time series variation set of the pixel position, and generate pixel-level time accumulation input data;

[0020] Calculating a total change value at each pixel position based on the pixel-level time-accumulated input data;

[0021] Based on the total change value, the sliding window method is used to calculate the mean and standard deviation of the total change in the neighborhood around each pixel, set a dynamic threshold, and screen pixel areas where the total change value is greater than the dynamic threshold to generate a dynamic change feature map of the surface of Chinese medicinal materials.

[0022] Preferably, the steps of acquiring the sparse three-dimensional point cloud and camera posture are:

[0023] Based on the multi-view image set, SIFT feature points are extracted from every two view images and feature descriptors are generated. Bidirectional matching is performed by calculating the Hamming distance between the feature descriptors. The RANSAC algorithm is used to fit the basic matrix and outlier matching point pairs with a projection error exceeding 0.5 pixels are eliminated to generate a set of matching pairs of feature points with the same name between the multi-view images.

[0024] Based on the set of matching pairs of feature points with the same name and the focal length, principal point coordinates, and radial distortion coefficient in the camera parameters, an epipolar geometry constraint equation is constructed. The camera rotation matrix and translation vector are solved by singular value decomposition of the essential matrix. The linear triangulation method is used to calculate the initial 3D spatial coordinates of the feature point pairs, generating the camera extrinsic parameter estimation results and a sparse 3D point coordinate set.

[0025] Based on the camera extrinsic parameter estimation results and the sparse 3D point coordinate set, the rotation matrix, translation vector and 3D point coordinates are iteratively adjusted with the reprojection error as the loss function until the average reprojection error converges to within 0.3 pixels. Abnormal 3D points with residuals exceeding 1.0 pixels after optimization are eliminated to generate a sparse 3D point cloud and camera pose.

[0026] Preferably, the steps of obtaining the three-dimensional model of the sample to be tested are:

[0027] Based on the sparse 3D point cloud and camera pose, the optical center distance between adjacent camera viewpoints, the mean 3D point reprojection error, and the surface curvature change gradient are extracted to establish an interpolation credibility assessment parameter set;

[0028] Calculating a dynamic connection distance threshold based on the interpolation credibility evaluation parameter group;

[0029] Based on the dynamic connection distance threshold, the wavefront advancing algorithm is used to perform regional growth on the point cloud, preferentially connecting adjacent point pairs with a spacing less than the dynamic connection distance threshold and a curvature difference less than 0.1, generating a set of triangular mesh faces with adaptive density, and establishing a three-dimensional model of the sample to be tested.

[0030] Preferably, the steps of obtaining the multi-dimensional defect indicator set are:

[0031] Based on the dynamic change characteristic map of the surface of the Chinese medicinal materials, the pixel point sets of all change areas were extracted. The judgment conditions for the moisture absorption defect feature were set as follows: the color value rising rate within 5 consecutive frames was ≥15% / second and the regional area expansion rate was ≥3% / second; the judgment conditions for the oiliness defect feature were set as follows: the texture direction consistency coefficient was ≤0.3 and the regional brightness mean decreasing gradient was ≥20 lumens / second;

[0032] Based on the three-dimensional model of the sample to be tested and the three-dimensional model of the same grade standard, a rough alignment is performed with the bottom surface of the model as the reference plane, and the model posture is iteratively adjusted until the mean Euclidean distance between the top of the medicinal material and the root bifurcation point is less than 0.5 mm and the maximum single point deviation does not exceed 1.2 mm, to generate a three-dimensional model comparison group after spatial registration;

[0033] Based on the three-dimensional model comparison group after spatial registration, the surface grid is divided into 10mm×10mm detection units. The surface curvature difference, normal vector deflection angle and volume expansion rate are synchronously calculated in each detection unit. The porosity change in the detection unit is superimposed on the hygroscopic defect area, and the surface gloss attenuation rate is superimposed on the oily defect area to generate a multi-dimensional defect indicator set.

[0034] Preferably, the steps for determining the defect category of the Chinese medicinal materials are as follows:

[0035] Based on the multi-dimensional defect indicator set, if the same inspection unit meets the conditions for both a moisture absorption defect candidate and an oil flooding defect candidate, it is preferentially determined as a moisture absorption defect. If the distance between defect candidate categories of different inspection units is ≤5mm, they are spatially adjacent and are merged into a composite defect area and upgraded to a severe defect category, generating a revised defect category determination result set.

[0036] Based on the corrected defect category determination result set, the area proportion and severity level of all defect areas are counted. If the total area proportion of hygroscopic defects is ≥5% or there is a serious defect category area, the sample is judged to be unqualified; if the total area proportion of defects is <5% and there is no serious defect category, a qualified mark is output and a defect category determination of the Chinese medicinal material is generated.

[0037] Compared with the prior art, the advantages and positive effects of the present invention are:

[0038] The present invention collects time-series images of Chinese medicinal material samples at regular intervals and performs pixel alignment to eliminate image offsets caused by ambient light fluctuations and device jitter, thereby ensuring the spatiotemporal consistency of dynamic change calculations. Multi-perspective images are structured and organized according to the shooting orientation and associated with camera parameters to establish geometric constraints between perspectives, thereby solving the problem of insufficient three-dimensional reconstruction accuracy caused by perspective isolation in traditional methods. Based on the registration of time-series image sequences, the color differences of pixels in adjacent frames are calculated, the total changes are accumulated along the time dimension, and a surface dynamic change feature map is generated, which can capture the gradual process of hygroscopic expansion and the diffusion path of oil-flooding areas that are difficult to identify with traditional static detection. Combining multi-perspective same-name feature point matching with camera parameter optimization, a sparse three-dimensional point cloud is constructed, and a surface geometric deformation model is established through triangulation of neighboring points to enhance the quantitative characterization capability of the depth of insect-eaten holes and the concave and convex features of oil-flooding areas. The dynamic change map is integrated with the spatial deformation parameters of the three-dimensional model, and multi-dimensional indicators such as color change rate, volume expansion rate, and curvature difference are simultaneously analyzed. A composite judgment mechanism based on logical rules and spatial distribution characteristics is established to improve the recognition accuracy of complex defects and reduce the risk of missed detection due to misjudgment of a single indicator. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Schematic diagram of the steps of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0041] See also Figure 1 The present invention provides a technical solution, a method for detecting defects in traditional Chinese medicine based on vision, comprising the following steps:

[0042] Collect time-series images of Chinese medicinal material samples at regular intervals, collect multi-view images, align the pixels of the time-series images, and structurally organize and associate camera parameters of the multi-view images according to the shooting orientation to generate a registered time-series image sequence and a multi-view image set;

[0043] Based on the registration of time-series image sequences, the color value differences of corresponding pixel positions between adjacent frame images in the sequence are calculated to obtain a frame-by-frame variation dataset. Based on the frame-by-frame variation dataset, the total variation value of each pixel position is accumulated along the time dimension to obtain the dynamic change feature map of the surface of Chinese medicinal materials;

[0044] Based on a multi-view image collection, a sparse 3D point cloud and camera pose are obtained by matching the feature points of the same name between images from different viewpoints and combining them with camera parameters. Based on the sparse 3D point cloud and camera pose, adjacent points in the point cloud are connected to construct triangular mesh patches to establish a 3D model of the sample to be tested.

[0045] Based on the dynamic change characteristic map of the surface of Chinese medicinal materials and the three-dimensional model of the sample to be tested, the pattern of the change area in the dynamic change characteristic map of the surface of Chinese medicinal materials is analyzed to see whether it conforms to the dynamic defect characteristics of moisture absorption or oil overflow, and a multi-dimensional defect indicator set is obtained. Based on the multi-dimensional defect indicator set, the qualification or defect category of the Chinese medicinal material sample is output, and a Chinese medicinal material defect category judgment is established.

[0046] The steps for obtaining the registration time sequence image sequence and multi-view image set are:

[0047] Time-series images and multi-view images of the same batch of Chinese medicinal materials are collected regularly. Pixel alignment is performed on the time-series images. SIFT feature points are extracted from adjacent frames and bidirectionally matched. The spatial Euclidean distance of matching point pairs is calculated. Outlier matching point pairs are removed using the RANSAC algorithm. The affine transformation matrix parameters are calculated and applied to the entire image to generate a sequence of registered time-series images.

[0048] Based on the registration of time-series image sequences and multi-view images, the multi-view images are structured and organized. The focal length, principal point coordinates, and radial distortion coefficient in the camera parameter file of each image are parsed. The mapping relationship between azimuth angle, pitch angle, and camera parameters is established to generate a multi-view image set.

[0049] Based on the registered temporal image sequence and multi-view image set, the effectiveness of temporal alignment is verified by calculating the SSIM structural similarity of adjacent frames in the temporal sequence, and the registered temporal image sequence and multi-view image set are generated.

[0050] Specifically, after regularly collecting the time-series images and multi-view images of the same batch of Chinese medicinal materials samples, pixel alignment is performed on the acquired time-series images. This process first traverses each pair of adjacent image frames in the time-series image sequence, for example, Frame and Frame, the scale-invariant feature transform (SIFT) algorithm is used to extract the key feature points and their corresponding 128-dimensional feature descriptors in the two frames of images respectively. Then, two-way matching is performed on the feature descriptors extracted from the two frames of images. Specifically, for the first Each feature point in the frame, In the frame, the nearest neighbor and the next nearest neighbor feature point are found based on the Euclidean distance of the feature descriptor, and the ratio of the nearest neighbor distance to the next nearest neighbor distance is calculated. If the ratio is less than the preset distance ratio threshold (for example, set to 0.75, which is an empirical value widely used in SIFT matching strategy and is suitable for distinguishing good matches from fuzzy matches), it is considered a potential match. Otherwise, for the first Each feature point in the frame, The same search and ratio test are performed in the frame. Only when two feature points are the best match for each other, they are considered as a valid initial matching point pair. Subsequently, the spatial Euclidean distance of all these initial matching point pairs in their respective image coordinate systems is calculated. This distance is not used for screening, but as input information or auxiliary judgment basis for subsequent algorithms. Then, the random sample consistency (RANSAC) algorithm is used to eliminate outlier matching point pairs caused by repeated textures or short-term occlusions. The execution process of the RANSAC algorithm is: randomly select 3 (the minimum number of point pairs required for affine transformation) initial matching point pairs, estimate an affine transformation matrix based on these 3 point pairs, and use this matrix to transform the first All matching points in the frame are transformed to the In the coordinate system of the frame, calculate the difference between the transformed point and the The pixel distance between corresponding matching points in the frame, point pairs with a distance less than a specific pixel threshold are considered "inliers". This pixel threshold is set according to the image resolution and the expected maximum pixel drift. For example, for a 1920x1080 resolution image, the expected maximum inter-frame misalignment is 2 pixels, then the threshold can be set to 1.5 pixels. This value is obtained by testing the stability of the imaging system, observing the maximum drift range of the feature points in the stable state and taking 75% of its upper limit. Repeat this process several times (for example, set the number of iterations to 1000 times, which is calculated based on the number of times to ensure a high probability of selecting a sample set that is all inliers, usually based on an estimate of the inlier rate), and select the set of affine transformation matrices that produces the most inliers as the optimal solution. Finally, apply the obtained optimal affine transformation matrix parameters (including 6 parameters: rotation, scaling, translation, and shear) to the first The pixel coordinates of the entire image of the frame are calculated by bilinear interpolation method to calculate the grayscale value or color value of the non-integer coordinate pixel after transformation, and the grayscale value or color value of the pixel after transformation is generated. Frame pixel alignment Frame image, repeat this alignment process for all adjacent frames in the entire temporal image sequence, and finally generate a registered temporal image sequence.

[0051] Based on the registered time-series image sequence generated in the previous step and the multi-view images collected, the disordered multi-view images need to be structured. The specific operation is to first read the camera parameter file associated with each multi-view image. The file is usually in text format, which records the internal parameters of the camera. By parsing the file content, the key parameter values ​​​​are extracted, including the focal length of the camera in the horizontal and vertical directions (respectively recorded as and , the unit is usually pixels), the coordinates of the main point of the image (denoted as , representing the intersection of the camera optical axis and the imaging plane, in pixels), and coefficients describing the radial distortion of the lens (usually including at least Low-order coefficients such as , used to correct the offset of pixels to the center or edge caused by the curvature of the lens). At the same time, it is necessary to obtain the spatial orientation information of the camera when shooting the multi-view image, that is, the azimuth (rotation angle around the vertical axis) and the pitch angle (rotation angle around the horizontal axis). These angle information are usually recorded directly by the encoder that controls the camera gimbal or robotic arm during acquisition, or are known according to a preset shooting point scheme. For example, the camera is installed on a rotating arm with a known radius and takes one picture every 10 degrees. The azimuth angle is an integer multiple of 10 degrees, and the pitch angle is determined according to the shooting level. Then, a data structure is established, such as a hash table or dictionary, and each group of obtained (azimuth, pitch angle) value pairs is used as a key, and the ( ) parameter set as the value, and store them in association. All multi-view images and their parameter files and orientation information are traversed to complete the construction of the mapping relationship. In this way, the scattered multi-view images and their corresponding precise shooting parameters and spatial orientations are integrated to generate a structured multi-view image set.

[0052] Based on the acquired registered temporal image sequence and the structured multi-view image set, in order to finally confirm the validity of the temporal image alignment, a verification process needs to be performed. This process calculates the accuracy of each pair of adjacent frames in the registered temporal image sequence (for example, the first Frame aligned with The specific steps of calculating the SSIM value are as follows: the two images to be compared are divided into several overlapping or non-overlapping calculation windows (for example, using a size of The sliding window of pixels, with a step size of 1 pixel, the choice of window size is based on the experience of balancing computational efficiency and sensitivity to local details). For each pair of windows and , calculate their pixel averages (representing brightness estimation ), pixel standard deviation (representing contrast estimation ) and covariance (representing a measure of structural similarity ), and then calculate the SSIM value of the window using the following formula: ,in, and is a stability constant introduced to prevent the denominator from approaching zero. Represents the dynamic range of the image pixel values ​​(for example, for an 8-bit grayscale image, ), and are two small constants, implemented according to the SSIM standard, usually set to and These values ​​are empirical values ​​selected based on the simulation of the perceptual characteristics of the human visual system. After calculating the SSIM values ​​of all windows, take the average of the SSIM values ​​of all windows to obtain the average SSIM (MSSIM) value between the two frames of images. Set an MSSIM threshold to determine whether the alignment is effective. The setting of this threshold is based on the evaluation of the actual alignment effect. For example, by observing the distribution of MSSIM values ​​of a large number of manually confirmed well-aligned image pairs, it is found that their values ​​are generally higher than 0.95, while the values ​​of poorly aligned image pairs are significantly lower than this value. Taking into account the possible small real changes on the surface of medicinal materials, a slightly lower threshold can be set, for example , if the calculated MSSIM values ​​of any adjacent frame pair are greater than or equal to , then the alignment of the entire time series is considered valid. After this verification, the obtained registered time series image sequence and multi-view image set are formally generated and confirmed for subsequent processing.

[0053] The steps for obtaining the frame-by-frame variation dataset are as follows:

[0054] Based on the registration time sequence image sequence, the RGB three-channel color value of each pixel position in the adjacent frame image is extracted, the current frame color value matrix and the previous frame color value matrix are established, and the adjacent frame color value matrix group is generated;

[0055] Based on the color value matrix group of adjacent frames, the color value difference of the corresponding pixel position is calculated. The calculation formula is:

[0056] ;

[0057] in, is the color value difference, The current frame Rank RGB component values ​​of column pixels, is the RGB component value of the corresponding position in the previous frame;

[0058] Based on the color value difference, all pixel positions and frame sequence indexes are traversed, and the color value differences are superimposed and arranged in frame sequence to generate a frame-by-frame change data set.

[0059] Specifically, based on the registered time-series image sequence generated in the previous step, each pair of temporally adjacent image frames in the sequence needs to be processed. Specifically, the currently processed frame is set as the first Frame (where Start from 2 and go up to the total number of frames in the sequence ), the previous frame is Frame, traverse the Each pixel position in the frame image is represented by its row coordinates and column coordinates The only one that determines From 1 to the image height , From 1 to image width ), for each pixel position , read and record it in the The intensity values ​​of the three color channels of red, green, and blue in the frame image are usually 8-bit unsigned integers ranging from 0 to 255. The RGB values ​​of all pixels in the frame are organized into a dimension of The data structure of the current frame color value matrix is ​​called , in the same way, read and record the Frame image at the corresponding pixel position The RGB color values ​​of the frame are organized into a color value matrix of the preceding frame, recorded as , these two matrices and It contains all the pixel-level color information required for subsequent color difference calculations. These two matrices are used as a processing unit to generate a matrix group of adjacent frame color values ​​for calculating the color difference between frames.

[0060] formula: The benefit of this formula is that, by calculating the Euclidean distance in the RGB color space, it can quantitatively and comprehensively reflect the overall color change of a single pixel between two adjacent moments, rather than being limited to changes in brightness. It is equally sensitive to changes in the red, green, and blue channels, making it suitable for capturing subtle or significant changes in the surface color of Chinese medicinal materials caused by moisture absorption (usually darkening of the color) or oiliness (usually darkening of the color or the appearance of oiliness leading to changes in reflectivity). This provides a fundamental, pixel-by-pixel change metric for the subsequent identification of these dynamic defects.

[0061] Calculation process: This formula calculates the straight-line distance between two points in the three-dimensional RGB color space (representing the colors of the same pixel in the previous and next frames respectively).

[0062] The calculation steps are as follows:

[0063] 1. Calculate the square of the red channel difference: ;

[0064] 2. Calculate the square of the green channel difference: ;

[0065] 3. Calculate the square of the blue channel difference: ;

[0066] 4. Add the three squared differences: ;

[0067] 5. Take the square root of the addition result: Let’s use the actual example: calculate the color value difference between the 4th frame and the 5th frame at the pixel position (500, 800) :

[0068] ;

[0069] ;

[0070] ;

[0071] ;

[0072] ;

[0073] The result shows that at the pixel position (500, 800), a color change of about 7.348 occurs from the 4th frame to the 5th frame. This value itself has no absolute physical meaning, but it relatively represents the amplitude of the color change. The larger the value, the more significant the color change. A value of 0 means that the color has not changed at all. The calculated color value difference is It will serve as basic data to construct the frame-by-frame change dataset required for subsequent analysis and is the key input information for identifying dynamically changing areas.

[0074] The color value difference calculated in the previous step, representing the color change of each pixel between adjacent frames , it is necessary to integrate the change information scattered between different frame conversions. The specific operation is, first, for the entire registration time series image sequence (a total of Frame), calculate the first frame to the second frame, the second frame to the third frame, ..., the Frame to The amount of difference in color values ​​at all pixel locations in the frame, which produces and image size ( ) the same difference matrix, where The matrix contains all pixels from the Frame to The amount of color value difference between frames , then, in order to construct the required dataset, it is necessary to traverse all pixel positions in the image (in From 1 to , From 1 to ), for each fixed pixel position , extract it from all The difference in color values ​​between frames, that is, the sequence This sequence records the frame-by-frame difference history of the pixel color over time. Then, this time series change of all pixel positions is organized together. A three-dimensional data structure (such as a three-dimensional array) can be used, whose dimensions are (image height , image width , frame conversion times ), where the stored value is the color value difference of the corresponding pixel at the corresponding frame conversion moment, or other equivalent data structures such as dictionary or list nesting are used to collect these color value differences arranged in frame order (or superimposed in time order) to generate a structured frame-by-frame change data set.

[0075] The steps for obtaining the characteristic map of dynamic changes on the surface of Chinese medicinal materials are as follows:

[0076] Based on the frame-by-frame variation data set, the color value difference sequence of each pixel position in all frames is extracted, the time series variation set of the pixel position is established, and the pixel-level time accumulation input data is generated;

[0077] Based on the pixel-level time-accumulated input data, the total change value of each pixel position is calculated using the following formula:

[0078] ;

[0079] in, is the total change value, For the Pixels in frame The amount of color value difference, For the The variance of the color value difference of the whole frame, is the total number of frames;

[0080] Based on the total change value, the sliding window method is used to calculate the The mean of the total variation in the neighborhood and standard deviation , set dynamic threshold , filter the pixel areas whose total change value is greater than the dynamic threshold, and generate the dynamic change characteristic map of the surface of Chinese medicinal materials.

[0081] Specifically, the frame-by-frame variation dataset generated in the previous step, which contains the color value differences of each pixel between adjacent frames, needs to be further processed to focus on the complete history of each pixel's own changes over time. The operation process is to traverse the dataset and for each specific pixel position in the image (in is the row index, (column index) extracts the color value difference calculated at all frame conversion moments corresponding to the pixel position from the data set. (in Represents the sequence number of frame conversion, starting from the second conversion To the last conversion , is the total number of frames), these extracted differences are sorted in chronological order (i.e. according to the frame conversion sequence number ) are organized into a sequence or list, e.g., for pixels , the sequence may be This sequence constitutes a set of time series changes at the pixel position, which records the complete dynamic process of the color change at that point. Repeat this extraction and organization process to bring together the time series change sets of all pixels to form a structured data set, which is the pixel-level time accumulation input data used to calculate the total change of each pixel in the next stage.

[0082] formula: The benefit of the formula is that it provides a method to measure the total amount of color change of pixels during the entire observation time, and introduces a normalization factor based on the global change variance. Taking the absolute value for accumulation ensures the cumulative nature of the change, and the denominator According to the intensity of the color change of the entire image at the moment of frame conversion (in terms of variance Measure) to adjust the contribution weight of the current frame difference. When the global change is drastic (such as sudden change in illumination causing is very large), the denominator increases, the contribution of the single point difference is suppressed, on the contrary, when the global change is stable ( Small), the denominator is close to 1, the contribution of the single point difference is more significant, this design makes the calculated total change It can better highlight those local pixel changes that are more abnormal relative to the global background changes of the same period, which helps to distinguish real local dynamics (such as defect development) from global interference;

[0083] Instructions for obtaining formula parameters:

[0084] Parameter representation and Frame-related pixels The color value difference of Frame to The color value difference of the frame, that is, ,in is obtained from the previously generated "frame-by-frame variation dataset". , since there is no preceding frame, define ,for arrive , read the corresponding pixel directly from the dataset In the Secondary frame conversion (i.e. ), for example, for pixels , the sequence of changes from frame 1 to frame 10 ( arrive ) could be: [5.2, 3.1, 4.5, 6.0, 5.8, 4.9, 7.1, 6.5, 5.5], plus ;

[0085] Parameter representation and The variance of the color value difference of the whole image associated with the frame, that is, Frame to The amount of color value difference between all pixels during frame conversion For example, the variance sequence of each frame conversion from frame 1 to frame 10 is calculated ( arrive ) is: [10.5, 8.2, 12.1, 15.0, 13.5, 11.8, 16.2, 14.8, 13.0], plus ;

[0086] Calculation process:

[0087] Counting pixels The total change in value ,use and the corresponding sequence sum Substitute the sequence into the calculation to get .

[0088] This result shows that the pixels The total color change accumulated during the entire observation period (10 frames) and normalized by the global change variance is approximately 13.371.

[0089] formula: The benefit of the formula is that it defines a dynamic threshold based on the statistical characteristics of the local neighborhood, which is used to filter out the pixel areas with significant changes from the total change value map, and calculate the local mean using a sliding window. and local standard deviation , so that the threshold It can adapt to the baseline level and fluctuation amplitude of background changes in different areas of the image. Compared with the global threshold, it can better handle the situation of uneven background changes, avoid missed detection in areas with gentle changes and false detection in areas with drastic changes, and is used to identify extreme values ​​or outliers relative to the local environment.

[0090] Instructions for obtaining formula parameters: The parameter represents the pixel Centered at The total change value of all pixels in the neighborhood (window) of The arithmetic mean of the total change value of all pixels needs to be extracted from the map of the corresponding window for calculation. The image boundary area is usually processed by zero padding, mirroring or repeating the boundary pixels to ensure that the window is always Size, for example, if , then calculate 5x5=25 pixels in the center The average of the values;

[0091] Parameters represent the same The total change value of all pixels in the neighborhood The standard deviation of the total change value, which measures the dispersion or volatility of the total change value in the neighborhood, also needs to be calculated by extracting data from the total change value map;

[0092] The parameter is the side length of the sliding window (number of pixels), which needs to be set in advance. Its selection affects the scale of detection and sensitivity to noise. Smaller Sensitive to small changes but easily affected by noise, larger The smoothing effect is good but may blur the boundaries of small features. Its value is usually set empirically based on the surface texture characteristics of the medicinal material and the expected defect size or selected through experimental optimization, for example, by testing different Value (such as pixels), observe which size can best segment the defect area while suppressing noise, and determine a suitable In this example, set ;

[0093] Substituting the parameters into the formula, we get 17.3. This result shows that for pixels and its neighborhood, its dynamic threshold is calculated as 17.3, and in the subsequent screening step, the total change value of the pixel itself will be (In the previous example, the neighborhood center value is 13.4) and compared with this threshold value of 17.3, if , then the pixel is marked as part of the dynamic change area, otherwise it is not marked. This calculation and comparison process is repeated for all pixels, and finally a binary image is obtained, in which the pixels with a value of 1 constitute the dynamic change feature map of the identified Chinese medicinal material surface.

[0094] The steps to obtain sparse 3D point cloud and camera pose are:

[0095] Based on a multi-view image set, SIFT feature points are extracted from every two-view images and feature descriptors are generated. Bidirectional matching is performed by calculating the Hamming distance between feature descriptors. The RANSAC algorithm is used to fit the basic matrix and eliminate outlier matching point pairs with projection errors exceeding 0.5 pixels. This generates a set of matching pairs of feature points with the same name between multi-view images.

[0096] Based on the set of matching pairs of feature points with the same name and the focal length, principal point coordinates, and radial distortion coefficient in the camera parameters, an epipolar geometry constraint equation is constructed. The camera rotation matrix and translation vector are solved by singular value decomposition of the essential matrix. The linear triangulation method is used to calculate the initial 3D spatial coordinates of the feature point pairs, generating the camera extrinsic parameter estimation results and a sparse 3D point coordinate set.

[0097] Based on the camera extrinsic parameter estimation results and the sparse 3D point coordinate set, the rotation matrix, translation vector and 3D point coordinates are iteratively adjusted with the reprojection error as the loss function until the average reprojection error converges to within 0.3 pixels. Abnormal 3D points with residuals exceeding 1.0 pixels after optimization are eliminated to generate a sparse 3D point cloud and camera pose.

[0098] Specifically, based on the multi-view image set generated in the previous step, we first need to find the corresponding feature points between images from different viewpoints. The specific operation is to select any two images from the set with different viewpoints, apply the scale-invariant feature transform (SIFT) algorithm to each image, detect key points at different scales by constructing a Gaussian difference pyramid, and calculate a 128-dimensional feature descriptor for each key point. Next, match the feature descriptor sets extracted from the two images, measure the similarity of the descriptors by calculating the Hamming distance between the feature descriptors, and use a two-way matching strategy to screen reliable matches: for each feature point in the first image, find the matching point with the smallest Hamming distance in the second image, and vice versa. Only when two feature points are the best match for each other are they considered as a candidate matching pair. Subsequently, the random sample consensus (RANSAC) algorithm is used to estimate the basic matrix between the two images, and further eliminate outlier matching point pairs that do not meet the epipolar geometry constraints. The RANSAC process is: repeatedly randomly select 8 pairs (the minimum number of point pairs required to calculate the basic matrix) of candidate matching point pairs, calculate a basic matrix model based on these 8 pairs of points, and then use the The model verifies all candidate matching point pairs and calculates the distance of each point to its corresponding epipolar line (i.e., projection error). If the distance is less than a preset pixel threshold, the point pair is considered an inlier. The pixel threshold here is set to 0.5 pixels. The basis for setting this value is: for a well-calibrated camera and accurate feature point extraction, the actual projection error of the matching point pair is usually expected to be at the sub-pixel level. Setting 0.5 pixels is a relatively strict standard, which aims to filter out large deviations caused by inaccurate feature positioning or incorrect matching, and ensure that the matching point pairs used for subsequent calculations have high accuracy. The threshold is usually obtained by performing a matching test on image pairs with known geometric relationships, analyzing the distribution of the inlier projection error, and selecting a boundary value that can effectively distinguish between inliers and outliers. The RANSAC algorithm will iterate (for example, setting 1000 iterations or until the inlier set is stable) and select the basic matrix that can obtain the largest number of inliers as the optimal estimate. Finally, all matching point pairs determined as inliers by the optimal basic matrix model constitute the matching pairs of feature points with the same name between the two view images. This process is repeated for all relevant image pairs in the multi-view image set to generate a set of matching pairs of feature points with the same name between the multi-view images.

[0099] Based on the matching pair set of feature points with the same name between the multi-view images obtained in the previous step, and combined with the camera intrinsic parameters (including focal length) corresponding to each image previously structured and stored in the multi-view image set , principal point coordinates and radial distortion coefficient ), start to estimate the relative posture (rotation and translation) between cameras and the 3D spatial position of feature points. First, for any pair of view images (such as view 1 and view 2) that have found matching feature points with the same name, use their camera intrinsic parameters (denoted as ) and the fundamental matrix (denoted as ), through the relationship Calculate the Essential Matrix Before applying this formula, it is necessary to use the radial distortion coefficient in the camera parameters to dedistort the two-dimensional image coordinates of the matching point pair to obtain the corrected coordinates and , the essential matrix Contains the rotation and translation information between the two camera coordinate systems and satisfies the epipolar constraints , next, the essential matrix Perform singular value decomposition, i.e. ,in is an orthogonal matrix, It is a diagonal matrix. Theoretically, for the essential matrix, its singular value form should be , four possible camera rotation matrices can be derived from the results of SVD and translation vectors Solution (usually fix the first camera pose to the identity matrix and the zero vector as the world coordinate system), in order to determine the only correct solution, a chirality check is required, that is, using these four possible sets of Construct the projection matrix with the camera internal parameters, and perform 3D reconstruction on at least one pair of feature points with the same name, and check whether the reconstructed 3D points are in front of the two cameras (that is, the depth value is positive). Only a pair that meets this condition is is a physically possible solution, choose the correct After that, they constitute the external parameter estimation result of camera 2 relative to camera 1. At the same time, for all matching pairs of feature points with the same name that meet the epipolar constraint, , using the linear triangulation method, using the known projection matrices of the two cameras (derived from the internal parameters and the external parameters solved) composition) and the corresponding two-dimensional point coordinates , by solving an overdetermined linear equation system to calculate the initial coordinates of each feature point pair in three-dimensional space , perform triangulation calculations on all matching point pairs, and collect all calculated and 3D point coordinates , generate the initial camera external parameter estimation results and sparse 3D point coordinate set.

[0100] Based on the initial camera external parameter estimation results obtained in the previous step (including the rotation matrix of each camera view relative to the reference view and translation vector T) and the corresponding sparse three-dimensional point coordinate set (including the initial three-dimensional spatial position of each feature point ), it is necessary to jointly optimize through bundle adjustment to improve the overall accuracy and consistency of camera pose and 3D point cloud. The core idea of ​​bundle adjustment is to minimize the reprojection error. The reprojection error is defined as: for a 3D space point , using the internal parameters of a camera 、External Reference Project it back to the image plane of the camera to get the predicted two-dimensional projection coordinates , the predicted coordinates and the two-dimensional feature point coordinates corresponding to the actually observed three-dimensional point The pixel distance between (usually Euclidean distance) is the reprojection error, expressed as , bundle adjustment constructs a large nonlinear least squares problem, whose loss function is the sum of the squares of the reprojection errors of all observed feature points in all camera views: ,in, Traverse all camera perspectives, Traverse all 3D points, It is The three-dimensional point in The observation coordinates in the camera, is a weight factor (usually 1, indicating that the point In camera The optimization process adjusts the rotation matrices of all cameras simultaneously through an iterative algorithm (such as the Levenberg-Marquardt algorithm). , translation vector And the coordinates of all 3D points , until the objective function converges. The convergence criterion is set as the average reprojection error of all valid observation points is less than 0.3 pixels. This 0.3 pixel threshold is set based on the experience of high-precision 3D reconstruction tasks. It aims to achieve higher reconstruction accuracy. Usually, a threshold with a good balance between accuracy and computational cost is determined by evaluating and comparing the accuracy of reconstruction results of multiple different scenes. When the average reprojection error meets the convergence condition, the optimization process stops. At this time, the optimized camera pose and 3D point coordinates are obtained. Finally, an outlier removal is required to check the final reprojection error (also called residual) of each 3D point in all cameras that observe it. If a 3D point If the post-optimization residual in at least one camera exceeds 1.0 pixel, the point is considered an abnormal 3D point (possibly caused by an initial matching error or a local minimum during the optimization process) and is removed from the 3D point set. This 1.0 pixel threshold is more relaxed than the convergence threshold (0.3 pixels) and is used to remove points that still show large inconsistencies even after optimization. Its setting is also based on experience, with the goal of retaining most of the correctly reconstructed points while removing obviously erroneous structural points. After optimization and elimination, an accurate sparse 3D point cloud and camera pose are ultimately generated.

[0101] The steps for obtaining the three-dimensional model of the sample to be tested are:

[0102] Based on the sparse 3D point cloud and camera pose, the optical center distance between adjacent camera viewpoints, the mean 3D point reprojection error, and the gradient of the surface curvature change rate are extracted to establish an interpolation credibility evaluation parameter set.

[0103] Based on the interpolation credibility evaluation parameter group, the dynamic connection distance threshold is calculated using the following formula:

[0104] ;

[0105] in, is the dynamic connection distance threshold, is the average distance between neighboring points in the sparse point cloud, is the distance between the optical centers of adjacent cameras, is the mean of the 3D point reprojection error, is the gradient modulus of the local curvature change rate of the point cloud;

[0106] Based on the dynamic connection distance threshold, the wavefront advancing algorithm is used to perform region growing on the point cloud. The adjacent point pairs with a spacing less than the dynamic connection distance threshold and a curvature difference less than 0.1 are preferentially connected to generate a set of triangular mesh faces with adaptive density to establish a three-dimensional model of the sample to be tested.

[0107] Specifically, based on the sparse 3D point cloud obtained after optimization in the previous steps and the precise posture of each camera (rotation matrix and translation vectors ), in order to guide the subsequent process of generating surface meshes from sparse point clouds, it is necessary to first evaluate the interpolation credibility of different regions in the point cloud. This is achieved by calculating a set of parameters that reflect the local data quality and geometric complexity. The specific operation is as follows: First, identify spatially adjacent camera viewpoint pairs, which can be calculated by calculating the optical centers of all cameras (the translation vectors in the camera poses Given), the distance between the two pairs is considered to be adjacent pairs if the distance is less than a certain range (for example, less than 1.5 times the average camera spacing). Then, for each pair of adjacent camera perspectives, the Euclidean distance between their optical centers is calculated and recorded as , which reflects the length of the observation baseline. Next, all the 3D points involved in the reconstruction of this pair of adjacent viewpoints are extracted, and the average reprojection error of these 3D points at these two specific viewpoints after the bundle adjustment is calculated, which is recorded as , which reflects the consistency of the positioning accuracy of these points under these two perspectives. Low error indicates high confidence. In addition, it is necessary to evaluate the local geometric complexity of the point cloud. For each point in the sparse 3D point cloud, the local surface curvature at the point is estimated by analyzing the distribution of its k nearest neighbor points (for example, k = 10 or 15) (for example, calculating the principal curvature). ), further calculate the speed of change of curvature in space, that is, the gradient of the rate of change of surface curvature, and take its modulus, recorded as , high gradient mode length means that the surface shape changes dramatically, combining these three calculated parameters (the distance between the optical centers of adjacent cameras , mean 3D point reprojection error , the gradient modulus of the local curvature change rate of the point cloud ), organize them to form a group of interpolation credibility assessment parameters associated with different regions of the point cloud or adjacent viewpoint pairs.

[0108] formula: The benefit of the formula is that it dynamically calculates a connection distance threshold based on the quality and geometric characteristics of the local data. , which is used to guide the subsequent surface mesh construction process, replacing the problem of using a global fixed threshold that may lead to overly conservative connections in high-quality areas and overly aggressive connections in low-quality or complex areas. The molecular part of the formula Average point spacing based on point cloud Scaling, scaling factor Taking into account the observation geometry (baseline ) and reconstruction accuracy (reprojection error ), when the baseline is long or the error is small, it allows connecting to a longer distance, indicating that the reconstruction of this area is more reliable; the denominator part The penalty for geometric complexity and observation conditions is introduced. When the surface curvature changes drastically ( Large) and a long observation baseline ( When the value is large, this item increases, resulting in Reduced, so that a more conservative connection strategy is adopted in areas with complex shapes or poor observation conditions, avoiding crossing features or generating wrong connections in sparse areas, The use of the function limits the growth of the denominator, making the penalty effect smoother and achieving adaptive adjustment of the connection distance overall;

[0109] Instructions for obtaining formula parameters:

[0110] The parameter represents the average distance between neighboring points in a sparse point cloud, reflecting the overall or local density of the point cloud. It is calculated by traversing each point in the point cloud and finding its nearest neighbors (e.g. ), calculate this point to here The average distance of neighbors, and finally the average neighbor distance of all points is averaged again to get the global For example, a sparse point cloud containing 10,000 points is calculated, and the average distance between neighboring points is mm;

[0111] The parameter is the distance between the optical centers of adjacent camera views. This value is extracted from the "Interpolation Credibility Evaluation Parameter Group" generated in the previous step and is related to the local area of ​​the point cloud corresponding to the current calculation threshold. For example, the current area is mainly composed of the distance between the optical centers. Observed by a pair of cameras with a resolution of mm;

[0112] The parameter is the average reprojection error (unit: pixel) of the 3D points related to the adjacent camera views. This value is also extracted from the "Interpolation Credibility Evaluation Parameter Group" and reflects the accuracy of the positions of these 3D points on the corresponding 2D image. For example, the average reprojection error corresponding to this area is Pixels;

[0113] The parameter is the gradient modulus of the surface curvature change rate of the point cloud in this local area, which quantifies the speed of the change of the surface curvature. This value is also extracted from the "Interpolation Credibility Evaluation Parameter Group". It is necessary to first estimate the curvature of each point in the point cloud (such as obtaining the normal vector and curvature through principal component analysis of the neighborhood covariance matrix), and then calculate the gradient size of the curvature in space. For example, the area is calculated ;

[0114] Calculation process: Use the parameter values ​​obtained above to calculate the dynamic connection distance threshold : mm, mm, Pixels, mm ;

[0115] Calculate the numerator:

[0116] ;

[0117] Evaluate the terms in the denominator:

[0118] ;

[0119] Calculate the denominator:

[0120] ;

[0121] calculate :

[0122] ;

[0123] The results show that in this specific point cloud area, based on its point cloud density, observation geometry, reconstruction accuracy and surface complexity evaluation, the calculated dynamic connection distance threshold is 25.29 mm. When the wavefront advancing algorithm is used to construct the triangulated mesh, only when the distance between two adjacent points is less than 25.29 mm and other conditions (such as curvature difference) are met, they will be given priority for connection into edges.

[0124] Based on the dynamic connection distance threshold calculated for different areas of the point cloud in the previous step , a surface reconstruction algorithm, such as the wavefront advancing algorithm, is used to generate triangular mesh patches from sparse 3D point clouds. The algorithm process is roughly as follows: First, an initial seed edge or seed triangle needs to be selected. Usually, any two points in the point cloud that are close to each other and have the same normal vector direction can be selected to form the initial edge, or three adjacent points that form a good shape (non-degenerate) triangle can be found as seeds. Then, the algorithm maintains an active boundary edge list (ie, "wavefront"), which initially contains the seed edge (or the three edges of the seed triangle). In each iteration, an edge is selected from the active boundary edge list (for example, denoted as ), within its neighborhood (e.g., a circle centered at the midpoint of an edge with a radius of Search for candidate points within the relevant spherical area) , the candidate point needs to meet the following conditions to be able to be connected to the edge Form a new triangle :First, point It cannot be a point already in the grid and The connection of points cannot cause the grid to intersect itself; secondly, the points With edge Two vertices of The distance must be less than the dynamic connection distance threshold corresponding to the area ; Third, in order to ensure the smoothness of the surface or the reasonable transition of features, point With vertex The normal vectors or curvatures of the triangles (or the newly formed triangles and the adjacent triangles) need to be similar, and one or more curvature values ​​(such as mean curvature or Gaussian curvature) need to be estimated for each point in advance, and then the potential connection points are compared. and Is the absolute value of the difference in curvature values ​​less than 0.1? This threshold of 0.1 is an empirical value obtained through experimental adjustment based on the expectation of the smoothness of the surface of the Chinese medicinal material sample and the need to retain detailed features. For example, the impact of values ​​such as 0.05, 0.1, and 0.2 on the reconstruction effect is tested, and 0.1 is selected as the optimal value for balancing smoothness and details. If the best candidate point that meets all conditions is found (e.g., the point that forms the triangle with the best angle), then the new triangle Add to the grid and add the two newly generated edges and Add them to the active boundary list (if they are not already on the active boundary) and move the edges Remove from the list, repeat this process, and continue to expand the grid boundary until the active boundary list is empty, that is, no suitable points can be found to expand the grid. It is calculated dynamically. The process can generate smaller triangular facets in areas with high point cloud density and good quality, and larger facets in sparse or complex areas, and finally form a set of triangular mesh facets with adaptive density. This set together constitutes the three-dimensional model of the sample to be tested.

[0125] The steps for obtaining the multi-dimensional defect indicator set are:

[0126] Based on the dynamic change characteristic map of the surface of traditional Chinese medicine, the pixel point set of all change areas was extracted. The judgment conditions for the moisture absorption defect feature were set as follows: the color value increase rate within 5 consecutive frames was ≥15% / second and the regional area expansion rate was ≥3% / second. The judgment conditions for the oiliness defect feature were set as follows: the texture direction consistency coefficient was ≤0.3 and the regional brightness mean decrease gradient was ≥20 lumens / second.

[0127] Based on the 3D model of the sample to be tested and the 3D model of the same grade standard, a rough alignment is performed with the bottom surface of the model as the reference plane. The model posture is iteratively adjusted until the mean Euclidean distance between the top of the medicinal material and the root bifurcation point is less than 0.5mm and the maximum single point deviation does not exceed 1.2mm. A comparison group of 3D models after spatial registration is generated;

[0128] Based on the three-dimensional model comparison group after spatial registration, the surface grid is divided into 10mm×10mm detection units. The surface curvature difference, normal vector deflection angle and volume expansion rate are simultaneously calculated in each detection unit. The porosity change in the detection unit is superimposed on the hygroscopic defect area, and the surface gloss attenuation rate is superimposed on the oily defect area to generate a multi-dimensional defect indicator set.

[0129] Specifically, based on the dynamic change feature map of the surface of the Chinese medicinal materials obtained in the previous step, spatially connected pixel sets are identified and extracted. Each set constitutes an independent change area. Then, for each identified change area, the original registered time-series image sequence and the frame-by-frame change data set are combined to analyze its dynamic behavior characteristics in the time dimension to determine whether it conforms to the pattern of a specific defect. For the determination of hygroscopic defects, it is necessary to examine the color change and area change trends of the area within a continuous time window (for example, set to 5 consecutive frames of images, the corresponding time length depends on the frame rate, if the frame rate is 2 frames / second, it is 2.5 seconds). The specific calculation method is: calculate the rate of change of the average color value of the pixels in the area (which can be calculated in RGB space or converted to a color space that is more in line with human perception, such as Lab) over time, for example, by calculating the linear growth slope of the average color value in the window or the average frame-by-frame growth percentage, and at the same time calculate the rate of change of the pixel area of ​​the area over time, which can also be obtained by linear fitting or average frame-by-frame expansion percentage. When the color value rising rate (usually manifested as a darkening of the color) reaches or exceeds 15% per second, and the regional area expansion rate reaches or exceeds 3% per second, the area is preliminarily judged to meet the characteristics of moisture absorption defects. These two thresholds (15% / second and 3% / second) is an empirical value obtained based on observation and data analysis of the moisture absorption process of a large number of actual Chinese medicinal materials, and represents the typical dynamic characteristics of moisture absorption. For the determination of oily defects, attention is paid to the texture changes and brightness changes of the region, and the texture direction consistency coefficient within the region is calculated. For example, by calculating the directional histogram of the image gradient in the region, the proportion of the main direction gradient energy to the total gradient energy is calculated, or by using the gray level co-occurrence matrix to calculate certain texture features (such as contrast and correlation) and comprehensively evaluate their degree of disorder, a coefficient between 0 (complete disorder) and 1 (high consistency) is obtained, and the average brightness value of the pixels in the region (for example, gray level co-occurrence matrix) is calculated. The gradient (i.e., rate of decline) of the luminance value (or brightness channel value) over time is calculated. When the texture directional consistency coefficient is less than or equal to 0.3, indicating that the original texture has been significantly damaged or become disordered, and when the average brightness gradient of the region reaches or exceeds 20 brightness units per second (brightness units correspond to the grayscale of an 8-bit image or the equivalent unit of lumens), the region is preliminarily judged to meet the characteristics of an oil flooding defect. The thresholds of 0.3 and 20 brightness units / second here are also set by quantitatively analyzing the texture and brightness changes of known oil flooding samples to set the limit values ​​that can effectively distinguish oil flooding from other changes, completing the feature calculation and condition judgment of all changed areas.

[0130] Based on the three-dimensional model of the sample to be tested generated in the previous step and the pre-prepared standard grade three-dimensional model representing the specification grade standard of the Chinese medicinal material (the standard model is usually constructed by averaging or selecting typical representatives after high-precision three-dimensional scanning of multiple batches of qualified samples of the same grade), the two three-dimensional models need to be accurately aligned in space for subsequent geometric comparison. The alignment process first performs a coarse alignment: identify and select the plane area on the two models that can represent their bottom or place the benchmark (this can be achieved through interactive selection or automatic detection algorithm based on the larger plane features at the bottom of the model), calculate an initial transformation (mainly rotation and translation) so that the bottom surface of the three-dimensional model of the sample to be tested is roughly parallel to the bottom surface of the standard grade three-dimensional model and close in height. After completing the coarse alignment, iterative adjustments for fine alignment are performed. This process relies on key anatomical landmarks that are pre-defined or automatically detected on the two models, such as the top point of the medicinal material (the farthest point along the main axis) and the main root bifurcation point (the location where the root structure branches) A one-to-one correspondence between these landmark points is established. A variant of the iterative closest point (ICP) algorithm or a landmark-based registration algorithm is used. In each iteration, a rigid transformation (containing three rotation parameters and three translation parameters) that minimizes the three-dimensional Euclidean distance between corresponding landmark points is calculated. This transformation is applied to the 3D model of the sample to be tested. Iterations continue until the preset convergence criteria are met: the average Euclidean distance between all corresponding landmark pairs is less than 0.5 mm, and the Euclidean distance between any pair of landmarks exceeds 1.2 mm. These two thresholds (0.5 mm average distance and 1.2 mm maximum distance) are determined based on the allowable deviation range for morphological dimensions in the quality inspection standards or pharmacopoeia regulations for this grade of Chinese medicinal materials. This ensures that the aligned models achieve sufficient matching accuracy in key morphological features. When the iteration meets the convergence criteria, adjustments are stopped. The 3D model of the sample to be tested and the 3D model of the standard grade now constitute the 3D model comparison set after spatial registration.

[0131] Based on the spatially aligned three-dimensional model of the sample to be tested and the specification-level three-dimensional model (i.e., the spatially aligned three-dimensional model comparison group), it is necessary to quantitatively evaluate the geometric deviation of the sample surface and the changes in physical properties associated with specific defects. First, the surface of the aligned specification-level three-dimensional model (or a reference surface wrapped thereon) is divided into regular inspection cells, for example, defined as a square grid area with a side length of 10 mm by 10 mm. This 10 mm × 10 mm size is set based on the conventional spatial resolution requirements of Chinese medicinal material defect detection and the convenience of subsequent statistical analysis, and can be adjusted according to the specific medicinal material type and size. Then, for each inspection cell, the geometric difference index between the surface area of ​​the three-dimensional model of the sample to be tested and the surface area of ​​the specification-level three-dimensional model corresponding to the cell is simultaneously calculated. Specifically, the difference in the average curvature (such as Gaussian curvature or mean curvature) of the two models within the cell is calculated; the deflection angle between the average surface normal vectors of the two models within the cell is calculated; and the change in the local volume enclosed by the surface of the sample to be tested relative to the specification-level surface within the cell is calculated and expressed as a volume expansion rate (for example, the average height difference or local volume relative to the cell area). In addition to these general geometric difference indicators, it is also necessary to combine the information of the potential defect area that has been identified and mapped to the surface of the three-dimensional model to perform targeted indicator superposition calculations: If a certain detection unit overlaps with the determined hygroscopic defect area, the porosity change is additionally calculated in the unit. This requires estimating its surface porosity by analyzing the local geometric features of the surface of the sample model to be tested in the unit (such as the number, depth and area distribution of tiny depressions), and comparing it with the porosity of the corresponding area of ​​the standard grade model to obtain the change; if the detection unit overlaps with the determined oil-blowing defect area, the porosity change is calculated in the unit. If the domains overlap, the surface gloss attenuation rate is additionally calculated. This calculation requires combining color and texture information, or using a specific lighting model and material reflection properties (BRDF) to estimate the glossiness of the surface of the sample to be tested in that unit, and compare it with the gloss reference value of the specification grade model to quantify its attenuation degree. After completing the calculation of these indicators (surface curvature difference, normal vector deflection angle, volume expansion rate, and porosity change or surface gloss attenuation rate based on the area type) for all inspection units, these quantitative indicators are summarized to form a multi-dimensional defect indicator set covering the entire sample surface.

[0132] The steps for determining the defect category of Chinese medicinal materials are as follows:

[0133] Based on a multi-dimensional defect indicator set, if the same inspection unit meets the criteria for both a moisture absorption defect candidate and an oil flooding defect candidate, it is prioritized as a moisture absorption defect. If the distance between defect candidate categories of different inspection units is ≤5mm, they are spatially adjacent and are merged into a composite defect area and upgraded to a severe defect category, generating a revised defect category determination result set.

[0134] Based on the corrected defect category determination result set, the area proportion and severity level of all defect areas are counted. If the total area proportion of hygroscopic defects is ≥5% or there is a severe defect category area, the sample is judged to be unqualified; if the total area proportion of defects is <5% and there is no severe defect category, a qualified mark is output and a defect category determination of the Chinese medicinal materials is generated.

[0135] Specifically, based on the multi-dimensional defect indicator set obtained in the previous step, which contains information on the geometric and physical property deviations of each inspection unit, the defect category of each unit needs to be finally confirmed and corrected. First, each 10 mm × 10 mm inspection unit is traversed and, based on the values ​​of the various indicators it contains (for example, color rise rate, area expansion rate, porosity change corresponding to moisture absorption; texture consistency coefficient, brightness decrease gradient, gloss decay rate corresponding to oil bleeding; and general indicators such as curvature difference, normal vector deflection angle, and volume expansion rate), a comprehensive assessment is made to determine whether the unit meets the previously set moisture absorption defect candidate conditions and oil bleeding defect candidate conditions. If a unit's indicators do strongly point to the possibility of both defects, a category determination is made according to preset rules: such units are preferentially determined as moisture absorption defects. The basis for this priority determination may be based on expert experience that moisture absorption is a more common or more basic form of deterioration, or that the visual or physical manifestation of moisture absorption is more likely to mask or confuse the oil bleeding characteristics, thereby directing the determination to a more specific or more critical category. After completing the preliminary category determination for all units, spatial correlation analysis is performed. The system then identifies all inspection units identified as having defects (whether due to moisture absorption or oily effluence) and merges spatially adjacent units (sharing edges or vertices) belonging to the same defect category to form a continuous defect region. The system then calculates the minimum distance between regions of different defect categories. For example, the shortest three-dimensional spatial distance between the boundary of a moisture absorption defect region and the boundary of the nearest oily effluence defect region is calculated. If this distance is less than or equal to 5 mm, the two different defect types are considered to have significant spatial proximity. This proximity may indicate mutual influence or a more complex deterioration process, thus requiring special treatment. The 5 mm threshold is based on research on deterioration patterns of traditional Chinese medicine, defining a distance range sufficient to indicate possible interaction or a common cause between defects. All defect regions of different categories (or the inspection units that comprise them) that meet this spatial proximity condition are merged into a larger "composite defect region," and the defect category of this region is upgraded to the "severe defect category." After the above conflict resolution, region merging, and severity upgrade, a revised set of defect category determination results is obtained.

[0136] Based on the revised defect category determination result set generated in the previous step (which contains the final defect category and severity level of each inspection unit or merged area), it is necessary to make a final qualified or unqualified judgment on the quality status of the entire Chinese medicinal material sample. First, using the result set and the three-dimensional model data of the sample, the total area of ​​all defect areas on the sample surface that are determined to be defects is statistically calculated, and the total area of ​​all areas determined to be "hygroscopic defects" is especially calculated. At the same time, the total surface area of ​​the sample is obtained (which can be directly calculated from the three-dimensional model). Then, the percentage of the total area of ​​hygroscopic defects to the total surface area of ​​the sample is calculated as: (total area of ​​hygroscopic defects / total surface area of ​​the sample) × 100%. Next, check whether there are any areas marked as "severe defect categories" in the revised defect category determination result set (that is, areas of different categories adjacent to each other in space). The final judgment is made according to the preset quality judgment standards: Standard 1, whether the calculated total area of ​​hygroscopic defects accounts for more than or equal to 5%; Standard 2, whether there is at least one serious defect category area. If this limit is exceeded, it is considered that the quality of the medicinal material is significantly affected. The existence of serious defect categories, regardless of the size of the area, usually represents more complex or more serious deterioration, and is therefore also regarded as a direct indicator of unqualified. If at least one of the above two standards is met (i.e., the area of ​​hygroscopic defects accounts for ≥5% or there are serious defect areas), the Chinese medicinal material sample is finally judged to be unqualified. If both standards are not met (i.e., the area of ​​hygroscopic defects accounts for <5% and there are no serious defect areas), the sample is judged to be qualified, and a qualified mark is output, and finally a defect category judgment (qualified or unqualified) for the Chinese medicinal material sample is generated.

[0137] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A vision-based method for detecting defects in Chinese medicinal materials, characterized in that: The following steps are involved: Collect time-series images of Chinese medicinal material samples at regular intervals, collect multi-view images, align the pixels of the time-series images, and structurally organize and associate camera parameters of the multi-view images according to the shooting orientation to generate a registered time-series image sequence and a multi-view image set; Based on the registered time-series image sequence, the color value difference between the corresponding pixel positions of adjacent frame images in the sequence is calculated to obtain a frame-by-frame variation data set. Based on the frame-by-frame variation data set, the total variation value of each pixel position is accumulated along the time dimension to obtain a dynamic change feature map of the surface of the Chinese medicinal material; Based on the multi-view image set, a sparse three-dimensional point cloud and a camera pose are obtained by matching feature points of the same name between images from different viewpoints and combining them with camera parameters. Based on the sparse three-dimensional point cloud and the camera pose, adjacent points in the point cloud are connected to construct triangular mesh patches to establish a three-dimensional model of the sample to be tested; Based on the dynamic change characteristic map of the Chinese medicinal material surface and the three-dimensional model of the sample to be tested, analyzing whether the pattern of the change area in the dynamic change characteristic map of the Chinese medicinal material surface conforms to the dynamic defect characteristics of moisture absorption or oiliness, obtaining a multi-dimensional defect index set, and outputting the qualification or defect category of the Chinese medicinal material sample based on the multi-dimensional defect index set, and establishing a Chinese medicinal material defect category determination; The steps for obtaining the multi-dimensional defect indicator set are: Based on the dynamic change characteristic map of the surface of the Chinese medicinal materials, the pixel point sets of all change areas were extracted. The judgment conditions for the moisture absorption defect feature were set as follows: the color value rising rate within 5 consecutive frames was ≥15% / second and the regional area expansion rate was ≥3% / second; the judgment conditions for the oiliness defect feature were set as follows: the texture direction consistency coefficient was ≤0.3 and the regional brightness mean decreasing gradient was ≥20 lumens / second; Based on the three-dimensional model of the sample to be tested and the three-dimensional model of the same grade standard, a rough alignment is performed with the bottom surface of the model as the reference plane, and the model posture is iteratively adjusted until the mean Euclidean distance between the top of the medicinal material and the root bifurcation point is less than 0.5 mm and the maximum single point deviation does not exceed 1.2 mm, to generate a three-dimensional model comparison group after spatial registration; Based on the three-dimensional model comparison group after spatial registration, the surface grid is divided into 10mm×10mm detection units. The surface curvature difference, normal vector deflection angle and volume expansion rate are synchronously calculated in each detection unit. The porosity change in the detection unit is superimposed on the hygroscopic defect area, and the surface gloss attenuation rate is superimposed on the oily defect area to generate a multi-dimensional defect indicator set.

2. The method for detecting defects in Chinese medicinal materials based on vision according to claim 1, characterized in that: The steps of obtaining the registration time sequence image sequence and the multi-view image set are: Time-series images and multi-view images of the same batch of Chinese medicinal materials are collected regularly. Pixel alignment is performed on the time-series images. SIFT feature points are extracted from adjacent frames and bidirectionally matched. The spatial Euclidean distance of matching point pairs is calculated. Outlier matching point pairs are removed using the RANSAC algorithm. The affine transformation matrix parameters are calculated and applied to the entire image to generate a sequence of registered time-series images. Based on the registered time-series image sequence and the multi-view images, the multi-view images are structurally organized, the focal length, principal point coordinates, and radial distortion coefficient in the camera parameter file of each image are parsed, a mapping relationship between the azimuth angle, the pitch angle, and the camera parameters is established, and a multi-view image set is generated; Based on the registered temporal image sequence and the multi-view image set, the validity of the temporal alignment is verified by calculating the SSIM structural similarity of adjacent frames in the temporal sequence, and the registered temporal image sequence and the multi-view image set are generated.

3. The method for detecting defects in Chinese medicinal materials based on vision according to claim 1, characterized in that: The steps for obtaining the frame-by-frame variation data set are: Based on the registered time-series image sequence, extracting the RGB three-channel color value of each pixel position in the adjacent frame image, establishing the current frame color value matrix and the previous frame color value matrix, and generating an adjacent frame color value matrix group; Calculating the color value difference of corresponding pixel positions based on the adjacent frame color value matrix group; Based on the color value difference, all pixel positions and frame sequence indexes are traversed, and the color value difference is superimposed and arranged in frame sequence to generate a frame-by-frame variation data set.

4. The method for detecting defects in Chinese medicinal materials based on vision according to claim 1, wherein: The steps for obtaining the characteristic map of dynamic changes on the surface of the Chinese medicinal material are as follows: Based on the frame-by-frame variation data set, extract the color value difference sequence of each pixel position in all frames, establish a time series variation set of the pixel position, and generate pixel-level time accumulation input data; Calculating a total change value at each pixel position based on the pixel-level time-accumulated input data; Based on the total change value, the sliding window method is used to calculate the mean and standard deviation of the total change in the neighborhood around each pixel, set a dynamic threshold, and screen pixel areas where the total change value is greater than the dynamic threshold to generate a dynamic change feature map of the surface of Chinese medicinal materials.

5. The method for detecting defects in Chinese medicinal materials based on vision according to claim 1, characterized in that: The steps for obtaining the sparse three-dimensional point cloud and camera posture are as follows: Based on the multi-view image set, SIFT feature points are extracted from every two view images and feature descriptors are generated. Bidirectional matching is performed by calculating the Hamming distance between the feature descriptors. The RANSAC algorithm is used to fit the basic matrix and outlier matching point pairs with a projection error exceeding 0.5 pixels are eliminated to generate a set of matching pairs of feature points with the same name between the multi-view images. Based on the set of matching pairs of feature points with the same name and the focal length, principal point coordinates, and radial distortion coefficient in the camera parameters, an epipolar geometry constraint equation is constructed. The camera rotation matrix and translation vector are solved by singular value decomposition of the essential matrix. The linear triangulation method is used to calculate the initial 3D spatial coordinates of the feature point pairs, generating the camera extrinsic parameter estimation results and a sparse 3D point coordinate set. Based on the camera extrinsic parameter estimation results and the sparse 3D point coordinate set, the rotation matrix, translation vector and 3D point coordinates are iteratively adjusted with the reprojection error as the loss function until the average reprojection error converges to within 0.3 pixels. Abnormal 3D points with residuals exceeding 1.0 pixels after optimization are eliminated to generate a sparse 3D point cloud and camera pose.

6. The method for detecting defects in Chinese medicinal materials based on vision according to claim 1, characterized in that: The steps for obtaining the three-dimensional model of the sample to be tested are: Based on the sparse 3D point cloud and camera pose, the optical center distance between adjacent camera viewpoints, the mean 3D point reprojection error, and the surface curvature change gradient are extracted to establish an interpolation credibility assessment parameter set; Calculating a dynamic connection distance threshold based on the interpolation credibility evaluation parameter group; Based on the dynamic connection distance threshold, the wavefront advancing algorithm is used to perform regional growth on the point cloud, preferentially connecting adjacent point pairs with a spacing less than the dynamic connection distance threshold and a curvature difference less than 0.1, generating a set of triangular mesh faces with adaptive density, and establishing a three-dimensional model of the sample to be tested.

7. The method for detecting defects in Chinese medicinal materials based on vision according to claim 1, characterized in that: The steps for determining the defect category of the Chinese medicinal materials are as follows: Based on the multi-dimensional defect indicator set, if the same inspection unit meets the conditions for both a moisture absorption defect candidate and an oil flooding defect candidate, it is preferentially determined as a moisture absorption defect. If the distance between defect candidate categories of different inspection units is ≤5mm, they are spatially adjacent and are merged into a composite defect area and upgraded to a severe defect category, generating a revised defect category determination result set. Based on the corrected defect category determination result set, the area percentage and severity level of all defect areas are calculated. If the total area percentage of moisture absorption defects is ≥5% or there is a severe defect category area, the sample is determined to be unqualified; If the total defect area accounts for less than 5% and there is no serious defect category, a qualified mark will be output and a Chinese medicinal material defect category judgment will be generated.

Citation Information

Patent Citations

  • Traditional Chinese medicine quality detection system based on image recognition

    CN118674954A

  • Traditional Chinese medicinal material production quality analysis system and method based on machine vision

    CN119693343A

Cited By

  • Tea form anomaly detection method and system based on visual fusion

    CN121438014A