Intelligent traditional Chinese medicine decoction piece screening method based on multi-modal data fusion

By using a multimodal data fusion method, visible light images, near-infrared spectra, and odor volatiles data of Chinese herbal medicine slices are collected and processed simultaneously, solving the problems of consistency and accuracy in the screening of Chinese herbal medicine slices. This enables comprehensive and precise screening of the quality of the slices, and is applicable to the modern Chinese medicine industry.

CN120831338AActive Publication Date: 2025-10-24WEIFANG NURSING VOCATIONAL COLLEGE +1

Patent Information

Application Number
CN202511332970.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-10-24
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing screening technologies for Chinese medicinal herbs are highly subjective, have poor stability, and are difficult to achieve consistency and reliability. Furthermore, a single detection technology cannot fully cover the key quality indicators of the herbs, and the data processing and integration process is insufficient, resulting in inaccurate screening results.

Method used

A multimodal data fusion method is adopted, which simultaneously acquires visible light image data, near-infrared spectral data and odor volatile data through a multi-source sensor array. Morphological cleaning, baseline drift correction and environmental interference removal are performed to extract morphological, chemical composition and volatile substance feature sets, construct a three-dimensional feature fusion space, and perform centralized search operation to determine the optimal screening decision vector.

Benefits of technology

It achieves comprehensive coverage of the quality of Chinese herbal medicine slices, reduces subjective bias caused by human experience, improves the objectivity and consistency of the screening process, and enhances the accuracy and efficiency of the screening results, making it suitable for the development of the modern Chinese medicine industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120831338A_ABST
    Figure CN120831338A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of traditional Chinese medicine decoction piece screening, and discloses an intelligent traditional Chinese medicine decoction piece screening method based on multi-modal data fusion. The method comprises the following steps: synchronously acquiring visible light image data, near infrared spectrum data and smell volatile matter data of a target traditional Chinese medicine decoction piece through a multi-source sensor array; respectively performing morphological cleaning, baseline drift correction and environmental interference elimination operations on the three types of data; extracting a form feature set, a chemical component feature set and a volatile substance feature set of the decoction pieces from the preprocessed data; calculating morphological stability factors, component stability factors and volatilization stability factors corresponding to the three types of feature sets; constructing a three-dimensional feature fusion space, and mapping the three types of feature sets into fusion feature vectors; executing centralized search in the space by taking the fusion feature vector as a starting point, and determining an optimal screening decision vector; therefore, more accurate grade classification and quality screening are realized, and the screening efficiency of the traditional Chinese medicine decoction pieces can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of traditional Chinese medicine decoction pieces screening, in particular to a multi-modal data fusion intelligent screening method for traditional Chinese medicine decoction pieces. BACKGROUND

[0002] As a key form of clinical application of traditional Chinese medicine, the quality of traditional Chinese medicine decoction pieces is directly related to the clinical efficacy and safety of drug use. In the current production and circulation links of traditional Chinese medicine decoction pieces, quality screening mainly relies on traditional manual identification methods, which are based on the experience of identification personnel to judge the appearance, odor and other characteristics of decoction pieces. However, this method has obvious problems of strong subjectivity and poor stability. The judgment results of different identification personnel or the same personnel in different states are prone to deviation, which makes it difficult to ensure the consistency and reliability of the screening results. With the development of technology, some fields have begun to try to use single detection technology for quality screening of traditional Chinese medicine decoction pieces, such as using visible light image technology to analyze the shape of decoction pieces, or using near-infrared spectroscopy technology to detect the chemical components of decoction pieces. However, single detection technology has obvious limitations. Only relying on visible light image technology cannot obtain the chemical component information inside the decoction pieces, and it is difficult to accurately judge the content of effective components and the internal quality of the decoction pieces. Although near-infrared spectroscopy technology can reflect the chemical component situation, it has weak recognition ability for the appearance defects of decoction pieces, such as insect damage and mold. In addition, the smell of traditional Chinese medicine decoction pieces is one of the important characteristics of its quality, and there is currently a lack of effective technical means to quantify it and integrate it into the screening process, resulting in the fact that the existing screening methods cannot comprehensively cover the key quality indicators of decoction pieces. In the existing attempts to use multiple technical means, there are obvious deficiencies in data processing and fusion. The data types obtained by different detection technologies are quite different, and the data formats and feature dimensions are different. Traditional data processing methods cannot effectively integrate different types of data, and often can only separately analyze and simply superimpose the results of various types of data, which cannot fully tap the internal correlation between different data, resulting in limitations in the accuracy and robustness of the screening model. These problems together result in the fact that the current traditional Chinese medicine decoction piece screening technology cannot meet the demand of modern traditional Chinese medicine industry for high-quality and high-efficiency screening, and there is an urgent need for a technical solution that can integrate multi-dimensional data and achieve comprehensive and accurate screening. SUMMARY

[0003] The purpose of the present application is to provide a multi-modal data fusion intelligent screening method for traditional Chinese medicine decoction pieces to solve the problems raised in the background art.

[0004] To achieve the above-mentioned purpose, the present application provides a multi-modal data fusion intelligent screening method for traditional Chinese medicine decoction pieces, which comprises: synchronously collecting visible light image data, near-infrared spectroscopy data and odor volatile data of the target traditional Chinese medicine decoction pieces by a multi-source sensor array; performing a morphological cleaning operation on the visible light image data, performing a baseline drift correction operation on the near-infrared spectrum data, and performing an environmental interference elimination operation on the odor volatile data; extracting a set of morphological features of the medicinal slices from the cleaned visible light image data, extracting a set of chemical component features from the corrected near-infrared spectrum data, and extracting a set of volatile substance features from the interference-eliminated odor volatile data; calculating a morphological stability factor of the set of morphological features of the medicinal slices, a component stability factor of the set of chemical component features, and a volatility stability factor of the set of volatile substance features; constructing a three-dimensional feature fusion space according to the morphological stability factor, the component stability factor, and the volatility stability factor, and mapping the set of morphological features of the medicinal slices, the set of chemical component features, and the set of volatile substance features into a fusion feature vector; performing a centralized search operation in the three-dimensional feature fusion space with the fusion feature vector as a starting point to determine an optimal screening decision vector; performing grade classification and quality screening on the target medicinal slices according to the optimal screening decision vector.

[0005] Preferably, the method for performing the morphological cleaning operation on the visible light image data is as follows: identifying a background pixel region in the visible light image data, and separating adhered medicinal slice contours by using a morphological opening operation; calculating a minimum circumscribed rectangle of each medicinal slice contour to generate a medicinal slice spatial distribution map; eliminating edge incomplete regions according to the medicinal slice spatial distribution map to output a complete medicinal slice image data set.

[0006] Preferably, the method for extracting the set of morphological features of the medicinal slices from the cleaned visible light image data is as follows: calculating a surface texture complexity, an edge curvature distribution histogram, and an aspect ratio feature of each medicinal slice based on the complete medicinal slice image data set; aggregating the surface texture complexity, the edge curvature distribution histogram, and the aspect ratio feature of all medicinal slices to form the set of morphological features of the medicinal slices.

[0007] Preferably, the method for calculating the morphological stability factor of the set of morphological features of the medicinal slices, the component stability factor of the set of chemical component features, and the volatility stability factor of the set of volatile substance features is as follows: statistically calculating a first quartile and a third quartile of all surface texture complexities in the set of morphological features of the medicinal slices; extending the first quartile downward by a preset texture offset as a lower bound of texture aggregation, and extending the third quartile upward by a preset texture offset as an upper bound of texture aggregation; calculating a variance value of the surface texture complexity between the lower bound of texture aggregation and the upper bound of texture aggregation, and setting the variance value as the morphological stability factor.

[0008] Preferably, the method for constructing a three-dimensional feature fusion space according to the morphological stability factor, the component stability factor and the volatile stability factor is: predefining three-dimensional space coordinate axes, wherein the X-axis corresponds to the morphological feature dimension of the decoction piece, the Y-axis corresponds to the chemical component feature dimension, and the Z-axis corresponds to the volatile substance feature dimension; mapping the morphological stability factor to the X-axis coordinate value, mapping the component stability factor to the Y-axis coordinate value, and mapping the volatile stability factor to the Z-axis coordinate value; generating a three-dimensional feature fusion space containing historical screening decision vectors with a three-dimensional coordinate point composed of the X-axis coordinate value, the Y-axis coordinate value and the Z-axis coordinate value as a center point.

[0009] Preferably, the method for performing a centralized search operation with the fusion feature vector as a starting point is: selecting a target subspace containing the fusion feature vector in the three-dimensional feature fusion space; constructing a spherical decision neighborhood with the fusion feature vector as the center and a preset decision step length as the radius; calculating the neighborhood decision density of the historical screening decision vectors in the spherical decision neighborhood.

[0010] Preferably, the method for determining the optimal screening decision vector is: randomly selecting a first candidate decision vector at the edge of the spherical decision neighborhood; calculating the first candidate neighborhood decision density of the first candidate decision vector; when the first candidate neighborhood decision density is greater than the neighborhood decision density, updating the first candidate decision vector to the current center and reconstructing the spherical decision neighborhood, and iteratively performing until a preset iteration number is met.

[0011] Preferably, the method further comprises: when the first candidate neighborhood decision density is less than the neighborhood decision density, accumulating the number of failed decision updates; if the number of failed decision updates exceeds a preset failure threshold, setting the decision vector corresponding to the current center as the optimal screening decision vector.

[0012] Preferably, the method for performing grade classification and quality screening on the target traditional Chinese medicine decoction piece according to the optimal screening decision vector is: performing normalization processing on the optimal screening decision vector to generate a standard screening decision value; outputting the quality grade label of the target traditional Chinese medicine decoction piece according to the distribution of the standard screening decision value in a preset grading threshold interval.

[0013] Preferably, the method for performing normalization processing on the optimal screening decision vector is: calculating the module length of the optimal screening decision vector in the three-dimensional feature fusion space; Divide the module length by the preset maximum radius of the decision space to obtain a standard screening decision value.

[0014] Compared with the prior art, the present application has the following advantages: The multi-modal data fusion intelligent screening method of traditional Chinese medicine decoction pieces can comprehensively obtain the appearance form, internal chemical components and volatile odor characteristics of the traditional Chinese medicine decoction pieces by synchronously collecting visible light image data, near-infrared spectrum data and odor volatile data through a multi-source sensor array, breaking the limitation of traditional single detection technology that can only obtain part of the quality information, and realizing all-round coverage of the decoction piece quality indicators. Compared with traditional manual identification, the method avoids subjective bias caused by manual experience through objective sensor data collection, makes the screening process more objective and consistent, and can effectively reduce the fluctuation of screening results caused by personnel experience differences. In the data preprocessing link, morphological cleaning, baseline drift correction and environmental interference removal operations are respectively performed according to the characteristics of different types of data, which can effectively remove noise and interference information in the data. The morphological cleaning of the visible light image can eliminate impurity pixels and background interference in the image, highlight the true morphological characteristics of the decoction piece, and improve the accuracy of subsequent morphological feature extraction; the baseline drift correction of the near-infrared spectrum data can eliminate the spectral baseline shift caused by factors such as instrument error and environmental temperature change, and ensure the authenticity and reliability of the chemical component information; the environmental interference removal of the odor volatile data can exclude the influence of other odor molecules in the external environment, accurately capture the volatile substance characteristics of the decoction piece itself, and provide a high-quality data basis for subsequent feature extraction. Respectively extracting the morphological feature set, the chemical component feature set and the volatile substance feature set from different types of preprocessed data, and calculating the corresponding stability factors, can convert each dimension of data into a feature index with clear physical meaning. The morphological stability factor can quantitatively reflect the integrity and consistency of the decoction piece morphology, the component stability factor can reflect the uniformity and effective component distribution of the chemical components, and the volatile stability factor can represent the stability and typicality of the odor characteristics. The introduction of these stability factors provides a clear and representative feature basis for subsequent data fusion, making the features of different types of data comparable and fusible. By constructing a three-dimensional feature fusion space, the three types of feature sets are mapped into a fusion feature vector, realizing the deep integration of different dimensions of data instead of simply superimposing the results. This fusion method can fully exploit the internal relationship between morphology, chemical components and odor characteristics, such as the potential connection between decoction piece morphology defects and chemical component changes, odor abnormalities, so that the fused feature vector can more comprehensively and accurately reflect the overall quality status of the decoction piece. Compared with the traditional data superposition method, this fusion strategy greatly improves the richness and effectiveness of feature information, providing a more reliable basis for subsequent screening decisions.

[0015] The centralized search operation is performed in the three-dimensional feature fusion space to determine the optimal screening decision vector, which can make accurate screening judgment based on multi-dimensional fusion features, and avoid misjudgment caused by single feature analysis. The search process can fully utilize the comprehensive information in the fusion feature vector, comprehensively consider multiple factors such as morphology, chemical composition and odor, and make overall evaluation on the quality of decoction pieces, so as to realize more accurate grade classification and quality screening. In addition, the whole process of the method is based on standardized technical operation, which is easy to realize automation and large-scale application, can significantly improve the efficiency of Chinese medicinal decoction piece screening, reduce the dependence on manual operation, meet the development needs of modern Chinese medicine industry, and help to promote the upgrading of Chinese medicinal decoction piece quality control technology and promote the standardized and modernized development of Chinese medicine industry. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A working principle diagram of the multi-modal data fusion intelligent screening method of Chinese medicinal decoction pieces is provided. Figure 2 A flowchart of the visible light image data morphological cleaning operation is provided. Figure 3 A flowchart for calculating three stability factors is provided. Figure 4 A flowchart for performing a centralized search operation is provided. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0018] Please refer to Figure 1 The present application provides a multi-modal data fusion intelligent screening method of Chinese medicinal decoction pieces, which comprises: Visible light image data, near-infrared spectral data, and odor volatile data of the target Chinese herbal medicine pieces are synchronously collected using a multi-source sensor array. The multi-source sensor array comprises a high-resolution camera, a near-infrared spectrometer, and a gas sensor, ensuring synchronized data acquisition. After acquisition, morphological cleaning is performed on the visible light image data, which involves image preprocessing to remove noise and separate adhering objects. Baseline drift correction is performed on the near-infrared spectral data to eliminate errors caused by instrument drift. Environmental interference removal is performed on the odor volatile data to filter out environmental factors such as temperature and humidity. Next, a morphological feature set of the herbal medicine piece, including surface texture and geometry, is extracted from the cleaned visible light image data. A chemical component feature set is extracted from the corrected near-infrared spectral data to reflect the chemical composition of the herbal medicine piece. A volatile compound feature set is extracted from the odor volatile data after interference removal to characterize the odor characteristics. Subsequently, the morphological stability factor of the herbal medicine piece morphological feature set, the component stability factor of the chemical component feature set, and the volatility stability factor of the volatile compound feature set are calculated to quantify the stability of the features. A three-dimensional feature fusion space is constructed based on the morphological stability factor, component stability factor, and volatility stability factor. The morphological, chemical component, and volatile substance feature sets of the medicinal slices are mapped into a fused feature vector. Within this three-dimensional feature fusion space, a centralized search operation is performed using the fused feature vector as the starting point to determine the optimal screening decision vector. Based on this optimal screening decision vector, the target Chinese medicinal slices are graded and quality-screened, and a quality grade label is output for the slices.

[0019] Example 1: See Figure 2 When implementing the morphological cleaning operation, visible light image data is collected by a high-resolution industrial camera with an image resolution of 4096×2160 pixels and a bit depth of 8 bits. The collected image is first converted into a grayscale image, and the grayscale value of each pixel is calculated using the weighted average method. The identification of background pixel areas is based on grayscale threshold segmentation. The threshold is set to a grayscale value of 30, and pixels below this value are identified as background. The morphological opening operation uses a circular structural element with a structural element radius of 5 pixels. The opening operation first performs an erosion operation on the image to eliminate small noise points and weak connections, and then performs an expansion operation to restore the main shape of the medicinal piece. The separation of the contours of the adhered medicinal pieces is achieved by calculating the connected domains, and each connected domain represents a potential medicinal piece object.

[0020] The calculation of the minimum bounding rectangle is based on the separated herb contour, using a rotating bounding box algorithm, traversing all points on the contour, calculating the bounding box in each direction, and selecting the smallest rectangle with the smallest area as the minimum bounding rectangle. The coordinates, width and height of each rectangle are recorded. The spatial distribution map of the herb is generated in the form of a two-dimensional array, with the same size as the original image, and each element stores the herb identifier or background marker at the corresponding position. The removal of the edge missing area is based on the position of the minimum bounding rectangle. If any side of the rectangle is less than 10 pixels away from the image boundary, the herb is considered to be edge missing and removed from the data set. The complete herb image data set is output as a series of cropped sub-images, each containing a complete herb, with a uniform size of 256x256 pixels.

[0021] The extraction of the herb morphology feature set is based on the complete herb image data set. The calculation of the surface texture complexity uses a gray level co-occurrence matrix, with the direction set to 0 degrees and the distance set to 1 pixel. The contrast and entropy features are extracted from the matrix. The contrast reflects the clarity of the texture, and the entropy represents the randomness of the texture. The surface texture complexity of each herb is represented by the weighted sum of the contrast and the entropy, with weights of 0.6 and 0.4 respectively. The edge curvature distribution histogram is obtained by first extracting the herb contour using the Canny edge detection algorithm, and then smoothing the contour point set using a B-spline curve. The curvature calculation is based on the fitted curve, and the reciprocal of the curvature radius is calculated at each contour point. The curvature value is normalized to the range of 0 to 1. The histogram divides the curvature value into 10 intervals, and the proportion of the number of points in each interval is calculated. The aspect ratio feature is directly calculated from the minimum bounding rectangle. The length is the long side of the rectangle, the width is the short side, and the ratio is the length divided by the width. All feature values are normalized to eliminate the dimension effect.

[0022] The composition of the herb morphology feature set is obtained by aggregating the feature values of all herbs. Each herb's features are represented as a vector, containing the surface texture complexity, the 10 interval values of the edge curvature distribution histogram, and the aspect ratio. The feature set is stored as a matrix, with rows corresponding to herb samples and columns corresponding to feature dimensions. During feature extraction, parallel computing is used to process each herb sub-image to improve processing efficiency. The final output of the herb morphology feature set is used for subsequent stability factor calculation and feature fusion. In specific applications, the collection environment of visible light image data needs to maintain uniform illumination to avoid shadow and glare interference. The effectiveness of the morphological cleaning operation depends on the size and shape selection of the structural elements, which need to be adjusted according to the type of herb. The accuracy of feature extraction is affected by image resolution and preprocessing effect, and high-resolution images can provide more detailed texture and edge information.

[0023] Example 2: see Figure 3In the process of calculating the morphological stability factor, the set of morphological characteristics of the decoction piece includes the surface texture complexity values of multiple decoction pieces. Assuming that a batch of 50 samples of Huangqi decoction pieces have surface texture complexity values distributed between 0.15 and 0.85 after normalization. First, sort all the values in ascending order, and determine the value at the 25% position as 0.32 (the first quartile) and the value at the 75% position as 0.68 (the third quartile). The preset texture offset is set to 0.05, which is determined based on historical data statistical analysis. The lower bound of texture aggregation is calculated as 0.32 minus 0.05, which is 0.27, and the upper bound of texture aggregation is calculated as 0.68 plus 0.05, which is 0.73. Select all surface texture complexity values within the range of 0.27 to 0.73, a total of 38 samples. Calculate the variance of these sample values: first, calculate the mean value 0.49, then calculate the sum of the squared differences between each value and the mean value, and finally divide by the number of samples to get the variance 0.021. This variance value is the morphological stability factor.

[0024] For the chemical component feature set, taking Danshen decoction pieces as an example, the feature set includes the concentration value of the main active ingredient Danshensan B. The concentration values of 50 samples are distributed between 8.2 mg / g and 22.7 mg / g. After sorting, the first quartile is determined to be 12.3 mg / g and the third quartile is determined to be 18.9 mg / g. The preset component offset is set to 1.2 mg / g, which is based on the concentration fluctuation range of similar decoction pieces. The lower bound of component aggregation is 12.3 minus 1.2, which is 11.1 mg / g, and the upper bound of component aggregation is 18.9 plus 1.2, which is 20.1 mg / g. Select 42 samples with concentration values between 11.1 mg / g and 20.1 mg / g, calculate the variance: after calculating the mean value 15.8 mg / g, calculate the sum of the squared deviations of each sample value from the mean value, and divide by the number of samples to get the variance 4.37. This value is the component stability factor.

[0025] In the processing of the volatile substance feature set, the feature set includes the volatile concentration value of menthol, and the 50 sample data is distributed between 105 ppm and 287 ppm. After sorting, the first quartile is determined to be 142 ppm and the third quartile is determined to be 218 ppm. The preset volatile offset is set to 25 ppm, which is determined according to the detection accuracy of the sensor. The lower bound of volatile aggregation is 142 minus 25, which is 117 ppm, and the upper bound of volatile aggregation is 218 plus 25, which is 243 ppm. Select 39 samples with concentration values between 117 ppm and 243 ppm, calculate the mean value 183 ppm, then calculate the sum of the squared differences between each sample value and the mean value, and divide by the number of samples to get the variance 1026. This variance value is the volatile stability factor.

[0026] The above calculation process needs to pay attention to three key links: first, the calculation of quartiles adopts linear interpolation method, and when the sample quantity cannot be divided, the position is determined by proportion. Second, the setting of the offset needs to consider the distribution characteristics of the characteristic value, and the texture offset is usually 0.3 to 0.7 times the standard deviation of the characteristic value, the chemical component offset refers to the standard error of the detection method, and the volatile offset combines the sensor accuracy parameters. Third, the variance calculation adopts the population variance formula to ensure the comparability of the stability factor. When implemented, a data processing pipeline is established: after the characteristic value is input, it is automatically sorted, the quartile calculation module is called, the aggregation interval is generated according to the preset offset, and finally the variance operation is performed. All parameter settings are recorded in the configuration file, including the offset coefficient, the variance calculation method, etc. For different medicinal piece varieties, the parameter combinations stored in the configuration file, such as the texture offset coefficient of Astragalus root medicinal piece is 0.5, the component offset of Salvia miltiorrhiza medicinal piece is a fixed value of 1.2, and the volatile offset of Mentha haplocalyx medicinal piece is a fixed value of 25. The system automatically loads the corresponding parameters according to the medicinal piece type.

[0027] The characteristic value abnormality processing mechanism includes two levels: when the characteristic value exceeds the historical data range, the data review process is started; when the sample proportion in the aggregation interval is less than 60%, the offset is automatically expanded by 10% to recalculate. All stability factor calculation results are attached with confidence labels, and the confidence is determined according to the sample proportion in the aggregation interval. The proportion higher than 80% is marked as high confidence, the proportion between 60% and 80% is marked as medium confidence, and the proportion lower than 60% is marked as low confidence.

[0028] The implementation mode shows parameter sensitivity in multiple batch tests. When the texture offset coefficient is in the interval of 0.4 to 0.6, the morphological stability factor fluctuates less than 5%; the component offset fixed value changes ±0.2 mg / g, causing the factor to fluctuate about 3%; and the volatile offset fixed value changes ±5 ppm, causing the factor to fluctuate about 7%. Therefore, the parameter setting needs to be calibrated by historical data, and at least 5 batches of data need to be accumulated for new medicinal piece varieties to determine the optimal parameters.

[0029] Example 3: see Figure 4 In the process of constructing the three-dimensional feature fusion space, a space coordinate system needs to be established first. Three orthogonal coordinate axes are defined: the X-axis represents the medicinal piece morphological feature dimension, the value range of which is determined by historical morphological stability factor data, and is usually normalized to the interval [0, 1]; the Y-axis corresponds to the chemical component feature dimension, which is calibrated based on the statistical distribution of the component stability factor; and the Z-axis represents the volatile substance feature dimension, whose scale is adjusted according to the maximum and minimum values of the volatile stability factor. The unit length of each axis represents the change amplitude of the stability factor, and the three dimensions together form a standardized cubic space.

[0030] The calculated morphological stability factor, component stability factor and volatility stability factor are mapped to spatial coordinate values. The mapping process employs a linear transformation: ; wherein: is the morphological stability factor of the current batch, and represent the minimum and maximum values of the morphological stability factor in the historical data, respectively; is the component stability factor of the current batch, and are its historical minimum and maximum values, respectively; is the volatility stability factor of the current batch, and are its historical minimum and maximum values, respectively. The mapping ensures that all coordinate values fall within the range [0, 1].

[0031] The historical screening decision vectors are loaded from the database, each vector containing three-dimensional coordinate values of past successful screening cases. These vectors are stored in the form of a point cloud in the feature fusion space, forming a decision sample distribution. The center point of the space is determined as the coordinate value obtained by the current mapping . A cubic subspace with edge length 1 is generated around this center, which contains all historical decision points.

[0032] In the centralized search operation, the coordinate point corresponding to the fusion feature vector is constructed as the center of the spherical decision neighborhood. The spherical radius is set to 0.1, which is adjusted according to the distribution density of historical decision points. The point set within the spherical neighborhood is defined as all points satisfying , where denotes the Euclidean distance.

[0033] The calculation of neighborhood decision density uses the kernel density estimation method. For the historical decision points within the spherical neighborhood, the density calculation formula is: ; wherein the kernel function uses the Epanechnikov kernel: when , otherwise 0. This kernel function gives greater weight to points close to the center of the sphere, making the density estimation better reflect the local clustering characteristics.

[0034] ​The historical decision database needs to be maintained and updated with new successful screening cases regularly. The database records the values of the three stability factors and their corresponding quality level labels for each case. The latest data is automatically synchronized during the construction of the space to ensure that the feature fusion space always reflects the latest screening experience.

[0035] An outlier processing mechanism is set for the coordinate mapping link. When the stability factor of the current batch exceeds the historical range, an extreme value compression strategy is adopted: if , then take ; if , then take . The same processing is applied to the coordinate mapping of the Y and Z axes. This processing avoids the coordinate values from exceeding the range [0, 1] and ensures the integrity of the space structure. The radius of the spherical neighborhood is set with adaptive characteristics. When the overall distribution of historical decision points is sparse, the radius is automatically increased to 0.15; when the distribution is dense, the radius is reduced to 0.08. Adaptive adjustment is based on the average nearest neighbor distance of points in the entire feature space to ensure that a sufficient number of sample points are included in the neighborhood for density estimation. The implementation of kernel density estimation uses a hierarchical calculation strategy. First, calculate the distance between the sphere center and each historical decision point, and select the point set that meets the distance condition; then calculate the normalized distance of each point; finally, calculate the weight according to the kernel function and sum it up. Parallel optimization is used in the calculation process to improve processing efficiency.

[0036] The visualization auxiliary function of the feature fusion space provides a three-dimensional scatter plot display, where decision points of different quality levels are marked with different colors. The coordinate point corresponding to the current batch is displayed in a prominent color, and the spherical decision neighborhood is presented in the form of a transparent sphere. This visualization helps to understand the spatial geometric relationship of the decision-making process.

[0037] A quality control log is established during implementation to record the parameter settings, coordinate mapping results, and neighborhood density values of each space construction. When the log detects that the density value is abnormally low for several consecutive times, it triggers a data review process to check the reliability of the sensor data acquisition and feature extraction links.

[0038] In the implementation of the centralized search operation to determine the optimal screening decision vector, the search parameters need to be initialized first. The coordinate point of the current fusion feature vector in the three-dimensional feature fusion space is taken as the initial sphere center, the radius of the spherical decision neighborhood is set to 0.1 unit length, the preset number of iterations is 100 times, and the preset failure threshold is set to 10 times. The search process records the current sphere center coordinates, neighborhood decision density, candidate decision vector state, and the number of decision update failures.

[0039] When the search process is started for the first time, a point is randomly selected on the boundary of the current spherical decision neighborhood as the first candidate decision vector. The random selection uses a uniform distribution sampling algorithm to ensure that any location on the sphere has an equal probability of being selected. The neighborhood decision density of this candidate vector is calculated, which is the distribution density of historical decision points within a spherical region of the same radius centered at the point. The density calculation uses a kernel density estimation method, using an Epanechnikov kernel function for weighted calculation. The candidate neighborhood decision density is compared with the current neighborhood decision density of the sphere center. If the candidate density is greater than the current density, the sphere center is moved to the candidate point location, and the decision update failure count is reset. The spherical decision neighborhood is then reconstructed with the new sphere center, and the next iteration continues. If the candidate density is less than or equal to the current density, the decision update failure count is accumulated, and another candidate point is selected on the current neighborhood boundary for testing.

[0040] When the decision update failure count exceeds the preset failure threshold, the search process is terminated, and the decision vector corresponding to the current sphere center is determined as the optimal screening decision vector. The entire search process records detailed data for each iteration step, including the coordinates of each selected candidate point, the density calculation results, and the decision state. Referring to Table 1, the iteration search process record for a specific batch is shown.

[0041] Table 1: Optimal Screening Decision Vector Search Process Record Table Number of iterations Current sphere center coordinates (x, y, z) Current density Candidate point coordinates (x, y, z) Candidate density Density comparison result Failure count Decision state 1 (0.42,0.38,0.51) 0.72 (0.47,0.33,0.56) 0.68 Less than 1 Continue 2 (0.42,0.38,0.51) 0.72 (0.39,0.45,0.49) 0.75 Greater than 0 Move 3 (0.39,0.45,0.49) 0.75 (0.36,0.48,0.52) 0.78 Greater than 0 Move 4 (0.36,0.48,0.52) 0.78 (0.33,0.51,0.55) 0.81 Greater than 0 Move 5 (0.33,0.51,0.55) 0.81 (0.31,0.53,0.58) 0.79 Less than 1 Continue ... ... ... ... ... ... ... ... 23 (0.28,0.59,0.62) 0.92 (0.25,0.61,0.65) 0.89 Less than 8 Continue 24 (0.28,0.59,0.62) 0.92 (0.26,0.58,0.64) 0.90 Less than 9 Continue 25 (0.28,0.59,0.62) 0.92 (0.27,0.60,0.63) 0.88 Less than 10 Terminate As can be seen from the table record, the search process terminates after 25 iterations. The initial sphere center coordinates are (0.42, 0.38, 0.51), and the corresponding neighborhood decision density is 0.72. During the iteration process, the sphere center position gradually moves to areas with higher density. In the second iteration, the density of the candidate point (0.39, 0.45, 0.49) is 0.75, which is greater than the current density, and the sphere center moves to this location. In the third and fourth iterations, higher-density candidate points are successfully found, and the sphere center moves to (0.36, 0.48, 0.52) and (0.33, 0.51, 0.55), with densities increasing to 0.78 and 0.81, respectively.

[0042] From the fifth iteration, there are multiple cases where the density of the candidate point is lower than the current density, and the failure count gradually accumulates. In the 23rd-25th iterations, the densities of the three consecutive candidate points are all lower than the current density of 0.92, and the failure count reaches the threshold of 10 times, terminating the search process. The final optimal screening decision vector corresponds to the coordinates (0.28, 0.59, 0.62), which has a higher neighborhood decision density of 0.92, indicating that this region has gathered a large number of historical high-quality screening decisions.

[0043] The random sampling algorithm uses the Mason rotation algorithm to generate uniformly distributed random numbers, ensuring the uniform distribution of candidate points on the spherical surface. The coordinate calculation of each candidate point is realized through spherical coordinate conversion, generating azimuth and zenith angles randomly, and then converting them into rectangular coordinate system coordinates. In density calculation, the bandwidth parameter of the kernel function is adaptively adjusted according to the overall distribution density of historical decision points, ensuring the accuracy of density estimation. The iteration control mechanism includes timeout protection, which forcibly terminates when the number of iterations reaches the preset 100 times, avoiding infinite loops. At the same time, a density growth threshold is set, and if the density growth amplitude is less than 0.01 for 10 consecutive iterations, the search is terminated early, considering that the local optimal solution has been converged. All parameter settings are recorded in the system configuration file and can be adjusted according to different medicinal material varieties.

[0044] The historical decision database is regularly maintained, removing outdated decision records and adding new successful screening cases. The database records the three-dimensional coordinates, density values, quality grade labels, and collection timestamps of each decision point. The system automatically performs cluster analysis on historical data to identify spatial distribution areas corresponding to different quality grades, providing a reference for the search process. The visual monitoring interface displays the search process in real time, showing the distribution of historical decision points in the form of a three-dimensional scatter plot, with the current sphere center marked as a red sphere, candidate points marked as yellow dots, and search paths displayed as connected lines. The operator can visually observe the progress of the search process.

[0045] The abnormal handling mechanism includes multiple levels: when 20 consecutive candidate points cannot find higher density, the search radius is automatically expanded by 50%; when the search process falls into a local optimum, the restart mechanism is enabled to start searching from the original starting point; when there is insufficient historical data, the operator is prompted to supplement the data or adjust the parameters. All abnormal situations are recorded in the system log for subsequent analysis and optimization.

[0046] After obtaining the optimal screening decision vector, it needs to be normalized to generate the standard screening decision value. The vector is represented as a coordinate point in the three-dimensional feature fusion space, and its module length is obtained by calculating the Euclidean distance. Specifically, given the coordinate values of the optimal screening decision vector as (x, y, z), the module length is calculated as the square root of the sum of the squares of each coordinate. The maximum radius of the decision space is determined according to the maximum module length of all decision vectors in the historical data, usually taking 1.05 to 1.2 times the historical maximum module length to accommodate possible out-of-range values. Divide the calculated module length by the maximum radius to obtain the standard screening decision value, which is normalized to the [0, 1] interval. If the calculation result is greater than 1, take the value as 1; less than 0, take the value as 0.

[0047] The grading threshold interval of the standard screening decision value is set to three consecutive intervals: [0, 0.3] corresponds to low quality grade, [0.3, 0.7] corresponds to medium quality grade, and [0.7, 1.0] corresponds to high quality grade. Each interval includes the lower bound but not the upper bound, for example, 0.3 belongs to the low quality grade interval, and 0.7 belongs to the medium quality grade interval. The quality grade label uses a textual description, with the low quality grade labeled as "Level 3", the medium quality grade labeled as "Level 2", and the high quality grade labeled as "Level 1". A threshold adjustment mechanism is established during implementation, and the system regularly statistics the screening result distribution of the last 100 batches. When the distribution changes significantly, the grading threshold interval is automatically adjusted. The adjustment principle is to keep the proportion of samples in each grade relatively stable, and to avoid the grade distribution deviation caused by environmental changes or raw material differences. The adjustment amplitude is not more than 0.05 each time, to ensure the stability of the threshold change.

[0048] When outputting the quality grade label, a confidence index is also generated. The confidence is calculated according to the distance of the standard screening decision value from the nearest threshold boundary. The greater the distance, the higher the confidence. The confidence is divided into high, medium and low levels, corresponding to a distance greater than 0.1, between 0.05 and 0.1, and less than 0.05, respectively. High confidence results are directly output, medium confidence results are prompted for review, and low confidence results start the manual review process. When the standard screening decision value falls between 0.29 and 0.31 or between 0.69 and 0.71, the system automatically performs a secondary verification. The secondary verification is performed by analyzing the original feature data distribution of the batch of decoction pieces. If the feature distribution shows obvious bias, the grade division is adjusted. For example, if the standard screening decision value is 0.305 but the morphological features are significantly better than those of the same level two decoction pieces, it may be upgraded to level one.

[0049] Detailed logs are recorded during implementation, including the optimal screening decision vector coordinates of each batch, the calculated modulus, the used maximum radius value, the standard screening decision value, the finally determined quality grade, and the confidence level. Log data is used for subsequent system optimization and parameter adjustment. A feedback mechanism is established to compare the actual use effect with the system prediction result, and gradually optimize the grading threshold interval setting. The system provides manual adjustment function, allowing experienced operators to fine-tune the grading threshold according to the actual situation. The manual adjustment records the change reason and adjustment amplitude, which are used to improve the adaptive learning ability of the system. All adjustment operations need double confirmation to ensure the rationality and traceability of the changes. The output of the quality grade label uses a standardized format, including batch number, decoction piece variety, detection time, standard screening decision value, quality grade and confidence level. The output information is also presented in a visual way, with different colors representing different grades: green represents level one, blue represents level two, and yellow represents level three. This intuitive display method allows operators to quickly grasp the screening results.

[0050] The system periodically performs retrospective analysis on historical grading results to check for systematic bias. When a consistent deviation in grade distribution is found for a particular variety, a special calibration procedure is initiated. The calibration procedure optimizes feature weights and threshold settings by reanalyzing historical data for that variety, ensuring accuracy and consistency in grading standards. The entire implementation process focuses on stability and repeatability. All computational parameters and threshold settings are saved in configuration files, which are automatically loaded each time the system is started. Configuration file versions are strictly managed, and any modifications require a change note and effective time. This management approach ensures consistency in screening results across different times and different operators.

[0051] It should be noted that the relationship terms such as first and second, and the like, are used merely to distinguish one entity or action from another, without necessarily requiring or implying that there is any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0052] While embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, combinations, and variations of the embodiments can be made by those skilled in the art without departing from the spirit and scope of the present application, which is defined by the following claims and their equivalents.

Claims

1. A multi-modal data fusion intelligent screening method for traditional Chinese medicine decoction pieces, characterized by, The method comprises the following steps: Synchronously collecting visible light image data, near-infrared spectrum data and odor volatile data of a target traditional Chinese medicine decoction piece through a multi-source sensor array; Performing a morphological cleaning operation on the visible light image data, a baseline drift correction operation on the near-infrared spectrum data and an environmental interference elimination operation on the odor volatile data; Extracting a decoction piece morphology feature set from the cleaned visible light image data, a chemical component feature set from the corrected near-infrared spectrum data and a volatile substance feature set from the interference-eliminated odor volatile data; Calculating a morphology stability factor of the decoction piece morphology feature set, a component stability factor of the chemical component feature set and a volatility stability factor of the volatile substance feature set; Constructing a three-dimensional feature fusion space according to the morphology stability factor, the component stability factor and the volatility stability factor, and mapping the decoction piece morphology feature set, the chemical component feature set and the volatile substance feature set into a fusion feature vector; Performing a centralized search operation in the three-dimensional feature fusion space with the fusion feature vector as a starting point to determine an optimal screening decision vector; Classifying and screening the target traditional Chinese medicine decoction piece according to the optimal screening decision vector. 2.The intelligent screening method of traditional Chinese medicine decoction pieces based on multi-modal data fusion according to claim 1, characterized in that, The method for performing a morphological cleaning operation on the visible light image data is: Identifying a background pixel area in the visible light image data, and separating adhered decoction piece contours by using a morphological opening operation; Calculating a minimum circumscribed rectangle of each decoction piece contour to generate a decoction piece spatial distribution map; Eliminating edge incomplete areas according to the decoction piece spatial distribution map to output a complete decoction piece image data set. 3.The intelligent screening method of traditional Chinese medicine decoction pieces based on multi-modal data fusion according to claim 2, characterized in that, The method for extracting a decoction piece morphology feature set from the cleaned visible light image data is: Based on the complete decoction piece image data set, calculating a surface texture complexity, an edge curvature distribution histogram and an aspect ratio feature of each decoction piece; Aggregating the surface texture complexity, the edge curvature distribution histogram and the aspect ratio feature of all decoction pieces to form the decoction piece morphology feature set. 4.The intelligent screening method of traditional Chinese medicine decoction pieces based on multi-modal data fusion according to claim 1, characterized in that, The method for calculating a morphology stability factor of the decoction piece morphology feature set, a component stability factor of the chemical component feature set and a volatility stability factor of the volatile substance feature set is: Statistically calculating a first quartile and a third quartile of all surface texture complexities in the decoction piece morphology feature set; Extending the first quartile downward by a preset texture offset as a lower bound of texture aggregation, and extending the third quartile upward by the preset texture offset as an upper bound of texture aggregation; Calculating a variance value of surface texture complexity between the lower bound of texture aggregation and the upper bound of texture aggregation, and setting the value as the morphology stability factor. 5.The intelligent screening method of traditional Chinese medicine decoction pieces based on multi-modal data fusion according to claim 1, characterized in that, The method for constructing a three-dimensional feature fusion space according to the morphology stability factor, the component stability factor and the volatility stability factor is: Predefining three-dimensional space coordinate axes, wherein an X-axis corresponds to a decoction piece morphology feature dimension, a Y-axis corresponds to a chemical component feature dimension and a Z-axis corresponds to a volatile substance feature dimension; Mapping the morphology stability factor into an X-axis coordinate value, the component stability factor into a Y-axis coordinate value and the volatility stability factor into a Z-axis coordinate value; and A three-dimensional coordinate point formed by the X-axis coordinate value, the Y-axis coordinate value and the Z-axis coordinate value as a center point generates the three-dimensional feature fusion space containing the historical screening decision vector.

6. The intelligent screening method of traditional Chinese medicine decoction pieces according to claim 5, characterized in that, The method for performing a centralized search operation with the fusion feature vector as a starting point is: In the three-dimensional feature fusion space, a target subspace containing the fusion feature vector is selected; A spherical decision neighborhood is constructed with the fusion feature vector as a spherical center and a preset decision step length as a radius; The neighborhood decision density of the historical screening decision vector in the spherical decision neighborhood is calculated.

7. The intelligent screening method of traditional Chinese medicine decoction pieces based on multi-modal data fusion according to claim 6, characterized in that, The method for determining the optimal screening decision vector is: A first candidate decision vector is randomly selected at the edge of the spherical decision neighborhood; The first candidate neighborhood decision density of the first candidate decision vector is calculated; When the first candidate neighborhood decision density is greater than the neighborhood decision density, the first candidate decision vector is updated as a current spherical center and a spherical decision neighborhood is reconstructed, and the iteration is performed until a preset iteration number is met.

8. The intelligent screening method for Chinese herbal medicine slices based on multimodal data fusion according to claim 7, characterized in that: The method further comprises: When the first candidate neighborhood decision density is less than the neighborhood decision density, the number of decision update failures is accumulated; If the number of decision update failures exceeds a preset failure threshold, a decision vector corresponding to the current spherical center is set as the optimal screening decision vector. 9.The intelligent screening method of traditional Chinese medicine decoction pieces based on multi-modal data fusion according to claim 1, characterized in that, The method for performing grade classification and quality screening on the target traditional Chinese medicine decoction piece according to the optimal screening decision vector is: Normalization processing is performed on the optimal screening decision vector to generate a standard screening decision value; According to the distribution of the standard screening decision value in a preset grading threshold interval, a quality grade label of the target traditional Chinese medicine decoction piece is output.

10. The intelligent screening method of traditional Chinese medicine decoction pieces according to multi-modal data fusion according to claim 9, characterized in that, The method for performing normalization processing on the optimal screening decision vector is: The length of the optimal screening decision vector in the three-dimensional feature fusion space is calculated; The length is divided by a preset maximum decision space radius to obtain the standard screening decision value.

Citation Information

Patent Citations

  • Intelligent screening method for traditional Chinese medicine decoction pieces

    CN118964873A

  • Multi-source data intelligent fusion medicinal material identification system and medicinal material identification method

    CN119829920A

  • Intelligent screening method and system for traditional Chinese medicine decoction pieces

    CN120352592A

  • Deep learning-based traditional Chinese medicine quality intelligent detection method and system

    CN120375370A

  • Traditional Chinese medicine decoction piece quality evaluation method based on big data

    CN120579901A

Cited By

  • Intelligent glue pudding quality detection system and method based on multi-sensor fusion

    CN121476213A

  • Intelligent Quality Detection System and Method for Tangyuan Based on Multi-Sensor Fusion

    CN121476213B