Construction method of medicinal material identification model based on image technology
By constructing a medicinal material identification map and reshaping texture features, an intelligent medicinal material identification model is generated, which solves the problems of low efficiency and insufficient accuracy of traditional medicinal material identification methods and realizes the automation and efficient management of medicinal material identification.
Patent Information
- Application Number
- CN202510198490.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-22
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-02-22
AI Technical Summary
Traditional medicinal material identification methods rely on manual experience, are inefficient and prone to errors, and are difficult to achieve high-precision and high-efficiency medicinal material identification when there are many types of medicinal materials with large morphological differences. Existing image technology-based methods lack robustness and accuracy in the face of differences and diversity in medicinal material image quality.
By acquiring medicinal material sample images for feature extraction, constructing a medicinal material identification map, collecting and reshaping texture features, generating a comprehensive feature representation, performing growth extension simulation and initial identification, performing result matching and edge comparison, reconstructing the central contour, and finally generating an intelligent medicinal material identification model.
It has achieved automation and improved accuracy in medicinal material identification, improved the efficiency and accuracy of medicinal material identification, enhanced the adaptability and robustness of the model, and supported the efficient management of medicinal materials.
Smart Images

Figure CN120126107B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of identification model construction, and in particular to a method for constructing a medicinal material identification model based on image technology. Background Art
[0002] Traditional medicinal material identification methods rely on manual experience and chemical detection methods. Although they can provide certain recognition effects, these methods require high operator experience and are time-consuming, resulting in low efficiency in large-scale applications. Especially when there are many types of medicinal materials with large morphological differences, these methods cannot guarantee the unity of high precision and high efficiency. In addition, manual operation is prone to introduce human errors, affecting the accuracy of recognition results, which in turn restricts the intelligent development of medicinal material identification. With the development of image technology and machine learning, medicinal material identification methods based on image technology have gradually attracted attention. In particular, deep processing of medicinal material images through techniques such as feature extraction and texture analysis has provided new ideas for medicinal material identification. However, due to the differences in image quality of medicinal material samples and the complexity and diversity of medicinal material images themselves, how to accurately and quickly extract effective image features and perform intelligent processing on them remains a difficult problem that needs to be solved urgently. In particular, when the texture features of medicinal material images are inconsistent or the number of samples is unbalanced, the robustness and accuracy of existing methods are challenged. Summary of the Invention
[0003] Based on this, it is necessary to provide a method for constructing a medicinal material identification model based on image technology to solve at least one of the above technical problems.
[0004] To achieve the above objectives, a method for constructing a medicinal material identification model based on image technology includes the following steps:
[0005] Step S1: Acquire a medicinal material sample image; perform feature extraction on the medicinal material sample image to obtain sample feature parameters; construct a knowledge graph for the medicinal material sample image based on the sample feature parameters to generate a medicinal material identification graph;
[0006] Step S2: Collecting original images of medicinal materials; reshaping the original images of medicinal materials with texture features to obtain standard original images of medicinal materials; performing multi-dimensional feature extraction and fusion modeling on the standard original images of medicinal materials to generate comprehensive feature representation;
[0007] Step S3: performing growth extension simulation on the original image of the standard medicinal material based on the comprehensive feature representation to generate candidate growth trajectories; performing initial category identification on the original image of the medicinal material based on the candidate growth trajectories and the medicinal material identification atlas to obtain an initial identification result;
[0008] Step S4: matching the medicinal material identification atlas based on the initial identification results to obtain a collection of matching medicinal material images; performing edge comparison on the collection of matching medicinal material images and the original image of the standard medicinal material to obtain edge comparison parameters;
[0009] Step S5: reconstructing the central contour of the original image of the medicinal material according to the edge contrast parameter to generate a reconstructed central contour; performing detail texture retrieval on the reconstructed central contour based on the matching image collection to generate a final matching result;
[0010] Step S6: Mark the final matching result and the original image of the medicinal material with a line to generate corresponding result data; use the corresponding result data to perform model training to generate an intelligent medicinal material identification model.
[0011] The present invention ensures the comprehensive capture of sample features by acquiring medicinal material sample images and performing feature extraction. The construction of the medicinal material identification atlas provides a knowledge basis for subsequent identification. The collection of original medicinal material images and the reconstruction of texture features optimize the image quality. The generated standard medicinal material original images provide a good foundation for multi-dimensional feature extraction. The generation of comprehensive feature representation improves the accuracy of medicinal material identification. The implementation of growth extension simulation provides a dynamic perspective for the category identification of medicinal materials. The generation of candidate growth trajectories enriches the identification information. The acquisition of initial identification results ensures rapid classification. The implementation of result matching enhances the practicality of the atlas. The formation of a collection of matching medicinal material images provides data support for edge contrast. The generation of edge contrast parameters improves the recognition ability of image details. The implementation of reconstructing the central contour ensures the accurate reproduction of the medicinal material morphology. The application of detail texture retrieval enhances the accuracy of the final matching result. The generation of connection marks provides a reliable data basis for model training. The construction of the intelligent medicinal material identification model improves the automation and accuracy of medicinal material identification, and ultimately achieves efficient medicinal material identification and management.
[0012] Preferably, step S1 includes the following steps:
[0013] Step S11: Acquire a medicinal material sample image; perform structure morphology spectrum mapping on the medicinal material sample image to obtain morphology mapping parameters;
[0014] Step S12: extracting features from the medicinal material sample image according to the morphological mapping parameters to obtain sample feature parameters;
[0015] Step S13: performing feature association analysis on the sample feature parameters to obtain a feature association network; performing hierarchical aggregation mapping on the feature association network to generate a category hierarchy matrix;
[0016] Step S14: reconstructing the feature distribution of the category hierarchy matrix to obtain classification feature data; mapping the classification feature data to an identification map to generate a medicinal material identification map.
[0017] The present invention can extract the morphological characteristics of medicinal materials by acquiring medicinal material sample images and performing structural morphological spectrum mapping, thereby providing a basis for subsequent feature extraction. The sample feature parameters obtained during the feature extraction process can comprehensively reflect the key characteristics of the medicinal materials. Feature association analysis can reveal the relationship between sample features and form a feature association network, which is helpful to understand the intrinsic connection of medicinal material features. Hierarchical aggregation mapping can effectively classify features by category and generate a category hierarchical matrix. Feature distribution reconstruction makes the classification feature data more accurate and can effectively improve the accuracy of classification. Identification map mapping converts the classification feature data into a visual medicinal material identification map, which is convenient for rapid identification and comparison of medicinal materials. The whole process not only improves the efficiency of medicinal material identification, but also enhances the scalability and adaptability of the model, providing a solid technical foundation for subsequent intelligent medicinal material identification.
[0018] Preferably, step S2 includes the following steps:
[0019] Step S21: collecting original images of medicinal materials; performing spectral tomography on the original images of medicinal materials to obtain spectral tomography data;
[0020] Step S22: performing superpixel segmentation on the spectral tomography data to generate medicinal material segmentation features; performing texture migration and reshaping on the medicinal material segmentation features to obtain a standard medicinal material original image;
[0021] Step S23: performing multi-scale gradient extraction on the original image of the standard medicinal material to obtain a gradient feature map; performing texture descriptor calculation on the gradient feature map to generate a texture feature vector;
[0022] Step S24: performing shape contour extraction on the texture feature vector to obtain a contour feature set; performing color matrix decomposition on the contour feature set to generate a color feature matrix;
[0023] Step S25: Perform spatial distribution statistics on the color feature matrix to obtain a spatial feature vector; perform feature fusion on the spatial feature vector to generate a comprehensive feature representation.
[0024] By collecting original images of medicinal materials and performing spectral tomography, the present invention can obtain comprehensive spectral information and provide rich medicinal material component characteristics. After superpixel segmentation, the spectral tomography data can achieve fine segmentation of the medicinal materials, ensuring the accuracy of feature extraction. Texture migration and reshaping further improves the standardization of the image, making subsequent processing more consistent. Multi-scale extraction of gradient feature maps helps to capture subtle changes in medicinal materials and enhances the robustness of features. The texture feature vectors generated by texture descriptor calculation effectively characterize the surface features of the medicinal materials. The extraction of contour feature sets can reveal the shape characteristics of the medicinal materials. Color matrix decomposition provides higher expression capabilities for color features. The analysis of the color feature matrix by spatial distribution statistics enables spatial feature vectors to have stronger distinguishing capabilities. The comprehensive feature representation generated after feature fusion gathers multi-dimensional information, provides a scientific basis for the comprehensive identification of medicinal materials, and improves the overall performance and accuracy of the identification model.
[0025] Preferably, step S3 includes the following steps:
[0026] Step S31: performing feature deformation simulation on the original image of the standard medicinal material based on the comprehensive feature representation to generate simulated deformation features;
[0027] Step S32: predicting the growth direction of the simulated deformation feature to obtain a candidate growth trajectory;
[0028] Step S33: performing feature point detection on the candidate growth trajectory to obtain key trajectory points; performing projection matching on the key trajectory points and the medicinal material identification map to generate a matching medicinal material map;
[0029] Step S34: performing initial identification of the category of the original medicinal material image according to the matched medicinal material atlas to obtain an initial identification result.
[0030] The present invention can generate a variety of simulated deformation features by simulating the feature deformation of the original image of the standard medicinal material, thereby enhancing the adaptability and flexibility of the model. The growth direction prediction provides directional information for the subsequent growth trajectory analysis. The generation of candidate growth trajectories helps to identify the growth law of the medicinal material. The feature point detection can accurately extract the key trajectory points, thereby improving the accuracy of information extraction. The projection matching of the key trajectory points and the medicinal material identification map realizes the effective integration of information. The generated matching medicinal material map provides a solid basis for the subsequent category identification. The initial category identification based on the matching medicinal material map can quickly identify the medicinal material type, thereby improving the identification efficiency and accuracy. The overall process enhances the intelligence level of the model, and provides technical support and guarantee for the automated identification of medicinal materials.
[0031] Preferably, step S32 includes the following steps:
[0032] Perform regional segmentation on the simulated deformation features to obtain a growth block map;
[0033] Perform boundary positioning on the growth block map to generate a boundary trajectory set; perform direction chain encoding on the boundary trajectory set to generate a trajectory encoding sequence;
[0034] Calculate the curvature of the trajectory coding sequence to obtain curvature feature data;
[0035] Predicting the growth direction based on the curvature feature data to obtain a set of predicted trajectories;
[0036] Perform deformation feature recognition on the predicted trajectory set to generate predicted deformation features;
[0037] Trajectory screening is performed based on the simulated deformation characteristics and the predicted deformation characteristics to generate candidate growth trajectories.
[0038] The present invention can effectively identify growth blocks by performing regional segmentation on simulated deformation features, thereby improving the efficiency of regional feature extraction. Boundary positioning provides clear boundary information for subsequent analysis, and the generated boundary trajectory set enhances the continuity of features. Directional chain encoding provides a systematic representation for trajectories. Curvature calculation can reveal subtle differences in morphological changes. The generated curvature feature data provides an important basis for growth direction prediction. The generation of predicted trajectory sets makes the analysis of growth trends more accurate. Deformation feature recognition can ensure the capture of important features. The trajectory screening process improves the accuracy of candidate growth trajectories. The overall process enhances the model's adaptability to medicinal material morphological changes, providing more comprehensive and reliable technical support for medicinal material identification.
[0039] Preferably, step S4 includes the following steps:
[0040] Step S41: constructing a graph index for the initial identification result to obtain index feature data; performing similarity measurement on the index feature data to generate similarity data;
[0041] Step S42: performing threshold screening on the medicinal material identification atlas based on the similarity data to obtain a collection of matching medicinal material images;
[0042] Step S43: performing image edge cutting on the standard medicinal material original image to obtain a medicinal material edge image; performing contrast enhancement processing on the medicinal material edge image to generate an enhanced edge image;
[0043] Step S44: performing edge contour tracing on the enhanced edge image to obtain the edge contour of the medicinal material; performing contour overlap comparison on the edge contour of the medicinal material based on the collection of matched medicinal material images to generate edge comparison parameters.
[0044] The present invention can generate systematic index feature data by constructing a graph index for the initial identification results, thereby enhancing the searchability of the data. The similarity measurement provides a quantitative basis for subsequent matching, and the generated similarity data can effectively reflect the similarity relationship between medicinal materials. The threshold screening process ensures the accuracy and relevance of the matching medicinal material image collection. The cutting of the medicinal material edge image improves the fineness of feature extraction, and the contrast enhancement process significantly improves the clarity of the image. The generated enhanced edge image provides a more distinct feature expression for subsequent analysis. The edge contour copying ensures the integrity of the edge feature. The contour overlap comparison can accurately identify similar features. The generated edge contrast parameters provide strong support for medicinal material identification. The overall process improves the recognition accuracy and efficiency of the model, and lays a solid technical foundation for the automated identification of medicinal materials.
[0045] Preferably, step S44 includes the following steps:
[0046] Extracting curvature features from the enhanced edge image to generate curvature distribution data; reconstructing the curvature distribution data through contour interpolation to obtain contour interpolation data;
[0047] Perform vectorization conversion processing on the contour interpolation data to obtain contour vector data; perform feature point positioning processing on the contour vector data to generate contour key point data;
[0048] Perform spline curve fitting on the key point data of the contour to obtain the edge contour of the medicinal material;
[0049] Perform contour registration on the edge contour of the medicinal material and the collection of matching medicinal material images to obtain registration contour data; perform overlap evaluation processing on the registration contour data to generate overlap evaluation data;
[0050] The overlap difference is quantified based on the overlap evaluation data to obtain the edge contrast parameter.
[0051] The present invention can obtain subtle changes in edge morphology by extracting curvature features from enhanced edge images. The generated curvature distribution data provides important information for contour morphology analysis. Contour interpolation reconstruction improves the smoothness and continuity of the contour. The obtained contour interpolation data provides a stable basis for subsequent processing. Vectorization conversion processing ensures the digitization and simplification of contour features. The generated contour vector data is convenient for subsequent analysis and calculation. Feature point positioning processing can accurately identify key points and enhance the expressiveness of the contour. Spline curve fitting realizes a refined description of the edge contour of medicinal materials. Registration processing standardizes the collection of matching medicinal material images to ensure the alignment and consistency of the contours. Overlap evaluation processing can quantify the similarity between contours. The generated overlap evaluation data provides a reliable basis for further feature analysis. Overlap difference quantification realizes in-depth comparison of edge features. The overall process improves the accuracy and efficiency of medicinal material identification and provides strong technical support for automated classification and recognition.
[0052] Preferably, step S5 includes the following steps:
[0053] Step S51: performing edge contour separation on the original image of the medicinal material according to the edge contrast parameter to obtain a separated edge contour; performing three-dimensional reconstruction on the original image of the medicinal material to generate a three-dimensional medicinal material model;
[0054] Step S52: performing de-edge positioning on the three-dimensional medicinal material model based on the separated edge contour to obtain a center point area; performing central contour cutting on the three-dimensional medicinal material model according to the center point area to generate a reconstructed center contour;
[0055] Step S53: extracting texture features from the reconstructed center contour to obtain texture feature parameters;
[0056] Step S54: performing detail texture retrieval on the texture feature parameters based on the matching image collection to generate a final matching result.
[0057] The present invention separates the edge contours of the original image of the medicinal material through edge contrast parameters, which can effectively extract the edge features of the medicinal material. The generated separated edge contours provide an accurate basis for subsequent analysis. The three-dimensional reconstruction of the medicinal material realizes the three-dimensional expression of the morphology. The generated three-dimensional medicinal material model enhances the comprehensive understanding of the shape of the medicinal material. The de-marginalization positioning ensures the accurate identification of the center point area. The centralized contour cutting provides a centralized perspective for the analysis of the three-dimensional model. The generation of the reconstructed center contour improves the representativeness of the model. The texture feature extraction can obtain the detailed information of the medicinal material surface. The generated texture feature parameters provide a quantitative basis for subsequent matching. The detailed texture retrieval realizes the depth comparison of the features based on the matching image collection. The generation of the final matching result ensures the accuracy and reliability of the medicinal material identification. The overall process improves the efficiency of the medicinal material identification and provides strong technical support for automatic identification.
[0058] Preferably, step S54 includes the following steps:
[0059] Performing scale normalization on texture feature parameters to obtain standard texture features; performing directionality analysis on standard texture features to generate texture direction vectors;
[0060] Partition encoding is performed on standard texture features based on texture direction vectors to obtain regional texture encoding;
[0061] Perform coding feature matching on the regional texture coding according to the matching image collection to generate a matching score matrix; sort the matching image collection based on the matching score matrix to obtain a sorted matching image;
[0062] The top search is performed based on the ranked matching images to generate the final matching results.
[0063] The present invention ensures the consistency and comparability of features by performing scale normalization processing on texture feature parameters. The generated standard texture features provide a unified basis for subsequent analysis. Directional analysis can reveal the directional features of the texture. The generated texture direction vector provides important information for the spatial distribution of the texture. The generation of regional texture coding realizes the systematic representation of texture features. The coding feature matching can effectively quantify the similarity between different images. The generated matching score matrix provides an accurate basis for subsequent sorting. Image sorting can be prioritized according to the degree of matching. The obtained sorted matching image is convenient for rapid identification and selection. The first-place retrieval ensures the efficiency and accuracy of the final matching result. The overall process improves the accuracy and efficiency of medicinal material identification and provides strong technical support for automated identification and classification.
[0064] Preferably, step S6 includes the following steps:
[0065] Step S61: extracting feature points from the final matching result to obtain feature point data; constructing line rules on the feature point data to obtain line rule data;
[0066] Step S62: Mark the final matching result and the original image of the medicinal material based on the connection rule data to generate corresponding result data;
[0067] Step S63: performing model parameterization processing based on the corresponding result data to obtain identification model parameters; iteratively optimizing the identification model parameters to generate iterative model data;
[0068] Step S64: Integrate the iterative model data to generate an intelligent medicinal material identification model.
[0069] The present invention can accurately identify key features by performing feature point extraction processing on the final matching results. The generated feature point data provides a basis for subsequent analysis. The connection rule construction realizes the systematic association between features. The obtained connection rule data enhances the visualization and understanding of features. Connection marking based on the connection rule data ensures the clear expression of matching results. The generated corresponding result data provides an important basis for model parameterization. The model parameterization processing can quantify the influence of features. The iterative optimization improves the accuracy and stability of the model. The generated iterative model data provides reliable support for subsequent applications. The integrated processing realizes the fusion of multiple models. The generated intelligent medicinal material identification model enhances the adaptability and practicality of the system, and provides a strong technical foundation for the automated identification of medicinal materials. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 A schematic diagram of the steps of the method for constructing a medicinal material identification model based on image technology;
[0071] Figure 2 Detailed implementation flow chart of step S2;
[0072] Figure 3 Detailed implementation flow chart of step S3;
[0073] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0074] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.
[0075] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.
[0076] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.
[0077] To achieve this, please refer to Figures 1 to 3 The method for constructing a medicinal material identification model based on image technology includes the following steps:
[0078] Step S1: Acquire a medicinal material sample image; perform feature extraction on the medicinal material sample image to obtain sample feature parameters; construct a knowledge graph for the medicinal material sample image based on the sample feature parameters to generate a medicinal material identification graph;
[0079] Step S2: Collecting original images of medicinal materials; reshaping the original images of medicinal materials with texture features to obtain standard original images of medicinal materials; performing multi-dimensional feature extraction and fusion modeling on the standard original images of medicinal materials to generate comprehensive feature representation;
[0080] Step S3: performing growth extension simulation on the original image of the standard medicinal material based on the comprehensive feature representation to generate candidate growth trajectories; performing initial category identification on the original image of the medicinal material based on the candidate growth trajectories and the medicinal material identification atlas to obtain an initial identification result;
[0081] Step S4: matching the medicinal material identification atlas based on the initial identification results to obtain a collection of matching medicinal material images; performing edge comparison on the collection of matching medicinal material images and the original image of the standard medicinal material to obtain edge comparison parameters;
[0082] Step S5: reconstructing the central contour of the original image of the medicinal material according to the edge contrast parameter to generate a reconstructed central contour; performing detail texture retrieval on the reconstructed central contour based on the matching image collection to generate a final matching result;
[0083] Step S6: Mark the final matching result and the original image of the medicinal material with a line to generate corresponding result data; use the corresponding result data to perform model training to generate an intelligent medicinal material identification model.
[0084] The present invention ensures the comprehensive capture of sample features by acquiring medicinal material sample images and performing feature extraction. The construction of the medicinal material identification atlas provides a knowledge basis for subsequent identification. The collection of original medicinal material images and the reconstruction of texture features optimize the image quality. The generated standard medicinal material original images provide a good foundation for multi-dimensional feature extraction. The generation of comprehensive feature representation improves the accuracy of medicinal material identification. The implementation of growth extension simulation provides a dynamic perspective for the category identification of medicinal materials. The generation of candidate growth trajectories enriches the identification information. The acquisition of initial identification results ensures rapid classification. The implementation of result matching enhances the practicality of the atlas. The formation of a collection of matching medicinal material images provides data support for edge contrast. The generation of edge contrast parameters improves the recognition ability of image details. The implementation of reconstructing the central contour ensures the accurate reproduction of the medicinal material morphology. The application of detail texture retrieval enhances the accuracy of the final matching result. The generation of connection marks provides a reliable data basis for model training. The construction of the intelligent medicinal material identification model improves the automation and accuracy of medicinal material identification, and ultimately achieves efficient medicinal material identification and management.
[0085] In an embodiment of the present invention, the method for constructing a medicinal material identification model based on image technology includes the following steps:
[0086] Step S1: Acquire a medicinal material sample image; perform feature extraction on the medicinal material sample image to obtain sample feature parameters; construct a knowledge graph for the medicinal material sample image based on the sample feature parameters to generate a medicinal material identification graph;
[0087] In this embodiment, when acquiring the medicinal material sample image, a high-resolution industrial camera is used for image acquisition. The industrial camera model is a 50-megapixel CMOS sensor camera equipped with a 50mm fixed-focus F1.8 lens to ensure that the acquired image has sufficient detail resolution. The image acquisition light source adopts a ring LED light source with a color temperature of 5000K to reduce the interference of ambient light on the image quality. During the acquisition process, the medicinal material sample is placed on a matte black background board to improve the edge contrast. The acquired sample image format is RAW format, and the resolution is set to 8192×5460 to ensure the accuracy of subsequent feature extraction. During feature extraction, a deep convolutional neural network (CNN) is used to extract features from the sample image. The extracted features include global color distribution, histogram of oriented gradients (HOG), local binary pattern (LBP), and the like. Pattern, LBP) (local binary pattern), etc. The dimension of the feature vector is set to 2048 dimensions to ensure that it contains sufficient image information. When constructing the knowledge graph using sample feature parameters, a graph database is used to store the sample feature data. During the construction process, each medicinal material sample corresponds to a node. The attributes of the node include the medicinal material name, main texture feature parameters, color histogram mean, shape contour feature value, etc. The edge is used to represent the similarity relationship between different samples. The similarity calculation uses cosine similarity, and the threshold is set to 0.85. The final knowledge graph is stored in the Neo4j database for subsequent rapid retrieval.
[0088] Step S2: Collecting original images of medicinal materials; reshaping the original images of medicinal materials with texture features to obtain standard original images of medicinal materials; performing multi-dimensional feature extraction and fusion modeling on the standard original images of medicinal materials to generate comprehensive feature representation;
[0089] In this embodiment, when collecting the original image of the medicinal material, the same industrial camera and light source configuration as the sample image is used to ensure data consistency. The image format is still set to RAW format. After the collection is completed, the image is denoised using Gaussian filtering, and the filter kernel size is set to 5×5. In the texture feature reconstruction process, the original image is first subjected to contrast limited adaptive histogram equalization (CLAHE) with a window size of 8×8 to enhance detail contrast. Then, wavelet transform is used for multi-scale texture decomposition. Daubechies4 (db4) wavelet basis is selected to perform four-level decomposition on the image. The low-frequency component is extracted for reconstructing the standard medicinal material original image. The final standard medicinal material original image is used for subsequent feature extraction. Multi-dimensional feature extraction uses ResNet-50 (deep residual network) for convolution feature extraction, and combined with principal component analysis (PCA) Principal Component Analysis (PCA) is used to reduce the feature dimensionality and finally obtain a comprehensive feature representation. The dimension of the comprehensive feature representation is set to 1024 to ensure the validity of the feature.
[0090] Step S3: performing growth extension simulation on the original image of the standard medicinal material based on the comprehensive feature representation to generate candidate growth trajectories; performing initial category identification on the original image of the medicinal material based on the candidate growth trajectories and the medicinal material identification atlas to obtain an initial identification result;
[0091] In this embodiment, when performing growth extension simulation on the original image of the standard medicinal material based on the comprehensive feature representation, a generative adversarial network (GAN) is used to simulate the medicinal material growth. The generative model adopts the StyleGAN architecture and uses the comprehensive feature representation as input to generate medicinal material images at different growth stages. During the training process, the discriminator uses PatchGAN (local discriminator) and the input size is set to 256×256 to ensure the consistency of local features. The generation of candidate growth trajectories is based on image sequence optical flow estimation (Optical Flow Estimation) and the RAFT (Recurrent All-Pairs Field Transforms) model is used for trajectory estimation to ultimately obtain candidate growth trajectories. During the initial category identification, the similarity network in the medicinal material identification map is used to match the candidate growth trajectories. The similarity is calculated using a graph embedding method. The t-SNE (t-distributed stochastic neighbor embedding) is used for visual dimensionality reduction. The initial identification results are stored in JSON format for subsequent processing.
[0092] Step S4: matching the medicinal material identification atlas based on the initial identification results to obtain a collection of matching medicinal material images; performing edge comparison on the collection of matching medicinal material images and the original image of the standard medicinal material to obtain edge comparison parameters;
[0093] In this embodiment, based on the initial identification results, when matching the results of the medicinal material identification map, the shortest path search algorithm (Dijkstra algorithm) is used to find the most matching medicinal material category in the knowledge map. The generation of the matching medicinal material image collection is based on KNN (K-nearest neighbor) classification, and the K value is set to 5 to ensure that the 5 samples with the highest similarity are selected. When performing edge comparison on the matching medicinal material image collection and the original image of the standard medicinal material, the Canny (Canny Edge Detection) algorithm is used to extract the edge, and the thresholds are set to 50 and 150. The extracted edge data uses Hausdorff distance (Hausdorff Distance) for shape similarity measurement, and the calculated edge comparison parameters are used for subsequent centralized contour reconstruction.
[0094] Step S5: reconstructing the central contour of the original image of the medicinal material according to the edge contrast parameter to generate a reconstructed central contour; performing detail texture retrieval on the reconstructed central contour based on the matching image collection to generate a final matching result;
[0095] In this embodiment, when reconstructing the central contour of the original medicinal material image based on the edge contrast parameter, morphological processing (Morphological Processing) combined with the active contour model (Active Contour Model, ACM) is used to optimize the contour. The morphological processing includes dilation and erosion. The size of the structural element is set to 3×3. The energy function of the ACM adopts the Chan-Vese model. The number of iterations is set to 500 to ensure contour convergence. After the reconstruction is completed, the local feature matching algorithm SIFT (Scale-Invariant Feature Transform) is used to retrieve the detail texture of the reconstructed central contour based on the matching image collection. During the matching process, FLANN (Fast Library for Approximate Nearest Neighbors) is used to perform nearest neighbor search to generate the final matching result.
[0096] Step S6: Mark the final matching result and the original image of the medicinal material with a line to generate corresponding result data; use the corresponding result data to perform model training to generate an intelligent medicinal material identification model.
[0097] In this embodiment, when the final matching results and the original images of the medicinal materials are marked with lines, Delaunay triangulation is used for spatial connection. During the triangulation process, the maximum side length threshold is set to 30 pixels to prevent excessive connection. After the line marking is completed, the corresponding result data is generated. The corresponding result data is constructed into a training set using the PyTorch framework. During the training process, the loss function uses cross entropy loss, the optimization algorithm uses Adam (Adaptive Moment Estimation), the initial learning rate is set to 0.001, and the number of training rounds is set to 100. Finally, an intelligent medicinal material identification model is generated. The model structure uses EfficientNet-B3 (efficient neural network). After the training is completed, the model parameters are stored in ONNX format for subsequent deployment.
[0098] Preferably, step S1 includes the following steps:
[0099] Step S11: Acquire a medicinal material sample image; perform structure morphology spectrum mapping on the medicinal material sample image to obtain morphology mapping parameters;
[0100] Step S12: extracting features from the medicinal material sample image according to the morphological mapping parameters to obtain sample feature parameters;
[0101] Step S13: performing feature association analysis on the sample feature parameters to obtain a feature association network; performing hierarchical aggregation mapping on the feature association network to generate a category hierarchy matrix;
[0102] Step S14: reconstructing the feature distribution of the category hierarchy matrix to obtain classification feature data; mapping the classification feature data to an identification map to generate a medicinal material identification map.
[0103] In this embodiment, when acquiring images of medicinal material samples, a CMOS industrial camera with a resolution of 6000×4000 pixels is used for acquisition, equipped with a fixed-focus lens of F2.0 with a focal length of 35mm to ensure that the image is clear and the details are complete. The light source uses a surface light source with a color temperature of 5500K to provide uniform lighting. The medicinal material samples are placed on a gray anti-reflective background plate to avoid light reflection interference. During the acquisition process, a robotic arm is used to automatically adjust the displacement of the sample. The robotic arm driven by a stepper motor has an accuracy of 0.02mm to achieve multi-angle acquisition. Six images at different angles are collected for each sample. The images are saved in TIFF format to ensure lossless storage. The structural morphology spectrum mapping is performed using Fourier transform (Fourier transform). Transform) is used to extract the frequency domain features of the image. The Python OpenCV library is used for implementation. After the image is grayed, a 2D Fourier transform is used to obtain the spectrum diagram, and the high-frequency and low-frequency components are separated. The intermediate-frequency components are retained by a bandpass filter. The filter bandwidth is set to 30-70 Hz. The obtained morphological mapping parameters include frequency distribution matrix, phase matrix and amplitude matrix. When extracting features according to the morphological mapping parameters, a method based on Gabor filter is used for image texture analysis. The Gabor filter parameters are set to 6 direction angles ranging from 0° to 150°, with a step size of 30°, a scale set to 4 scales, and a frequency parameter range of 0.05-0.4. The extracted texture features include statistics such as energy, mean and standard deviation. Combined with the color histogram features, the HSV color space is used to quantize each channel into 16 levels to form a 48-dimensional color feature vector. The feature fusion algorithm is used. The algorithm uses a weighted averaging method to perform a weighted fusion of texture and color features, with a weight ratio of 0.7:0.3. This results in a 2048-dimensional sample feature parameter vector. Mutual information analysis is used to calculate the correlation between features using the mutual_info_classif function in the scikit-learn library. The discretization parameter is set to 10 intervals. Mutual information values are calculated for each pair of features to generate a feature correlation matrix. A threshold of 0.15 is set, and feature pairs below the threshold are considered irrelevant. The feature correlation matrix is then fed into a graph convolutional network (GCN) for association learning. The GCN uses a three-layer structure with 512, 256, and 128 hidden nodes per layer. ReLU (rectified linear unit) activation function is used, and the learning rate is set to 0.0005, the number of iterations is 200, the output feature association network is stored in the graph structure data format, the hierarchical aggregation mapping is implemented by an algorithm based on hierarchical clustering, the Ward link method is used for clustering, the distance metric uses the Euclidean distance, the clustering results are mapped, and the generated category hierarchy matrix is an N×N matrix, where N is the number of medicinal material categories. When reconstructing the feature distribution of the category hierarchy matrix, the principal component analysis (PCA) method is used for matrix dimensionality reduction, and the cumulative variance explanation rate of the retained principal component is set to 95%. The obtained low-dimensional representation matrix is input into the autoencoder for reconstruction. The autoencoder adopts a symmetrical structure. The encoder and decoder each contain 2 hidden layers, with 128 and 64 nodes in each layer respectively. The activation function uses the Sigmoid function, and the mean squared error loss function (Mean Squared Error Loss), the number of training rounds was set to 500, and the learning rate was 0.001. The resulting reconstructed matrix was represented as categorical feature data. When mapping the categorical feature data to the identification map, the k-means clustering algorithm was used for classification. The k value was set to the number of medicinal material categories. The clustering process used randomly initialized centers, the maximum number of iterations was 300, and the convergence threshold was set to 1e-4. The final medicinal material identification map was stored in the Neo4j graph database. The map nodes represented the individual medicinal material categories, the edges represented the feature similarity between categories, and the edge weights were the distances between cluster centers.
[0104] Preferably, step S2 includes the following steps:
[0105] Step S21: collecting original images of medicinal materials; performing spectral tomography on the original images of medicinal materials to obtain spectral tomography data;
[0106] Step S22: performing superpixel segmentation on the spectral tomography data to generate medicinal material segmentation features; performing texture migration and reshaping on the medicinal material segmentation features to obtain a standard medicinal material original image;
[0107] Step S23: performing multi-scale gradient extraction on the original image of the standard medicinal material to obtain a gradient feature map; performing texture descriptor calculation on the gradient feature map to generate a texture feature vector;
[0108] Step S24: performing shape contour extraction on the texture feature vector to obtain a contour feature set; performing color matrix decomposition on the contour feature set to generate a color feature matrix;
[0109] Step S25: Perform spatial distribution statistics on the color feature matrix to obtain a spatial feature vector; perform feature fusion on the spatial feature vector to generate a comprehensive feature representation.
[0110] In this embodiment, when collecting the original image of the medicinal material, a CCD industrial camera with a resolution of 8000×6000 pixels and a lens of 50mm is used. An F1.8 fixed-focus lens ensures sufficient light flux and detail resolution. The light source is a ring LED light with a color temperature of 5000K to ensure uniform lighting. The medicinal material samples are placed on a black flocked background plate with a reflectivity of less than 2%. A mechanical slide driven by a stepper motor with an accuracy of 0.01mm is used to achieve multi-angle automatic shooting. A total of 7 original images are collected for each sample at angles of 0°, 30°, 60°, 90°, 120°, 150° and 180°. The images are losslessly saved in TIFF format. Spectral tomography uses a hyperspectral imaging device with a spectral range of 400nm to 1000nm and a spectral resolution of 5nm. Spectral preprocessing is performed using ENVI (Environmental Visualization Image Analysis System) to remove noise and background. A band selection algorithm is used to retain 20 spectral bands with identification value. When performing superpixel segmentation on the spectral tomography data, the SLIC superpixel segmentation algorithm (Simple Linear Iterative Segmentation) is used. Clustering) was implemented using the Python scikit-image library. The number of superpixels was set to 500, and the compactness parameter of each superpixel was 10 to ensure clear boundaries. The generated medicinal material segmentation features were stored in the form of pixel index matrices. Texture migration and reshaping used a deep convolutional generative adversarial network (DCGAN). The generator and discriminator each used a 4-layer convolution structure, the convolution kernel size was 3×3, the stride was 1, and the batch normalization used the BatchNorm layer. The batch size for each training round was 64, the number of training iterations was 5000, and the learning rate was 0.0002. When performing multi-scale gradient extraction on the original image of the standard medicinal material, the Laplacian pyramid algorithm (Laplacian Pyramid), the image was downsampled four times, each time reducing the resolution by half. A Gaussian pyramid was constructed using the OpenCV pyrDown function, and a Laplacian pyramid was obtained by subtracting the Gaussian pyramids. The gradient map at each scale was extracted using the Sobel operator, with a convolution kernel size of 3×3 and a step size of 1. The extracted gradient feature map was stored in matrix form. The texture descriptor was calculated using the LBP algorithm (local binary pattern) with a radius of 2 and a neighborhood of 16 points. The calculated texture feature vector was 256-dimensional. When extracting shape contours from the texture feature vector, the Canny edge detection algorithm was used with thresholds set to 100 and 200. The extracted edges were contour-fitted using polygonal approximation with an approximation accuracy of 0.01. The resulting contour feature set is the coordinate set of each contour point. Principal component analysis is used for color matrix decomposition. The three-channel color matrix in the HSV color space is reduced in dimensionality, retaining 99% of the cumulative variance to obtain the color feature matrix. When performing spatial distribution statistics on the color feature matrix, the KDE (kernel density estimation) method is used to calculate the spatial probability density of the color distribution. A Gaussian kernel is used as the kernel function, and the bandwidth parameter is set to 0.5. The calculated spatial feature vector is 64-dimensional and saved in HDF5 format. Feature fusion uses a weighted fusion method, with the weights of the texture feature vector, contour feature set, and spatial feature vector set to 0.4, 0.3, and 0.3, respectively. The fused comprehensive feature is represented as a 512-dimensional vector.
[0111] Preferably, step S3 includes the following steps:
[0112] Step S31: performing feature deformation simulation on the original image of the standard medicinal material based on the comprehensive feature representation to generate simulated deformation features;
[0113] Step S32: predicting the growth direction of the simulated deformation feature to obtain a candidate growth trajectory;
[0114] Step S33: performing feature point detection on the candidate growth trajectory to obtain key trajectory points; performing projection matching on the key trajectory points and the medicinal material identification map to generate a matching medicinal material map;
[0115] Step S34: performing initial identification of the category of the original medicinal material image according to the matched medicinal material atlas to obtain an initial identification result.
[0116] In this embodiment, when simulating feature deformation of the original image of the standard medicinal material based on the comprehensive feature representation, the Thin Plate Spline deformation method is adopted, which is implemented through the interpolate module of the scipy library. The input comprehensive feature representation is a 512-dimensional vector, 64 of which are randomly selected as control points. The displacement range of each control point is between -5 and 5 pixels. The deformation weights between the control points are calculated by the RBF (Radial Basis Function) kernel function, and the output is the deformed image coordinate mapping matrix. The mapping matrix is applied to remap the coordinates of the original image of the standard medicinal material pixel by pixel. Bilinear interpolation is selected as the interpolation method. When predicting the growth direction of the simulated deformation feature, a long short-term memory network (LSTM) is used. The network structure is a 3-layer LSTM layer with 128 hidden units in each layer. The input is a time series representation of the simulated deformation feature with a time step of 10 and an input feature dimension of 64. The output is a displacement vector for each time step. The loss function uses the mean square error. The Adam optimizer is used for training with a learning rate of 0.001, a training batch size of 32, and 3000 training iterations. The output candidate growth trajectory is a three-dimensional coordinate sequence stored in JSON format, containing the X, Y, and Z values of each coordinate and the corresponding time step. When detecting feature points on the candidate growth trajectory, the SIFT algorithm (scale-invariant feature transform) is used. , using the cv.SIFT_create() function in the OpenCV library, setting the contrast threshold to 0.04 and the edge threshold to 10, the extracted key trajectory points are the coordinates, scale, direction and descriptor of each feature point, the descriptor length is 128 dimensions, the key trajectory point data is stored in XML format, the projection matching adopts the homography matrix estimation algorithm based on RANSAC (random sampling consistency), the key trajectory points are matched with the feature points in the medicinal material identification atlas by FLANN (fast nearest neighbor search), the index algorithm is set to KD tree, the number of trees is 5, the number of neighboring points returned for each query is 2, and the ratio is used. The value test threshold of 0.75 was used to screen matching points. When the original medicinal material image was initially identified according to the matching medicinal material atlas, a convolutional neural network (CNN) was used. The network structure consisted of 5 convolutional layers with convolution kernel sizes of 7×7, 5×5, 3×3, 3×3 and 1×1, respectively. Each convolution layer was followed by a ReLU activation function and a maximum pooling layer with a pooling kernel size of 2×2 and a stride of 2. The last two layers were fully connected layers with 512 and 256 neurons, respectively. The output layer was a Softmax layer with 50 classifications. The input medicinal material original image was normalized, and each pixel value was subtracted from the mean 128 and divided by the standard deviation 64.
[0117] Preferably, step S32 includes the following steps:
[0118] Perform regional segmentation on the simulated deformation features to obtain a growth block map;
[0119] Perform boundary positioning on the growth block map to generate a boundary trajectory set; perform direction chain encoding on the boundary trajectory set to generate a trajectory encoding sequence;
[0120] Calculate the curvature of the trajectory coding sequence to obtain curvature feature data;
[0121] Predicting the growth direction based on the curvature feature data to obtain a set of predicted trajectories;
[0122] Perform deformation feature recognition on the predicted trajectory set to generate predicted deformation features;
[0123] Trajectory screening is performed based on the simulated deformation characteristics and the predicted deformation characteristics to generate candidate growth trajectories.
[0124] In this embodiment, when performing regional segmentation on the simulated deformation features, a superpixel segmentation algorithm based on simple linear iterative clustering (SLIC) is adopted. The input simulated deformation features are three-channel images of 512×512 pixels. The slice function in the skimage library is used, the number of superpixels is set to 500, the compactness parameter of each superpixel region is 10, the number of iterations is 50, each pixel is clustered according to its color and spatial position, and the output growth block map is a two-dimensional matrix with each superpixel label marked, which is stored as a PNG format image file with a resolution of 300dpi. When performing boundary positioning on the growth block map, a boundary extraction method based on the Canny edge detection algorithm is adopted. The input growth block map is a two-dimensional matrix with each superpixel label marked. The block map was preprocessed with Gaussian blur using the cv2.GaussianBlur() function with a kernel size of 5×5 and a standard deviation of 1.0. The low threshold of the Canny algorithm was set to 50 and the high threshold was set to 150. The output boundary trajectory was a set of edge pixel coordinates stored in JSON format. Each trajectory contained at least 30 pixels. The boundary trajectory set was extracted using the cv2.findContours() function of the OpenCV library. The contour approximation accuracy was set to 0.02 times the contour length. When performing direction chain coding on the boundary trajectory set, the Freeman chain code coding method was used. The input boundary trajectory was scanned in a clockwise direction, and the chain code direction was divided into There are 8, corresponding to 0°, 45°, 90°, 135°, 180°, 225°, 270° and 315° respectively. NumPy array is used to store each direction code in sequence. The chain code sequence length is 512. When calculating the curvature of the trajectory code sequence, the differential geometry method is used to calculate the curvature of each trajectory point by three-point difference. The input trajectory code sequence length is 512. The numpy.gradient() function of the scipy library is used to perform first-order difference on the trajectory point coordinates. The difference interval is 1 pixel. The output curvature feature data is a one-dimensional array containing 512 curvature values. When predicting the growth direction based on the curvature feature data, the direction prediction based on Bayesian estimation is used. The model uses normalized curvature feature data with a mean of 0 and a standard deviation of 1. Training is performed using the GaussianNB() function in the scikit-learn library. The training data consists of 1000 curvature samples, each with a length of 512. The output prediction trajectory is a coordinate sequence of 200 trajectory points, each consisting of three coordinate values: X, Y, and Z. Deformation feature recognition is performed on the predicted trajectory using a convolutional neural network (ResNet-50). The input prediction trajectory is augmented with random rotations, translations, and scaling. The rotation angle is ±15°, the translation range is ±10 pixels, and the scaling factor is 0.9 to 1.1. The training data volume is 5000 trajectory images, each with a size of 224×224 pixels. The output predicted deformation features are 2048-dimensional feature vectors. When screening trajectories based on simulated deformation features and predicted deformation features, the cosine similarity calculation method is used. The length of the two input feature vectors is 2048. The cosine function of the scipy library is used for similarity calculation. The threshold is set to 0.8. Trajectories with similarity greater than 0.8 are retained. The selected candidate growth trajectories are stored in PLY format, which contains the coordinate sequence and corresponding feature vector of each trajectory.
[0125] Preferably, step S4 includes the following steps:
[0126] Step S41: constructing a graph index for the initial identification result to obtain index feature data; performing similarity measurement on the index feature data to generate similarity data;
[0127] Step S42: performing threshold screening on the medicinal material identification atlas based on the similarity data to obtain a collection of matching medicinal material images;
[0128] Step S43: performing image edge cutting on the standard medicinal material original image to obtain a medicinal material edge image; performing contrast enhancement processing on the medicinal material edge image to generate an enhanced edge image;
[0129] Step S44: performing edge contour tracing on the enhanced edge image to obtain the edge contour of the medicinal material; performing contour overlap comparison on the edge contour of the medicinal material based on the collection of matched medicinal material images to generate edge comparison parameters.
[0130] In this embodiment, when constructing the atlas index for the initial identification result, the KD-Tree (K-dimensional tree) indexing algorithm is used. The input initial identification result is a medicinal material atlas dataset containing 512-dimensional feature vectors. The scipy.spatial.KDTree function in the scipy library is used for index construction. The balance factor is set to 0.5, each index node contains no more than 10 feature vectors, the index tree depth is 16 layers, and the output index feature data is a KD-Tree index structure stored in a pickle file format. Each leaf node corresponds to a medicinal material atlas feature set. The L2 distance is used for node splitting during the index construction process. The splitting process is balanced by calculating the mean of the feature vector. The element value of each eigenvector is normalized to the interval [0,1]. The index construction time is about 30 seconds. When measuring the similarity of the index feature data, the cosine similarity measurement method is used. The input index feature data is read in batches, and 128 eigenvectors are read each time. The cosine function of the scipy library is used to calculate the similarity. The threshold is set to 0.85. The similarity calculation is implemented by matrix multiplication. The two input matrices are 128×512 and 512×512. The output similarity data is a 128×512 matrix, which is stored as a CSV file. Each matrix element represents the similarity between the index feature data and the initial identification result. The calculation process is accelerated by GPU, and the device is NVIDIA RTX 3080, the calculation takes about 5 seconds. When threshold screening is performed on the medicinal material identification atlas based on similarity data, the threshold selection method based on Otsu's method is adopted. The input similarity data is statistically analyzed by histogram, and the number of histogram intervals is 256. The cv2.threshold() function is used for automatic threshold segmentation, and the output threshold is 0.78. The collection of matched medicinal material images after screening is a JPEG format file containing 256 medicinal material images. The size of each image is 512×512 pixels, the resolution is 300dpi, and the image storage path is set to " / mnt / data / matching_herb_images / ". The file of each image is The item name is named by the medicinal material ID plus a timestamp. The image collection is indexed and managed using the DataFrame of the Pandas library. Each index record contains the image path, similarity value, and medicinal material category label. The Canny edge detection algorithm is used to perform image edge segmentation on the original standard medicinal material images. The input standard medicinal material original image is in TIFF format and has a size of 1024×1024 pixels. The cv2.Canny() function is used for edge detection. The low threshold is set to 100 and the high threshold is set to 200. The output medicinal material edge image is a binary image, with white representing edge pixels and black representing background. The image is stored in PNG format. The edge segmentation process is performed using cv2.The findContours() function extracts edge contours, and the minimum contour area is set to 500 pixels. When contrast enhancement is performed on the edge image of medicinal materials, the histogram equalization method is used. The input medicinal material edge image is an 8-bit grayscale image. The cv2.equalizeHist() function is used for histogram equalization. The output enhanced edge image is an edge image with a more uniform distribution of gray levels. The image size is 512×512 pixels. When tracing the edge contour of the enhanced edge image, the method based on Active Contour is used. The Snake algorithm of the active contour model is used. The input enhanced edge image is processed by Gaussian smoothing. The smoothing kernel size is 7×7 and the standard deviation is 1.5. The scipy.ndimage.gaussian_filter() function is used for smoothing. The number of iterations of the Snake algorithm is 500 times, the elasticity parameter is 0.5, and the smoothing parameter is 0.3. The output medicinal material edge contour is a NumPy array containing the coordinates of each contour point. The array length is 1024. When the edge contour of the medicinal material is overlapped and compared based on the matching medicinal material image collection, Hausd is used. Contour matching is performed using the Hausdorff distance. The input medicinal material edge contour is a NumPy array containing 1024 coordinate points. The number of contour points of each image in the matching medicinal material image collection is 512. The scipy.spatial.distance.directed_hausdorff() function is used for distance calculation. The output edge comparison parameter is the Hausdorff distance matrix between each pair of contours. The matrix size is 256×1 and is stored as an HDF5 file. Each distance value represents the degree of edge matching between the corresponding matching medicinal material image and the original standard medicinal material image.
[0131] Preferably, step S44 includes the following steps:
[0132] Extracting curvature features from the enhanced edge image to generate curvature distribution data; reconstructing the curvature distribution data through contour interpolation to obtain contour interpolation data;
[0133] Perform vectorization conversion processing on the contour interpolation data to obtain contour vector data; perform feature point positioning processing on the contour vector data to generate contour key point data;
[0134] Perform spline curve fitting on the key point data of the contour to obtain the edge contour of the medicinal material;
[0135] Perform contour registration on the edge contour of the medicinal material and the collection of matching medicinal material images to obtain registration contour data; perform overlap evaluation processing on the registration contour data to generate overlap evaluation data;
[0136] The overlap difference is quantified based on the overlap evaluation data to obtain the edge contrast parameter.
[0137] In this embodiment, when the curvature feature is extracted from the enhanced edge image, the second-order derivative calculation method is used. The input enhanced edge image is a grayscale image of 512×512 pixels. The cv2.Laplacian() function in the OpenCV library is used to extract the second-order edge information of the image. The convolution kernel size is 3×3, and the output curvature image is a 32-bit floating-point matrix. Each pixel value represents the local curvature of the point. The x-axis and y-axis gradients of the image are calculated by the gradient function of the NumPy library. The curvature calculation adopts the discrete approximation of κ=(x′y″-y′x″) / (x′2+y′2)^(3 / 2), and finally generates Curvature distribution data, when performing contour interpolation reconstruction on curvature distribution data, the bicubic interpolation algorithm is used. The input curvature distribution data is processed by the scipy.ndimage.map_coordinates() function. During the interpolation process, a 64×64 grid is used to sample the 512×512 original curvature data. The output contour interpolation data is a 256×256 two-dimensional array. Each value represents the curvature value after interpolation. The interpolation weight is calculated by the weighted average of 16 neighborhood points. The cyclic boundary condition is used in the interpolation reconstruction process. When performing vector conversion processing on the contour interpolation data, the Marching algorithm is used. Squares algorithm, the input contour interpolation data uses scipy.ndimage.measurements.find_objects() function to detect connected domains, the minimum area threshold of the connected domain is set to 50 pixels, and the skimage.measure.find_contours() function is used to extract contour lines, and the contour level is set to 0.5. The output contour vector data is a GeoJSON file containing 1024 coordinate points. The x and y values of each coordinate point are retained to 4 decimal places, and the coordinate data unit is pixel. When the contour vector data is processed for feature point positioning, the Harris corner detection algorithm is used. The input contour vector data is processed by shap The LineString object from the scipy library was geometrically transformed and the corner response values were calculated using the cv2.cornerHarris() function. The window size was set to 3×3, the aperture parameter of the Sobel operator was set to 3, and the Harris corner detection parameter k was set to 0.04. The output contour keypoint data was a CSV file containing 128 keypoints. The coordinates and corner response value of each keypoint were recorded. Points with a response value greater than 0.01 were considered valid feature points. When fitting the contour keypoint data with a spline curve, a cubic B-spline fitting method was used. The input contour keypoint data was parameterized using the scipy.interpolate.splprep() function, and the fitting parameter s was set to 0.5, k is set to 3, the output medicinal material edge contour is parameterized B-spline curve data, the number of curve segments is 64, and the iterative closest point algorithm (ICP) is used to perform contour registration on the medicinal material edge contour and the matching medicinal material image collection. The input medicinal material edge contour is an array of 1024 coordinate points, and the number of contour points of each image in the matching medicinal material image collection is 512. The o3d.registration.registration_icp() function of the Open3D library is used for rigid transformation registration. The maximum number of iterations is 200, the convergence threshold is 1e-6, and the output registration contour data is the coordinate array after rotation and translation transformation. When performing overlap evaluation on the registration contour data, the Jaccard similarity coefficient calculation method is used. The input registration contour data and the contour data of the matching medicinal material image collection are processed by cv2.fi. The llPoly() function generates a binary mask image with a size of 512×512. White pixels represent contour areas. The intersection and union are calculated using the bitwise_and() and bitwise_or() functions of the NumPy library. The output overlap assessment data is a floating-point array. When quantifying the overlap difference of the overlap assessment data, a pixel difference statistical method is used. The input overlap assessment data is passed through the count_nonzero() function of the NumPy library to calculate the number of difference pixels. The structural similarity index is calculated using the skimage.metrics.structural_similarity() function. The difference quantification parameters include pixel difference rate, edge offset distance, and structural similarity. The output edge contrast parameters are a TXT file containing three quantitative indicators, each with three decimal places.
[0138] Preferably, step S5 includes the following steps:
[0139] Step S51: performing edge contour separation on the original image of the medicinal material according to the edge contrast parameter to obtain a separated edge contour; performing three-dimensional reconstruction on the original image of the medicinal material to generate a three-dimensional medicinal material model;
[0140] Step S52: performing de-edge positioning on the three-dimensional medicinal material model based on the separated edge contour to obtain a center point area; performing central contour cutting on the three-dimensional medicinal material model according to the center point area to generate a reconstructed center contour;
[0141] Step S53: extracting texture features from the reconstructed center contour to obtain texture feature parameters;
[0142] Step S54: performing detail texture retrieval on the texture feature parameters based on the matching image collection to generate a final matching result.
[0143] In this embodiment, when the edge contour of the original image of the medicinal material is separated according to the edge contrast parameter, the Canny edge detection algorithm is used. The input original image of the medicinal material is an RGB image of 1024×1024 pixels. First, the cv2.cvtColor() function is used to convert it into a grayscale image, and the cv2.GaussianBlur() function is used for Gaussian smoothing. The convolution kernel size is set to 5×5 and the standard deviation is set to 1.4. Then, the cv2.Canny() function is called for edge detection, the low threshold is set to 50, the high threshold is set to 150, and the extracted edge image is subjected to contour extraction by the cv2.findContours() function. Take, the contour type is set to cv2.RETR_EXTERNAL, the approximation method is set to cv2.CHAIN_APPROX_SIMPLE, the edge offset distance recorded in the edge contrast parameter is used, and the offset area is marked on the original image through the cv2.drawContours() function. The offset distance is 5 pixels. When the original image of the medicinal material is reconstructed into 3D, the multi-view stereo vision (MVS, Multi-ViewStereo) method is used. The input original image is reconstructed into 3D through the COLMAP tool. First, SIFT (Scale-Invariant Feature Transform, The feature transform detector was used to extract key points. The number of feature points was set to 20,000. A robust estimation algorithm based on RANSAC was used for feature matching to generate camera pose data between views. A sparse point cloud was generated using the sparse reconstruction module. The point cloud density was set to 500 points per square millimeter. A continuous three-dimensional surface was generated using the surface reconstruction module based on the Poisson surface reconstruction algorithm. The octree depth was set to 10, and the number of smoothing iterations was set to 3. When de-marginalizing the three-dimensional medicinal material model based on the separation edge contour, the MeshLab tool was used. The input separation edge contour and the three-dimensional medicinal material model were loaded into MeshLab and aligned using the ICP algorithm. The maximum number of iterations was set to 100, and the error threshold was set to 1e-6. After alignment, the edge area defined by the separation edge contour was selected using the "Selection" function of MeshLab. The selected edge area was deleted using the "Delete Selected Faces" tool to obtain the de-marginalized three-dimensional model. The 3D medicinal material model was then de-marginalized using cv2.The moments() function calculates the center point of the de-marginalized model. The output center point area is a JSON file containing x, y, and z coordinates. When the three-dimensional medicinal material model is cut according to the center point area, the Boolean operation tool in Blender software is used. The input three-dimensional medicinal material model is plane-cut using Blender's "Knife Project" tool. The cutting plane is defined by the center point area. The reconstructed center contour after cutting is output as an STL file. When the texture feature of the reconstructed center contour is extracted, the gray-level co-occurrence matrix (GLCM) is used. Matrix) method, the input reconstructed center contour is converted into a grayscale image by the cv2.cvtColor() function, the image size is 256×256 pixels, the grayscale co-occurrence matrix is calculated using the skimage.feature.graycomatrix() function, the grayscale level is set to 64, the calculation directions are 0°, 45°, 90° and 135°, and the distance is set to 1 pixel. Then the skimage.feature.graycoprops() function is called to extract the four texture feature parameters of contrast, correlation, energy and homogeneity. The calculation result of each parameter is retained to three decimal places. When performing detailed texture retrieval based on texture feature parameters from a collection of matching images, the k-nearest neighbor (k-NN) algorithm is used. The input texture feature parameters are compared with the texture feature database of the matching image collection. Training is performed using the KNeighborsClassifier() function in the scikit-learn library. During training, the k value is set to 5, the distance metric is Euclidean distance, and the ratio of the training set to the test set is 8:2. After training, the texture feature parameters are input for classification. The final matching results are output as the matching herbal name, matching probability, and matching image index.
[0144] Preferably, step S54 includes the following steps:
[0145] Performing scale normalization on texture feature parameters to obtain standard texture features; performing directionality analysis on standard texture features to generate texture direction vectors;
[0146] Partition encoding is performed on standard texture features based on texture direction vectors to obtain regional texture encoding;
[0147] Perform coding feature matching on the regional texture coding according to the matching image collection to generate a matching score matrix; sort the matching image collection based on the matching score matrix to obtain a sorted matching image;
[0148] The top search is performed based on the ranked matching images to generate the final matching results.
[0149] In this embodiment, when performing scale normalization processing on texture feature parameters, bilinear interpolation is used to rescale the input high-dimensional texture feature matrix, mapping all texture feature vectors to a uniform numerical range of 0 to 1. In the specific implementation process, the maximum and minimum values of the texture feature matrix are selected as normalization boundaries. The normalization operation is completed by traversing each pixel point of the texture matrix and applying a linear transformation formula. A standard size of 128×128 is selected as the spatial scale of the output texture matrix. The bilinear interpolation algorithm is called through the resize function of the OpenCV library to complete matrix resampling. This ensures that the output standard texture feature matrix has a uniform scale specification while maintaining texture detail features. The output data is in the form of a 128×128 floating-point matrix with a numerical range of 0 to 1. When performing directional analysis on the standard texture features, a Gabor filter is used to extract the directional characteristics of the texture. Eight Gabor kernels with different directions are selected for convolution operation. The parameters of each Gabor kernel are set to have a frequency of 0.2, a directional spacing of π / 8, a standard deviation of 2.0, and an aspect ratio of 0.5. The convolution response is calculated pixel by pixel to obtain the response value of each pixel in 8 directions. The maximum response principle is used to select the direction with the largest response value for each pixel as its main direction. The final texture direction vector is an integer matrix of 128×128 dimensions, where the direction encoding value of each pixel is an integer between 0 and 7, corresponding to 8 different directions. When partitioning the standard texture features based on the texture direction vector, the 128×128 standard texture feature matrix is first divided into 16 8×8 sub-regions, and the texture direction vectors in each sub-region are statistically analyzed. By counting the number of pixels in 8 directions in each sub-region, an 8-dimensional histogram is formed as the direction coding feature of the sub-region. After the value of each histogram is accumulated and calculated, L2 normalization is performed to eliminate the feature amplitude difference. Finally, 16 8-dimensional regional texture coding vectors are obtained, and All regional vectors are merged into a 128-dimensional global regional texture code. When the regional texture code is matched according to the matching image collection, the cosine similarity measurement method (CosineSimilarity) is used to calculate the similarity between the input regional texture code and the corresponding regional texture code of each image in the matching image collection. The regional texture code is extracted by traversing each image in the matching image collection image by image. The dot function of the NumPy library is used to calculate the dot product of the input code and each image code, and then divided by the L2 norm product of the two to obtain the cosine similarity. Finally, a matching score matrix containing the similarity values of all matching images is generated. The rows of the matrix represent the input images, the columns represent the matching images, and the matrix elements represent the similarity scores of the corresponding image pairs. When sorting the matching image collection based on the matching score matrix, the quick sorting algorithm (Quick sorting algorithm) is used. Sort) sorts the similarity score vectors corresponding to each input image in descending order. Use the Python standard library sorted function with the reverse = True parameter to sort the match score vectors, generating a sorted matching image index array. This index array is then used to rearrange the matching image collection so that the matching images corresponding to each input image in the collection are arranged in descending order of similarity score. When performing a top-rank search based on the sorted matching images, the top-ranked matching image for each input image in the sorted matching images is extracted as the final matching result. By iterating over the matching image collection index array and selecting the first element of each index array, Python's index method is called to obtain the corresponding image's position in the collection. This image is then extracted to generate the final matching result set. The final output matching result set is an image list containing the best matching images for all input medicinal material images.
[0150] Preferably, step S6 includes the following steps:
[0151] Step S61: extracting feature points from the final matching result to obtain feature point data; constructing line rules on the feature point data to obtain line rule data;
[0152] Step S62: Mark the final matching result and the original image of the medicinal material based on the connection rule data to generate corresponding result data;
[0153] Step S63: performing model parameterization processing based on the corresponding result data to obtain identification model parameters; iteratively optimizing the identification model parameters to generate iterative model data;
[0154] Step S64: Integrate the iterative model data to generate an intelligent medicinal material identification model.
[0155] In this embodiment, when the final matching result is subjected to feature point extraction processing, the SIFT algorithm (Scale-Invariant Feature Transform) is used to extract key feature points in the image, the cv2.SIFT_create() method in the OpenCV library is used to generate a SIFT feature extractor, the final matching result is input into the detectAndCompute function, the key point positions of the image at different scales are calculated by Gaussian Pyramid and Difference of Gaussian (DoG), the key points with weak edge response are screened out using the principal curvature ratio, and finally feature point data including key point coordinates, scales, directions and local feature descriptors are obtained, and then the Delaunay triangulation algorithm (DelaunayTriangulation) is used to construct the line rules for the feature point data, and the scipy.spatial.Delaunay function is used to generate a triangulation structure with no sharp angles and uniform side lengths based on the feature point coordinates, and the line rules are constructed by traversing the triangle edges to generate line rule data including edge coordinates, lengths and connection relationships, and the final matching result is compared based on the line rule data. When marking lines on the original image of medicinal materials, the cv2.line function of OpenCV is used for annotation processing. First, the coordinates of the feature points corresponding to each edge in the line rule data are extracted, and the cv2.line function is called to draw the corresponding line segment on the image. The line segment color is set to red (RGB value is 255,0,0), and the line width is set to 2 pixels. At the same time, the cv2.circle function is called at the endpoint position of each line to draw a circular mark with a diameter of 5 pixels for endpoint identification. The final result data is an image matrix containing all the line and feature point marks. When the model parameterization is performed based on the corresponding result data, the principal component analysis algorithm (PCA, Principal Component Analysis) performs dimensionality reduction on high-dimensional features such as line length, line direction, and feature point distribution. The sklearn.decomposition.PCA function is called with a set number of principal components of 64. The principal component matrix is obtained through covariance matrix decomposition. The first 64 principal components are extracted as model parameters, forming a discriminant model parameter vector containing 64 values, including the mean line length, direction variance, and feature point density. The discriminant model parameters are then iteratively optimized using the Adam optimization algorithm (Adaptive Moment Estimation) based on the torch.optim.Adam function for parameter training. The model parameters are then fed into the PyTorch neural network training framework, with a learning rate set to 0.001, with a batch size of 32 and 1000 training iterations per round. Backpropagation is used to calculate gradients and update parameters, generating iterative model data containing optimized parameter values and training weights. When integrating the iterative model data, the Random Forest algorithm, an ensemble learning method, is used to integrate multiple trained models. The sklearn.ensemble.RandomForestClassifier function is called, and the number of base classifiers is set to 50. Each sub-model in the iterative model data is trained and predicted, and the prediction results of the 50 sub-models are counted. The final output label is determined by majority voting. The resulting intelligent medicinal material identification model includes the trained random forest model and the final ensemble classifier.
[0156] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.
[0157] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing a medicinal material identification model based on image technology, characterized in that: The following steps are involved: Step S1: Acquire medicinal material sample images; Extract features from medicinal material sample images to obtain sample feature parameters; Construct a knowledge graph of medicinal material sample images based on sample feature parameters to generate a medicinal material identification graph; Step S2: Collecting original images of medicinal materials; reshaping the original images of medicinal materials with texture features to obtain standard medicinal materials original images; performing multi-dimensional feature extraction and fusion modeling on the standard medicinal materials original images to generate comprehensive feature representation; Step S3: Based on the comprehensive feature representation, a growth extension simulation is performed on the original image of the standard medicinal material to generate a candidate growth trajectory; based on the candidate growth trajectory and the medicinal material identification map, an initial category identification is performed on the original image of the medicinal material to obtain an initial identification result; Step S4: matching the medicinal material identification atlas based on the initial identification results to obtain a collection of matching medicinal material images; performing edge comparison on the collection of matching medicinal material images and the original image of the standard medicinal material to obtain edge comparison parameters; Step S5: reconstructing the central contour of the original image of the medicinal material according to the edge contrast parameter to generate a reconstructed central contour; performing detail texture retrieval on the reconstructed central contour based on the matching image collection to generate a final matching result; Step S6: Mark the final matching result and the original image of the medicinal material with a line to generate corresponding result data; use the corresponding result data to perform model training to generate an intelligent medicinal material identification model.
2. The method for constructing a medicinal material identification model based on image technology according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: Acquire a medicinal material sample image; perform structure morphology spectrum mapping on the medicinal material sample image to obtain morphology mapping parameters; Step S12: extracting features from the medicinal material sample image according to the morphological mapping parameters to obtain sample feature parameters; Step S13: performing feature association analysis on the sample feature parameters to obtain a feature association network; performing hierarchical aggregation mapping on the feature association network to generate a category hierarchy matrix; Step S14: reconstructing the feature distribution of the category hierarchy matrix to obtain classification feature data; mapping the classification feature data to an identification map to generate a medicinal material identification map.
3. The method for constructing a medicinal material identification model based on image technology according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: collecting original images of medicinal materials; performing spectral tomography on the original images of medicinal materials to obtain spectral tomography data; Step S22: performing superpixel segmentation on the spectral tomography data to generate medicinal material segmentation features; performing texture migration and reshaping on the medicinal material segmentation features to obtain a standard medicinal material original image; Step S23: performing multi-scale gradient extraction on the original image of the standard medicinal material to obtain a gradient feature map; performing texture descriptor calculation on the gradient feature map to generate a texture feature vector; Step S24: performing shape contour extraction on the texture feature vector to obtain a contour feature set; performing color matrix decomposition on the contour feature set to generate a color feature matrix; Step S25: Perform spatial distribution statistics on the color feature matrix to obtain a spatial feature vector; perform feature fusion on the spatial feature vector to generate a comprehensive feature representation.
4. The method for constructing a medicinal material identification model based on image technology according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: performing feature deformation simulation on the original image of the standard medicinal material based on the comprehensive feature representation to generate simulated deformation features; Step S32: predicting the growth direction of the simulated deformation feature to obtain a candidate growth trajectory; Step S33: performing feature point detection on the candidate growth trajectory to obtain key trajectory points; performing projection matching on the key trajectory points and the medicinal material identification map to generate a matching medicinal material map; Step S34: performing initial identification of the category of the original medicinal material image according to the matched medicinal material atlas to obtain an initial identification result.
5. The method for constructing a medicinal material identification model based on image technology according to claim 4, characterized in that: Step S32 includes the following steps: Perform regional segmentation on the simulated deformation features to obtain a growth block map; Perform boundary positioning on the growth block map to generate a boundary trajectory set; perform direction chain encoding on the boundary trajectory set to generate a trajectory encoding sequence; Calculate the curvature of the trajectory coding sequence to obtain curvature feature data; Predicting the growth direction based on the curvature feature data to obtain a set of predicted trajectories; Perform deformation feature recognition on the predicted trajectory set to generate predicted deformation features; Trajectory screening is performed based on the simulated deformation characteristics and the predicted deformation characteristics to generate candidate growth trajectories.
6. The method for constructing a medicinal material identification model based on image technology according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: constructing a graph index for the initial identification result to obtain index feature data; performing similarity measurement on the index feature data to generate similarity data; Step S42: performing threshold screening on the medicinal material identification atlas based on the similarity data to obtain a collection of matching medicinal material images; Step S43: performing image edge cutting on the standard medicinal material original image to obtain a medicinal material edge image; performing contrast enhancement processing on the medicinal material edge image to generate an enhanced edge image; Step S44: performing edge contour tracing on the enhanced edge image to obtain the edge contour of the medicinal material; performing contour overlap comparison on the edge contour of the medicinal material based on the collection of matched medicinal material images to generate edge comparison parameters.
7. The method for constructing a medicinal material identification model based on image technology according to claim 6, characterized in that: Step S44 includes the following steps: Extracting curvature features from the enhanced edge image to generate curvature distribution data; reconstructing the curvature distribution data through contour interpolation to obtain contour interpolation data; Perform vectorization conversion processing on the contour interpolation data to obtain contour vector data; perform feature point positioning processing on the contour vector data to generate contour key point data; Perform spline curve fitting on the key point data of the contour to obtain the edge contour of the medicinal material; Perform contour registration on the edge contour of the medicinal material and the collection of matching medicinal material images to obtain registration contour data; perform overlap evaluation processing on the registration contour data to generate overlap evaluation data; The overlap difference is quantified based on the overlap evaluation data to obtain the edge contrast parameter.
8. The method for constructing a medicinal material identification model based on image technology according to claim 1, characterized in that: Step S5 includes the following steps: Step S51: performing edge contour separation on the original image of the medicinal material according to the edge contrast parameter to obtain a separated edge contour; performing three-dimensional reconstruction on the original image of the medicinal material to generate a three-dimensional medicinal material model; Step S52: performing de-edge positioning on the three-dimensional medicinal material model based on the separated edge contour to obtain a center point area; performing central contour cutting on the three-dimensional medicinal material model according to the center point area to generate a reconstructed center contour; Step S53: extracting texture features from the reconstructed center contour to obtain texture feature parameters; Step S54: performing detail texture retrieval on the texture feature parameters based on the matching image collection to generate a final matching result.
9. The method for constructing a medicinal material identification model based on image technology according to claim 8, characterized in that: Step S54 includes the following steps: Performing scale normalization on texture feature parameters to obtain standard texture features; performing directionality analysis on standard texture features to generate texture direction vectors; Partition encoding is performed on standard texture features based on texture direction vectors to obtain regional texture encoding; Perform coding feature matching on the regional texture coding according to the matching image collection to generate a matching score matrix; sort the matching image collection based on the matching score matrix to obtain a sorted matching image; The top search is performed based on the ranked matching images to generate the final matching results.
10. The method for constructing a medicinal material identification model based on image technology according to claim 1, characterized in that: Step S6 includes the following steps: Step S61: extracting feature points from the final matching result to obtain feature point data; constructing line rules on the feature point data to obtain line rule data; Step S62: Mark the final matching result and the original image of the medicinal material based on the connection rule data to generate corresponding result data; Step S63: performing model parameterization processing based on the corresponding result data to obtain identification model parameters; iteratively optimizing the identification model parameters to generate iterative model data; Step S64: Integrate the iterative model data to generate an intelligent medicinal material identification model.
Citation Information
Patent Citations
Traditional Chinese medicinal material quality control map equipment, traditional Chinese medicinal material quality control map construction method, electronic equipment and computer program product
CN118609143A
Drug repositioning method and system fusing multi-source knowledge graph
WO2024138803A1