Structured feature selection method and system for authentic medicinal materials based on Raman peak intensity information
By generating two-dimensional chemical composition distribution data and Raman peak intensity difference maps of medicinal material slices, combined with spatial topology analysis and feature screening, the problems of insufficient feature extraction accuracy and model generalization ability in the identification of authentic medicinal materials are solved, and high-precision medicinal material identification is achieved.
Patent Information
- Application Number
- CN202510812514.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing technology for identifying authentic medicinal materials has problems such as insufficient accuracy in extracting the spatial distribution features of chemical components, light scattering interference affecting reconstruction accuracy, and insufficient model generalization ability, which leads to limited identification accuracy and robustness.
A scanning device is used to generate two-dimensional distribution data of chemical component concentration gradients, and a dual-wavelength laser source is used to generate a spatial response comparison map of Raman peak intensity differences. Combined with spatial topological analysis, the target concentration area and its boundary intersection points are identified, and a spatial correlation model is constructed to screen feature combinations that match the distribution of authentic medicinal materials, and then input into the discriminant model for identification.
It improves the ability to capture the spatial heterogeneity of chemical composition, enhances the feature resolution ability, optimizes the input quality of the discrimination model, and improves the reliability and interpretability of the identification results.
Smart Images

Figure CN120352410B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of structured feature selection, and in particular to a method and system for selecting structured features of authentic medicinal materials based on Raman peak intensity information. Background Art
[0002] In the area of quality control for traditional Chinese medicine (TCM), authentic medicinal materials, due to the ecological and environmental differences in their specific production areas, develop unique spatial distribution patterns of active ingredients (e.g., gradient distribution, regional enrichment, etc.). While traditional physical and chemical testing can determine component content, it cannot characterize the spatial heterogeneity of components, resulting in insufficient identification accuracy. Currently, there is an urgent need for a technology that can non-destructively and rapidly acquire the three-dimensional spatial distribution characteristics of multiple components within and on the surface of medicinal materials, and to establish a quantifiable spatial pattern discrimination model to support the accurate identification and traceability of authentic medicinal materials.
[0003] The current mainstream approach uses near-infrared hyperspectral imaging (NIR-HSI) combined with chemometrics. This technology collects hyperspectral cube data from medicinal material samples within a specific near-infrared shortwave spectral range. Using partial least squares discriminant analysis (PLS-DA) or support vector machines (SVM), it extracts characteristic bands and constructs heat maps of chemical composition distribution in different regions. By comparing the spatial distribution differences between samples from authentic and non-authentic production areas (e.g., the annular enrichment pattern of flavonoids in rhizome cross-sections), identification is achieved through spatial pattern matching.
[0004] However, this technology has three limitations: First, near-infrared spectroscopy has poor resolution for structurally similar components (such as isomers), limiting the accuracy of spatial distribution feature extraction. Second, high humidity on the surface of medicinal materials easily generates light scattering interference, affecting the accuracy of three-dimensional spatial reconstruction. Third, existing models rely on a large number of labeled samples to construct a "standard spatial atlas library," but the high cost of obtaining authentic medicinal samples and the existence of regional protection barriers lead to insufficient model generalization. These shortcomings restrict the practical application value of this technology in the identification of rare medicinal materials. Summary of the Invention
[0005] The present application provides a method and system for selecting structured features of authentic medicinal materials based on Raman peak intensity information, in order to solve the problem of poor structured feature selection effect in the prior art.
[0006] In a first aspect, the present application provides a method for selecting structured features of authentic medicinal materials based on Raman peak intensity information, comprising:
[0007] The surface area of the medicinal material slice is spatially sampled by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components;
[0008] Alternately projecting excitation light generated by a dual-wavelength laser source onto the same sampling area, generating a spatial response comparison map based on the difference in Raman peak intensity obtained at the two wavelengths, and forming a three-dimensional feature stack based on the spatial response comparison map and the two-dimensional distribution data;
[0009] performing spatial topological analysis on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections;
[0010] Constructing a spatial correlation model between the target concentration areas, and analyzing the distribution density and directional vector relationship of the boundary intersection points in each target concentration area through the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of the medicinal material components;
[0011] A structured feature selection operation is performed to retain, in the graph structure data based on constraints, feature combinations that match the prior knowledge of the spatial distribution of authentic medicinal materials, and the screened feature combinations are input as structured features into an authenticity discrimination model, so that the discrimination model outputs an identification result based on the joint weights of the feature combinations in the spatial dimension and the spectral dimension.
[0012] Optionally, retaining a feature combination matching the prior knowledge of spatial distribution of authentic medicinal materials in the graph structure data through constraint conditions includes:
[0013] Extracting characteristic distribution patterns of the graph structure data from a historical authentic medicinal material sample library, and constructing a constraint rule set based on the characteristic distribution patterns;
[0014] Traversing all feature combinations in the graph structure data, screening out feature combinations that meet preset conditions as a primary screening set; wherein the feature combination that meets the preset conditions refers to a feature combination with multiple constraint tags;
[0015] For each feature combination in the primary screening set, the product value of the regional area dispersion and the intersection density is calculated, and the feature combinations whose product values are ranked within a predefined ratio range are retained as the screened feature combinations.
[0016] Optionally, the filtered feature combination is input as a structured feature into an authenticity discrimination model, so that the discrimination model outputs an identification result based on the joint weight of the feature combination in the spatial dimension and the spectral dimension, including:
[0017] The filtered feature combination is input into the authenticity discrimination model as a discrete unit with spatial coordinate labels;
[0018] A dual-channel calculation path is established within the authenticity discrimination model. The first channel maps the two-dimensional coordinate values of each discrete unit to a virtual plane divided by an orthogonal grid. The standard deviation of the spectral response amplitude difference of all discrete units within each orthogonal grid is calculated to form a spatial fluctuation coefficient per grid unit. The second channel extracts the straight-line distance between each discrete unit and the intersection point of the target concentration region to which it belongs, and combines this with the directional correlation strength to generate an angle correction factor.
[0019] generating a joint weight value according to the spatial fluctuation coefficient and the angle correction factor in an asymmetric superposition manner;
[0020] Construct a computing network with a ring connection structure, inject the joint weight value of each discrete unit into its corresponding ring node, and achieve spatial transmission of the weight value through the impedance matching relationship between adjacent ring nodes. The output end of the computing network can receive the joint weight value output by each port, and the sum of the joint weight values of each port is compared with a preset azimuth reference value to generate a phase comparison result.
[0021] An azimuth offset map is generated based on the phase comparison results, and the spatial distribution morphology of the offset extreme value areas and the target concentration areas that appear continuously in the azimuth offset map is overlapped and analyzed. When the boundary of the offset extreme value area forms a specific angle relationship with the geometric center of at least three target concentration areas, the confirmation mechanism of the authenticity discrimination model is triggered, and the identification result is output.
[0022] Optionally, analyzing the distribution density and directional vector relationship of boundary intersections in each target concentration region by the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of medicinal material components includes:
[0023] An initial topological network is established using the spatial correlation model, with the geometric centers of each target concentration area as nodes and boundary intersections as connection hubs;
[0024] In the initial topological network, the number of intersections within a preset radius around each boundary intersection is calculated as a density distribution parameter, and the direction angle of the line connecting each boundary intersection and the center of gravity of the target concentration area to which it belongs is recorded as a direction relationship parameter;
[0025] Vector synthesis is performed on boundary intersections with overlapping density distribution parameters to generate weighted edges representing the interaction strength between regions;
[0026] The density distribution parameter and the direction relationship parameter are used as node attributes, and the vector synthesis direction of the weighted edge is used as an edge attribute to construct graph structure data containing multi-level spatial association features.
[0027] Optionally, the surface area of the medicinal material slice is spatially sampled by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components, including:
[0028] The surface area is irradiated point by point along a preset grid path by a scanning device, and each grid node corresponds to a sampling position;
[0029] At each sampling position, monochromatic excitation light is vertically incident on the surface of the medicinal material, and the Raman scattering signal reflected at the sampling position is synchronously received, and the intensity value of the preset wavelength range in the Raman scattering signal is extracted as the feature value of the grid node;
[0030] According to the spatial coordinate arrangement of the grid nodes, the feature quantities of all grid nodes are filled into a two-dimensional matrix in row-column order, and the difference between the feature quantities of adjacent grid nodes is calculated. If the difference exceeds a set threshold, an interpolation node is inserted in the corresponding row-column gap to supplement the feature quantities of the interpolation nodes to form a continuously distributed two-dimensional data plane;
[0031] The characteristic quantity of each sampling position in the two-dimensional data plane is mapped into a grayscale value to generate two-dimensional distribution data with spatial coordinates as horizontal and vertical axes and grayscale values representing chemical component concentrations.
[0032] Optionally, the excitation light generated by the dual-wavelength laser source is alternately projected onto the same sampling area, a spatial response comparison graph is generated based on the difference in Raman peak intensities obtained at the two wavelengths, and a three-dimensional feature stack is formed based on the spatial response comparison graph and the two-dimensional distribution data, including:
[0033] The output wavelength of the dual-wavelength laser source is alternately switched by a synchronous controller, and a complete spectrum acquisition is performed on the same sampling area at each wavelength;
[0034] The characteristic peak intensity value of the Raman signal collected at the first wavelength is marked as a first response value, and the characteristic peak intensity value of the Raman signal collected at the second wavelength is marked as a second response value;
[0035] Calculating the difference between the first response value and the second response value for each spatial sampling point, and mapping the difference to the same spatial coordinate system as the two-dimensional distribution data to generate a spatial response comparison graph;
[0036] The spatial response comparison map is used as an additional signal layer and is stacked and combined with the two-dimensional distribution data to form a three-dimensional feature stack with three layers of signal intensity.
[0037] Optionally, performing spatial topological analysis on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections includes:
[0038] A segmentation operation is performed on each signal layer in the three-dimensional feature stack, and continuous areas with signal intensity values higher than a preset threshold are extracted as candidate concentration areas;
[0039] Performing spatial position comparison on candidate concentration regions in multiple signal layers, and retaining candidate regions whose spatial overlap in at least two signal layers exceeds a set ratio as target concentration regions;
[0040] Detect the boundary curves of each target concentration area and calculate the minimum spacing and curvature change rate between the boundary curves corresponding to adjacent target concentration areas;
[0041] When the minimum spacing between boundary curves corresponding to adjacent target concentration regions is less than a preset ratio of the average width of the regions and the curvature change rate exceeds a preset mutation threshold, it is determined that there is an intersection between the boundaries corresponding to adjacent target concentration regions.
[0042] In a second aspect, the present application provides a structured feature selection system for authentic medicinal materials based on Raman peak intensity information, comprising:
[0043] The first generation module is used to perform spatial sampling on the surface area of the medicinal material slice by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components;
[0044] A second generation module is configured to alternately project excitation light generated by a dual-wavelength laser source onto the same sampling area, generate a spatial response comparison map based on the difference in Raman peak intensities obtained at the two wavelengths, and form a three-dimensional feature stack based on the spatial response comparison map and the two-dimensional distribution data;
[0045] an identification module for performing spatial topological analysis on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections;
[0046] A third generation module is used to construct a spatial correlation model between the target concentration areas, and analyze the distribution density and direction vector relationship of the boundary intersection points in each target concentration area through the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of the medicinal material components;
[0047] The output module is used to perform a structured feature selection operation, retaining feature combinations that match the prior knowledge of the spatial distribution of authentic medicinal materials in the graph structure data through constraints, and inputting the screened feature combinations as structured features into the authenticity discrimination model, so as to output the identification results through the discrimination model based on the joint weights of the feature combinations in the spatial dimension and the spectral dimension.
[0048] In a third aspect, the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a method for selecting structured features of authentic medicinal materials based on Raman peak intensity information as described in the first aspect above.
[0049] In a fourth aspect, the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements a method for selecting structured features of authentic medicinal materials based on Raman peak intensity information as described in the first aspect.
[0050] The present application performs spatial sampling of the surface area of a medicinal material slice by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components; the excitation light generated by a dual-wavelength laser source is alternately projected onto the same sampling area, a spatial response comparison map is generated based on the difference in Raman peak intensity obtained at the two wavelengths, and a three-dimensional feature stack is formed based on the spatial response comparison map and the two-dimensional distribution data; spatial topology analysis is performed on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections; a spatial correlation model is constructed between the target concentration regions, and the distribution density and direction vector relationship of the boundary intersections in each target concentration region are analyzed through the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of the medicinal material components; a structured feature selection operation is performed, and through constraints, a feature combination that matches the prior knowledge of the spatial distribution of authentic medicinal materials is retained in the graph structure data, and the screened feature combination is input as a structured feature into the authenticity discrimination model, so that the discrimination model outputs an identification result based on the joint weight of the feature combination in the spatial dimension and the spectral dimension.
[0051] This application has the following beneficial effects:
[0052] A scanning device performs high-precision spatial sampling of the surface of medicinal material slices, acquiring two-dimensional distribution data of chemical component concentration gradients. This provides multi-dimensional foundational data for subsequent analysis and ensures that spatial heterogeneity of chemical composition is fully captured. Alternating excitation light from a dual-wavelength laser source generates spatial response comparison maps of Raman peak intensity differences. Combined with the two-dimensional distribution data, these maps form a three-dimensional feature stack, enhancing sensitivity to spatial-spectral synergistic changes in chemical composition and improving feature resolution. Topological analysis precisely locates target concentration regions and their boundary intersections, revealing the macroscopic distribution patterns of chemical components within the medicinal material and providing key node data for spatial correlation modeling. By analyzing the distribution density and directional vector relationships at boundary intersections, a graph-structured data representation of component arrangement patterns is constructed. Complex spatial relationships are converted into computable topological networks, enabling a digital representation of the medicinal material structure. Constraints are used to screen feature combinations that match prior knowledge of authentic medicinal materials, eliminating redundant information while retaining highly discriminative features in the joint spatial-spectral dimension, thereby optimizing the input quality of the discriminant model.
[0053] Furthermore, the present application constructs a constraint rule set by extracting feature distribution laws from a historical sample library, and combines the screening mechanism of the product value of regional area discreteness and intersection density to achieve hierarchical optimization of feature combinations; establishes a dual-channel calculation path in the discriminant model, and uses the asymmetric superposition of orthogonal grid space fluctuation coefficients and angle correction factors to generate joint weights, and realizes weight space transfer through impedance matching of the ring connection network, and finally triggers the identification mechanism based on the geometric relationship between the azimuth offset map and the target concentration area.
[0054] This application accurately retains the spatial distribution characteristics that characterize authenticity through dynamic constraint rules and multi-dimensional feature quantification; combines dual-channel feature fusion and the spatial transmission mechanism of the ring network to enhance the model's ability to analyze the spatial arrangement pattern of medicinal ingredients, and achieves highly robust authenticity identification through phase comparison of azimuth offset maps, significantly improving the reliability and interpretability of the identification results.
[0055] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0057] Figure 1 A flowchart of a method for selecting structured features of authentic medicinal materials based on Raman peak intensity information provided by the present application is shown;
[0058] Figure 2 A schematic diagram of the structure of a system for selecting structured features of authentic medicinal materials based on Raman peak intensity information provided by the present application is shown;
[0059] Figure 3 A schematic structural diagram of a computing device provided by the present application is shown. DETAILED DESCRIPTION
[0060] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0061] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.
[0062] Currently, traditional Raman spectroscopy is widely used in the field of authentic medicinal material identification. Its core flaws are reflected in two aspects. On the one hand, existing methods mostly focus on the static analysis of single-point spectral information, which can only obtain qualitative data on local chemical components, but cannot reveal the spatial heterogeneity of the distribution of chemical components within the medicinal materials, resulting in a lack of effective characterization of the gradient concentration distribution pattern unique to authentic medicinal materials. On the other hand, traditional feature selection strategies overly rely on the screening of single indicators in the spectral dimension, ignoring the synergistic correlation between spatial distribution characteristics and spectral response, resulting in the loss of key discriminant information. In addition, existing technologies lack the ability to quantitatively analyze the topological structure of multiple concentration regions in medicinal material slices and their spatial correlation, making it difficult to construct an interpretable distribution model. Ultimately, the identification model is susceptible to interference from nonspecific features, and its accuracy and robustness are limited.
[0063] To address the above technical bottlenecks, the present invention proposes a structured feature selection method for authentic medicinal materials based on Raman peak intensity information. This method constructs a three-dimensional feature stack through spatial sampling and dual-wavelength excitation technology, integrating the multidimensional information of two-dimensional distribution data and spatial response comparison maps, breaking through the traditional single spectral analysis dimension; further combining spatial topology analysis with graph structure modeling, accurately extracting the distribution density and directional vector relationship of the boundary intersection points of the target concentration area, and quantitatively characterizing the spatial arrangement pattern of the medicinal material components; finally, through a feature screening mechanism constrained by prior knowledge, retaining feature combinations that match the spatial structure-activity laws of authentic medicinal materials, and achieving high-specificity identification based on a spatial-spectral joint weight model. Compared with traditional methods, this scheme, through three-dimensional feature fusion and spatial correlation modeling, for the first time achieves the coordinated analysis of the spatial heterogeneity distribution of chemical components and spectral response, effectively overcoming the problem of neglecting gradient distribution features in single-point analysis; at the same time, the structured feature selection mechanism significantly improves the extraction efficiency of key discriminant features, reduces noise interference through joint weight calculation, ultimately improving the accuracy of authenticity identification, and giving the model the ability to interpret and verify the spatial structure-activity relationship of medicinal materials.
[0064] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0065] Figure 1 The present invention provides a flowchart of a method for selecting structured features of authentic medicinal materials based on Raman peak intensity information. Figure 1 As shown, the method includes:
[0066] 101. The surface area of the medicinal material slice is spatially sampled by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components.
[0067] Optionally, the step 101 of spatially sampling the surface area of the medicinal material slice by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components may specifically include:
[0068] 1011. Irradiate the surface area point by point along a preset grid path using a scanning device, where each grid node corresponds to a sampling position;
[0069] 1012. At each sampling position, use monochromatic excitation light to vertically impinge on the surface of the medicinal material, synchronously receive the Raman scattering signal reflected at the sampling position, and extract the intensity value of the preset wavelength range in the Raman scattering signal as the feature value of the grid node;
[0070] 1013. According to the spatial coordinate arrangement of the grid nodes, the feature quantities of all grid nodes are filled into a two-dimensional matrix in row-column order, and the difference between the feature quantities of adjacent grid nodes is calculated. If the difference exceeds a set threshold, an interpolation node is inserted in the corresponding row-column gap to supplement the feature quantities of the interpolation nodes to form a continuously distributed two-dimensional data plane;
[0071] 1014. Map the feature quantity of each sampling position in the two-dimensional data plane into a grayscale value to generate two-dimensional distribution data with spatial coordinates as horizontal and vertical axes and grayscale values representing chemical component concentrations.
[0072] In the above scheme, a scanning device refers to a device that uses optical or electronic technology to detect the surface of an object point by point, capable of collecting physical or chemical signals related to spatial locations. A grid node refers to the intersection of a preset grid path. Each node corresponds to a sampling point on the surface of a medicinal material slice and is used to locate the location for Raman signal collection. Monochromatic excitation light refers to an incident light source with a single wavelength, which is used to excite the medicinal material surface to produce a Raman scattering signal. Its wavelength is selected to match the Raman characteristic peak of the target chemical component. The Raman scattering signal is the inelastic scattered light signal generated by the molecules on the medicinal material surface under monochromatic light excitation. Its wavelength shift is related to the molecular vibration mode and can be used for chemical component analysis. The characteristic value is the intensity value within a preset wavelength range extracted from the Raman signal, which represents the relative concentration of the chemical component corresponding to the sampling location. Interpolation nodes are virtual nodes inserted between the original grid nodes. Their characteristic values are calculated using an interpolation algorithm and are used to fill in the sudden changes in the concentration gradient between adjacent nodes. The two-dimensional data plane is a continuous two-dimensional matrix composed of the characteristic values of the grid nodes and the interpolation nodes, which reflects the spatial distribution of the chemical component concentration. A two-dimensional matrix refers to a matrix in which the row index and the column index correspond to the horizontal and vertical physical positions of the medicinal material surface, respectively.
[0073] In the embodiment of the present application, first, step 1011 controls the scanning device to follow a preset fine grid path, for example, on a 1cmx1cm medicinal material slice surface, with a grid step size of 100um, to systematically spatially sample the surface area. The scanning device, such as a confocal Raman microscope, drives the sample stage or laser focus to precisely move to the sampling position corresponding to each grid node, for example, starting from coordinates (0,0) and scanning line by line to (10mm,10mm), generating a total of 100x100 sampling points, laying the foundation for subsequent point-to-point measurement.
[0074] Secondly, through step 1012, at each sampling position, the laser source emits monochromatic excitation light, for example, 532nm wavelength, which is vertically incident on the surface of the medicinal material. A spectrometer and a CCD detector are used synchronously to receive the reflected Raman scattering signal. Signal processing uses a spectral analysis algorithm, such as fast Fourier transform to extract the intensity value assuming a preset wavelength range of 500-600nm, and the intensity value is stored as a feature value. The feature value calculation is achieved by integrating the signal intensity in the preset wavelength range. For example, from the coordinate list obtained in step 1011, each position is processed in turn. After laser irradiation, the detector captures the signal, and the software extracts the intensity value and associates it with the corresponding grid node. This step outputs the feature value data set for each grid node.
[0075] Then, step 1013 is performed to arrange the spatial coordinates of the grid nodes and fill the feature quantities output in step 1012 into the two-dimensional matrix in row-column order. The difference calculation uses a simple subtraction algorithm to compare the feature quantities of adjacent nodes, where the simple subtraction algorithm is for upper, lower, left, and right neighbors. If the absolute value of the difference exceeds the set threshold, for example, the set threshold = 10, then an interpolation node is inserted in the corresponding gap, such as between rows or columns. Interpolation uses a bilinear interpolation algorithm to calculate the feature quantity of the new node based on the feature quantities and coordinates of the adjacent grid nodes. The formula is (in, is a constant term, and They are and The linear coefficient of for and The interaction coefficients are calculated. All interpolated nodes and the original nodes together form an expanded regular grid, for example, from 100x100 to 199x199, ultimately generating a continuously distributed two-dimensional data plane. For example, first fill the matrix, iterate over all adjacent pairs, insert nodes when the difference is too large, and update the matrix.
[0076] Finally, step 1014 normalizes the feature quantity of the two-dimensional data plane (including the original and interpolated nodes) output in step 1013 to the range of 0-255, using a linear mapping algorithm (where feature is the feature value of the current node, and The global minimum The maximum value) is converted to grayscale value. The horizontal axis is the spatial coordinate. and vertical axis , using an image processing library (such as OpenCV) to generate a grayscale image. For example, after inputting a feature matrix, the software automatically calculates global extreme values and completes grayscale mapping, ultimately outputting a two-dimensional distribution image file.
[0077] Taking the detection of saponin components in ginseng slices as an example, the grid spacing is 0.05mm, the excitation light wavelength is 532nm, and the threshold is 2000.
[0078] Specifically, ginseng slices were fixed to a stage, and a scanning device generated a grid with a 0.05mm spacing, covering the slice surface. Using a 785nm laser for excitation, the Raman signal at each node was collected, and the intensity values of the saponin characteristic peaks were extracted. If the intensity difference between adjacent nodes in a certain area exceeded a threshold, interpolated nodes were inserted to smooth the transition. A grayscale distribution map of saponin concentration was generated, showing a gradient characteristic of saponin decreasing from the center of the slice outward.
[0079] This step achieves high-resolution spatial distribution detection of chemical components on the surface of medicinal materials through point-by-point scanning and Raman signal extraction; through difference threshold judgment and interpolation node insertion, the continuity and visualization of the concentration gradient are enhanced; the final grayscale distribution data plane can intuitively reflect the enrichment area and diffusion trend of chemical components, providing a reliable basis for medicinal material quality evaluation.
[0080] 102. Alternately project the excitation light generated by the dual-wavelength laser source onto the same sampling area, generate a spatial response comparison graph based on the difference in Raman peak intensity obtained at the two wavelengths, and form a three-dimensional feature stack based on the spatial response comparison graph and the two-dimensional distribution data.
[0081] Optionally, in step 102, alternately projecting the excitation light generated by the dual-wavelength laser source onto the same sampling area, generating a spatial response comparison graph based on the difference in Raman peak intensities obtained at the two wavelengths, and forming a three-dimensional feature stack based on the spatial response comparison graph and the two-dimensional distribution data may specifically include:
[0082] 1021. Alternately switch the output wavelength of the dual-wavelength laser source through a synchronous controller, and perform complete spectrum acquisition on the same sampling area at each wavelength;
[0083] 1022. Mark the characteristic peak intensity value of the Raman signal collected at the first wavelength as a first response value, and mark the characteristic peak intensity value of the Raman signal collected at the second wavelength as a second response value;
[0084] 1023. Calculate the difference between the first response value and the second response value for each spatial sampling point, and map the difference to the same spatial coordinate system as the two-dimensional distribution data to generate a spatial response comparison graph;
[0085] 1024. Use the spatial response comparison map as an additional signal layer and stack it with the two-dimensional distribution data to form a three-dimensional feature stack with three layers of signal strength.
[0086] In the above scheme, a dual-wavelength laser source refers to a laser device that can output two fixed wavelengths, such as 532nm and 785nm, and stimulates the differential response of the target chemical components by alternating wavelength switching. The synchronization controller refers to an electronic module used to coordinate the laser wavelength switching and the spectral acquisition timing to ensure that the Raman signals at the two wavelengths are accurately collected in chronological order. The Raman peak intensity difference refers to the difference in the intensity of the Raman characteristic peak produced by the same chemical component under excitation at different wavelengths, reflecting the effect of wavelength selectivity on the detection sensitivity. The spatial response comparison map refers to a two-dimensional matrix with spatial coordinates as the reference and difference intensity as the value, which is used to visualize the response differences of chemical components under excitation at different wavelengths. The three-dimensional feature stacking refers to a three-dimensional data structure formed by superimposing the two-dimensional distribution data and the spatial response comparison map in layers, which contains the original concentration distribution and wavelength response difference information.
[0087] In an embodiment of the present application, a timing pulse signal is first sent by the synchronization controller in step 1021 to drive the dual-wavelength laser source to switch alternately between two preset wavelengths. The switching frequency is usually in the order of kHz to ensure sampling synchronization. At each wavelength, a confocal Raman microscope is used to perform point-by-point spectrum acquisition on the same sampling area. The acquisition process uses a spectrometer to record the full-band Raman spectrum of each spatial point, and the scanning path is controlled by software to ensure coverage of the entire two-dimensional sampling area. In terms of algorithm, spatial registration technology is used to ensure that the coordinates of the sampling points at different wavelengths are consistent. The output is two complete sets of spectral data sets.
[0088] Next, in step 1022, the Raman characteristic peak intensity at each sampling point at the first wavelength is marked as a first response value, and the corresponding peak intensity at the second wavelength is marked as a second response value. For example, if the characteristic peak intensity at a sampling point is 8000 photons under 532nm excitation and 5000 photons under 785nm excitation, the first and second response values are recorded as I1 and I2, respectively.
[0089] Then, step 1023 is used to calculate the difference in dual-wavelength response for each sampling point output from step 1022, using the first wavelength response value I1 and the second wavelength response value I2. The absolute difference algorithm ΔI=|I1-I2| is used to quantify the response difference. For example, at the sampling point (2,3), I1=850 and I2=420, then ΔI=|850-420|=430. Next, the spatial coordinate arrangement established in step 101 is strictly inherited, including the exact coordinates of the original grid nodes and all interpolation nodes to create an empty matrix M_ΔI with exactly the same size as the two-dimensional distribution data. For example, a 5×5 grid corresponds to a 5×5 matrix. Each sampling point is traversed by the coordinate index, and the ΔI value of the point is accurately filled into the corresponding position of M_ΔI. For example, ΔI=430 at the coordinate (2,3) is filled into M_ΔI[2][3]. For the interpolation nodes inserted in step 101, their ΔI values are calculated and filled synchronously to ensure full spatial coverage. Finally, through the matrix-to-image conversion algorithm, the M_ΔI matrix is mapped into a spatial response comparison map, where the horizontal / vertical axes correspond to the spatial coordinates and the pixel grayscale value represents the ΔI intensity. For example, ΔI=0 is displayed as black and ΔI≥500 is displayed as white, generating a spatial response comparison map that is strictly aligned with the original concentration distribution.
[0090] Finally, in step 1024, the spatial response comparison map is recorded as a matrix M_ΔI, also of size m×n, where each element M_ΔI[i][j] represents the difference in dual-wavelength response at the same location. The two-dimensional distribution data obtained in step 101 is recorded as a matrix M_gray, of size m×n, where each element M_gray[i][j] represents the grayscale value of the chemical component concentration at the spatial coordinate (i, j). Next, a three-dimensional array data structure of dimension m×n×2 is created. M_gray is used as the first signal layer, layer Z=0, and is completely copied to the D[:,:,0] channel of the three-dimensional array. That is, for each spatial location (i, j), D[i][j][0]=M_gray[i][j]. Then, M_ΔI is used as the second signal layer, layer Z=1, and is copied to the D[:,:,1] channel. That is, for each location (i, j), D[i][j][1]=M_ΔI[i][j]. Finally, if there is a need for expansion, such as adding a third signal layer, additional feature data can be stacked at layer Z = 2 to form a complete three-dimensional feature stacking data structure. For example, for a 3×3 grid system, the original concentration matrix is [[180, 200, 190], [210, 220, 200], [190, 180, 210]], and the response contrast matrix is [[310, 290, 300], [280, 270, 290], [290, 310, 280]], then the three-dimensional feature at position (1, 1) after stacking is [220, 270].
[0091] Taking the detection of ferulic acid components in angelica slices as an example, specifically, dual-wavelength laser sources alternately irradiate the same area of the angelica slices, and collect Raman signals separately. To extract the characteristic peak intensity of ferulic acid, the first response value is recorded at 532nm, and the second response value is recorded at 830nm. When the difference is calculated, it is found that the response value of 532nm in the epidermal area is 2000 photons higher than that of 830nm, and 1500 photons lower in the medulla area. The generated spatial response comparison map shows that the epidermis is a high-difference area. The comparison map is superimposed with the ferulic acid concentration distribution map, and the three-dimensional stack shows that the epidermal area has both high concentration and high wavelength response differences.
[0092] This step enhances the ability to identify the spatial response characteristics of specific chemical components through dual-wavelength excitation and difference comparison; three-dimensional feature stacking combines concentration distribution and wavelength selectivity information, which can reveal the response patterns of chemical components under different excitation conditions, providing a data basis for multimodal analysis while avoiding the limitations of single wavelength detection.
[0093] 103. Perform spatial topological analysis on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections.
[0094] Optionally, performing spatial topological analysis on the three-dimensional feature stack in step 103 to identify at least three target concentration regions and their boundary intersections may specifically include:
[0095] 1031. Perform a segmentation operation on each signal layer in the three-dimensional feature stack, and extract continuous regions with signal intensity values higher than a preset threshold as candidate concentration regions;
[0096] 1032. Perform spatial position comparison on the candidate concentration regions in the multiple signal layers, and retain the candidate regions whose spatial overlap in at least two signal layers exceeds a set ratio as target concentration regions;
[0097] 1033. Detect the boundary curves of each target concentration region, and calculate the minimum spacing and curvature change rate between the boundary curves corresponding to adjacent target concentration regions;
[0098] 1034. When the minimum spacing between boundary curves corresponding to adjacent target concentration regions is less than a preset ratio of the average width of the region, and the curvature change rate exceeds a preset mutation threshold, it is determined that there is an intersection between the boundaries corresponding to the adjacent target concentration regions.
[0099] In the above scheme, a three-dimensional feature stack refers to a three-dimensional data structure formed by stacking multiple layers of signal data, such as medical images, remote sensing data, or sensor network data, along a spatial dimension. Each layer contains the spatial distribution of signal intensity values. A candidate concentration region refers to a continuous region extracted through a segmentation operation in a single layer of signal data, whose signal intensity value exceeds a preset threshold, such as the grayscale threshold of a lesion region in a medical image. Spatial overlap refers to the proportion of overlapping area of candidate regions in different signal layers when projected onto the same spatial coordinate system, and is used to measure the consistency of cross-layer regions. The boundary curve refers to the outer contour of the target concentration region, a continuous curve fitted by discrete points that characterizes the region shape. The average region width refers to the average span of the target concentration region in the main extension direction, and is used to normalize the boundary spacing judgment criterion. The curvature change rate refers to the derivative of the local curvature of the boundary curve, reflecting the degree of abrupt change in the boundary shape, such as a sharp corner or inflection point. Boundary intersections are determined by detecting when the gradient change rate of adjacent regions exceeds a set change amplitude.
[0100] In the embodiment of the present application, first, each signal layer of the three-dimensional feature stack is independently segmented through step 1031. The Otsu algorithm is used to automatically calculate the optimal threshold for the first signal layer, traverse the grayscale levels of 0-255, select the threshold T0 that maximizes the inter-class variance σ²=ω0ω1(μ0-μ1)², generate a binary mask, and then perform 8-neighborhood connected area marking, and output the candidate concentration area set C0. Synchronously set the medicinal material-specific fixed threshold T1 for the second signal layer, mark the connected areas after denoising by the morphological opening operation of the 3×3 structural element, and output the candidate concentration area set C1. For example, the concentration layer of the angelica slices segments the ferulic acid enrichment area R1, and the difference layer segments the dual-wavelength response area S1.
[0101] Secondly, according to the spatial coordinate arrangement and spatial response comparison diagram, it is ensured that the C0 and C1 regions are in the same spatial reference system through step 1032. The affine transformation matrix is used to correct the position offset to perform spatial position comparison, where the formula is is the coordinate before transformation, is the transformed coordinate, is the rotation and scaling matrix, is the translation vector; calculate the Jaccard overlap of each pair of regions, the formula is ,in is the number of overlapping pixels, is the number of pixels in the union. The calculation results retain only regions with an overlap greater than 0.6, forming the target concentration region set T. For example, if the overlap between R1 and S1 is 0.75, the target region T1 is generated.
[0102] Next, the boundary of each target area is extracted in step 1033. Moore's boundary tracing algorithm is used to search the 8 neighborhoods in a counterclockwise direction starting from the upper left corner pixel to generate a closed boundary curve. For adjacent target areas (e.g. and ), calculate the Hausdorff distance between boundary curves (in, are the coordinates of the boundary points, Indicates a point To the curve the shortest Euclidean distance of ), and calculate the mean curvature change rate (in, For the The tangent angle of the boundary point, is the total number of boundary points). For example, the distance between the scutellaria saponin region and the polysaccharide region , curvature change rate radian .
[0103] Finally, the average width of the target area is calculated in step 1034, and the volume equivalent formula is used for the target area. Conclusion (in, represents the average width, and the volume is the total volume of the target area). Set the judgment condition. If the minimum spacing Less than And the curvature change rate Exceeding the mutation threshold , then mark the boundary crossing. In this case and , so in three-dimensional coordinates The intersection is confirmed as a boundary crossing point (where is the minimum spacing, is the curvature change rate, represents the rate of change of curvature along the path).
[0104] Taking the 5×5 mm slice of authentic medicinal material A as an example, specifically, the three-dimensional feature stacking The pixel concentration layer and the difference layer are processed separately:
[0105] The concentration layer is segmented using Otsu adaptive threshold (threshold ), extract area , whose coordinate range is (10-30, 20-40); the difference layer is fixed by the threshold Segmentation, get region , the coordinate range is (15-32, 22-42); calculate the overlap of the two regions , generate the target area ; For adjacent target areas , calculate the boundary parameters in step 1033: minimum Hausdorff distance ; Mean curvature change rate radian ; Judgment condition: spacing condition ;Curvature condition ; Finally, in the three-dimensional coordinates Identify as a boundary intersection.
[0106] This step achieves precise identification of 3D target regions through layer-by-layer segmentation and spatial registration, effectively locating intersections based on boundary geometry. Its advantage lies in its integration of multi-level signal correlations, avoiding the limitations of single-layer analysis. Furthermore, it enhances the robustness of boundary topology analysis through curvature variation and spacing criteria, making it suitable for automated analysis of complex 3D structures.
[0107] 104. Construct a spatial correlation model between the target concentration areas, and analyze the distribution density and directional vector relationship of the boundary intersection points in each target concentration area through the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of the medicinal material components.
[0108] Optionally, the step 104 of analyzing the distribution density and direction vector relationship of the boundary intersection points in each target concentration region by using the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of the medicinal material components may specifically include:
[0109] 1041. Establish an initial topological network using the spatial correlation model, with the geometric centers of each target concentration area as nodes and boundary intersections as connection hubs;
[0110] 1042. In the initial topological network, calculate the number of intersections within a preset radius around each boundary intersection as a density distribution parameter, and record the direction angle of the line connecting each boundary intersection and the center of gravity of the target concentration area to which it belongs as a direction relationship parameter;
[0111] 1043. Perform vector synthesis on boundary intersections with overlapping density distribution parameters to generate weighted edges representing the interaction strength between regions;
[0112] 1044. Use the density distribution parameter and the direction relationship parameter as node attributes, and use the vector synthesis direction of the weighted edge as an edge attribute to construct graph structure data containing multi-level spatial association features.
[0113] In the above scheme, target concentration regions refer to areas of medicinal ingredient concentration distribution delineated through image analysis or chemical testing. Each region represents a high concentration of a specific chemical component and includes attributes such as geometry, boundaries, and concentration gradients. Boundary intersections are the intersection points of the boundaries of different target concentration regions and are used to characterize the spatial contact relationship between regions. The geometric centroid is the geometric center of the target concentration region, obtained by calculating the coordinates of the polygonal vertices of the region's outline and serving as a node in the topological network. The initial topological network is an undirected graph with geometric centroids as nodes and boundary intersections as connecting hubs, describing the adjacency relationship between regions. The density distribution parameter is the number of other intersections within a preset radius centered on the boundary intersection, reflecting the density of the region boundary. The directional relationship parameter is the directional angle between the line connecting the boundary intersection and the geometric centroid of the region to which it belongs, quantifying the direction of the intersection's contribution to the region. Weighted edges are edges generated through vector synthesis, with their weights representing the strength of the interaction between regions and their directions determined by the angle of the synthesized vectors. Graph structure data refers to a topological network containing node attributes such as density and directional angle, and edge attributes such as weight and direction, used to describe the spatial arrangement pattern of medicinal ingredients.
[0114] In the embodiment of the present application, first, step 1041 is used to calculate the geometric center of gravity coordinates of each target concentration area set based on the target concentration area set and the boundary intersection point set output in step 103, wherein the formula for calculating the center of gravity coordinates is: ,in is the geometric center coordinate of the kth target concentration area, n is the total number of boundary intersections contained in the area, (xi,yi) is the spatial coordinate of the i-th boundary intersection. With the center of gravity as the node and the intersection as the hub, an initial topological network is established, where the node set V={ }, hub set H={ }, edge set ,in Indicates area For example, the centroids of the three saponin-rich regions G1-G3 and the boundary intersections P1-P5 in Astragalus slices were used to construct a network with 3 nodes and 5 hubs.
[0115] Next, at step 1042, the density distribution parameter of each boundary intersection is calculated. ,by Draw a circle with a preset radius r as the center, for example r = 0.5mm, and count the number of other intersections in the circle .in For intersection The density value at are the coordinates of other boundary intersection points. Synchronously calculate the direction relationship parameter θm and obtain the vector The direction angle of , ,in is the center of gravity of the region to which the boundary intersection set belongs. Calculate the center of gravity and direction angle of each boundary intersection according to the above steps, and record the direction angle of the line connecting each boundary intersection and the center of gravity of the target concentration region to which it belongs as the direction relationship parameter.
[0116] Then, in step 1043, points with the same or similar density distribution parameters are identified, and the intersection points with overlapping density distributions are identified. >0 performs vector synthesis. Extract the direction vector of all boundary intersections in the same area. The formula is , weighted composite region representative vector , where the weight , Owned by the same region Generate weighted edges e between the geometric center nodes of different target concentration areas ij , weight . Weighted edge e ij The weight w ij The interaction strength between target concentration regions is directly quantified to generate weighted edges representing the interaction strength between regions. For example, the angle between the composite vectors of the saponin region and the polysaccharide region is 60°, resulting in an edge with a weight of 0.87.
[0117] Finally, in step 1044, for each target concentration area, the density distribution parameters of all associated boundary intersections (area) are summarized and their arithmetic mean is calculated as the node density attribute to quantify the complexity of the boundary structure of the area and represent the dense distribution of multiple intersections; at the same time, the regional representative vector generated in step 1043 is extracted, and its direction angle is used as the node direction attribute to characterize the overall distribution trend of the chemical composition in space. For each weighted edge, the angle between the two regional representative vectors calculated in step 1043 is used as the edge direction attribute, and the weight is directly inherited as the edge strength attribute. The two together quantify the spatial interaction pattern. Finally, the three-layer association features are integrated to construct graph structure data, where the first-layer regional features are node attributes that describe the structural properties of the target concentration area itself; the second-layer interaction features are edge attributes that capture the spatial interaction relationship between adjacent areas; and the third-layer topological features are network connection structures that reveal the spatial arrangement framework of multiple areas, forming graph structure data of a multi-level spatial association model.
[0118] Taking the slices of Scutellaria baicalensis as an example, the spatial distribution of flavonoids in the slices was analyzed.
[0119] Specifically, five high-flavonoid concentration regions were delineated through microscopic imaging and threshold segmentation. The coordinates of their centroids were calculated, 12 boundary intersections were detected, and a topological network consisting of five nodes and 12 hubs was constructed. A radius of 100 μm was set, and 2 to five neighboring points around each intersection were counted, with the azimuth angle recorded. For example, the azimuth angle between intersection A and the centroid of region 1 was 45°. For a cluster of intersections with a density of 3, vectors with azimuth angles of 30°, 90°, and 150° were synthesized, generating weighted edges with a weight of 2.6 and a direction of 60°. Finally, a graph consisting of five nodes and eight weighted edges was generated, which was used to identify the "radial-cluster" distribution pattern of flavonoid components.
[0120] This step constructs a multi-level spatial correlation model by quantifying the density and direction characteristics of boundary intersections. It can intuitively reveal the spatial interaction pattern of medicinal ingredients, improve the efficiency of analyzing the distribution patterns of complex ingredients, and provide structured input for ingredient prediction based on graph neural networks.
[0121] 105. Perform a structured feature selection operation to retain, in the graph structure data based on constraints, a feature combination that matches the prior knowledge of the spatial distribution of authentic medicinal materials.
[0122] Optionally, the feature combination that matches the prior knowledge of the spatial distribution of authentic medicinal materials in the graph structure data through the constraint conditions in step 105 may specifically include:
[0123] 1051. Extracting characteristic distribution patterns of the graph structure data from a historical authentic medicinal material sample library, and constructing a constraint rule set based on the characteristic distribution patterns;
[0124] 1052. Traverse all feature combinations in the graph structure data and select feature combinations that meet preset conditions as a primary screening set; wherein the feature combinations that meet the preset conditions are feature combinations with multiple constraint tags;
[0125] 1053. For each feature combination in the primary screening set, calculate the product value of the area dispersion and the intersection density, and retain the feature combinations whose product values are ranked within a predefined ratio range as the screened feature combinations.
[0126] In the above scheme, graph-structured data refers to a feature network represented in a graph format. Nodes represent feature attributes, such as the climate and soil properties of a medicinal material distribution area, and edges represent relationships between features. A feature combination refers to a parameter consisting of at least two of the following: the number of boundary intersections, the area ratio of target concentration regions, and the angle between directional vectors. A feature distribution pattern refers to the spatial distribution pattern of feature combinations in historical authentic medicinal material samples, such as the correlation between high-altitude areas and specific precipitation levels. A constraint rule set includes logical rules for regional area dispersion thresholds, intersection density intervals, and vector angle distribution patterns, such as "authentic medicinal materials must meet both the requirements of an altitude greater than 1000 meters and precipitation greater than 800 mm." A constraint tag is an identifier indicating that a feature combination satisfies a particular constraint rule, such as Rule A or Rule B. Regional area dispersion refers to the degree of dispersion of the spatial area distribution corresponding to a feature combination, such as the standard deviation or entropy value, reflecting the degree of regional area variability. Intersection density refers to the ratio of the number of topological intersections corresponding to a feature combination to the regional area, reflecting the complexity of the boundary.
[0127] In an embodiment of the present application, first, step 1051 is used to extract the characteristic distribution pattern of the graph structure data generated in step 1044 from the historical authentic medicinal material sample library; Gaussian mixture model clustering is performed on the density mean of the node attributes, the angle between the reverse angle and the edge attributes, and the weight of all samples in the historical authentic medicinal material sample library to form a characteristic matrix, for example, the covariance matrix type is full. The probability distribution of each characteristic matrix is calculated. For each characteristic matrix, the arithmetic mean of the eigenvalues is taken as the confidence interval to generate constraint rules. Feature matrices with strong correlation are merged to generate a set of constraint rules.
[0128] Then, all possible feature combinations of the graph structure data of the current medicinal material are traversed in step 1052. Each feature combination includes at least one density mean or direction angle of a node attribute and the angle α or weight of the associated edge attribute, such as the G1 attribute of the saponin region node. The edge e12 attribute α12=42∘ connected to the polysaccharide region forms a combination c1; then the constraint rule set is used to perform rule matching on each combination, and the feature combination that matches both rules is screened, where the counter count=0 is initialized. If each combination satisfies rule 1, the count is increased by 1, and if it satisfies rule 2, the count is increased by 1. For example, in c1 ∈[2.8,3.4] triggers rule 1, and α12 = 42∘ < 50∘ triggers rule 2, resulting in count = 2. Each combination is marked as valid only when count ≥ 2. Finally, all combinations with a count ≥ 2 are integrated, such as c1. A primary screening set is constructed. For example, 15 dual-rule matching combinations were screened from Astragalus slices.
[0129] Finally, in step 1053, the area dispersion of each feature combination in the primary screening set is calculated, and the formula is: ,in δ k is the area dispersion of the feature combination, is the actual measured area of a certain feature combination, is the predefined standard reference area, The actual measurement have to and intersection density ,in have to , find the product value The top 15% of the combinations are retained and the final output is the filtered feature combination that matches the prior knowledge of the spatial distribution of authentic medicinal materials.
[0130] Take the Chinese medicinal herb Astragalus as an example.
[0131] Specifically, we identified characteristic combinations from historical data that define authentic Astragalus production areas, such as "annual average temperature 6-10°C, soil pH 6.5-7.5, and slope <15°." We then screened for combinations that simultaneously met these temperature, pH, and slope criteria, eliminating combinations that only met one of these criteria. We then calculated the regional area dispersion and intersection density of the retained combinations, retaining those with low dispersion and high density.
[0132] This step achieves precise optimization of feature combinations through historical constraints and geometric feature screening. Its advantage lies in integrating prior knowledge to enhance feature relevance. It also combines regional distribution stability with topological complexity to effectively distinguish key features from redundant ones, enhancing the interpretability and reliability of the spatial distribution model of authentic medicinal materials.
[0133] 106. Input the filtered feature combination as a structured feature into an authenticity discrimination model, so that the discrimination model outputs an identification result based on the joint weight of the feature combination in the spatial dimension and the spectral dimension.
[0134] Optionally, in step 106, inputting the filtered feature combination as a structured feature into the authenticity discrimination model, so as to output the identification result according to the joint weight of the feature combination in the spatial dimension and the spectral dimension through the discrimination model, may specifically include:
[0135] 1061. Input the filtered feature combination into the authenticity discrimination model as a discrete unit with spatial coordinate labels;
[0136] 1062. A dual-channel calculation path is established within the authenticity discrimination model. The first channel maps the two-dimensional coordinate values of each discrete unit to a virtual plane divided by an orthogonal grid. The standard deviation of the spectral response amplitude difference of all discrete units within each orthogonal grid is calculated to form a spatial fluctuation coefficient per grid unit. The second channel extracts the straight-line distance between each discrete unit and the intersection point of the boundary of the target concentration region to which it belongs, and combines this with the directional correlation strength to generate an angle correction factor.
[0137] 1063. Generate a joint weight value according to the spatial fluctuation coefficient and the angle correction factor in an asymmetric superposition manner;
[0138] 1064. Construct a computing network including a ring connection structure, inject the joint weight value of each discrete unit into its corresponding ring node, and achieve spatial transmission of the weight value through the impedance matching relationship between adjacent ring nodes, so that the output end of the computing network can receive the joint weight value output by each port, and perform a phase comparison between the sum of the joint weight values of each port and a preset azimuth reference value to generate a phase comparison result;
[0139] 1065. Generate an azimuth offset map based on the phase comparison result, and perform overlapping analysis on the spatial distribution patterns of the offset extreme value regions and the target concentration regions that appear continuously in the azimuth offset map. When the boundary of the offset extreme value region forms a specific angle relationship with the geometric center of at least three target concentration regions, trigger the confirmation mechanism of the authenticity discrimination model and output the identification result.
[0140] In the above scheme, a discrete unit refers to a selected feature combination. Each unit contains spatial coordinates, such as two-dimensional latitude and longitude or image pixel location, and corresponding spectral feature data, such as spectral response values in different bands. An orthogonal grid is a mathematical method that divides a virtual plane into a grid of equally spaced rows and columns, with grid boundaries formed by perpendicularly intersecting straight lines. This method is used to calculate spatial distribution characteristics. The spatial fluctuation coefficient is the standard deviation calculated based on the difference in spectral response amplitudes of discrete units within the orthogonal grid, reflecting the stability or variability of spectral characteristics within a region. The angle correction factor is a parameter generated by combining the straight-line distance from the discrete unit to the boundary of the target concentration region and the directional correlation strength, used to correct directional bias in weight calculation. The combined weight value is a comprehensive weight obtained by asymmetric superposition of the spatial fluctuation coefficient and the angle correction factor, representing the importance of the feature combination in the spatial and spectral dimensions. A ring connection structure is a network topology in which each node corresponding to a discrete unit is connected to its adjacent nodes through a ring path, and weight values are transmitted between nodes through impedance matching. The azimuth offset map is a two-dimensional map generated based on phase comparison that displays the offset distribution within a region. Extreme value areas indicate abnormal spatial distribution or areas with significant features. Impedance matching means that the impedance value is inversely proportional to the cosine component of the directional correlation strength between units. Phase comparison involves setting up a directionally selective receiving array at the output of the computing network. This array contains acquisition ports distributed at equal angles along eight azimuths. Each port only receives weighted values transmitted from ring nodes at a specific azimuth angle, and then performs a phase comparison between the accumulated weighted values of each port and a preset azimuth reference value.
[0141] In the embodiment of the present application, the filtered feature combinations, such as the node attributes of the saponin region in the Astragalus slice that meet the spatial distribution rules, are first combined with the associated edge attributes in step 1061. The spatial position information carried by the combination, i.e., the surface coordinates of the medicinal material corresponding to the feature combination, such as the two-dimensional grid positions x = 0.2 mm and y = 0.3 mm recorded during Raman scanning, are marked as inherent attributes of the discrete units and input into the authenticity discrimination model.
[0142] Next, a dual-channel calculation path is established through step 1062. The first channel maps the discrete unit coordinates to a virtual plane divided by an orthogonal grid, and for each discrete unit in the grid, calculates the standard deviation of its spectral response amplitude difference, which is expressed as follows: Get the spatial fluctuation coefficient, where is the standard deviation of the spectral response amplitude difference, For the The spectral response amplitude of a discrete unit, is the mean of the spectral response amplitudes of all discrete units in the grid, is the total number of discrete units in the grid. The second channel measures the straight-line distance from each discrete unit to the nearest boundary point of its region. For example, the distance between the unit in the saponin region and the boundary point is 0.05 mm. The angle correction factor is obtained by directional consistency adjustment based on the boundary association weight of the boundary point and the angle between the direction of the discrete unit pointing to the boundary point and the direction of the principal component distribution.
[0143] Next, in step 1063, the spatial fluctuation coefficient is obtained as the stability reference value, and (1 + λ) is used as the multiplication coefficient, where λ is the angle correction factor. The two calculation results are multiplied together to obtain the final joint weight.
[0144] Then, a ring computing network is constructed in step 1064, so that each discrete unit, for example, the coordinates of the saponin area unit (0.2mm, 0.3mm) are mapped to a ring node, and the initial weight value of the node is the joint weight value calculated in step 1063. Then the dynamic impedance value between adjacent nodes is calculated, and the formula is impedance , w A With w B is the weight value of nodes A and B. For example, if the weight of node A is 1.54 and the weight of adjacent node B is 1.30, then , the weight is converted into a transfer signal through the current transfer model, for example, node A outputs current to the adjacent node Ampere; when the network is running, the signal converges from the boundary node to the center, and finally the total current value of each port is accumulated at the output end. For example, the current of the three output ends is [2.1∠30°, 1.8∠45°, 3.4∠20°]. The vector phase of this synthetic current vector is compared with the authentic medicinal material orientation reference value. The formula is phase difference , where arg() is the complex vector phase angle function. For example, the phase difference ΔΦ = 15° between the actual synthesized phase of 30° and the reference phase of 15° generates a phase comparison result.
[0145] Finally, in step 1065, the data from the phase comparison results, such as the deviation ΔΦ = 14° between the network output phase of 32° and the reference phase of 18°, are combined with the spatial coordinates of each discrete unit and the phase difference to generate an azimuth offset map. The phase offset values of all units are marked within the coordinate system. For example, areas with ΔΦ > 10° are rendered as red thermal areas, forming a two-dimensional offset matrix containing spatial distribution information. Spatial registration of the azimuth offset map with the target concentration area is then performed, and the azimuth offset map is superimposed on the two-dimensional distribution data of the medicinal material slice. The continuously appearing extreme offset areas in the map, such as red areas with an area greater than 0.1 mm², are extracted as targets to be analyzed. Next, a geometric relationship check is performed to calculate the boundary contour line of the offset extreme value area and measure the angular relationship between it and the geometric center points of at least three target concentration areas. A triangular network is constructed using the center point group, and the angle θ formed by the foot of the perpendicular from the offset area boundary line to the lines connecting the centers of gravity is calculated. For example, the foot angle of the perpendicular from the boundary point to the AB line is 118°, the foot angle of the perpendicular to the BC line is 122°, and the foot angle of the perpendicular to the AC line is 120°. Finally, when the angle relationship matches a specific authentic structural rule, such as 115°≤θ≤125° and the extreme value area covers the midpoint of the center of gravity line, the authenticity discrimination model confirmation mechanism is triggered. For example, the authenticity identification result of Astragalus slices is output because they simultaneously meet the triangular angle of 120°±2° and the offset area coverage rate reaches 85%.
[0146] Taking the authenticity identification of Chinese medicinal materials as an example, a 10×10 pixel grid, a spatial fluctuation coefficient weight of 0.6, and an angle correction factor weight of 0.4 are used.
[0147] Specifically, in the hyperspectral image data of Chinese medicinal materials in a certain area, each pixel contains spatial coordinates and spectral characteristics. The pixel points are input into the model as discrete units, and the coordinates correspond to the image positions. The orthogonal grid is divided into 0×10 pixels, and the standard deviation of the spectral response amplitude in each grid is calculated; at the same time, the distance and direction angle from the pixel to the boundary of the preset authentic production area are calculated. The spatial fluctuation coefficient with a weight of 0.6 and the angle correction factor with a weight of 0.4 are superimposed to generate a joint weight. A ring network is constructed to transfer the weights, and the ports with larger output phase differences correspond to abnormal areas. It is found that the extreme offset area forms an equilateral triangle angle with the centers of the three authentic production areas, and it is determined to be authentic medicinal materials.
[0148] This step fuses spatial and spectral information through a dual-channel computational path, incorporates the spatial transfer characteristics of a ring network, and dynamically adjusts the joint weights. Ultimately, this method achieves highly robust authenticity identification through phase comparison and morphological overlap analysis. Its advantages lie in its ability to adapt to differences in feature distribution across regions, reduce noise interference, and verify the reliability of results through geometric relationships, significantly improving identification accuracy and interpretability.
[0149] The following is a complete example of steps 101 to 106:
[0150] In a local medicinal material identification system based on spatial-spectral fusion, 10 samples of local medicinal materials from area A and 10 samples of common medicinal materials from area B were first selected. A precision microtome was used to prepare transverse slices with a thickness of 50 μm. These slices were placed on a microscopic confocal Raman scanning platform for spatial sampling. A 5×5 mm² area was scanned with a step accuracy of 0.5 μm. The original number of sampling points was downsampled to a 500×500 grid during actual processing, with each pixel representing a 10 μm×10 μm area, to generate a two-dimensional chemical composition distribution matrix.
[0151] Next, a dual-wavelength laser module with wavelengths of 830 nm and 785 nm was configured. A fiber coupler was used to alternately project the two excitation beams onto the same sampling area. When the 830 nm laser was detected, the system recorded the intensity of the characteristic Raman peak at 1600 cm⁻¹; when the laser was switched to 785 nm, the intensity of the characteristic peak at 1250 cm⁻¹ was captured. A spatial response comparison map was generated through difference calculation. The pixel intensity values of this map were used as the third dimension (Z3) and were aligned and stacked with the previous two-dimensional distribution data, forming a 256×256×3 three-dimensional feature stack consisting of spatial coordinates and the three eigenvalues.
[0152] A density clustering algorithm was then used to perform spatial topological analysis of the three-dimensional stack. By setting constraints of a flavonoid concentration threshold >2.5 mg / g, a saponin concentration gradient change rate <15%, and a Raman response difference >300, target concentration regions α (1.2 mm in diameter), β (0.8 mm in diameter), and γ (0.5 mm in diameter) were identified in the northwest quadrant of the slice. A Voronoi diagram was used to identify the boundary intersections of the three regions, resulting in the localization of five characteristic intersections.
[0153] Next, a spatial association model based on graph theory was constructed. Each intersection was used as a node, and the Euclidean distance between nodes was calculated as the edge weight. The node density in the northwest quadrant was also analyzed, reaching 18 nodes per mm². Directional vector analysis revealed that nodes in region α exhibited a 30° radial distribution, while nodes in region β exhibited a 60° circular arrangement. These topological features were quantified into an adjacency matrix and stored as graph-structured data.
[0154] Subsequently, feature selection criteria were established based on the "Spatial Distribution Standards for Authentic Medicinal Materials": ① Node density threshold ≥ 15 nodes / mm²; ② Radial angle deviation < 5°; ③ Circular arrangement completeness > 80%. Using these three constraints, 12 feature combinations meeting the requirements were selected from the 28 features in the original graph structure data, including key metrics such as spatial distribution dispersion (0.32) and feature angle matching.
[0155] Finally, the filtered structured features were fed into a pre-trained authenticity discrimination model. This authenticity model, using a combined spatial dimension weight of 0.6 and a spectral dimension weight of 0.4, determined that the sample's spatial distribution pattern matched authentic medicinal materials from site A by 87.3%, while its match with common medicinal materials from site B was only 41.2%. The model ultimately concluded that the sample met the characteristics of authentic medicinal materials.
[0156] This example demonstrates the complete technical process from data collection to result interpretation. All technical parameters are virtual settings, without specific brand or region information, to comply with confidentiality requirements. Each step includes specific operation methods, data processing procedures, and decision-making logic, forming a closed-loop technical verification system.
[0157] Figure 2 The present invention provides a schematic diagram of a structured feature selection system for authentic medicinal materials based on Raman peak intensity information. Figure 2 As shown, the system includes:
[0158] The first generating module 21 is used to perform spatial sampling on the surface area of the medicinal material slice by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components;
[0159] A second generating module 22 is configured to alternately project the excitation light generated by the dual-wavelength laser source onto the same sampling area, generate a spatial response comparison graph based on the difference in Raman peak intensities obtained at the two wavelengths, and form a three-dimensional feature stack based on the spatial response comparison graph and the two-dimensional distribution data;
[0160] an identification module 23 for performing spatial topological analysis on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections;
[0161] The third generation module 24 is used to construct a spatial correlation model between the target concentration areas, and analyze the distribution density and direction vector relationship of the boundary intersection points in each target concentration area through the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of the medicinal material components;
[0162] The output module 25 is used to perform a structured feature selection operation, retaining feature combinations that match the prior knowledge of the spatial distribution of authentic medicinal materials in the graph structure data through constraint conditions, and inputting the screened feature combinations as structured features into the authenticity discrimination model, so as to output the identification results through the discrimination model based on the joint weights of the feature combinations in the spatial dimension and the spectral dimension.
[0163] Figure 2 The structured feature selection system for authentic medicinal materials based on Raman peak intensity information can be performed Figure 1The implementation principles and technical effects of the method for selecting authentic medicinal materials based on Raman peak intensity information in the illustrated embodiment are not further elaborated. The specific manner in which each module and unit performs operations in the system for selecting authentic medicinal materials based on Raman peak intensity information in the above-mentioned embodiment has been described in detail in the embodiments of the method and will not be further elaborated here.
[0164] In one possible design, Figure 2 The structured feature selection system for authentic medicinal materials based on Raman peak intensity information of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;
[0165] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .
[0166] The processing component 32 is used for the above Figure 2 The embodiment provides a method for selecting structured features of authentic medicinal materials based on Raman peak intensity information.
[0167] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.
[0168] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0169] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.
[0170] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.
[0171] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0172] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0173] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment shows a method for selecting structured features of authentic medicinal materials based on Raman peak intensity information.
[0174] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0175] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0176] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for selecting structured features of authentic medicinal materials based on Raman peak intensity information, characterized in that: include: The surface area of the medicinal material slice is spatially sampled by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components; Alternately projecting excitation light generated by a dual-wavelength laser source onto the same sampling area, generating a spatial response comparison map based on the difference in Raman peak intensity obtained at the two wavelengths, and forming a three-dimensional feature stack based on the spatial response comparison map and the two-dimensional distribution data; performing spatial topological analysis on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections; Constructing a spatial correlation model between the target concentration areas, and analyzing the distribution density and directional vector relationship of the boundary intersection points in each target concentration area through the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of the medicinal material components; A structured feature selection operation is performed to retain, in the graph structure data based on constraints, feature combinations that match the prior knowledge of the spatial distribution of authentic medicinal materials, and the screened feature combinations are input as structured features into an authenticity discrimination model, so that the discrimination model outputs an identification result based on the joint weights of the feature combinations in the spatial dimension and the spectral dimension.
2. The method according to claim 1, characterized in that The method of retaining a feature combination matching the prior knowledge of the spatial distribution of authentic medicinal materials in the graph structure data through constraint conditions includes: Extracting characteristic distribution patterns of the graph structure data from a historical authentic medicinal material sample library, and constructing a constraint rule set based on the characteristic distribution patterns; Traversing all feature combinations in the graph structure data, screening out feature combinations that meet preset conditions as a primary screening set; wherein the feature combination that meets the preset conditions refers to a feature combination with multiple constraint tags; For each feature combination in the primary screening set, the product value of the regional area dispersion and the intersection density is calculated, and the feature combinations whose product values are ranked within a predefined ratio range are retained as the screened feature combinations.
3. The method according to claim 1, characterized in that The filtered feature combination is input as a structured feature into the authenticity discrimination model, so that the discrimination model outputs the identification result according to the joint weight of the feature combination in the spatial dimension and the spectral dimension, including: The filtered feature combination is input into the authenticity discrimination model as a discrete unit with spatial coordinate labels; A dual-channel calculation path is established within the authenticity discrimination model. The first channel maps the two-dimensional coordinate values of each discrete unit to a virtual plane divided by an orthogonal grid. The standard deviation of the spectral response amplitude difference of all discrete units within each orthogonal grid is calculated to form a spatial fluctuation coefficient per grid unit. The second channel extracts the straight-line distance between each discrete unit and the intersection point of the target concentration region to which it belongs, and combines this with the directional correlation strength to generate an angle correction factor. generating a joint weight value according to the spatial fluctuation coefficient and the angle correction factor in an asymmetric superposition manner; Construct a computing network with a ring connection structure, inject the joint weight value of each discrete unit into its corresponding ring node, and achieve spatial transmission of the weight value through the impedance matching relationship between adjacent ring nodes. The output end of the computing network can receive the joint weight value output by each port, and the sum of the joint weight values of each port is compared with a preset azimuth reference value to generate a phase comparison result. An azimuth offset map is generated based on the phase comparison results, and the spatial distribution morphology of the offset extreme value areas and the target concentration areas that appear continuously in the azimuth offset map is overlapped and analyzed. When the boundary of the offset extreme value area forms a specific angle relationship with the geometric center of at least three target concentration areas, the confirmation mechanism of the authenticity discrimination model is triggered, and the identification result is output.
4. The method according to claim 1, wherein The spatial correlation model is used to analyze the distribution density and directional vector relationship of the boundary intersection points in each target concentration area to generate graph structure data representing the spatial arrangement pattern of the medicinal material components, including: An initial topological network is established using the spatial correlation model, with the geometric centers of each target concentration area as nodes and boundary intersections as connection hubs; In the initial topological network, the number of intersections within a preset radius around each boundary intersection is calculated as a density distribution parameter, and the direction angle of the line connecting each boundary intersection and the center of gravity of the target concentration area to which it belongs is recorded as a direction relationship parameter; Vector synthesis is performed on boundary intersections with overlapping density distribution parameters to generate weighted edges representing the interaction strength between regions; The density distribution parameter and the direction relationship parameter are used as node attributes, and the vector synthesis direction of the weighted edge is used as an edge attribute to construct graph structure data containing multi-level spatial association features.
5. The method according to claim 1, characterized in that The surface area of the medicinal material slice is spatially sampled by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components, including: The surface area is irradiated point by point along a preset grid path by a scanning device, and each grid node corresponds to a sampling position; At each sampling position, monochromatic excitation light is vertically incident on the surface of the medicinal material, and the Raman scattering signal reflected at the sampling position is synchronously received, and the intensity value of the preset wavelength range in the Raman scattering signal is extracted as the feature value of the grid node; According to the spatial coordinate arrangement of the grid nodes, the feature quantities of all grid nodes are filled into a two-dimensional matrix in row-column order, and the difference between the feature quantities of adjacent grid nodes is calculated. If the difference exceeds a set threshold, an interpolation node is inserted in the corresponding row-column gap to supplement the feature quantities of the interpolation nodes to form a continuously distributed two-dimensional data plane; The characteristic quantity of each sampling position in the two-dimensional data plane is mapped into a grayscale value to generate two-dimensional distribution data with spatial coordinates as horizontal and vertical axes and grayscale values representing chemical component concentrations.
6. The method according to claim 1, characterized in that The method comprises alternately projecting the excitation light generated by the dual-wavelength laser source onto the same sampling area, generating a spatial response comparison graph based on the difference in Raman peak intensity obtained at the two wavelengths, and forming a three-dimensional feature stack based on the spatial response comparison graph and the two-dimensional distribution data, including: The output wavelength of the dual-wavelength laser source is alternately switched by a synchronous controller, and a complete spectrum acquisition is performed on the same sampling area at each wavelength; The characteristic peak intensity value of the Raman signal collected at the first wavelength is marked as a first response value, and the characteristic peak intensity value of the Raman signal collected at the second wavelength is marked as a second response value; Calculating the difference between the first response value and the second response value for each spatial sampling point, and mapping the difference to the same spatial coordinate system as the two-dimensional distribution data to generate a spatial response comparison graph; The spatial response comparison map is used as an additional signal layer and is stacked and combined with the two-dimensional distribution data to form a three-dimensional feature stack with three layers of signal intensity.
7. The method according to claim 1, wherein The performing spatial topological analysis on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections includes: A segmentation operation is performed on each signal layer in the three-dimensional feature stack, and continuous areas with signal intensity values higher than a preset threshold are extracted as candidate concentration areas; Performing spatial position comparison on candidate concentration regions in multiple signal layers, and retaining candidate regions whose spatial overlap in at least two signal layers exceeds a set ratio as target concentration regions; Detect the boundary curves of each target concentration area and calculate the minimum spacing and curvature change rate between the boundary curves corresponding to adjacent target concentration areas; When the minimum spacing between boundary curves corresponding to adjacent target concentration regions is less than a preset ratio of the average width of the regions and the curvature change rate exceeds a preset mutation threshold, it is determined that there is an intersection between the boundaries corresponding to adjacent target concentration regions.
8. A structured feature selection system for authentic medicinal materials based on Raman peak intensity information, characterized in that: include: The first generation module is used to perform spatial sampling on the surface area of the medicinal material slice by a scanning device to generate two-dimensional distribution data containing concentration gradients of different chemical components; A second generation module is configured to alternately project excitation light generated by a dual-wavelength laser source onto the same sampling area, generate a spatial response comparison map based on the difference in Raman peak intensities obtained at the two wavelengths, and form a three-dimensional feature stack based on the spatial response comparison map and the two-dimensional distribution data; an identification module for performing spatial topological analysis on the three-dimensional feature stack to identify at least three target concentration regions and their boundary intersections; A third generation module is used to construct a spatial correlation model between the target concentration areas, and analyze the distribution density and direction vector relationship of the boundary intersection points in each target concentration area through the spatial correlation model to generate graph structure data representing the spatial arrangement pattern of the medicinal material components; The output module is used to perform a structured feature selection operation, retaining feature combinations that match the prior knowledge of the spatial distribution of authentic medicinal materials in the graph structure data through constraints, and inputting the screened feature combinations as structured features into the authenticity discrimination model, so as to output the identification results through the discrimination model based on the joint weights of the feature combinations in the spatial dimension and the spectral dimension.
9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the method for selecting structured features of authentic medicinal materials based on Raman peak intensity information as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, the method for selecting structured features of authentic medicinal materials based on Raman peak intensity information according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-parameter nondestructive in-situ detector for cross-border goods
CN113310965A
Self-mixing interferometry
US20250052663A1