Intelligent distinguishing and detecting system for components of phellinus igniarius from different habitats

By using an intelligent differentiation and detection system, which utilizes spectral data analysis and clustering algorithms, the problem of inaccurate regional identification caused by overlapping components of Sanghuang has been solved, and accurate regional identification of Sanghuang samples has been achieved.

CN120253724BActive Publication Date: 2025-12-05HENAN SANSE PIGEON DAIRY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510411407.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-12-05
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

In existing technologies, the main components of Sanghuang may overlap in certain wavelength ranges, leading to overlapping spectral signals, which increases the difficulty of component differentiation and results in inaccurate identification of Sanghuang producing areas.

Method used

An intelligent detection system for differentiating components of Phellinus linteus from different origins was adopted, including a data acquisition module, a similarity analysis module, a clustering module, and an origin determination module. By acquiring spectral data, the system analyzes the similarity of spectral shapes and the dissimilarity of the same component content. The origin of Phellinus linteus samples is determined by using the K-means clustering algorithm and principal component analysis.

Benefits of technology

It effectively reduced the influence of component overlap in the spectral data of Sanghuang samples, improved the accuracy of Sanghuang production area differentiation, and ensured accurate clustering and matching of samples with unknown production areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120253724B_ABST
    Figure CN120253724B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent sensor, in particular to a kind of different origin Phellinus igniarius component intelligent distinguishing detection system, the system includes: data acquisition module, for obtaining the spectral data of several origin unknown Phellinus igniarius samples and known origin representative Phellinus igniarius sample;Similarity analysis module, for determining the spectral shape similarity between the spectral data of two different Phellinus igniarius samples and the same component content dissimilarity;Clustering module, for determining distance index according to spectral shape similarity and same component content dissimilarity, based on distance index, the clustering cluster of the spectral data of all origin unknown Phellinus igniarius samples is obtained;Origin determination module, for matching the spectral data in each clustering cluster with the spectral data of all known origin representative Phellinus igniarius sample based on distance index, and determining the origin of the Phellinus igniarius sample in each clustering cluster according to matching result.The present application effectively improves the accuracy of Phellinus igniarius origin identification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent sensor, in particular to a system for intelligently distinguishing and detecting components of Phellinus igniarius from different producing areas. BACKGROUND

[0002] Phellinus igniarius is a medicinal fungus, and the components of Phellinus igniarius from different producing areas may not only differ in types, but also differ in contents even if the types are the same. Therefore, intelligently distinguishing and detecting the producing areas of Phellinus igniarius is crucial for ensuring the efficacy and quality of Phellinus igniarius.

[0003] In the prior art, the method for distinguishing Phellinus igniarius from different producing areas by using spectrum detection is a rapid and non-destructive analysis method. However, since the main components of Phellinus igniarius such as polysaccharides, flavonoids and triterpenoids may overlap in certain wavelength ranges, that is, the spectrum signals of different components overlap, which increases the difficulty of component distinction, and causes inaccurate comparison and analysis between the spectrum graphs of Phellinus igniarius, thereby leading to inaccurate distinction of the producing areas of Phellinus igniarius. SUMMARY

[0004] In order to solve the technical problem of inaccurate distinction of the producing areas of Phellinus igniarius due to the overlapping of spectrum signals of different components, the present application aims to provide a system for intelligently distinguishing and detecting components of Phellinus igniarius from different producing areas, and the technical solution adopted is as follows:

[0005] In the first aspect, the present application provides a system for intelligently distinguishing and detecting components of Phellinus igniarius from different producing areas, and the system comprises:

[0006] a data acquisition module, configured to acquire spectrum data of Phellinus igniarius samples from unknown producing areas and representative Phellinus igniarius samples from known producing areas, wherein the spectrum data comprises absorbance at different wavelengths within a set wavelength range;

[0007] a similarity analysis module, configured to analyze the spectrum data to determine the spectral shape similarity and the same-component content dissimilarity between the spectrum data of any two different Phellinus igniarius samples after reducing the influence of component overlap;

[0008] a clustering module, configured to determine a distance index between the spectrum data of any two different Phellinus igniarius samples according to the spectral shape similarity and the same-component content dissimilarity after reducing the influence of component overlap, and to cluster the spectrum data of all Phellinus igniarius samples from unknown producing areas based on the distance index to obtain a plurality of clustering clusters;

[0009] a producing area determination module, configured to match the spectrum data in each clustering cluster with the spectrum data of all representative Phellinus igniarius samples from known producing areas based on the distance index, and to determine the producing area of the Phellinus igniarius sample corresponding to each clustering cluster according to the matching result.

[0010] With reference to the first aspect, in some possible implementation manners, the similarity analysis module comprises a first analysis module, and the first analysis module comprises:

[0011] a first wavelength segment acquisition unit, configured to acquire a plurality of first wavelength segments corresponding to wavelength ranges in which spectral shapes of spectral data of any two different Phellinus igniarius samples are different;

[0012] a principal component analysis unit, configured to acquire first spectral data segments and second spectral data segments of the spectral data of any two different Phellinus igniarius samples in each first wavelength segment respectively, and perform principal component analysis on the first spectral data segments and the second spectral data segments to acquire a plurality of principal components and characteristic values of the principal components of the first spectral data segments and the second spectral data segments;

[0013] a matching analysis unit, configured to perform one-to-one matching on the principal components of the first spectral data segments and the second spectral data segments, and determine spectral shape similarity between the spectral data of any two different Phellinus igniarius samples after reducing the influence of component overlap according to a matching result and the characteristic values of the principal components.

[0014] With reference to the first aspect, in some possible implementation manners, the first wavelength segment acquisition unit comprises:

[0015] a difference sequence acquisition unit, configured to determine a spectral shape difference sequence of any two different Phellinus igniarius samples according to a difference in variation trends of absorbance of the spectral data of the any two different Phellinus igniarius samples at the same wavelength, the spectral shape difference sequence reflecting the difference in variation trends of the spectral data of the any two different Phellinus igniarius samples at the same wavelength;

[0016] a first clustering unit, configured to cluster target elements in the spectral shape difference sequence to obtain a plurality of clusters, and if variation trends of the spectral data of any two different Phellinus igniarius samples at a wavelength are the same, an element corresponding to the wavelength in the spectral shape difference sequence is taken as a target element;

[0017] the first wavelength segment acquisition unit, configured to take a wavelength range formed by a maximum wavelength and a minimum wavelength corresponding to the target elements in each cluster as a first wavelength segment.

[0018] With reference to the first aspect, in some possible implementation manners, the difference sequence acquisition unit comprises:

[0019] a wavelength label obtaining unit, configured to determine a wavelength label corresponding to each wavelength in the spectral data, set the wavelength label corresponding to a certain wavelength as a first value if absorbance at the certain wavelength is less than absorbance at a next wavelength, set the wavelength label corresponding to the certain wavelength as a second value if absorbance at the certain wavelength is equal to absorbance at the next wavelength, and set the wavelength label corresponding to the certain wavelength as a third value if absorbance at the certain wavelength is greater than absorbance at the next wavelength;

[0020] a first sequence obtaining unit, configured to arrange all the wavelength labels in a wavelength order to obtain a spectral shape sequence corresponding to the spectral data;

[0021] a second sequence obtaining unit, configured to perform an AND operation on wavelength labels at the same wavelength in the spectral shape sequences corresponding to the spectral data of any two different Phellinus baumii samples to obtain a spectral shape difference sequence of the any two different Phellinus baumii samples.

[0022] In combination with the first aspect, in some possible implementation manners, the matching analysis unit includes:

[0023] a matching unit, configured to perform one-to-one matching on the principal components of the first spectral data segment and the second spectral data segment to obtain a plurality of one-to-one matching pairs and matching edge values of the one-to-one matching pairs;

[0024] a same-component-degree obtaining unit, configured to determine a same-component-degree between the first spectral data segment and the second spectral data segment according to a quantity difference between the principal components of the first spectral data segment and the second spectral data segment, the matching edge values of the one-to-one matching pairs, and eigenvalues of the two principal components in the one-to-one matching pairs;

[0025] a spectral shape similarity obtaining unit, configured to determine a spectral shape similarity between the spectral data of any two different Phellinus baumii samples after reducing the influence of component overlap according to a wavelength range proportion of the first wavelength segments corresponding to the first spectral data segment and the second spectral data segment in a set wavelength range, the same-component-degree between the first spectral data segment and the second spectral data segment, and a wavelength range proportion of a total wavelength range other than all the first wavelength segments in the set wavelength range in the set wavelength range.

[0026] In combination with the first aspect, in some possible implementation manners, the similarity analysis module includes a second analysis module, and the second analysis module includes:

[0027] a second wavelength segment obtaining unit, configured to obtain a plurality of second wavelength segments corresponding to a wavelength range in which the spectral shapes of the spectral data of any two different Phellinus baumii samples are the same in a set wavelength range;

[0028] The first content difference acquisition unit is configured to acquire third spectral data segments and fourth spectral data segments of the spectral data of any two different Phellinus samples in each second wavelength segment, and determine the first same-component content difference and the first maximum-component content difference according to absorbance differences at the same wavelengths in the third spectral data segments and the fourth spectral data segments.

[0029] The second content difference acquisition unit is configured to determine the second same-component content difference and the second maximum-component content difference according to absorbance differences at the same wavelengths in the one-to-one matching pairs of the first spectral data segments and the second spectral data segments.

[0030] The dissimilarity acquisition unit is configured to determine the same-component content dissimilarity between the spectral data of any two different Phellinus samples according to all the first same-component content differences and the first maximum-component content differences and all the second same-component content differences and the second maximum-component content differences corresponding to the two Phellinus samples.

[0031] In combination with the first aspect, in some possible implementation manners, the first content difference acquisition unit comprises:

[0032] The first same-component content difference acquisition unit is configured to determine a sum of absolute values of absorbance differences at the same wavelengths in the third spectral data segments and the fourth spectral data segments, to obtain the first same-component content difference.

[0033] The first maximum-component content difference acquisition unit is configured to determine a maximum value of absolute values of absorbance differences at the same wavelengths in the third spectral data segments and the fourth spectral data segments, to obtain the first maximum-component content difference.

[0034] In combination with the first aspect, in some possible implementation manners, the dissimilarity acquisition unit comprises:

[0035] The comprehensive same-component content difference determination unit is configured to determine a sum of all the first same-component content differences and the second same-component content differences corresponding to any two different Phellinus samples, to obtain a comprehensive same-component content difference.

[0036] The maximum-component content difference variance determination unit is configured to determine a sum of variances and a maximum value of all the first maximum-component content differences and the second maximum-component content differences corresponding to any two different Phellinus samples, to obtain a maximum-component content difference variance and a maximum-component content difference maximum value.

[0037] The dissimilarity determination unit is configured to determine the same-component content dissimilarity between the spectral data of any two different Phellinus samples according to the comprehensive same-component content difference, the maximum-component content difference variance, and the maximum-component content difference maximum value.

[0038] In combination with the first aspect, in some possible implementation manners, the clustering module comprises:

[0039] a negative correlation mapping processing unit configured to perform negative correlation mapping processing on the spectral shape similarity to obtain a spectral shape similarity processing value;

[0040] a distance index obtaining unit configured to calculate a product of the spectral shape similarity processing value and the same-component content dissimilarity as a distance index;

[0041] a second clustering unit configured to perform clustering on the spectral data of all the unknown-origin Phellinus igniarius samples based on the distance index by using a K-means clustering algorithm to obtain a plurality of clustering clusters.

[0042] In some possible implementation manners of the first aspect, the origin determination module comprises:

[0043] a first distance index obtaining unit configured to determine an average value of the similarity index between each clustering cluster and the spectral data of the representative Phellinus igniarius sample of each same origin to obtain an average distance index;

[0044] a second distance index obtaining unit configured to determine a minimum value in the average distance index between each clustering cluster and the spectral data of the representative Phellinus igniarius sample of different origins to obtain a minimum average distance index;

[0045] an origin determination unit configured to determine the origin corresponding to the minimum average distance index as the origin of the Phellinus igniarius sample corresponding to each clustering cluster.

[0046] In a second aspect, the present application further provides a method for intelligently distinguishing and detecting different-origin Phellinus igniarius components, which comprises the following steps:

[0047] obtaining spectral data of a plurality of unknown-origin Phellinus igniarius samples and representative Phellinus igniarius samples of known origins, the spectral data comprising absorbance at different wavelengths within a set wavelength range;

[0048] analyzing the spectral data to determine spectral shape similarity and same-component content dissimilarity between spectral data of any two different Phellinus igniarius samples after reducing the influence of component overlap;

[0049] determining a distance index between spectral data of any two different Phellinus igniarius samples according to the spectral shape similarity and the same-component content dissimilarity after reducing the influence of component overlap, and performing clustering on the spectral data of all the unknown-origin Phellinus igniarius samples based on the distance index to obtain a plurality of clustering clusters;

[0050] based on the distance index, matching the spectral data in each clustering cluster with the spectral data of the representative Phellinus igniarius samples of all the known origins, and determining the origin of the Phellinus igniarius sample corresponding to each clustering cluster according to the matching result.

[0051] In a third aspect, the present application further provides a device for intelligently distinguishing and detecting components of Phellinus igniarius from different producing areas, comprising a memory, a processor, and computer program code stored in the memory, the processor being configured to call and run the computer program code from the memory, so that the device executes the method in the first aspect or any possible implementation manner of the first aspect.

[0052] In a fourth aspect, the present application further provides a computer program product, which comprises computer program code, when the computer program code is run on a computer, so that the computer executes the method in the first aspect or any possible implementation manner of the first aspect.

[0053] In a fifth aspect, the present application further provides a computer readable storage medium, which stores computer program code, when the computer program code is run on a computer, so that the computer executes the method in the first aspect or any possible implementation manner of the first aspect.

[0054] The present application has the following beneficial effects: the present application obtains spectral data of a plurality of Phellinus igniarius samples from unknown producing areas and representative Phellinus igniarius samples from known producing areas, analyzes the spectral data, determines the spectral shape similarity and the in-component content dissimilarity between the spectral data of any two different Phellinus igniarius samples after reducing the influence of the overlapping of spectral signals of different components, and then accurately determines the distance index between the spectral data of any two different Phellinus igniarius samples, so as to realize accurate clustering of the spectral data of all Phellinus igniarius samples from unknown producing areas based on the distance index, and obtain a plurality of clustering clusters. Then, based on the distance index, the spectral data in each clustering cluster is accurately matched with the spectral data of all representative Phellinus igniarius samples from known producing areas, and the producing area of the Phellinus igniarius sample corresponding to each clustering cluster is determined according to the matching result. The present application effectively reduces the influence of the overlapping of spectral signals of different components in the spectral data of the Phellinus igniarius sample, ensures the classification accuracy of the spectral data of the Phellinus igniarius sample, and finally accurately matches the clustering cluster of the Phellinus igniarius sample from unknown producing areas with the spectral data of the Phellinus igniarius sample from known producing areas, thereby effectively improving the accuracy of the Phellinus igniarius producing area distinction. BRIEF DESCRIPTION OF DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, and the advantages thereof, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0056] Figure 1 It is a structural schematic diagram of a different producing area Phellinus igniarius component intelligent distinguishing and detecting system according to an embodiment of the present application.

[0057] Figure 2 A step flow chart of a different origin Phellinus igniarius component intelligent differentiation and detection method of an embodiment of the present application;

[0058] Figure 3 A schematic diagram of absorbance of a Phellinus igniarius sample 1 of an embodiment of the present application;

[0059] Figure 4 A schematic diagram of absorbance of a Phellinus igniarius sample 2 of an embodiment of the present application;

[0060] Figure 5 A structural schematic diagram of a different origin Phellinus igniarius component intelligent differentiation and detection device of an embodiment of the present application. DETAILED DESCRIPTION

[0061] To clearly illustrate the technical features of the present scheme, the present application will be described in detail below with specific embodiments and in conjunction with the accompanying drawings.

[0062] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, rather these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are for exemplary purposes only, and are not intended to limit the scope of protection of the present application.

[0063] It should be understood that each step described in the method embodiments of the present application can be performed in different order, and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0064] The term "comprising" and variations thereof as used herein are open-ended, that is "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.

[0065] It should be noted that the concepts of "first", "second", etc. mentioned in the present application are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0066] Although the operations or steps are described in a particular order in the drawings, it should not be understood that the order described is the only order in which the operations or steps can be performed, or that all illustrated operations or steps are necessary for all embodiments. The operations or steps can be performed in serial, in parallel, or some combination of the two. Some operations or steps can not be required in all embodiments.

[0067] Meanwhile, it can be understood that the data involved in the technical solutions of the present application (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and relevant provisions. Unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as understood by those skilled in the art to which the present application belongs, and all parameters or indicators in the formulas involved in the present application are normalized values that eliminate the influence of dimensions.

[0068] In order to solve the technical problem of inaccurate differentiation of Phellinus baumii due to overlapping of spectral signals of different components, the embodiments of the present application provide a kind of intelligent differentiation detection system for different components of Phellinus baumii from different producing areas. The system is essentially a software system, which is composed of various modules that realize corresponding functions, and the corresponding structural diagram is as shown in Figure 1 The core of the system is to realize an intelligent differentiation detection method for different components of Phellinus baumii from different producing areas. Each module in the system corresponds to each step in the method, and the corresponding flow chart is as shown in Figure 2 The modules of the system will be described in detail below in combination with the specific steps in the method.

[0069] The data acquisition module is used to acquire spectral data of a plurality of Phellinus baumii samples from unknown producing areas and representative Phellinus baumii samples from known producing areas, and the spectral data includes absorbance at different wavelengths within a set wavelength range.

[0070] Specifically, the producing areas of Phellinus baumii are distributed in different provinces and regions. In order to intelligently differentiate and detect the producing areas of different Phellinus baumii, a plurality of Phellinus baumii samples from unknown producing areas are first acquired, for example, 30 Phellinus baumii samples from unknown producing areas can be acquired. There are many reasons for the unknown producing areas of Phellinus baumii, for example, Phellinus baumii may be transferred several times, and the original producing area information may be lost or tampered with during the transfer process, or the label of Phellinus baumii may be incorrectly marked due to negligence or misunderstanding, resulting in inaccurate producing area information.

[0071] The specific process of acquiring a plurality of Phellinus baumii samples from unknown producing areas is as follows: (1) collecting samples: collecting representative samples from all Phellinus baumii from unknown producing areas; (2) pretreatment: appropriately pretreating the Phellinus baumii samples, such as drying, crushing, sieving, etc., to ensure the consistency and representativeness of the samples.

[0072] Then, the spectral data of the above several unknown origin Phellinus samples are obtained. The specific process of obtaining the spectral data is as follows: (1) selecting a suitable spectrometer: in this embodiment, an infrared spectrometer is selected. (2) setting instrument parameters: the set wavelength range of the infrared spectrum (IR) is 4000-400 wavenumers (cm-1), which is used to detect the vibration absorption of functional groups such as hydroxyl and carbonyl in Phellinus. (3) sample loading: placing the pretreated Phellinus sample on the sample cell of the spectrometer. (4) spectral scanning: starting the spectrometer to scan the Phellinus sample and recording the absorbance at different wavelengths.

[0073] The spectral data obtained from the above several unknown origin Phellinus samples include infrared spectra of different Phellinus samples from the same origin and different origins. The horizontal axis of the infrared spectrum is the wavelength, and the vertical axis is the absorbance.

[0074] In the same way as described above for obtaining the spectral data of the several unknown origin Phellinus samples, the spectral data of at least one representative Phellinus sample from each origin of all Phellinus origins can be obtained, and the origin of each representative Phellinus sample is marked.

[0075] The similarity analysis module is used to analyze the spectral data and determine the spectral shape similarity and the same component content dissimilarity between the spectral data of any two different Phellinus samples after reducing the overlapping influence of components.

[0076] Specifically, due to the influence of geographical environment, climate conditions, soil types, cultivation techniques and other factors, the components of Phellinus from different origins may not only differ in type, but even if the component types are the same, their contents may also differ significantly. Therefore, in order to distinguish the origins of the above several unknown origin Phellinus samples, it is necessary to first distinguish the Phellinus samples with different components, and then distinguish the Phellinus samples with the same components but different contents.

[0077] For the infrared spectrum, the size of the absorbance reflects the content of the components in the sample. The larger the absorbance, the higher the content of the component in the sample. At the same time, the shape of the spectral curve can be used to identify specific components in the sample. Different chemical substances will have different absorption characteristics at a specific wavelength. If the spectral curve shapes of two samples are similar, they are likely to contain the same components.

[0078] Since the main components of Sanghuang such as polysaccharides, flavonoids and triterpenoids may overlap in certain wavelength ranges, which may cause different shapes of spectral curves in the overlapping wavelength range with different contents of the same component, and it is easy to be mistaken as different components. For ease of understanding, for example: in the overlapping wavelength range [1001, 1004] of components A and B, the content of component A of any two different Sanghuang samples 1 and 2 is the same, and the absorbance is {5, 10, 11, 12}, that is, the spectral curve shape of component A is the same, and the absorbance is the same, while the content of component B of Sanghuang samples 1 and 2 is different, and the absorbance is {1, 2, 1, 2} and {1, 4, 1, 4}, that is, the spectral curve shape of component B is the same, and the single absorbance is different, then the absorbance of Sanghuang samples 1 and 2 is {6, 12, 12, 14} and {6, 14, 12, 16} respectively. At this time, the schematic diagram of the absorbance of Sanghuang samples 1 and 2 is shown in Figure 3 and Figure 4 According to Figure 3 and Figure 4 It can be seen that the spectral curve shapes of Sanghuang samples 1 and 2 with the same component but different contents are different in the overlapping wavelength range of components A and B. Therefore, it is necessary to first divide the range with different curve shapes in the infrared spectrum, and further analyze whether the range is caused by different components or by the same component but with overlapping components and different contents.

[0079] Further, the above similarity analysis module includes a first analysis module, and the first analysis module includes: a first wavelength segment acquisition unit, configured to acquire a plurality of first wavelength segments corresponding to wavelength ranges with different spectral shapes of spectral data of any two different Sanghuang samples in a set wavelength range; a principal component analysis unit, configured to acquire first spectral data segments and second spectral data segments of the spectral data of any two different Sanghuang samples in each first wavelength segment, and perform principal component analysis on the first spectral data segments and the second spectral data segments to acquire a plurality of principal components and characteristic values of the first spectral data segments and the second spectral data segments; a matching analysis unit, configured to one-to-one match the principal components of the first spectral data segments and the second spectral data segments, and determine the spectral shape similarity between the spectral data of any two different Sanghuang samples after reducing the influence of component overlap according to the matching result and the characteristic values of the principal components.

[0080] Specifically, for any two different Phellinus baumii samples 1 and 2 (the Phellinus baumii samples 1 and 2 can be Phellinus baumii samples of unknown origin or representative Phellinus baumii samples of known origin), the spectral shape difference between the two spectral data of the two samples can be analyzed to determine a plurality of first wavelength ranges corresponding to the wavelength ranges in which the spectral shapes of the two spectral data are different. Further, in some possible implementation manners, the first wavelength range obtaining unit includes: a difference sequence obtaining unit, configured to determine a spectral shape difference sequence of any two different Phellinus baumii samples according to the difference in the change trend of the absorbance of the spectral data of the two samples at the same wavelength, the spectral shape difference sequence reflecting the difference in the change trend of the spectral data of the two samples at the same wavelength; a first clustering unit, configured to cluster target elements in the spectral shape difference sequence to obtain a plurality of clusters, and if the change trend of the spectral data of any two different Phellinus baumii samples at a certain wavelength is the same, the element corresponding to the wavelength in the spectral shape difference sequence is taken as a target element; and a first wavelength range obtaining unit, configured to take a wavelength range formed by the maximum wavelength and the minimum wavelength corresponding to the target elements in each cluster as a first wavelength range.

[0081] In order to obtain the plurality of first wavelength ranges corresponding to the wavelength ranges in which the spectral shapes of the two spectral data of any two different Phellinus baumii samples 1 and 2 are different, first, the change trend corresponding to each wavelength in each spectral data can be determined according to the change of the absorbance at adjacent wavelengths in the spectral data, and then the change trends corresponding to the same wavelength of the two spectral data are compared, so as to obtain the spectral shape difference sequence corresponding to the two spectral data. Finally, the wavelength range is filtered based on the spectral shape difference sequence, that is, each wavelength range corresponding to the wavelength ranges in which the spectral shapes of the two spectral data of any two different Phellinus baumii samples 1 and 2 are different can be determined.

[0082] Further, in some possible implementation manners, the difference sequence obtaining unit includes: a wavelength label obtaining unit, configured to determine a wavelength label corresponding to each wavelength in the spectral data, if the absorbance at a certain wavelength in the spectral data is less than the absorbance at the next wavelength, the wavelength label corresponding to the wavelength is set to a first value, if the absorbance at a certain wavelength is equal to the absorbance at the next wavelength, the wavelength label corresponding to the wavelength is set to a second value, and if the absorbance at a certain wavelength is less than the absorbance at the next wavelength, the wavelength label corresponding to the wavelength is set to a third value; a first sequence obtaining unit, configured to arrange all the wavelength labels in the order of wavelengths to obtain a spectral shape sequence corresponding to the spectral data; and a second sequence obtaining unit, configured to perform an exclusive OR operation on the wavelength labels at the same wavelength in the spectral shape sequences corresponding to the spectral data of any two different Phellinus baumii samples to obtain a spectral shape difference sequence of the two samples.

[0083] In this embodiment, for any sample of *Phellinus linteus*, i.e., its infrared spectrum, the absorbance at the i-th and (i+1)-th wavelengths is obtained. and ;when When the i-th wavelength is assigned a wavelength label with a value of 1, it indicates that the shape of the spectral curve is rising; when When the i-th wavelength is assigned a wavelength label with a value of 0, it indicates that the shape of the spectral curve remains unchanged; when When the i-th wavelength is assigned a wavelength label with a value of -1, it indicates that the spectral curve shape is decreasing. In this way, wavelength labels corresponding to all wavelengths in the spectral data of any Sanghuang sample can be obtained. Arranging all wavelength labels in ascending order of wavelengths in the spectral data yields the spectral shape sequence corresponding to the spectral data of any Sanghuang sample. It should be understood that for the last wavelength in the spectral data of a Sanghuang sample, since there is no subsequent wavelength, a wavelength label can be left unassigned, or the label of the second-to-last wavelength can be directly assigned to the last wavelength.

[0084] Let the spectral data of any two different Sanghuang samples 1 and 2 be denoted as A1 and A2, respectively. In the spectral shape sequences of spectral data A1 and A2, if the values ​​at the same index (i.e., the element values ​​corresponding to the same wavelength) are the same (i.e., the spectral curve shapes are the same), then the label value corresponding to that index is recorded as 1; if the values ​​at the same index (i.e., the element values ​​corresponding to the same wavelength) are different (i.e., the spectral curve shapes are different), then the label value corresponding to that index is recorded as 0. Arrange all the label values ​​in ascending order of index, and thus obtain the spectral shape difference sequence of any two different Sanghuang samples 1 and 2. For example, if the spectral shape sequence of spectral data A1 is {0, 0, 1} and the spectral shape sequence of spectral data A2 is {-1, 0, 1}, then the spectral shape difference sequence corresponding to spectral data A1 and A2 is {0, 1, 1}.

[0085] For the spectral shape difference sequence corresponding to spectral data A1 and A2, since the element 0 in the sequence indicates that the spectral shapes corresponding to spectral data A1 and A2 are the same, that is, they have the same trend of change at a certain wavelength, the wavelength interval between any two 0s in the sequence (the wavelength interval between adjacent 0s is 0) is used as the clustering distance. The K-means clustering algorithm is used to cluster all 0s in the spectral shape difference sequence, thereby clustering 0s with smaller wavelength intervals into one class, obtaining multiple clusters. The K value in the K-means clustering algorithm can be obtained using the elbow method, and the specific acquisition process is a well-known technique and will not be elaborated here.

[0086] A minimum wavelength corresponding to all 0s in any one cluster class to a maximum wavelength to form a wavelength range with a different curve shape, thereby obtaining a plurality of wavelength ranges with different curve shapes, and each wavelength range is taken as a first wavelength segment. It should be understood that if there is only one 0 in a certain cluster class, the cluster class is not analyzed. Further, for each first wavelength segment, the first spectral data segment and the second spectral data segment of the spectral data A1 and A2 in each first wavelength segment can be obtained, and are denoted as A11 and A22 respectively.

[0087] The PCA principal component analysis method is used to perform principal component analysis on the first spectral data segment A11 and the second spectral data segment A22 respectively, to obtain a plurality of principal components decomposed from the first spectral data segment A11 and a plurality of principal components decomposed from the second spectral data segment A22, and a characteristic value corresponding to each principal component. The greater the characteristic value, the more information and data changes the principal component contains. In infrared spectrum analysis, this means that the principal component may be closely related to the main chemical composition or structural feature in the sample. For infrared spectrum data with three component overlaps, three principal components can be obtained in theory after decomposition by the PCA principal component analysis method. Each spectral data segment and its decomposed principal components have the same length.

[0088] The principal components of the first spectral data segment A11 and the second spectral data segment A22 are matched one by one to obtain a matching result, and then the spectral shape similarity between the spectral data of any two different Sanghuang samples after reducing the influence of component overlap is determined according to the matching result and the characteristic value of the principal component. Further, in some possible implementation manners, the matching analysis unit includes: a matching unit, configured to match the principal components of the first spectral data segment and the second spectral data segment one by one to obtain a plurality of one-to-one matching pairs and matching edge values thereof; a component same degree acquisition unit, configured to determine the component same degree between the first spectral data segment and the second spectral data segment according to the quantity difference between the principal components of the first spectral data segment and the second spectral data segment, the matching edge values of the one-to-one matching pairs, and the characteristic values of the two principal components in the one-to-one matching pairs; and a spectral shape similarity acquisition unit, configured to determine the spectral shape similarity between the spectral data of any two different Sanghuang samples after reducing the influence of component overlap according to the wavelength range proportion of the first wavelength segment corresponding to each first spectral data segment and the second spectral data segment in the set wavelength range, the component same degree between the first spectral data segment and the second spectral data segment, and the wavelength range proportion of the total wavelength range in the set wavelength range except all the first wavelength segments.

[0089] In the embodiment of the present application, for the first spectral data segment A11 of the spectral data A1 and A2 in each first wavelength segment

[0090] For the second spectral data segment A22, using the principal components corresponding to A11 as left nodes and the principal components corresponding to A22 as right nodes, the Kuhn-Munkres matching algorithm is used to perform one-to-one matching between the left and right nodes, obtaining several one-to-one matching pairs and the matching edge value of each pair. Each matching pair contains one left node and one right node, and each matching pair is unique; there are no one-to-many cases. The Kuhn-Munkres matching algorithm, also known as the Hungarian algorithm, is an efficient algorithm for solving assignment problems. This is a well-known technique. Specifically, each left node is connected to all right nodes by an edge. The matching edge value on the connecting edge is the Pearson correlation coefficient of the two principal components corresponding to the two nodes. The Pearson correlation coefficient is a well-known calculation method; the larger the correlation coefficient, the more similar the changing trends of the two principal components. Then, through Kuhn-Munkres matching, following the principle of maximum matching, one-to-one matching pairs between left and right nodes are obtained. A one-to-one matching pair indicates that the components decomposed in A11 are highly likely to be the same as those decomposed in A22.

[0091] Therefore, based on the quantitative differences between the principal components of the first spectral data segment A11 and the second spectral data segment A22 in each first wavelength band, the matching boundary values ​​of one-to-one matching pairs, and the eigenvalues ​​of the two principal components in each one-to-one matching pair, the degree of component similarity between the first spectral data segment A11 and the second spectral data segment A22 in each first wavelength band of spectral data A1 and A2 is determined:

[0092] ;

[0093] In the formula: For any two different Sanghuang samples 1 and 2, the spectral data A1 and A2 are in the first... The degree of compositional similarity between the first spectral data segment A11 and the second spectral data segment A22 under the first wavelength band; The spectral data A1 of Phellinus linteus sample 1 are in The number of principal components decomposed in the first spectral data segment A11 under the first wavelength band. For the spectral data A2 of Phellinus linteus sample 2 The number of principal components decomposed in segment A22 of the second spectral curve under the first wavelength band; It is an exponential function with the natural constant e as its base; For any two different Sanghuang samples 1 and 2, the spectral data A1 and A2 are in the first... The number of one-to-one matching pairs between the first spectral data segment A11 and the second spectral data segment A22 in the first wavelength band; Let be the matching edge value of the t-th one-to-one matching pair, that is, the Pearson correlation coefficient between the principal components corresponding to the left and right nodes in the t-th one-to-one matching pair; and , respectively, are the eigenvalues ​​of the principal components corresponding to the left and right nodes in the t-th one-to-one matching pair.

[0094] In the above formula, when When the value is 0, it indicates that the number of overlapping chemical components in the first spectral data segment A11 and the second spectral data segment A22 is the same. The larger the value of the degree of component similarity between the first spectral data segment A11 and the second spectral data segment A22, the more accurate the result will be. Therefore, the exponential function can be used to... right Negative correlation normalization was performed. Meanwhile, in each one-to-one matching pair between the left and right nodes of the principal components in the first spectral data segment A11 and the second spectral data segment A22, the larger the matching edge weight, the more similar the changing trends of the two principal components. Conversely, the larger the eigenvalues ​​of the two principal components, the closer the principal components are to the main chemical components in the *Phellinus linteus* sample. Therefore, [the following is used as a reference:] The matching edge weights of the two principal components The weights are then calculated, and the weights of the matching edges between the two principal components in all one-to-one matching pairs are summed in a weighted manner to obtain the weighted cumulative value of the matching edge weights. The weighted matching edge weight accumulation value The larger the value, the more consistent the overlapping chemical compositions are between the first spectral data segment A11 and the second spectral data segment A22.

[0095] Furthermore, for the spectral data A1 and A2 of any two different Sanghuang samples 1 and 2, the similarity of the spectral shape after reducing the influence of component overlap between the spectral data A1 and A2 of any two different Sanghuang samples 1 and 2 is determined based on the proportion of each first wavelength segment within the set wavelength range, the degree of component similarity between the first spectral data segment A11 and the second spectral data segment A22 of spectral data A1 and A2 in each first wavelength segment, and the proportion of the total wavelength range excluding all first wavelength segments within the set wavelength range. :

[0096] ;

[0097] In the formula: The wavelength range for the spectral data is set, which is the size of the wavelength range of the infrared spectrum (4000-400). Let A1 and A2 be the number of the first wavelength bands in the spectral data of any two different Sanghuang samples 1 and 2; For any two different Sanghuang samples 1 and 2, the spectral data A1 and A2 are... The wavelength range size of the first wavelength band; For any two different Sanghuang samples 1 and 2, the spectral data A1 and A2 are in the first... a same component degree between the first spectral data segment A11 and the second spectral data segment A22 under the first wavelength segment; to set a total wavelength range size in the wavelength range except all the first wavelength segments; to be a linear normalization function.

[0098] In the above formula, the same component degree is greater, the spectral data A1 and A2 under the wavelength range of the first wavelength segment are more similar in shape after reducing the influence of the same component, that is The wavelength range should be the wavelength range with the same spectral shape, so the closer the spectral shape similarity H is to 1, the more similar the spectral data A1 and A2 of any two different Phellinus baumii samples 1 and 2 are in shape.

[0099] The above analyzes the spectral data shape similarity after reducing the influence of the same component, and due to the influence of factors such as geographical environment, climate condition, soil type, and cultivation technology, the content of the same component of Phellinus baumii in different producing areas may have significant differences. In addition, even if the Phellinus baumii in the same producing area, the content of the same component of Phellinus baumii may also have certain differences, but the difference is smaller than that of different producing areas. Therefore, it is also necessary to analyze the spectral data of different Phellinus baumii samples to determine the difference between the same component content between the spectral data of any two different Phellinus baumii samples.

[0100] Further, in some possible implementation manners, the similarity analysis module further includes a second analysis module, and the second analysis module includes: a second wavelength segment acquisition unit, configured to acquire a plurality of second wavelength segments corresponding to the wavelength range with the same spectral shape of the spectral data of any two different Phellinus baumii samples in the set wavelength range; a first content difference acquisition unit, configured to acquire third spectral data segments and fourth spectral data segments of the spectral data of any two different Phellinus baumii samples under each second wavelength segment respectively, and determine a first same component content difference and a first maximum component content difference according to the absorbance difference of the same wavelength in the third spectral data segment and the fourth spectral data segment; a second content difference acquisition unit, configured to determine a second same component content difference and a second maximum component content difference according to the absorbance difference of the same wavelength in the one-to-one matching pair of the first spectral data segment and the second spectral data segment; and an dissimilarity acquisition unit, configured to determine the same component content dissimilarity between the spectral data of any two different Phellinus baumii samples according to all the first same component content differences and the first maximum component content differences and the second same component content differences and the second maximum component content differences corresponding to any two different Phellinus baumii samples.

[0101] ​Specifically, for the spectral data Al and A2 of any two different Phellinus igniarius samples 1 and 2, a plurality of second wavelength segments corresponding to the wavelength range in which the spectral shapes of the spectral data Al and A2 in the set wavelength range are the same are obtained, that is, the continuous wavelength segments between every two adjacent first wavelength segments in the set wavelength range of the spectral data Al and A2 and the continuous wavelength segments before the first first wavelength segment and after the last first wavelength segment are taken as the respective second wavelength segments of the spectral data Al and A2 in the set wavelength range. It should be understood that if there are also continuous wavelength segments before the first first wavelength segment and after the last first wavelength segment, the continuous wavelength segments are also taken as the second wavelength segments of the spectral data Al and A2. For each second wavelength segment, the third spectral data segment and the fourth spectral data segment of the spectral data Al and A2 in each second wavelength segment can be obtained and are denoted as A33 and A44, respectively.

[0102] Further, according to the absorbance difference of the third spectral data segment A33 and the fourth spectral data segment A44 of the spectral data Al and A2 in each second wavelength segment at the same wavelength, the first same-component content difference and the first maximum-component content difference of the spectral data Al and A2 in each second wavelength segment are determined. In some possible implementation manners, the first content difference obtaining unit further includes a first same-component content difference obtaining unit configured to determine the sum of the absolute values of the absorbance differences of the third spectral data segment and the fourth spectral data segment at the same wavelength to obtain the first same-component content difference, and a first maximum-component content difference obtaining unit configured to determine the maximum value of the absolute values of the absorbance differences of the third spectral data segment and the fourth spectral data segment at the same wavelength to obtain the first maximum-component content difference.

[0103] In this embodiment, for the third spectral data segment A33 and the fourth spectral data segment A44 of the spectral data Al and A2 in the yth second wavelength segment, the absolute values of the absorbance differences at the same wavelength are obtained, the sum of the absolute values of the absorbance differences at the same wavelength is taken as the first same-component content difference of the yth second wavelength segment, and the maximum value of the absolute values of the absorbance differences at the same wavelength is taken as the first maximum-component content difference of the yth second wavelength segment.

[0104] Meanwhile, for each one-to-one matching pair of the principal components decomposed from the first spectral data segment A11 and the second spectral data segment A22 of the spectral data A1 and A2 under each first wavelength segment, according to the absorbance difference of the two principal components under the same wavelength in each one-to-one matching pair, in the same way, the second same-component content difference and the second maximum-component content difference of each one-to-one matching pair are obtained. That is, for each one-to-one matching pair, the absolute value of the difference of the data values under the same sequence number in the principal components corresponding to the left and right nodes is calculated, the sum of the absolute values of the differences of the data values under the same sequence number is taken as the second same-component content difference of each one-to-one matching pair, and the maximum of the absolute values of the differences of the data values under the same sequence number is taken as the second maximum-component content difference of each one-to-one matching pair.

[0105] Further, based on all the first same-component content differences and the first maximum-component content differences and the second same-component content differences and the second maximum-component content differences corresponding to the spectral data A1 and A2, the same-component content dissimilarity between the spectral data A1 and A2 can be determined. Further, in some possible implementation manners, the above dissimilarity obtaining unit comprises: a comprehensive same-component content difference determining unit, configured to determine the sum of all the first same-component content differences and the second same-component content differences corresponding to any two different Sanghuang samples, to obtain a comprehensive same-component content difference; a maximum-component content difference variance determining unit, configured to determine the variance and the maximum value of all the first maximum-component content differences and the second maximum-component content differences corresponding to any two different Sanghuang samples, to obtain a maximum-component content difference variance and a maximum-component content difference maximum value; and a dissimilarity determining unit, configured to determine the same-component content dissimilarity between the spectral data of any two different Sanghuang samples according to the comprehensive same-component content difference, the maximum-component content difference variance and the maximum-component content difference maximum value.

[0106] In the embodiment of the present application, for the spectral data A1 and A2, all the first same-component content differences of the second wavelength segments and all the second same-component content differences under the first wavelength segments form a same-component content difference set; similarly, all the first maximum-component content differences of the second wavelength segments and all the second maximum-component content differences under the first wavelength segments form a maximum-component content difference set. At this time, based on the same-component content difference set and the maximum-component content difference set, the same-component content dissimilarity between the spectral data A1 and A2 is determined.

[0107] ;

[0108] In the formula, the comprehensive same-component content difference is the sum of all the same-component content differences in the same-component content difference set, and the maximum-component content difference variance is the variance of all the maximum-component content differences in the maximum-component content difference set. ​​​a maximum value of all the maximum component content differences in the maximum component content difference set, i.e., a maximum component content difference maximum value; is a linear normalization function.

[0109] In the above formula, The greater, the greater the difference in the same component content between the spectral data A1 and A2. At the same time, since the component content of the same origin Phellinus igniarius may also have some differences, but relatively small, i.e. The smaller, and the difference between the maximum component content differences of different components is also smaller, i.e. The smaller; and the content of Phellinus igniarius of different origins may have significant differences, resulting in The greater, the greater the difference between the maximum component content differences of different components in the spectral data A1 and A2, and there is a component with a large content difference, and the spectral data A1 and A2 corresponding to the Phellinus igniarius are more dissimilar in the same component content. The greater, the greater the difference between the maximum component content differences of different components in the spectral data A1 and A2, and there is a component with a large content difference, and the spectral data A1 and A2 corresponding to the Phellinus igniarius are more dissimilar in the same component content. The greater, the greater the difference between the maximum component content differences of different components in the spectral data A1 and A2, and there is a component with a large content difference, and the spectral data A1 and A2 corresponding to the Phellinus igniarius are more dissimilar in the same component content. The greater, the greater the difference between the maximum component content differences of different components in the spectral data A1 and A2, and there is a component with a large content difference, and the spectral data A1 and A2 corresponding to the Phellinus igniarius are more dissimilar in the same component content. The greater, the greater the difference between the maximum component content differences of different components in the spectral data A1 and A2, and there is a component with a large content difference, and the spectral data A1 and A2 corresponding to the Phellinus igniarius are more dissimilar in the same component content. The greater, the greater the difference between the maximum component content differences of different components in the spectral data A1 and A2, and there is a component with a large content difference, and the spectral data A1 and A2 corresponding to the Phellinus igniarius are more dissimilar in the same component content.

[0110] At this point, according to the above manner, the spectral shape similarity and the same component content dissimilarity between the spectral data of any two different Phellinus igniarius samples after reducing the influence of component overlap can be determined.

[0111] The clustering module is configured to determine a distance index between the spectral data of any two different Phellinus igniarius samples according to the spectral shape similarity and the same component content dissimilarity after reducing the influence of component overlap, and cluster the spectral data of all Phellinus igniarius samples with unknown origins based on the distance index to obtain a plurality of clustering clusters.

[0112] Specifically, based on the spectral shape similarity and the same component content dissimilarity between the spectral data of any two different Phellinus igniarius samples after reducing the influence of component overlap, a distance index between the spectral data of any two different Phellinus igniarius samples is determined. The smaller the value of the spectral shape similarity , and the greater the value of the same component content dissimilarity , the more dissimilar the spectral data of any two different Phellinus igniarius samples, and the greater the corresponding distance index.

[0113] Furthermore, in some possible implementations, the clustering module includes: a negative correlation mapping processing unit, used to perform negative correlation mapping processing on spectral shape similarity to obtain a spectral shape similarity processing value; a distance index acquisition unit, used to calculate the product of the spectral shape similarity processing value and the dissimilarity of the same component content as a distance index; and a second clustering unit, used to cluster the spectral data of all Sanghuang samples with unknown origins based on the distance index using the K-means clustering algorithm to obtain several clusters.

[0114] In this embodiment, the spectral shape similarity between spectral data A1 and A2 of any two different Sanghuang samples 1 and 2 after reducing the influence of component overlap is used. Dissimilarity in content of the same component Determine the distance index between the spectral data A1 and A2 of any two different Sanghuang samples 1 and 2:

[0115] ;

[0116] In the formula: Let A1 and A2 be the distance index between the spectral data of any two different Sanghuang samples 1 and 2. The spectral shape similarity value is obtained by performing a negative correlation mapping on the spectral data shape similarity H of A1 and A2 of the two different Sanghuang samples 1 and 2 after reducing the influence of component overlap. The value is processed using this spectral shape similarity. Dissimilarity of the same component content between spectral data A1 and A2 After adjustments, the final distance index between spectral data A1 and A2 is obtained. The higher the value of the shape similarity H of the spectral data A1 and A2 after reducing the influence of component overlap, the more similar the spectral data A1 and A2 are, and the higher the distance index. The smaller the value, the better; simultaneously, when the content dissimilarity of the same component between spectral data A1 and A2 is... The larger the value, the less similar the spectral data A1 and A2 are, and the greater the distance index. The farther away, the better.

[0117] Using the distance between the spectral data of two different Sanghuang samples from unknown origins as the clustering distance, the K-means clustering algorithm is applied to cluster the spectral data of all Sanghuang samples from unknown origins based on all clustering distances, resulting in multiple clusters. Each cluster corresponds to one or more Sanghuang samples. The K value in the K-means clustering algorithm can be obtained using the elbow method; the specific acquisition process is a well-known technique and will not be elaborated here.

[0118] The origin determination module is configured to match the spectral data in each cluster with the spectral data of the representative Phellinus samples of all known origins based on the distance index, and determine the origin of the Phellinus sample corresponding to each cluster according to the matching result.

[0119] Specifically, the spectral data in each cluster is matched with the spectral data of the representative Phellinus samples of all known origins based on the distance index, the spectral data of the representative Phellinus sample of a certain origin that is most matched with the spectral data in each cluster is determined, and the certain origin is taken as the origin of the Phellinus sample corresponding to each cluster.

[0120] Further, in some possible implementation manners, the origin determination module comprises: a first distance index acquisition unit configured to determine the average value of the similarity indexes between each cluster and the spectral data of the representative Phellinus samples of each same origin, to obtain the average distance index; a second distance index acquisition unit configured to determine the minimum value in the average distance indexes between each cluster and the spectral data of the representative Phellinus samples of different origins, to obtain the minimum average distance index; and an origin determination unit configured to take the origin corresponding to the minimum average distance index as the origin of the Phellinus sample corresponding to each cluster.

[0121] In the embodiment, for the spectral data of the plurality of representative Phellinus samples of the origin A province, the distance index between each representative Phellinus sample and each spectral data in each cluster can be determined in the manner of determining the distance index between the spectral data A1 and A2 of any two different Phellinus samples 1 and 2. Further, the average distance index between each cluster and the spectral data of the representative Phellinus samples of each same origin can be determined by determining the average value of all distance indexes between the spectral data of all representative Phellinus samples and each spectral data in each cluster. The minimum value in all average similarity indexes corresponding to each cluster is determined to obtain the minimum average distance index. The origin corresponding to the minimum average distance index is the most matched origin of each cluster, and thus the origin corresponding to the minimum average distance index is taken as the origin of the Phellinus sample corresponding to each cluster. Therefore, the origin identification of the Phellinus sample with unknown origin can be realized.

[0122] Based on the same inventive concept, as shown in Figure 2 The present embodiment further provides a method for intelligently distinguishing and detecting components of Phellinus samples from different origins, which comprises the following steps:

[0123] Obtaining the spectral data of a plurality of Phellinus samples with unknown origins and representative Phellinus samples of known origins, wherein the spectral data comprises absorbance at different wavelengths in a set wavelength range;

[0124] Spectral data were analyzed to determine the similarity of spectral shapes and the dissimilarity of the same component content between any two different Sanghuang samples after reducing the influence of component overlap.

[0125] Based on the similarity of spectral shape after reducing the influence of component overlap and the dissimilarity of the same component content, the distance index between the spectral data of any two different Sanghuang samples is determined. Based on the distance index, the spectral data of all Sanghuang samples with unknown origins are clustered to obtain several clusters.

[0126] Based on the distance index, the spectral data in each cluster is matched with the spectral data of representative Sanghuang samples from all known origins, and the origin of the Sanghuang sample corresponding to each cluster is determined according to the matching results.

[0127] Based on the same inventive concept, embodiments of the present invention also provide an intelligent detection device for differentiating and detecting components of Phellinus linteus from different origins, such as... Figure 5 As shown, the device includes: a memory 501, a processor 502, and computer program code 503 stored in the memory 501 and running on the processor 502. When the processor 502 executes the computer program code 503, the device can perform the module implementation steps of any of the intelligent differentiation and detection systems for different origins of Phellinus linteus that were described above.

[0128] In this embodiment of the invention, the device can be divided into functional modules based on the module implementation steps in the above system. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0129] Based on the same inventive concept, embodiments of the present invention also provide a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to execute the module implementation steps in any of the aforementioned intelligent differentiation and detection systems for components of Phellinus linteus from different origins.

[0130] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the module implementation steps in any of the aforementioned intelligent differentiation and detection systems for components of Phellinus linteus from different origins.

[0131] It should be noted that the above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the foregoing embodiments of the present application have been described in detail, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent ones; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An intelligent distinguishing and detecting system for components of different origins of Phellinus igniarius, characterized in that, The system comprises: a data acquisition module, configured to acquire spectral data of a plurality of Phellinus igniarius samples with unknown origins and representative Phellinus igniarius samples with known origins, the spectral data comprising absorbance at different wavelengths within a set wavelength range; a similarity analysis module, configured to analyze the spectral data, and determine spectral shape similarity and same-component content dissimilarity between the spectral data of any two different Phellinus igniarius samples after reducing the influence of component overlap; a clustering module, configured to determine a distance index between the spectral data of any two different Phellinus igniarius samples according to the spectral shape similarity and the same-component content dissimilarity, and cluster the spectral data of all Phellinus igniarius samples with unknown origins based on the distance index to obtain a plurality of clustering clusters; an origin determination module, configured to match the spectral data in each clustering cluster with the spectral data of all representative Phellinus igniarius samples with known origins based on the distance index, and determine the origin of the Phellinus igniarius sample corresponding to each clustering cluster according to a matching result; the similarity analysis module comprises a first analysis module, and the first analysis module comprises: a first wavelength segment acquisition unit, configured to acquire a plurality of first wavelength segments corresponding to a spectral shape different wavelength range of the spectral data of any two different Phellinus igniarius samples within the set wavelength range; a principal component analysis unit, configured to acquire first spectral data segments and second spectral data segments of the spectral data of any two different Phellinus igniarius samples within each first wavelength segment respectively, and perform principal component analysis on the first spectral data segments and the second spectral data segments to acquire a plurality of principal components and characteristic values of the principal components of the first spectral data segments and the second spectral data segments; a matching analysis unit, configured to perform one-to-one matching on the principal components of the first spectral data segments and the second spectral data segments, and determine spectral shape similarity between the spectral data of any two different Phellinus igniarius samples after reducing the influence of component overlap according to a matching result and the characteristic values of the principal components; the matching analysis unit comprises: a matching unit, configured to perform one-to-one matching on each principal component of the first spectral data segments and the second spectral data segments to obtain a plurality of one-to-one matching pairs and matching edge values of the one-to-one matching pairs; a component same degree acquisition unit, configured to determine a component same degree between the first spectral data segments and the second spectral data segments according to a quantity difference between each principal component of the first spectral data segments and the second spectral data segments, the matching edge values of the one-to-one matching pairs, and characteristic values of two principal components in the one-to-one matching pairs; a spectral shape similarity acquisition unit, configured to determine the spectral shape similarity between the spectral data of any two different Phellinus igniarius samples after reducing the influence of component overlap according to a wavelength range proportion of a first wavelength segment corresponding to each first spectral data segment and second spectral data segment within the set wavelength range, the component same degree between the first spectral data segments and the second spectral data segments, and a wavelength range proportion of a total wavelength range excluding all first wavelength segments within the set wavelength range within the set wavelength range.

2. The system according to claim 1, wherein, the first wavelength segment acquisition unit comprises: The difference sequence acquisition unit is configured to determine a spectrum shape difference sequence of the spectrum data of any two different Phellinus igniarius samples according to a difference in a change trend of the absorbance of the spectrum data of any two different Phellinus igniarius samples at the same wavelength, and the spectrum shape difference sequence reflects the difference in the change trend of the spectrum data of any two different Phellinus igniarius samples at the same wavelength. The first clustering unit is configured to cluster target elements in the spectrum shape difference sequence to obtain a plurality of clusters, and if the change trends of the spectrum data of any two different Phellinus igniarius samples at a wavelength are the same, an element corresponding to the wavelength in the spectrum shape difference sequence is taken as a target element. The first wavelength segment acquisition unit is configured to take a wavelength range formed by a maximum wavelength and a minimum wavelength of the target elements in each cluster as a first wavelength segment.

3. The system according to claim 2, wherein, The difference sequence acquisition unit includes: The wavelength label acquisition unit is configured to determine a wavelength label corresponding to each wavelength in the spectrum data, set the wavelength label corresponding to a wavelength as a first value if the absorbance at the wavelength is less than the absorbance at a next wavelength, set the wavelength label corresponding to the wavelength as a second value if the absorbance at the wavelength is equal to the absorbance at the next wavelength, and set the wavelength label corresponding to the wavelength as a third value if the absorbance at the wavelength is greater than the absorbance at the next wavelength. The first sequence acquisition unit is configured to arrange all the wavelength labels in a wavelength order to obtain a spectrum shape sequence corresponding to the spectrum data. The second sequence acquisition unit is configured to perform an AND operation on wavelength labels at the same wavelength in the spectrum shape sequences corresponding to the spectrum data of any two different Phellinus igniarius samples to obtain a spectrum shape difference sequence of the spectrum data of any two different Phellinus igniarius samples.

4. The system according to claim 1, wherein, The similarity analysis module includes a second analysis module, and the second analysis module includes: The second wavelength segment acquisition unit is configured to obtain a plurality of second wavelength segments corresponding to wavelength ranges in which the spectrum shape of the spectrum data of any two different Phellinus igniarius samples is the same in a set wavelength range. The first content difference acquisition unit is configured to obtain a third spectrum data segment and a fourth spectrum data segment of the spectrum data of any two different Phellinus igniarius samples at each second wavelength segment, respectively, and determine a first same-component content difference and a first maximum-component content difference according to a difference in the absorbance at the same wavelength in the third spectrum data segment and the fourth spectrum data segment. The second content difference acquisition unit is configured to determine a second same-component content difference and a second maximum-component content difference according to a difference in the absorbance at the same wavelength in a one-to-one matching pair of the first spectrum data segment and the second spectrum data segment. The dissimilarity acquisition unit is configured to determine a same-component content dissimilarity between the spectrum data of any two different Phellinus igniarius samples according to all the first same-component content differences and the first maximum-component content differences and all the second same-component content differences and the second maximum-component content differences corresponding to the any two different Phellinus igniarius samples.

5. The system according to claim 4, wherein the system is characterized by, The first content difference acquisition unit includes: The first same-component-content-difference obtaining unit is configured to determine a sum of absolute values of differences in absorbance at the same wavelength in the third spectral data segment and the fourth spectral data segment, and obtain a first same-component-content-difference. The first maximum-component-content-difference obtaining unit is configured to determine a maximum value of absolute values of differences in absorbance at the same wavelength in the third spectral data segment and the fourth spectral data segment, and obtain a first maximum-component-content-difference.

6. The system according to claim 4, wherein the system is characterized by, The dissimilarity obtaining unit comprises: The comprehensive same-component-content-difference determining unit is configured to determine a sum of all first same-component-content-differences and second same-component-content-differences corresponding to any two different Phellinus baumii samples, and obtain a comprehensive same-component-content-difference. The maximum-component-content-difference variance determining unit is configured to determine a sum of variances of all first maximum-component-content-differences and second maximum-component-content-differences corresponding to any two different Phellinus baumii samples, and obtain a maximum-component-content-difference variance. The dissimilarity determining unit is configured to determine a same-component-content-dissimilarity between spectral data of any two different Phellinus baumii samples according to the comprehensive same-component-content-difference, the maximum-component-content-difference variance, and the maximum-component-content-difference maximum value.

7. The system according to claim 1, wherein the system is characterized by, The clustering module comprises: The negative correlation mapping processing unit is configured to perform negative correlation mapping processing on the spectral shape similarity, and obtain a spectral shape similarity processing value. The distance index obtaining unit is configured to calculate a product of the spectral shape similarity processing value and the same-component-content-dissimilarity as a distance index. The second clustering unit is configured to perform clustering on spectral data of all Phellinus baumii samples with unknown origins based on the distance index and using a K-means clustering algorithm, and obtain a plurality of clustering clusters.

8. The system according to claim 1, wherein the system is characterized by, The origin determining module comprises: The first distance index obtaining unit is configured to determine an average value of similarity indexes between each clustering cluster and spectral data of representative Phellinus baumii samples of each same origin, and obtain an average distance index. The second distance index obtaining unit is configured to determine a minimum value in average distance indexes between each clustering cluster and spectral data of representative Phellinus baumii samples of different origins, and obtain a minimum average distance index. The origin determining unit is configured to determine an origin corresponding to the minimum average distance index as an origin of each clustering cluster of Phellinus baumii samples.

Citation Information

Patent Citations

  • Method for identifying production place of dwarf lilyturf root by using near infrared spectrum technology

    CN102243170A

  • Ore grade and element detection method and system

    CN118641528A