Morphological character-based cactus plant germplasm identification method and system and medium
By using standardized measurements and improved principal component analysis to screen key morphological traits and construct an identification reference database, the accuracy problems existing in traditional methods are solved, enabling rapid and accurate identification of cactus species.
Patent Information
- Application Number
- CN202511433722.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-01-02
AI Technical Summary
Traditional morphological identification methods for cacti suffer from low accuracy, high subjectivity, poor repeatability, and insufficient standardization, making it difficult to accurately delineate interspecific boundaries and affecting the development of medicinal resources and taxonomic research.
A morphological trait data matrix was constructed using standardized measurement rules. Key morphological traits were screened using an improved principal component analysis method, and standardized processing was performed to construct an identification reference database. Germplasm was determined by spatial similarity calculation.
It improves the accuracy and standardization of cactus identification, enables rapid and accurate germplasm identification, reduces reliance on the experience of identification personnel, and is suitable for automated identification of large numbers of samples.
Smart Images

Figure CN121256346A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of species identification, in particular to a morphological trait-based Opuntia germplasm identification method, system and medium. BACKGROUND
[0002] Among natural medicine resources, Opuntia plants are of great concern due to their unique biological characteristics and extensive medicinal value. Some Opuntia plants are widely used in traditional medicine due to their pharmacological activities such as anti-inflammatory, hemostatic, detoxification, and treatment of respiratory and digestive system diseases.
[0003] Currently, traditional morphological identification techniques play a fundamental role in the management, classification and preliminary identification of plant germplasm resources, especially in field surveys and herbarium work due to their convenience and intuitiveness. Traditional Opuntia identification methods mainly rely on the experience-based identification of macroscopic morphological traits such as spines, pads, fruits, etc. However, Opuntia plants have high morphological plasticity (great morphological differences in the same species under different environments), frequent interspecific hybridization (leading to mixed traits), and ambiguous intraspecific and interspecific boundaries, which greatly increase the difficulty of traditional morphological identification methods. Specifically, the difficulty of experience-based morphological identification of Opuntia plants mainly comes from two aspects: first, the contradictory relationship between morphological traits (such as spines, small nests, seed characteristics) in different Opuntia species, reflecting strong functional plasticity and phenotypic convergence; second, external morphological characteristics such as spine density and small nest distribution, which are easily driven by environmental adaptation and natural selection, may independently evolve similar phenotypes in different phylogenetic backgrounds, further exacerbating the disconnection between classification and true phylogenetic relationships. Single or few morphological traits fail to capture all variations, making it difficult for traditional morphological experience-based identification methods to accurately delineate interspecific boundaries. Therefore, for such a group of Opuntia with strong phenotypic plasticity and ambiguous interspecific boundaries, purely relying on morphological experience-based identification methods in the application of Opuntia will inevitably have technical problems such as low accuracy, strong subjectivity, poor repeatability, and insufficient standardization, which seriously restricts the effective development of Opuntia medicinal resources, taxonomic research, and market supervision. SUMMARY
[0004] To solve the above technical problems, the present application provides a morphological trait-based Opuntia germplasm identification method, system and medium, aiming to optimize the morphological identification technical system and provide a scientific basis for rapid, accurate and standardized identification of Opuntia plants.
[0005] In a first aspect, the present application provides a morphological trait-based Opuntia germplasm identification method, comprising the following steps:
[0006] S1, constructing a morphological trait data matrix of different Opuntia plants by using a standardized measurement rule;
[0007] S2, screening key morphological traits by using an improved principal component analysis method based on the morphological trait data matrix;
[0008] S3, performing standardized processing on each key morphological trait to obtain a standardized key morphological vector;
[0009] S4, constructing an identification reference database containing a standard species vector based on the standardized key morphological vector;
[0010] S5, collecting measurement data of key morphological traits of a plant sample to be measured and performing standardized processing to obtain a standardized morphological vector to be measured;
[0011] S6, calculating the spatial similarity between the morphological vector to be measured and each standard species vector in the identification reference database, and determining the Opuntia plant germplasm corresponding to the plant sample to be measured based on the calculation result of the spatial similarity.
[0012] In some embodiments, the morphological trait data matrix of different Opuntia plants is constructed by using a standardized measurement rule, including:
[0013] Selecting quantitative morphological traits of Opuntia plants, the quantitative morphological traits including the number of spines, the distance between small cavities, the height of the plant, the number of small cavities in each row, the length of spines, the length of palm pieces, the width of palm pieces, the diameter of seeds, the width of fruits, and the length of fruits;
[0014] Setting a corresponding measurement standard for each quantitative morphological trait to constitute a standardized measurement rule;
[0015] Based on the standardized measurement rule, collecting quantitative morphological trait measurement data of several individuals of each Opuntia plant and calculating the average value to obtain the average data of each quantitative morphological trait of each Opuntia plant;
[0016] The collection of average data of quantitative morphological traits of all Opuntia plants constitutes a morphological trait data matrix.
[0017] In some embodiments, the key morphological traits are screened by using an improved principal component analysis method based on the morphological trait data matrix, including:
[0018] Based on the morphological trait data matrix, a corresponding covariance matrix is calculated;
[0019] Solving the covariance matrix to obtain the eigenvalues of each principal component and the corresponding eigenvectors;
[0020] Based on the eigenvalues of each principal component, the contribution rate of each principal component is calculated, which represents the percentage of the variance explained by the principal component in the total variance;
[0021] Randomly combine the contribution rates of each principal component, each combination containing k principal components, obtain the cumulative contribution rate corresponding to each combination, and select the principal components in the combination corresponding to the maximum cumulative contribution rate as important principal components;
[0022] Based on the eigenvalues of the important principal components and the corresponding eigenvectors, the load values of each important principal component are calculated, which represent the contribution degree of quantitative morphological traits to important principal components;
[0023] Sort the load values of each important principal component from large to small, and select the quantitative morphological traits corresponding to the first few load values as key morphological traits.
[0024] In some embodiments, based on the eigenvalues of the important principal components and the corresponding eigenvectors, the load values of each important principal component are calculated, including:
[0025] Calculate the square root of the eigenvalue of the important principal component, then multiply it by the corresponding eigenvector, and take the absolute value of the product obtained, which is taken as the load value of the corresponding important principal component.
[0026] In some embodiments, the key morphological traits are standardized to obtain a standardized key morphological vector, including:
[0027] Calculate the average value and standard deviation of each key morphological trait measurement data in all Opuntia plants;
[0028] Subtract each key morphological trait measurement data from the corresponding average value, and then divide the difference by the corresponding standard deviation to obtain the standardized key morphological vector.
[0029] In some embodiments, based on the standardized key morphological vector, an identification reference database containing standard species vectors is constructed, including:
[0030] Combine the standardized key morphological vector corresponding to each Opuntia plant into a standard species vector, and the standard species vector is used to uniquely determine the coordinate position of the corresponding Opuntia plant in the multi-dimensional morphological space;
[0031] Associate all standard species vectors with the corresponding Opuntia plants and numbers to form an identification reference database.
[0032] In some embodiments, the spatial similarity between the measured morphological vector and each standard species vector in the identification reference database is calculated, and the corresponding Opuntia plant germplasm of the measured plant sample is determined based on the calculation result of the spatial similarity, including:
[0033] calculating the difference between the morphological vector to be tested and each standard species vector in each key morphological dimension, and calculating the square of the difference;
[0034] summing the square of the difference between the morphological vector to be tested and each standard species vector in all key morphological dimensions to obtain the spatial similarity between the morphological vector to be tested and each standard species vector;
[0035] selecting the standard species vector corresponding to the minimum spatial similarity, and determining that the plant sample to be tested is the Opuntia plant germplasm corresponding to the standard species vector.
[0036] In some embodiments, after determining the Opuntia plant germplasm corresponding to the plant sample to be tested based on the calculation result of the spatial similarity, the method further comprises:
[0037] calculating the spatial similarity between each pair of standard species vectors in the identification reference database to form a similarity matrix;
[0038] extracting all non-zero similarity values from the similarity matrix to form a similarity dataset;
[0039] determining a similarity threshold value based on the similarity dataset using the percentile method;
[0040] comparing the minimum value in the calculation result of the spatial similarity with the similarity threshold value, if the minimum value in the calculation result of the spatial similarity is less than or equal to the similarity threshold value, then the Opuntia plant germplasm corresponding to the plant sample to be tested determined in S6 is the final identification result, if the minimum value in the calculation result of the spatial similarity is greater than the similarity threshold value, then triggering the manual review operation.
[0041] In a second aspect, the present application provides an Opuntia plant germplasm identification system based on morphological traits, comprising:
[0042] a standardization measurement module for constructing morphological trait data matrices of different Opuntia plants using standardization measurement rules;
[0043] a morphological trait screening module for screening key morphological traits based on the morphological trait data matrices using an improved principal component analysis method;
[0044] a standardization processing module for standardizing each key morphological trait to obtain a standardized key morphological vector;
[0045] a database construction module for constructing an identification reference database containing standard species vectors based on the standardized key morphological vector;
[0046] The sample processing module is used for collecting and standardizing measurement data of key morphological traits of the plant sample to be tested, and obtaining a standardized morphological vector of the plant sample to be tested;
[0047] The species identification module is used for calculating spatial similarity between the morphological vector to be tested and each standard species vector in the identification reference database, and determining the Cereus plant germplasm corresponding to the plant sample to be tested based on the calculation result of the spatial similarity.
[0048] In a third aspect, a computer readable storage medium has a computer program stored thereon, and the computer program is executed by a processor to implement the Cereus plant germplasm identification method based on morphological traits as described above.
[0049] The beneficial technical effects of the present application at least include:
[0050] 1. The Cereus plant germplasm identification method, system and medium based on morphological traits are adopted, and through the synergistic effect of the ring-by-ring design, the technical problems of low accuracy, strong subjectivity, poor repeatability and insufficient standardization in the application of the existing morphological experience identification method to the Cereus genus are systematically solved. Specifically: first, the ambiguous morphological description is converted into accurate numerical values by strictly defining the quantitative morphological traits and measurement standards, and the subjective bias caused by the personal experience of the identifier is eliminated; second, the improved principal component analysis method is adopted to effectively screen out the key morphological trait combination with the largest contribution to the Cereus species division from the numerous quantitative morphological traits, and the redundant and interfering information is discarded, thereby directly improving the identification accuracy and concentrating resources on the most effective identification features; then, the standardization processing of each key morphological trait eliminates the dimensional difference between the key morphological traits, and realizes the quantifiable comparison of the key morphological traits; finally, the spatial similarity calculation provides an accurate and quantifiable similarity index, and the identification problem is converted into a mathematical problem of "finding the nearest neighbor in the morphological space". The complete technical solution successfully constructs a standardized, streamlined and highly operable morphological identification system. The expert knowledge is solidified into specific operation processes and algorithms, which significantly reduces the dependence on the experience of the identifier, and realizes a high identification accuracy when facing the highly complex Cereus genus, which is significantly better than the traditional morphological experience identification method. The present application not only provides a new and effective solution for the rapid and low-cost identification of medicinal plant resources, but also provides a feasible solution for the classification dilemma caused by the morphological plasticity and genetic complexity of the Cereus genus.
[0051] 2、The application uses improved principal component analysis for dimension reduction. First, the principal components generated by linear transformation are orthogonal to each other, which completely solves the problem of multicollinearity, thereby realizing the extraction of pure and independent identification dimensions. The key morphological trait combination screened by the application has the lowest redundancy of classification information carried, making the subsequent identification reference database more stable and efficient, and thus significantly improving the accuracy of identification. Second, important principal components are selected according to eigenvalues and cumulative contribution rates to achieve the optimal balance between information retention and operational efficiency, capturing as much classification information as possible with the least variables. Finally, the contribution of each quantitative morphological trait to species differentiation is objectively quantified by analyzing the load value, thereby greatly improving the pertinence of identification efficiency. The final identification system is lightweight and maintains high discrimination ability (as the key morphological traits carry the most core variation information), making it possible to achieve rapid and accurate on-site identification.
[0052] 3、The morphological characteristics of Cactus species are converted into coordinate points in a mathematical space, thereby transforming the fuzzy, experience-based "morphological similarity" judgment into precise, calculable "spatial distance" comparison, injecting modern data science ideas into traditional morphological identification. Specifically, traditional morphological identification methods rely on text-based morphological characteristics (such as "many spines" and "large leaves"), but when multiple traits need to be considered, the human brain cannot accurately quantify the weight of each trait and perform comprehensive calculations, and is easily disturbed by individual prominent traits. Therefore, the application constructs a "vector space model" that abstracts each Cactus species as a point in a multi-dimensional morphological space, i.e., converts morphological characteristics into specific coordinates in a multi-dimensional morphological space. In this "vector space model", the species identification problem is transformed into a nearest neighbor search problem, achieving lossless quantization of information while ensuring that the sample determination is based on the similarity of its overall morphological profile, avoiding misjudgment caused by a single or few morphological traits. Even if a key morphological trait varies greatly due to environmental reasons, as long as other key morphological traits are stable, the overall coordinate position will not shift dramatically, thereby effectively improving the accuracy and standardization level of Cactus germplasm resource morphological identification.
[0053] 4、The application converts a complex taxonomic problem into a standardized and efficient computational problem, requiring only one spatial similarity operation (calculating the spatial similarity between each standard species vector in the identification reference database) to identify an unknown plant sample, enabling automatic identification of batch samples with much higher efficiency than manual methods, making it applicable to practical scenarios that require rapid screening of a large number of samples. The identification method is simple and intuitive, and the differences in multiple key morphological dimensions are cumulative, ultimately outputting an objective, accurate, and quantifiable identification conclusion.
[0054] Other features and advantages of the present application will be disclosed in the following detailed description, drawings. BRIEF DESCRIPTION OF DRAWINGS
[0055] The present application will be further described below with reference to the drawings:
[0056] Figure 1 The flow chart of the cactus plant germplasm identification method based on morphological traits of the embodiments of the present application is shown.
[0057] Figure 2 The structural schematic diagram of the cactus plant germplasm identification system based on morphological traits of the embodiments of the present application is shown. DETAILED DESCRIPTION
[0058] The technical solutions of the embodiments of the present application will be explained and described below with reference to the drawings of the embodiments of the present application. However, the following embodiments are only preferred embodiments of the present application, not all. Based on the embodiments in the embodiments, other embodiments obtained by those skilled in the art without creative labor also belong to the protection scope of the present application.
[0059] In the following description, the orientation or position relationship appearing in the terms such as "inner", "outer", "upper", "lower", "left", "right" and the like is only for the convenience of describing the embodiments and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.
[0060] Embodiment one:
[0061] Please refer to the drawings Figure 1 , Figure 1 The flow schematic diagram of the cactus plant germplasm identification method based on morphological traits provided by one embodiment of the present specification is shown.
[0062] As Figure 1 shown, the cactus plant germplasm identification method based on morphological traits can at least include the following steps:
[0063] S1, constructing morphological trait data matrix of different cactus plants by using standardized measurement rules.
[0064] It can be understood that the inventive concept of step S1 is to realize that high-quality data input is a prerequisite for the success of any identification system. In the face of the traditional difficulties of cactus identification, the solution of the present embodiment is not to abandon morphological identification, but to optimize it through the precise thinking of engineering, and to change the original qualitative technology relying on "eyes and hands" into a set of objective technology based on quantitative measurement, strict standards and statistical processing.
[0065] Specifically, in this embodiment, standardized measurement rules are used to construct morphological trait data matrices of different Opuntia plants, including:
[0066] S11, quantitative morphological traits of Opuntia plants are selected, including spine number, areole spacing, plant height, number of areoles per row, spine length, cladode length, cladode width, seed diameter, fruit width, and fruit length.
[0067] In this embodiment, 10 main quantitative morphological traits in Opuntia are selected by extensively reviewing existing Opuntia plant specimens in the collection, namely, spine number (Spine Number), areole spacing (Areole Spacing, unit: cm), plant height (Plant Height, unit: m), number of areoles per row (Areoles per Row), spine length (Spine Length, unit: cm), cladode length (Cladode Length, unit: cm), cladode width (Cladode Width, unit: cm), seed diameter (Seed Diameter, unit: mm), fruit width (Fruit Width, unit: cm), and fruit length (Fruit Length, unit: cm).
[0068] S12, measurement standards corresponding to each quantitative morphological trait are set to form standardized measurement rules.
[0069] In this embodiment, precise and repeatable measurement standards are set for each quantitative morphological trait. The specific measurement standards can be set according to actual identification conditions, for example, the measurement standard of the quantitative morphological trait “spine length” is set as the average length of the longest 5 spines on the Opuntia plant, the measurement standard of the quantitative morphological trait “cladode length” is set as the vertical distance from the base to the top, the measurement of cladode length and width is only performed on mature stem segments, and the measurement of fruits and seeds strictly selects fully mature samples. This embodiment does not limit this.
[0070] S13, based on the standardized measurement rules, quantitative morphological trait measurement data of several individuals of each Opuntia plant are collected and the average value is calculated to obtain the average data of each quantitative morphological trait of each Opuntia plant.
[0071] S14, the collection of average data of quantitative morphological traits of all Opuntia plants forms a morphological trait data matrix.
[0072] For example, assuming that m Opuntia species are studied in this embodiment, n=10 quantitative morphological traits are measured for each species, and after average calculation, an m×n morphological trait data matrix Z is obtained, where Z[i,j] represents the average data of the jth quantitative morphological trait of the ith species.
[0073] Understandably, measuring multiple individuals and calculating the average value can better represent the stable characteristics of a species. This can minimize measurement errors and filter out accidental variations in individual plants, approximating the "true value" of the trait for that species. As a result, a stable and representative data matrix of species morphological traits can be obtained, effectively reducing the risk of misjudgment due to individual differences.
[0074] S2, based on the morphological trait data matrix, uses an improved principal component analysis method to screen out key morphological traits.
[0075] Understandably, given the complex and redundant nature of cactus traits, this embodiment does not rely on experience to select and directly use the original morphological traits. Instead, it uses an improved principal component analysis method to automatically extract the most important and independent key morphological trait combinations among cactus species. The purpose is to objectively and data-drivenally screen out the core morphological trait combinations with the greatest classification and identification value from numerous quantitative morphological traits, thereby achieving dimensionality reduction and noise reduction.
[0076] Specifically, in this embodiment, based on the morphological trait data matrix, a modified principal component analysis method is used to screen and obtain key morphological traits, including:
[0077] S21, based on the morphological data matrix, the corresponding covariance matrix is calculated.
[0078] It is understandable that this embodiment calculates the covariance matrix of the morphological trait data matrix to explore the degree of linear correlation between each pair of 10 quantitative morphological traits.
[0079] The calculation method for the covariance matrix corresponding to the morphological data matrix is similar to that in existing technologies, and will not be elaborated upon in this embodiment. The covariance matrix calculated from the morphological data matrix Z is denoted as Σ, where Σ is a... The matrix is Σ[j,j], where the diagonal element Σ[j,h] is the variance of the j-th quantitative morphological trait, and the off-diagonal element Σ[j,h] is the covariance between the j-th and h-th quantitative morphological traits. Its value is between -1 and 1, reflecting the linear correlation between the two quantitative morphological traits. A positive value indicates a positive correlation, a negative value indicates a negative correlation, and 0 indicates no linear correlation.
[0080] S22, solve the covariance matrix to obtain the eigenvalues of each principal component and the corresponding eigenvectors.
[0081] Among them, eigenvalues ( ) is a set of scalars, each principal component has an eigenvalue, the size of each eigenvalue represents the importance of the corresponding principal component of the variation direction (that is, the size of the variance corresponding to the principal component), the larger the eigenvalue, the more significant the variation in that direction.
[0082] wherein the eigenvector (v ) is a direction vector with a length of 1, and each eigenvalue corresponds to an eigenvector, which defines a new coordinate axis direction in the original 10-dimensional quantitative morphological trait space, that is, the direction of the principal component.
[0083] wherein the principal component , the principal component Each element in the vector is the coordinate value of each species in the i-th principal component direction, which is called the principal component score, that is, the projection of each species onto the straight line defined by the new direction vector , and the projection length is calculated, which tells us where the data point "looks" in the new perspective.
[0084] S23, based on the eigenvalues of each principal component, the contribution rate of each principal component is calculated, which represents the percentage of the variance explained by the principal component in the total variance.
[0085] wherein the contribution rate of the i-th principal component It can be understood that the importance of each principal component is intuitively reflected by the contribution rate.
[0086] S24, randomly combine the contribution rates of each principal component, each combination contains k principal components, obtain the cumulative contribution rate corresponding to each combination, and select the principal components in the combination corresponding to the maximum cumulative contribution rate as important principal components.
[0087] wherein, according to the eigenvalues from large to small, the principal components are sorted, and the number of k in this embodiment is less than the number of quantitative morphological traits, that is, , the first k principal components are selected in turn, so that their cumulative contribution rate reaches a high level.
[0088] S25, based on the eigenvalues of the important principal components and the corresponding eigenvectors, the load value of each important principal component is calculated, which represents the contribution degree of the quantitative morphological trait to the important principal component.
[0089] It can be understood that the higher the absolute value of the load value of a quantitative morphological trait on a principal component, the greater the contribution of the quantitative morphological trait to the composition of the principal component, that is, the quantitative morphological trait is the main driving force in the variation direction represented by the principal component.
[0090] Further, in this embodiment, based on the eigenvalue of the important principal component and the corresponding eigenvector, the load value of each important principal component is calculated, including:
[0091] The square root of the eigenvalue of the important principal component is calculated, and then multiplied by the corresponding eigenvector, and the absolute value of the obtained product is taken as the load value of the corresponding important principal component, which can be expressed as:
[0092] ,
[0093] Wherein, represents the load value between the jth quantitative morphological trait and the ith principal component, represents the jth component of the ith eigenvector, which can be understood as the initial weight of the quantitative morphological trait j in the "formula" constituting the principal component . The obtained is a value between-1 and 1, and the intensity is represented by the absolute value.
[0094] It can be understood that this embodiment is only the load value calculation process of one of the important principal components, and the load value calculation methods of other important principal components can refer to this embodiment.
[0095] It can be understood that the eigenvector Itself is a direction unit vector. Multiply it by , which is equivalent to scaling the direction vector. The vector coordinates after scaling are the load values, which directly represent the projection length of the original variable axis on the new principal component axis. The longer the projection, the closer the relationship.
[0096] S26, the load values of each important principal component are sorted from large to small, and the quantitative morphological traits corresponding to the first few load values are selected as the key morphological traits.
[0097] For example, the improved principal component analysis is carried out on the basis of 10 quantitative morphological traits of 36 species of cactus genus, and the principal component analysis results are shown in Table 1 and Table 2. Principal component ordering Eigenvalue Contribution rate (%) Cumulative contribution rate (%) PC1 2.9415 29.4149 29.4149 PC2 1.5762 15.7615 45.1764 PC3 1.3079 13.0789 58.2553 PC4 0.9929 9.9287 68.1840 PC5 0.9363 9.3628 77.5468 PC6 0.7749 7.7491 85.2959 PC7 0.5161 5.1606 90.4564 PC8 0.4410 4.4095 94.8659 PC9 0.2991 2.9914 97.8573 PC10 0.2143 2.1427 100 .
[0098] Table 1 Principal component analysis results of quantitative morphological traits of 36 species of cactus genus Quantitative morphological traits PC1 PC2 PC3 Palm length 6.3153 0.7688 0.3327 Fruit width 6.2446 -0.2134 0.1622 Plant height 5.5293 0.1125 -0.7923 Fruit length 5.3897 -0.7524 -0.485 Cup distance 5.1349 -1.5356 1.513 Palm width 4.3964 1.4798 1.5732 Seed diameter 1.0641 1.0619 2.5637 Cup number -0.1107 3.1842 -1.0886 Spine number -1.5064 1.4643 2.0615 Spine length -3.2134 -0.8457 3.2825 .
[0099] Table 2 Load of quantitative morphological traits in the first three principal components
[0100] From Table 1 and Table 2, it can be seen that when the first five principal components are introduced, the cumulative contribution rate can only exceed 70%, while the first seven principal components can cumulatively explain more than 90% of the total variation, reflecting the multidimensionality and dispersion of the morphological variation of Cactaceae plants. Among the first three principal components, PC1 has the highest eigenvalue (2.94) and the largest contribution rate (29.41%), mainly with significant positive loadings of the length of the palm piece (6.3153), the width of the fruit (6.2446), the height of the plant (5.5293), the length of the fruit (5.3897), the distance between the small pits (5.1349), and the width of the palm piece (4.3964), and negative loadings of the length of the spine (-3.2134) and the number of spines (-1.5064) (Table 2). This shows that PC1 mainly reflects the resource allocation trade-off between the overall structure size of the plant and the defense structure characteristics. PC2 explains 15.76% of the total variation, and the positive loadings are mainly concentrated in the number of small pits (3.1842), the width of the palm piece (1.4798), and the number of spines (1.4643), and the negative loadings are mainly the distance between the small pits (-1.5356), reflecting the morphological differentiation of the plant surface defense structure (small pit and spine density). PC3 explains 13.08% of the total variation, and on this principal component, the length of the spine (3.2825), the diameter of the seed (2.5637), and the number of spines (2.0615) have higher positive loadings, indicating the characteristics of the coordinated variation of the spine morphology and the reproductive structure (seed size), while the height of the plant (-0.7923) and the number of small pits (-1.0886) have negative loadings. Therefore, in summary, the first three principal components correspond to the main morphological variation directions of Cactaceae plants in plant structure, defense characteristics, and reproductive characteristics.
[0101] It can be understood that, on the one hand, the traditional method has the technical problems of trait redundancy and noise interference, that is, there is often correlation between the multiple morphological traits used by the traditional method (for example, the palm piece length and the palm piece width may be positively correlated), which leads to information overlap. Including all these high-correlation traits in the identification system not only increases unnecessary workload, but also introduces problems such as multicollinearity in mathematics, which interferes with the accuracy and stability of the identification model. On the one hand, the traditional method is difficult to scientifically weigh the information loss and efficiency, that is, if all traits are used for comprehensiveness, the efficiency will be reduced; if the traits are arbitrarily reduced for efficiency, key classification information may be lost. Therefore, the improved principal component analysis method is used for dimension reduction in the embodiment. First, the principal components (new variables) generated by linear transformation are orthogonal (uncorrelated) to each other, which completely solves the multicollinearity problem, thereby realizing the extraction of pure and independent identification dimensions. The combination of key morphological traits screened by the embodiment has the lowest classification information redundancy, making the subsequent identification reference database more stable and efficient, because the identification reference database is determined based on a group of “independent” rather than “entangled” prediction variables, thereby significantly improving the accuracy of identification. Secondly, important principal components are screened according to eigenvalues and cumulative contribution rates to realize the optimal balance between information retention and operational efficiency. Through this method, we neither blindly use all traits nor arbitrarily discard traits, but in a reliable way, we capture as much classification information as possible with the least variables. Finally, key morphological traits are objectively screened by analyzing the load values. By analyzing the load values of each quantitative morphological trait on the important principal components, the contribution of each quantitative morphological trait to species differentiation is objectively quantified, thereby ensuring that the screened key morphological traits are the characteristics that are mathematically verified and contribute most to species differentiation, so that the identification work can focus on the most effective features, avoiding wasting effort on low identification power traits, thereby greatly improving the pertinence of identification efficiency. Further, the finally established identification system is not only lightweight, but also maintains high discrimination ability (because the key morphological traits carry the most core variation information), which makes it possible to realize fast and accurate on-site identification.
[0102] In summary, step S2 is the key turning point of the entire method from “traditional description” to “modern data-driven”. S2 not only solves the subjectivity and redundancy of the traditional method, but more importantly, it provides a mathematically optimized and optimized input feature set for the entire identification system. As a core step, it combines with the traditional morphology to build an efficient identification morphological space.
[0103] S3, standardizing each key morphological trait to obtain a standardized key morphological vector.
[0104] Specifically, in this embodiment, each key morphological trait is standardized to obtain a standardized key morphological vector, including:
[0105] S31, calculate the average value and standard deviation of each key morphological trait measurement data in all Cactus plants.
[0106] S32, subtract each key morphological trait measurement data from the corresponding average value, and then divide the obtained difference value by the corresponding standard deviation to obtain a standardized key morphological vector.
[0107] It can be understood that the original measurement data of each key morphological trait selected in this embodiment is standardized to eliminate the deviation caused by the difference in dimension and numerical range of different key morphological traits, and to ensure that each key morphological trait is in the same scale and coordinate system in subsequent analysis.
[0108] S4, based on the standardized key morphological vector, an identification reference database containing a standard species vector is constructed.
[0109] Specifically, in this embodiment, based on the standardized key morphological vector, an identification reference database containing a standard species vector is constructed, including:
[0110] S41, the standardized key morphological vector corresponding to each Cactus plant is combined into a standard species vector, and the standard species vector is used to uniquely determine a coordinate position of the corresponding Cactus plant in a multi-dimensional morphological space.
[0111] It can be understood that for each Cactus plant, several standardized key morphological vectors are combined in a fixed order to form an ordered array, i.e. a standard species vector, which can be understood as a multi-dimensional morphological vector. The standard species vector uniquely determines a point (coordinate position) of the Cactus plant in a multi-dimensional morphological space, and each dimension in the multi-dimensional morphological space represents a key morphological trait. In theory, different species have different coordinate positions in this multi-dimensional morphological space due to their morphological differences.
[0112] S42, associate all standard species vectors with the corresponding Cactus plants and numbers to form an identification reference database.
[0113] It can be understood that the present embodiment associates the "multi-dimensional morphological vector" of each Cactus plant with its accurate species name, number, and other information, and stores and manages them in the form of a database, which is the identification reference database. The essence of the identification reference database is an m x q lookup matrix (m represents the number of Cactus species, and q represents the number of key morphological traits, i.e., the number of standardized key morphological vectors), each row represents a Cactus species, and each column represents a standardized key morphological vector.
[0114] It can be understood that the inventive concept of the present embodiment is derived from a core idea: converting the morphological characteristics of Cactus species into coordinate points in a mathematical space, thereby transforming the fuzzy, experience-based "morphological similarity" judgment into precise, calculable "spatial distance" comparison, and injecting modern data science ideas into traditional morphological identification. Specifically, the traditional morphological identification method relies on textually described morphological characteristics (such as "many spines" and "large leaves"), but when multiple traits need to be considered simultaneously, the human brain has difficulty accurately quantifying the weight of each trait and performing comprehensive calculations, and is easily disturbed by individual prominent traits. Therefore, the present embodiment constructs a "vector space model" that abstracts each Cactus species as a point in a multi-dimensional morphological space, i.e., converts morphological characteristics into specific coordinates in a multi-dimensional morphological space. In this "vector space model", the species identification problem is transformed into a nearest neighbor search problem, achieving lossless quantization of information while ensuring that the judgment of the sample is based on the similarity of its overall morphological profile, avoiding misjudgment caused by a single or a few morphological traits. Even if a key morphological trait varies greatly due to environmental reasons, as long as other key morphological traits are stable, the overall coordinate position will not shift dramatically, thereby effectively improving the accuracy and standardization level of Cactus plant germplasm resource morphological identification.
[0115] S5, collecting key morphological trait measurement data of the test plant sample and performing standardization processing to obtain a standardized test morphological vector.
[0116] It can be understood that step S5 ensures that the data of the test plant sample and the data in the identification reference database are in the same scale and coordinate system, making the subsequent spatial similarity calculation have mathematical meaning and comparability.
[0117] S6, calculating the spatial similarity between the test morphological vector and each standard species vector in the identification reference database, and determining the corresponding Cactus plant germplasm of the test plant sample based on the calculation result of the spatial similarity.
[0118] Specifically, in this embodiment, the spatial similarity between the to-be-tested morphological vector and each standard species vector in the identification reference database is calculated, and the to-be-tested plant sample is determined to correspond to a Cereus plant germplasm based on the calculation result of the spatial similarity, including:
[0119] S61, the difference between the to-be-tested morphological vector and each standard species vector in each key morphological dimension is calculated, and the square of the difference is calculated.
[0120] It can be understood that the square operation of the difference in this embodiment ensures that the difference is positive and amplifies the contribution of larger differences.
[0121] S62, the sum of the square of the difference between the to-be-tested morphological vector and each standard species vector in all key morphological dimensions is calculated, to obtain the spatial similarity between the to-be-tested morphological vector and each standard species vector.
[0122] It can be understood that the sum of the square of the difference between the to-be-tested morphological vector and each standard species vector in all key morphological dimensions in this embodiment synthesizes the differences in all morphological characteristics, and there is no bias in artificially assigning weights.
[0123] S63, the standard species vector corresponding to the minimum spatial similarity is selected, and the to-be-tested plant sample is determined to be the Cereus plant germplasm corresponding to the standard species vector.
[0124] It can be understood that this embodiment converts a complex taxonomic problem into a standardized and efficient calculation problem, and only one spatial similarity operation (calculating the spatial similarity between each standard species vector in the identification reference database) is required to identify an unknown plant sample. Batch samples can be automatically identified, which is much more efficient than manual methods, allowing it to be applied to practical scenarios that require rapid screening of a large number of samples. At the same time, the identification method is simple and intuitive, and the differences in multiple key morphological dimensions are cumulative, resulting in an objective, accurate, and quantifiable identification result.
[0125] In summary, the present embodiment does not simply stack technical units, but through the synergistic effect of a series of interlocking designs, it systematically solves the technical problems of low accuracy, strong subjectivity, poor repeatability, and insufficient standardization of existing morphological experience identification methods in the application of Cereus. Specifically:
[0126] First, by strictly defining quantitative morphological characteristics and measurement standards, ambiguous morphological descriptions are converted into precise numerical values, eliminating subjective bias caused by the personal experience of the identifier, and achieving a transition from "subjective experience" to "objective data". At the same time, it provides a reliable and consistent data basis for all subsequent statistical analysis;
[0127] Secondly, by employing an improved principal component analysis method, the key morphological trait combinations that contribute most to the differentiation of cactus species were effectively screened from numerous quantitative morphological traits. Redundant and interfering information was eliminated, essentially mimicking the ability of identification experts to "focus on key points," but the process was completely objective and data-driven, achieving a shift from "overall observation" to "feature extraction." Simultaneously, the selected key morphological traits greatly simplified the complexity of subsequent calculations and directly improved the accuracy of identification, allowing resources to be concentrated on the most effective distinguishing features.
[0128] Next, by standardizing the key morphological traits, the dimensional differences between them were eliminated. This allowed data in different units, such as "fruit width (cm)" and "plant height (m)," to be fairly compared and calculated within the same multidimensional morphological space, achieving a transformation from "incomparable" to "quantifiable comparison." This is also a prerequisite for subsequent spatial similarity calculations. Without standardization, the calculated spatial similarity would be meaningless.
[0129] Finally, spatial similarity calculation provides a precise and quantifiable similarity index. This transforms the identification problem into a mathematical problem of "finding the nearest neighbor in morphological space," shifting from "similarity judgment" to "precise measurement," ultimately achieving accurate and repeatable identification. Simultaneously, spatial similarity calculation integrates all the outputs of the preceding steps (i.e., the standardized key morphological characteristic vectors), becoming the core of the entire identification scheme's decision-making process.
[0130] Therefore, this complete technical solution successfully constructs a standardized, streamlined, and highly operable morphological identification system. It solidifies expert knowledge into specific operational procedures and algorithms, significantly reducing reliance on the experience of identification personnel. This allows non-experts to conduct relatively reliable preliminary identifications, achieving a high accuracy rate even when dealing with the highly complex cactus genus. This is significantly superior to traditional morphological experience-based identification methods, providing a novel and effective solution for the rapid and low-cost identification of medicinal plant resources. It also offers a feasible approach to addressing the classification dilemmas arising from the morphological plasticity and genetic complexity of cacti.
[0131] Example 2:
[0132] This embodiment only applies to comparisons with... Figure 1 The differences between the two embodiments will be described in the following descriptions. The technical concepts of the remaining designs are similar to those of the first embodiment, and will not be repeated here.
[0133] Furthermore, in this embodiment, after "S6, determining the cactus germplasm corresponding to the plant sample to be tested based on the spatial similarity calculation results", it may also include:
[0134] S71, calculate the spatial similarity between each pair of standard species vectors in the identification reference database to form a similarity matrix.
[0135] Understandably, for all known species in the identification reference database (e.g., 36 cactus species), the spatial similarity between each pair of species is calculated, generating a symmetric similarity matrix containing the spatial similarity values of all species pairs. This similarity matrix is m×m (where m is the number of cactus species), with diagonal elements being 0 (spatial similarity within the same cactus genus).
[0136] S72, extract all non-zero similarity values (i.e., spatial similarity between different cactus species pairs) from the similarity matrix to form a similarity dataset.
[0137] S73, based on the similarity dataset, uses the percentile method to determine the similarity threshold.
[0138] Specifically, in this embodiment, the percentile method is used to determine the similarity threshold. The similarity threshold is set to a certain percentile of the similarity dataset, such as the 95th percentile (meaning that the spatial similarity of 95% of cactus species pairs is less than this threshold), so as to ensure that the similarity threshold can cover the normal variation between most known species.
[0139] S74 compares the minimum value in the spatial similarity calculation results with the similarity threshold. If the minimum value in the spatial similarity calculation results is less than or equal to the similarity threshold, then the cactus germplasm corresponding to the plant sample to be tested identified in S6 is the final identification result, indicating that the plant sample to be tested falls within the multidimensional morphological space of known cactus species. If the minimum value in the spatial similarity calculation results is greater than the similarity threshold, then a manual review operation is triggered, indicating that the plant sample to be tested may be an abnormal sample, a hybrid, or a new species not included in the identification reference database.
[0140] For example, the way to trigger the manual review operation can be: prompting "cannot be reliably identified" or "may be an unknown species", and listing detailed comparative data of these plant samples to be tested and their key morphological traits, and forwarding them to identification experts for final manual judgment, thereby further ensuring the reliability of the final identification results.
[0141] Understandably, the inventive concept of threshold verification proposed in this embodiment stems from addressing the inherent uncertainty in morphological identification. Traditional identification methods rely on subjective judgment, which is difficult to standardize and replicate, and cannot effectively handle abnormal samples or unknown species, easily leading to misjudgments. This embodiment, by introducing a statistically based similarity threshold, transforms subjective experience into an objective standard. It then utilizes known spatial similarity distribution data among species to automatically set a similarity threshold, ensuring that decisions are based on overall data characteristics rather than individual experiences. Through threshold filtering, it identifies and processes plant samples exceeding the normal range of variation, balancing automation and accuracy, and effectively improving adaptability to complex real-world situations.
[0142] Example 3:
[0143] Please see the appendix Figure 2 , Figure 2 This is a schematic diagram of a morphological trait-based germplasm identification system for cacti provided in one embodiment of this specification.
[0144] like Figure 2 As shown, this morphological trait-based germplasm identification system for cacti may include at least:
[0145] Standardized measurement module 1 is used to construct a morphological trait data matrix of different cactus species using standardized measurement rules;
[0146] Morphological trait screening module 2 is used to screen key morphological traits based on the morphological trait data matrix using an improved principal component analysis method.
[0147] Standardization module 3 is used to standardize each key morphological property to obtain the standardized key morphological vector.
[0148] Database construction module 4 is used to construct an identification reference database containing standard species vectors based on standardized key morphological vectors;
[0149] The sample processing module 5 is used to collect key morphological data of the plant sample to be tested and perform standardization processing to obtain the standardized morphological vector of the sample to be tested.
[0150] Species identification module 6 is used to calculate the spatial similarity between the morphological vector to be tested and the standard species vectors in the identification reference database, and to determine the cactus germplasm corresponding to the plant sample to be tested based on the spatial similarity calculation results.
[0151] It is understood that the technical concept of the morphological trait-based cactus germplasm identification system provided in this embodiment is similar to the technical concept of the aforementioned morphological trait-based cactus germplasm identification method, and will not be repeated here.
[0152] Example 4:
[0153] Another embodiment of this specification provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps in the above-described method embodiments. The constituent modules of the above-described electronic device, if implemented as software functional units and used as independent downstream task predictions or applications, can be stored in the computer-readable storage medium.
[0154] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).
[0155] The above description is merely a preferred embodiment disclosed in this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of protection involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0156] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
Claims
1. A method for germplasm identification of *Cactus* species based on morphological traits, characterized in that, Includes the following steps: S1, a morphological trait data matrix of different cactus species was constructed using standardized measurement rules; S2, based on the morphological trait data matrix, the key morphological traits were screened using an improved principal component analysis method; S3, standardize each key morphological property to obtain the standardized key morphological vector; S4. Based on the standardized key morphological vectors, construct an identification reference database containing standard species vectors; S5. Collect key morphological data of the plant sample to be tested and standardize them to obtain the standardized morphological vector of the sample to be tested. S6. Calculate the spatial similarity between the morphological vector to be tested and the vectors of each standard species in the identification reference database. Based on the calculation results of spatial similarity, determine the cactus germplasm corresponding to the plant sample to be tested.
2. The method for germplasm identification of cacti based on morphological traits as described in claim 1, characterized in that, A morphological trait data matrix for different cactus species was constructed using standardized measurement rules, including: Quantitative morphological traits of cactus species were selected, including the number of spines, the spacing between nests, the plant height, the number of nests per row, the spine length, the palm length, the palm width, the seed diameter, the fruit width, and the fruit length. Establish corresponding measurement standards for each quantitative morphological trait to form standardized measurement rules; Based on standardized measurement rules, quantitative morphological trait measurement data of several individuals of each cactus species were collected and the average value was calculated to obtain the mean data of each quantitative morphological trait of each cactus species. The set of mean data of quantitative morphological traits of all cactus species constitutes a morphological trait data matrix.
3. The method for germplasm identification of cacti based on morphological traits as described in claim 1, characterized in that, Based on the morphological trait data matrix, a modified principal component analysis method was used to screen out key morphological traits, including: Based on the morphological data matrix, the corresponding covariance matrix is calculated; Solve the covariance matrix to obtain the eigenvalues and corresponding eigenvectors of each principal component; Based on the eigenvalues of each principal component, the contribution rate of each principal component is calculated. The contribution rate represents the percentage of variance explained by the principal component relative to the total variance. The contribution rates of each principal component are randomly combined, and each combination contains k principal components. The cumulative contribution rate of each combination is obtained, and the principal component in the combination with the largest cumulative contribution rate is selected as the important principal component. Based on the eigenvalues and corresponding eigenvectors of the important principal components, the loading values of each important principal component are calculated. The loading values represent the degree of contribution of quantitative morphological traits to the important principal components. The loading values of each important principal component were sorted from largest to smallest, and the quantitative morphological traits corresponding to the top few loading values were selected as key morphological traits.
4. The method for germplasm identification of cacti based on morphological traits as described in claim 3, characterized in that, Based on the eigenvalues and corresponding eigenvectors of the important principal components, the loading values of each important principal component are calculated, including: Calculate the square root of the eigenvalues of the important principal components, multiply it by the corresponding eigenvectors, and take the absolute value of the product. Use this absolute value as the loading value of the corresponding important principal components.
5. The method for germplasm identification of cacti based on morphological traits as described in claim 2, characterized in that, The key morphological traits are standardized to obtain standardized key morphological vectors, including: Calculate the mean and standard deviation of the measurement data of each key morphological trait in all cactus species; Subtract the measurement data of each key morphological trait from the corresponding average value, and then quotient the difference with the corresponding standard deviation to obtain the standardized key morphological vector.
6. The method for germplasm identification of cacti based on morphological traits as described in claim 1, characterized in that, Based on standardized key morphological vectors, an identification reference database containing standard species vectors is constructed, including: The standardized key morphological vectors corresponding to each cactus species are combined into a standard species vector, which is used to uniquely determine the coordinate position of the corresponding cactus species in the multidimensional morphological space. All standard species vectors are associated with their corresponding cactus genus and numbers to form an identification reference database.
7. The method for germplasm identification of cacti based on morphological traits as described in claim 6, characterized in that, Calculate the spatial similarity between the morphological vector of the test sample and the vectors of each standard species in the identification reference database. Based on the spatial similarity calculation results, determine the cactus germplasm corresponding to the test plant sample, including: Calculate the difference between the morphological vector to be tested and the vector of each standard species in each key morphological dimension, and then calculate the square of the difference. The spatial similarity between the morphological vector to be tested and each standard species vector is obtained by summing the squared differences in all key morphological dimensions across all key morphological dimensions. The standard species vector corresponding to the minimum spatial similarity is selected to determine that the plant sample to be tested is the cactus germplasm corresponding to the standard species vector.
8. The method for germplasm identification of cacti based on morphological traits as described in claim 1, characterized in that, After determining the cactus germplasm corresponding to the tested plant sample based on spatial similarity calculation results, the following steps are also included: Calculate the spatial similarity between each pair of standard species vectors in the identification reference database to form a similarity matrix; Extract all non-zero similarity values from the similarity matrix to form a similarity dataset; Based on the similarity dataset, the percentile method is used to determine the similarity threshold; The minimum value in the spatial similarity calculation result is compared with the similarity threshold. If the minimum value in the spatial similarity calculation result is less than or equal to the similarity threshold, the cactus germplasm corresponding to the plant sample to be tested in S6 is the final identification result. If the minimum value in the spatial similarity calculation result is greater than the similarity threshold, a manual review operation is triggered.
9. A germplasm identification system for cacti based on morphological traits, characterized in that, include: The standardized measurement module is used to construct morphological trait data matrices for different cactus species using standardized measurement rules; The morphological trait screening module is used to screen key morphological traits based on the morphological trait data matrix using an improved principal component analysis method. The standardization module is used to standardize each key morphological property to obtain the standardized key morphological vector. The database construction module is used to build an identification reference database containing standard species vectors based on standardized key morphological vectors; The sample processing module is used to collect key morphological data of the plant sample to be tested and perform standardization processing to obtain the standardized morphological vector of the sample. The species identification module is used to calculate the spatial similarity between the morphological vector of the test sample and the standard species vectors in the identification reference database. Based on the spatial similarity calculation results, the genetic material of the cactus species corresponding to the test plant sample is determined.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for germplasm identification of cactus plants based on morphological traits as described in any one of claims 1 to 8.