Identification method of paphiopedilum harzianum leaf back anthocyanin
By using image processing and K-means clustering algorithms from OpenCV and scikit-learn libraries, the problem of non-destructive, rapid, and low-cost anthocyanin detection in Paphiopedilum leaves was solved. This method enables accurate identification and area calculation of anthocyanin regions and is applicable to the detection of different species.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies cannot achieve non-destructive, rapid, and low-cost detection of anthocyanins in Paphiopedilum leaves, and the models have poor generalization ability, making them difficult to apply to different species.
Using OpenCV and scikit-learn libraries, anthocyanins on the underside of Paphiopedilum leaves were identified through image preprocessing and K-means clustering algorithm. Valid pixels were selected using RGB and HSV color spaces, and the colors were clustered into anthocyanin regions using K-means algorithm to calculate the anthocyanin area ratio.
It enables non-destructive, rapid, and accurate identification of anthocyanins on the underside of Paphiopedilum leaves, reducing technical barriers and costs, making it suitable for widespread application in ordinary laboratories, and providing intuitive and highly adaptable results.
Smart Images

Figure CN121747084A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of plant biotechnology, specifically to a method for identifying anthocyanins on the underside of Paphiopedilum leaves. Background Technology
[0002] Sclerophyllum oxypetalum ( Paphiopedilum micranthum As a national first-class protected plant unique to China, *Paphiopedilum sclerotium* possesses both extremely high ornamental value and medicinal potential. Its leaves are rich in anthocyanins, a key metabolite determining its ornamental traits and associated with the biosynthesis of various flavonoid medicinal active compounds. However, current research on *Paphiopedilum sclerotium* focuses primarily on molecular biology (such as specific primer development), resource conservation, and cultivation, with relatively little research on the association between its phenotypic characteristics and metabolites. Methods for rapidly evaluating associated metabolites using phenotypic characteristics with ornamental value are still scarce, and existing technologies have significant limitations, restricting the scientific conservation and development of this rare species.
[0003] Currently, non-destructive testing methods for plant anthocyanins mainly rely on image processing techniques or specialized instrument analysis. For example: 1. A scheme based on image processing and mathematical models CN106124506A (Preparatory Document 1) discloses a method for measuring anthocyanins in leaves by image processing and establishing a mathematical model. The method includes the following steps: 1) Establish a relationship model between leaf image parameters and leaf anthocyanin content Y=f(X), where Y represents leaf anthocyanin content and X represents the ratio of blue value to green value in the selected area of the leaf; 2) Measure the X value of the leaf to be tested, and calculate the anthocyanin content of the leaf to be tested by Y=f(X). The relationship model between the leaf image parameters and the leaf anthocyanin content is: Y=aX+bX2+c. The values of a, b, and c are calculated by measuring the Y value and the X value of different samples. The specific steps for measuring the X value of different samples include: selecting the leaf analysis area; acquiring the RGB image of the leaf using an RGB vision sensor; acquiring data of six parameters of the image features of the selected area of the leaf using Matlab software: red (R), green (G), blue (B), hue (H), saturation (S), and intensity (I); and calculating the ratio of the blue value to the green value of the selected area; the leaf is a lettuce leaf.
[0004] However, looking at this technical solution, on the one hand, the premise for establishing the model Y = f(X) is that the actual anthocyanin content (Y value) of a large number of samples must first be measured using chemical methods (such as ultraviolet spectrophotometry). This process is inherently destructive, requiring the grinding and extraction of leaves. Therefore, this method is essentially a "semi-non-destructive detection based on destructive calibration." For rare and endangered plants like Paphiopedilum sclerotium with small sample sizes, the large amount of destructive sampling required for the initial modeling violates my country's legal regulations for nationally protected plants and is therefore impossible to implement.
[0005] On the other hand, this model has poor generalization ability and limited versatility. The parameters a, b, and c in its model Y = aX + bX² + c are obtained by fitting data from a specific plant (such as lettuce). Different species have differences in leaf structure, color characteristics, and anthocyanin types, requiring the model to be refitted multiple times for species differences. It is a non-universal model that depends on a specific object and is difficult to apply directly to Paphiopedilum sclerotium.
[0006] 2. Hyperspectral instrument-based solution Publication No. CN105241822A (Prior Document 2) proposes a "Method for Determining Anthocyanin Content in Peony Leaves Based on Hyperspectral Analysis." This method integrates peony leaf reflectance spectral data with anthocyanin content data, determines the sensitive wavelength band for anthocyanin content through correlation analysis, constructs a spectral monitoring model for anthocyanin content based on the characteristic wavelength band in selected software, and then uses a hyperspectral radiometer to measure the reflectance spectral data of the peony leaves in the characteristic wavelength band. The data is then imported into the spectral monitoring model to calculate the anthocyanin content of the peony leaves. Although this scheme has high accuracy, it has the following problems: 1. The core of this solution is the use of a hyperspectral radiometer. Such equipment is expensive (usually tens of thousands to hundreds of thousands of RMB), complex to operate, and requires professional personnel. This strictly limits its application to professional research institutions with corresponding financial and technical strength, making it difficult to promote in ordinary laboratories, nurseries, or horticultural enterprises.
[0007] 2. In the first step of establishing this hyperspectral model, "sampling and data acquisition," it is clearly stated that "the anthocyanin content of each peony leaf needs to be measured separately." The method used to measure this content is inevitably a traditional, destructive chemical method (such as spectrophotometry). Therefore, this "non-destructive" detection method is itself based on large-scale destructive sampling and chemical analysis. For rare plants like *Paphiopedilum sclerotium*, the initial modeling itself is a formidable obstacle.
[0008] 3. This model is constructed based on spectral data and chemical measurements of peony leaves. The leaf tissue structure, epidermal wax, cell morphology, and specific types of anthocyanins vary among different plants, all of which affect their spectral characteristics. The same technical approach has been applied to anthocyanin detection in the leaves of various plants, including corn, wheat, rice, and apples, demonstrating that this model remains a non-universal model dependent on a specific species. Species changes inevitably lead to model reconstruction; for example, using it for *Paphiopedilum sclerotium* would result in significant prediction errors and a lack of reliability. Developing a separate hyperspectral model for *Paphiopedilum sclerotium* would then return to the issues of high cost and destructive sampling.
[0009] In summary, existing methods for detecting anthocyanins in plant leaves, whether based on mathematical models using RGB images (as in Comparative File 1) or on monitoring models based on hyperspectral imaging (as in Comparative File 2), all suffer from the following common and insurmountable drawbacks: First, they all belong to "pseudo-nondestructive" testing, and their core models must be calibrated through destructive sampling and chemical analysis, which makes them completely unsuitable for endangered protected plants with precious samples such as Paphiopedilum sclerotium.
[0010] Secondly, they rely heavily on fixed mathematical models for specific species, have poor generalization ability, and are difficult to transfer directly to different species (such as from lettuce / peony to hard-leaved slipper orchid), thus lacking versatility.
[0011] Therefore, for the specific protected species Paphiopedilum sclerotium and its research needs, developing a method that can achieve truly non-destructive testing without any prior destructive calibration, eliminate dependence on fixed mathematical models, and is simple to operate, low in cost, and suitable for widespread application has always been a technical problem that needs to be solved by those in the field. Summary of the Invention
[0012] The purpose of this invention is to provide a method for identifying anthocyanins on the underside of Paphiopedilum leaves, which can simply, clearly, and accurately observe the accumulation of anthocyanins on the underside of Paphiopedilum leaves, providing relatively strong technical support for the research of ornamental plants with variegated leaf characteristics such as Paphiopedilum, especially rare and endangered plants that are strictly protected by national regulations.
[0013] The Python code logic in this invention is as follows: First, the `cv2.imread` function provided by the open-source OpenCV library is used to read the input image file. The code is executed as "img = cv2.imread(image_path)". OpenCV reads images in BGR (Blue-Green-Red) format by default. To accommodate different processing needs, the code converts the image to two color spaces: 1. RGB color space: used for subsequent visualization display. "img_rgb = cv2.cvtColor(img,cv2.COLOR_BGR2RGB)" This code converts BGR to RGB to ensure that the image colors are displayed correctly; 2. HSV Color Space: "img_hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)" is used for color clustering and anthocyanin region identification. This code presents HSV (Hue, Saturation, Value) in a way that is more in line with human eye perception of color. The H component is directly related to the color type and is suitable for distinguishing regions such as purple and green. It is the core color space for subsequent anthocyanin identification.
[0014] Next, the HSV image undergoes channel separation and effective pixel filtering using the code "h, s, v = cv2.split(img_hsv)". This code aims to remove low-saturation background areas such as white and gray. A saturation threshold of "saturation_threshold=30" is then set, and an effective pixel mask is constructed: "valid_mask = s>saturation_threshold". In this layer of code, only pixels that might belong to anthocyanin-colored regions are assigned the value True to exclude background interference. Subsequently, the HSV pixels filtered through the mask are reshaped into a two-dimensional array "pixels = img_hsv[valid_mask].reshape(-1, 3)" to meet the subsequent K-means input format requirements.
[0015] Based on the KMeans clustering algorithm from the open-source scikit-learn library, "kmeans = KMeans(n_clusters=k, random_state=42), labels = kmeans.fit_predict(pixels), centers =kmeans.cluster_centers_", pixels in the color space are divided into k categories, making the pixels within each category as close as possible to their cluster centers in the HSV space. In this invention, experiments have verified that setting 20 cluster centers effectively distinguishes anthocyanin regions, green leaf mesophyll regions, white backgrounds, and transitional tones for the color distribution of images of the underside of *Paphiopedilum sclerotium* leaves. Therefore, the code is set as "def extract_anthocyanin_ratio(image_path, k=20, saturation_threshold=30):", which splits purple and green of different shades, brightness, and saturation into multiple clusters; multiple clusters can be flexibly selected and merged into anthocyanin categories, reducing the situation where different color components are mixed in the same cluster.
[0016] After clustering is completed, first analyze the HSV values of each cluster center to determine which clusters may represent anthocyanin colors. The code is "anthocyanin_clusters = []; for i, c in enumerate(centers): hue = c[0]; if 120 < hue < 220 or hue < 6: anthocyanin_clusters.append(i)". Verified by the experiments of the present invention, anthocyanin color development mainly concentrates in the vicinity of red: Hue < 6 and the blue-violet to purple region: 120 < Hue < 220. After numbering the anthocyanin cluster numbers identified by the retrieval, construct the corresponding mask on the entire image "label_map = np.full(valid_mask.shape, -1, dtype=np.int8); label_map[valid_mask]=labels; anthocyanin_mask=np.isin(label_map, anthocyanin_clusters).astype(np.uint8) * 255", and further generate a red overlay visualization image based on this mask "red_overlay = np.zeros_like(img); red_overlay[anthocyanin_mask == 255]= [0, 0, 255]; pure_red_highlight =np.where(anthocyanin_mask[..., None].repeat(3, axis=2),red_overlay,img)".
[0017] After completing the marking of the anthocyanin region, calculate the proportion of the anthocyanin area by counting the number of pixels in the mask "anthocyanin_pixels = np.sum(anthocyanin_mask == 255); total_pixels = np.sum(valid_mask); ratio = anthocyanin_pixels / total_pixels if total_pixels>0 else 0". Finally, output the result "return ratio, img_rgb, pure_red_highlight", and the displayed result shows the two images side by side, and the title of the annotation map is accompanied by the anthocyanin ratio.
[0018] The above specific flowchart is shown in the following figure:
[0019] Figure 13 Furthermore, the method for identifying anthocyanins on the underside of Paphiopedilum leaves in this application includes the following steps: Step 1: Preload all the open-source environments required for Anaconda to run on your computer; Step 2: Image acquisition and preprocessing, specifically: Take a picture of the complete back of a Paphiopedilum leaf in the center of the photo, crop out the excess part in the photo editing software, and fill all areas except the back of the leaf with white to obtain the processed image and name and save it. Step 3: Enable the Anaconda PowerShell prompt and load the anthocyanin identification code file. Specifically: Place the Python code file for anthocyanin recognition in the computer's user root directory, then enable AnacondaPowerShell prompt by typing `conda activate`; subsequently, type `jupyter notebook` to start the JupyterNotebook service. Step 4: In the Jupyter Notebook, open the Python code file for anthocyanin recognition and enter the code execution interface; Step 5: Set the number of color clusters, specifically: In the code block 'def extract_anthocyanin_ratio(image_path, k=, saturation_threshold=30):' in the Python code file, set the number of color clusters k to 10; Step Six: Load the image and run the code, specifically: Enter the filename of the processed image saved in step two into the image path variable specified in the code, click the run button to execute the code and obtain the result of the proportion of anthocyanin to the leaf underside area.
[0020] Furthermore, the specific value of the number of color clusters k in step five is 10, which is the best and optimal choice for recognition effect.
[0021] Furthermore, the output of step six also includes a visualization image in which the anthocyanin recognition area is highlighted in a prominent color. Beneficial effects
[0022] 1. Accurate, Non-destructive, and Convenient Recognition: Based on the color clustering algorithm (K-means), the preprocessed leaf underside image is analyzed, achieving non-destructive and rapid recognition of anthocyanin regions. Preprocessing removes invalid background, effectively eliminating interference from non-target areas and resulting in more accurate recognition.
[0023] 2. High adaptability and intuitive results: This application avoids complex model training and has high human controllability. It directly outputs the anthocyanin area percentage, providing a quantitative and intuitive result.
[0024] 3. Low cost and easy to promote: The entire solution is based on an open-source software environment and Python code, without relying on expensive and complex special instruments. The operation process is simple, which greatly reduces the technical threshold and usage cost. It is particularly suitable for promotion and application in ordinary laboratories and has positive practical significance for the scientific research and protection of Paphiopedilum sclerotium. Attached Figure Description
[0025] Figure 1 Image showing the underside of a Paphiopedilum leaf.
[0026] Figure 2 Image preprocessing demonstration.
[0027] Figure 3 The location of the user's root directory where the images and recognition Python code files used in Example 1 are stored.
[0028] Figure 4 In Example 1, after the Anaconda PowerShell prompt is launched, a preloaded anthocyanin recognition Python code file startup item is loaded.
[0029] Figure 5 : A diagram showing the location of Python code files in the notebook page in Example 1.
[0030] Figure 6 Example 1 sets the number and range of color clusters in the item column.
[0031] Figure 7 Example 1: Pasting images and launching Python code files in the project panel.
[0032] Figure 8 When the number of color clusters K = 10, the final recognition effect diagram of Example 1 is shown.
[0033] Figure 9 When the number of color clusters K = 14, refer to the final recognition result diagram in Example 1.
[0034] Figure 10 When the number of color clusters K=3, refer to the final recognition result diagram in Example 2.
[0035] Figure 11 When the number of color clusters K = 8, refer to the final recognition result diagram in Example 3.
[0036] Figure 12 When the number of color clusters K = 20, refer to the final recognition result diagram in Example 4.
[0037] Figure 13 The flowcharts show the key algorithm processes, judgment conditions, and parameter selection criteria used in the code of this application. Detailed Implementation
[0038] The distinctive purplish-red spots on the underside of *Paphiopedilum sclerotium* leaves, caused by anthocyanin accumulation, possess not only high ornamental value but also serve as a crucial feature for studying the physiological and ecological adaptation mechanisms of this orchid. This unique phenotype is highly representative among all variegated-leaf orchids. As a typical variegated-leaf orchid, the accumulation and aggregation of anthocyanins on the underside of *Paphiopedilum sclerotium* leaves can change due to variations in the growth environment or external stimuli. Researchers can use this intuitive phenotypic information to correlate changes in anthocyanin content, enabling rapid predictions of the orchid's growth status and reducing the cost of fully orthogonal experiments without background information. This readily observable phenotypic characteristic makes it the best material for studying anthocyanin metabolism changes in orchids.
[0039] In the embodiments of this invention, *Paphiopedilum sclerotium*, native to the river valleys of Wenshan City, Malipo County, and Xichou County under the jurisdiction of Wenshan Prefecture, Yunnan Province, was selected as the research object. As the only native habitat of *Paphiopedilum sclerotium* in Yunnan Province (*Paphiopedilum sclerotium* in my country is only distributed in Wenshan City, Malipo County, and Xichou County, Wenshan Prefecture, Yunnan Province), this avoids the errors caused by ex-situ / near-situ conservation in shaping plant phenotypes and ensures that the method has good operability for *Paphiopedilum sclerotium* growing under all different conservation conditions. The invention will be further described in detail below with specific examples. Example 1:
[0040] The method for identifying anthocyanins on the underside of Paphiopedilum leaves in this embodiment includes the following steps: Step 1: Preload all the open-source environments required for Anaconda to run on your computer; Step 2: Image acquisition and preprocessing, specifically: Select a healthy, growing Paphiopedilum sclerotium plant. Choose mature leaves with intact shape and relatively flat surfaces, based on the plant's growth. With the leaf back facing up, ensure the leaf tip is on the left and the leaf base on the right before taking a photo. The entire underside of the Paphiopedilum leaf should be centered in the photo (e.g.,...). Figure 1 (As shown in the image). Note that when taking photos, ensure that the back of the leaf is clear and intact, and that there is no light or extraneous objects interfering with the image.
[0041] In photo editing software, crop out the excess parts of the photograph and fill in all areas except the underside of the leaf with white (e.g., ...). Figure 2 As shown), name the processed photo "example.png" and save it in the user's root directory (e.g., Figure 3 (as shown) Step 3: Enable the Anaconda PowerShell prompt and load the anthocyanin identification code file. Specifically: Place the Python code file for anthocyanin recognition in the computer's user root directory, then enable AnacondaPowerShell prompt by typing `conda activate`; subsequently, type `jupyter notebook` to start the JupyterNotebook service (e.g., ...). Figure 4 (as shown) It should be noted that the complete source code file required for the implementation of this invention can be found in the electronic attachment "ga.ipynb" of this application, which is submitted electronically along with the patent application.
[0042] Step 4: In the Jupyter Notebook, open the Python code file for anthocyanin recognition (e.g., ...). Figure 5 As shown), enter the code execution interface; Step 5: Set the number of color clusters, specifically: In the Python code file, in the code block 'def extract_anthocyanin_ratio(image_path, k=, saturation_threshold=30):', change the value after k= to 10 (e.g., Figure 6 (as shown) Step Six: Load the image and run the code, specifically: Copy the photo processed as required in step two, and enter a name after the "r'. / " character. In this example, the running image is entered as '(r'. / example.png')', and the result is as shown in the example. Figure 7 As shown.
[0043] Click the "▶" symbol in the Python code file interface to start running the Python code file, and wait for the Python code file to run and calculate before outputting the results.
[0044] Output results: Anthocyanin pixels account for 72.36%. Simultaneously, an overlay image is generated and displayed, where areas identified as anthocyanins are marked with a red semi-transparent overlay, highly overlapping with the actual purplish-red area on the underside of the leaf.
[0045] The standardized requirements for accurately identifying and determining the specific results of anthocyanin identification are as follows: (1) The identified image is complete and clearly shows the leaf veins.
[0046] (2) After the identification is completed, there are no unidentified anthocyanin regions.
[0047] (3) After the identification is completed, there is no case where anthocyanin-free regions are identified.
[0048] (4) The red-covered area highly coincides with the actual anthocyanin location.
[0049] If all of the above conditions are met, and the anthocyanin recognition area is accurate and complete, then the test result is considered reliable.
[0050] In this embodiment, the identified image is complete and clearly shows the leaf veins. After identification, there are no unidentified anthocyanins, and no areas without anthocyanins are identified. The red-covered areas highly overlap with the actual anthocyanin locations, as shown in the example. Figure 8 As shown. Compare with Example 1:
[0051] Repeat steps one through four of Example 1; Step 5: In the code block of the Python code file, 'def extract_anthocyanin_ratio(image_path, k=, saturation_threshold=30):', change the value after k= to 14; Repeat step six in step 1.
[0052] Output result: When k=14 in step six, the anthocyanin pixel ratio is 66.22%. This can be seen from the recognized image. Figure 9 As shown in the figure, after the identification is completed, the red covered area generally does not overlap with the actual anthocyanin location. There are unidentified anthocyanins and areas without anthocyanins are identified. Compare with Example 2:
[0053] Repeat steps one through four of Example 1; Step 6: In the code block of the Python code file, 'def extract_anthocyanin_ratio(image_path, k=, saturation_threshold=30):', change the value after k= to 3; Repeat step six in step 1.
[0054] Output result: When k=3 in step six, the proportion of anthocyanin pixels is 16.98%, as can be seen from the recognized image. Figure 10As shown in the figure, the red-covered area has a very low degree of overlap with the actual anthocyanin location, with a large number of unidentified anthocyanins and some areas without anthocyanins being identified. Compare with Example 3:
[0055] Repeat steps one through four of Example 1; Step 6: In the code block of the Python code file, 'def extract_anthocyanin_ratio(image_path, k=, saturation_threshold=30):', change the value after k= to 8; Repeat step six in step 1.
[0056] Output result: When k=8 in step six, the proportion of anthocyanin pixels is 56.63%, as can be seen from the recognized image. Figure 11 As shown in the figure, the red-covered area does not overlap much with the actual anthocyanin location. There are unidentified anthocyanins in the upper right and lower middle parts of the leaf back, while areas without anthocyanins are identified on the left side of the leaf back and at the leaf tip. Compare with Example 4:
[0057] Repeat steps one through four of Example 1; Step 6: In the code block of the Python code file, 'def extract_anthocyanin_ratio(image_path, k=, saturation_threshold=30):', change the value after k= to 20; Repeat step six in step 1.
[0058] Output result: When k=20 in step six, the anthocyanin pixel ratio is 20.88%, which can be seen from the recognized image. Figure 12 As shown in the figure, after the identification is completed, the red covered area has a low degree of overlap with the actual anthocyanin location, and there are a large number of unidentified anthocyanins. There are also cases where areas without anthocyanins in the lower part of the leaf are identified.
[0059] The above examples and comparative examples demonstrate that the method of the present invention can effectively identify anthocyanins on the underside of *Paphiopedilum sclerotium* leaves. By adjusting the number of color clusters, k, it can adapt to different image conditions and recognition accuracy requirements. Specifically, k=10 achieved the best recognition performance under the experimental conditions.
[0060] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for identifying anthocyanins on the underside of Paphiopedilum leaves, characterized in that: Includes the following steps: Step 1: Preload all the open-source environments required for Anaconda to run on your computer; Step 2: Image acquisition and preprocessing, specifically: Take a picture of the complete back of a Paphiopedilum leaf in the center of the photo, crop out the excess part in the photo editing software, and fill all areas except the back of the leaf with white to obtain the processed image and name and save it. Step 3: Enable the Anaconda PowerShell prompt and load the anthocyanin identification code file. Specifically: Place the Python code file for anthocyanin recognition in the computer's user root directory, then enable AnacondaPowerShell prompt by typing `conda activate`; subsequently, type `jupyter notebook` to start the JupyterNotebook service. Step 4: In the Jupyter Notebook, open the Python code file for anthocyanin recognition and enter the code execution interface; Step 5: Set the number of color clusters, specifically: In the code block 'def extract_anthocyanin_ratio(image_path, k=, saturation_threshold=30):' in the Python code file, set the number of color clusters k to 10; Step Six: Load the image and run the code, specifically: Enter the filename of the processed image saved in step two into the image path variable specified in the code, click the run button to execute the code and obtain the result of the proportion of anthocyanin to the leaf underside area.
2. The method for identifying anthocyanins on the underside of Paphiopedilum leaves according to claim 1, characterized in that: The output of step six also includes a visualization image in which the anthocyanin-recognized regions are highlighted in a prominent color.
Citation Information
Patent Citations
Measurement method of content of anthocyanin in leaves of peony on the basis of hyperspectrum
CN105241822A
Leaf anthocyanin content measuring method
CN106124506A