Plant raw material identification method

By combining liquid chromatography-mass spectrometry (LC-MS) with principal component analysis (PCA), an anthocyanin fingerprint model for bilberry, blueberry, and cranberry anthocyanins was established. This solved the problem of identifying the source of anthocyanins in bilberry, blueberry, and cranberry anthocyanins, enabling accurate identification of these anthocyanins and ensuring the authenticity and accuracy of dietary supplement ingredients.

CN121027391APending Publication Date: 2025-11-28HEILONGJIANG FEIHE DAIRY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511129317.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-06-06
Filing Date
2025-08-12
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies are insufficient to accurately identify anthocyanins from plant sources such as blueberries, bilberries, and cranberries, leading to adulteration and labeling errors, which pose health risks and economic losses.

Method used

An anthocyanin fingerprint model was established using liquid chromatography-mass spectrometry combined with principal component analysis (PCA) and Mahalanobis distance classification. Through normalization and visualization, anthocyanin components from different plant sources were distinguished.

Benefits of technology

It enables accurate identification of anthocyanins from blueberries, bilberries, and cranberries, ensuring the authenticity and accuracy of dietary supplement ingredients and reducing the risk of adulteration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121027391A_ABST
    Figure CN121027391A_ABST
Patent Text Reader

Abstract

The invention relates to a plant source identification method, in particular to a chemometrics method based on combination of a liquid chromatography-mass spectrometry anthocyanin fingerprint spectrum, principal component analysis (PCA) and a mahalanobis distance classification method, and is applied to the field of food.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for identifying plant origin, which is a chemometric method based on liquid chromatography-mass spectrometry of anthocyanin fingerprint combined with principal component analysis (PCA) and Mahalanobis distance classification, used in the field of food. BACKGROUND

[0002] Vaccinium is a genus of terrestrial shrubs widely distributed in the Northern Hemisphere, comprising more than 400 species. Among them, V. myrtillus L., V. corymbosum L. and V. macrocarpon Aiton have high economic value. Vaccinium berries are rich in polyphenolic compounds, especially anthocyanins, which give them strong antioxidant capacity. Therefore, Vaccinium has a long history of use in traditional medicine and is widely consumed as food. Modern clinical studies support the multiple therapeutic effects of Vaccinium plants in preventing or treating cardiovascular diseases, diabetes, obesity, cancer, urinary tract infections and senile diseases.

[0003] The recent surge in demand for dietary ingredients and dietary supplements containing Vaccinium plants has led to quality problems. In many cases, blueberry products are intentionally or unintentionally sold as Vaccinium. Cranberry juice or extract is completely or partially replaced by cheaper anthocyanin sources such as grape seed, black rice and mulberry extract.

[0004] In a 2016 study, it was reported that more than 30% of the Vaccinium products surveyed did not match the claimed anthocyanin content. It was also reported that 45% of Vaccinium-containing dietary supplements failed to pass certification tests. Several researchers pointed out that up to 24% of extracts and 70% of Vaccinium supplements had label errors in species composition or active ingredient dosage.

[0005] Although most adulteration or label errors have limited impact on consumer health, some cases can pose serious health risks due to the presence of toxic or allergenic substances and can result in losses. Allergic reactions are a common problem with berry adulteration, as consumers can react to certain berries, but not all berries.

[0006] In several studies, adulteration or label errors of Vaccinium products were found by comparing the anthocyanin profile in the sample with that of the standard material. Anthocyanins are naturally occurring water-soluble phenolic compounds that are the source of red, blue and purple color in many fruits, vegetables and cereals. The chemical structure of anthocyanins varies depending on the type of aglycone (anthocyanidin), as well as the type, number and location of glycosylation and acylation.

[0007] To date, over 700 different anthocyanins have been identified. The differences in anthocyanin composition between blueberry and non-blueberry species are mainly attributed to their inherent genetic differences. For example, bilberry contains a variety of anthocyanins, including delphinidin, cyanidin, petunidin, peonidin, and malvidin derivatives. In contrast, cranberry is dominated by peonidin, cyanidin, and malvidin derivatives. In addition, non-blueberry species such as Rubus subg. rubus, Aronia, and others also contain high levels of a variety of anthocyanins. In addition to interspecies differences, intraspecies variation has also been observed, which is influenced by factors such as cultivar, ripeness, and geographic region. For example, the amount of anthocyanin accumulation in blueberries increases as they ripen, particularly the content of proanthocyanidins and anthocyanins. Wild blueberries can have 1.7 times more total anthocyanin content per fruit than cultivated blueberries. Anthocyanin profiles are also affected by adverse environmental conditions. For example, the content of delphinidin, cyanidin, and petunidin derivatives in blueberries increases after mechanical damage or insect infestation, while the content of malvidin derivatives decreases. In addition, various manufacturing processes, such as extraction and purification, can lead to the hydrolysis and degradation of anthocyanins. Therefore, relying solely on the baseline anthocyanin profile in certified reference materials may not be sufficient to reliably identify bilberry products, and considering seasonal, temporal, and geographic variations is crucial for conducting robust authentication studies.

[0008] One strategy to address this challenge is to employ a combination of analytical techniques or an orthogonal approach to gain a more comprehensive understanding of the target sample's chemical composition. This synergistic approach can provide powerful insights into the chemical relationships within a species. One orthogonal approach is to first generate a computer-simulated database containing anthocyanin profiles from different real bilberry samples that cover different cultivars, growing regions, and environmental conditions. Then, new samples can be compared against the reference library rather than a single reference sample to determine their similarity to known profiles.

[0009] Principal component analysis (PCA) is an existing mature unsupervised chemometric method that helps identify samples that may be mislabeled by analyzing the chemical variation between and within species after dimensionality reduction.

[0010] To date, there have been limited studies on detecting adulteration in dietary supplements containing bilberry, and no universal analytical method has been published to identify different types of bilberry. SUMMARY

[0011] Problems to be solved by the invention

[0012] To address the problem of difficult identification of anthocyanin sources, the goal of the present invention is to develop a chemometric method for distinguishing between blueberry, cranberry species or other species of anthocyanin profiles and to assess the anthocyanin sources of existing anthocyanin-containing nutritional supplements.

[0013] In combination with the Mahalanobis distance decision boundary, the present invention applies principal component analysis (PCA) to three standard samples (blueberry, cranberry and bilberry) and commercial nutritional supplements that use or may use anthocyanins, and visualizes the characterization and comparison, thereby avoiding the influence of false adulteration or misidentification.

[0014] The present study is based on liquid chromatography-tandem mass spectrometry (LC-MS / MS) technology, using known standard samples as a training data set to define and optimize computer-simulated regional classification, and further testing the test samples to distinguish whether the real samples have potential adulteration risks.

[0015] Solution for solving the problem

[0016] The above technical problems can be solved by implementing the following technical solutions:

[0017] The plant raw material identification method of the present invention is a principal component analysis identification method based on mass spectrometry analysis and combined with Mahalanobis distance analysis, which comprises:

[0018] The steps of model establishment and the steps of identification,

[0019] The step of model establishment comprises:

[0020] i. Taking blueberry, cranberry and bilberry as three standard samples, and extracting the anthocyanin components contained in each standard sample to obtain three standard sample extracts;

[0021] ii. For each of the three standard sample extracts, the content of the selected anthocyanin components in the standard sample is analyzed by liquid chromatography-mass spectrometry, and further, the content of each selected anthocyanin component is normalized based on the total content of the selected anthocyanin components to obtain the mass spectrometry peak area ratio of each selected anthocyanin component relative to the total selected anthocyanin components;

[0022] The selected anthocyanin components at least include anthocyanin components classified as cyanidin (Cy), delphinidin (Dp), petunidin (Pt), peonidin (Pn), pelargonidin (Pg) and malvidin (Mv),

[0023] iii. performing principal component analysis on the three sets of data obtained from the normalization process, and establishing three non-overlapping distribution areas corresponding to the components of each set of data, wherein the principal component analysis uses anthocyanins and the ratio of cyanidin (Cy) to malvidin (Mv) as variables, and the distribution areas are established based on Hotelling's T2 analysis with a 95% confidence level, wherein the distribution area derived from the cranberry extract is designated as distribution area A, the distribution area derived from the blueberry extract is designated as distribution area B, and the distribution area derived from the raspberry extract is designated as distribution area C

[0024] The step of identifying comprises:

[0025] i'. selecting the object to be identified, and extracting the anthocyanin components contained therein;

[0026] ii'. for each of the three standard sample extracts, detecting the content of the selected anthocyanin components in the standard sample by a detection system, which comprises liquid chromatography-mass spectrometry, and further normalizing the content of each of the selected anthocyanin components based on the total content of the selected anthocyanin components, to obtain the ratio of each selected anthocyanin component relative to the total selected anthocyanin components,

[0027] The selected anthocyanin components are the same as, different from, or partially the same as the anthocyanin components in step ii;

[0028] iii'. the data obtained from the normalization process are processed in the same manner as step iii to obtain the distribution area D of the object to be identified.

[0029] Further, the positional relationship of the distribution area D is compared with the distribution areas A, B, and C.

[0030] In some specific embodiments, the extraction is performed in an alcohol solution with an acid.

[0031] In some specific embodiments, the detection system comprises a high-performance thin-layer chromatography (HPTLC) system with a photodiode array detector (PDA).

[0032] In some specific embodiments, the liquid chromatography-mass spectrometry is performed in a multiple reaction monitoring mode (MRM), and data are collected and processed using software.

[0033] In some specific embodiments, the anthocyanins in step ii include at least the following 18 anthocyanins: Dp-3-galactoside, Dp-3-glutamic acid, Dp-3-arabinoside, Cy-3-galactoside, Cy-3-glutamic acid, Pt-3-galactoside, Pt-3-glutamic acid, Pg-3-galactoside, Cy-3-arabinoside, Pg-3-glutamic acid, Pt-3-arabinoside, Pn-3-galactoside, Mv-3-galactoside, Pg-3-arabinoside, Pn-3-glutamic acid, Mv-3-glutamic acid, Pn-3-arabinoside, and Mv-3-arabinoside.

[0034] In some specific embodiments, after the normalization process, the normalized data is outputted or stored in a form of machine-readable data used in the principal component analysis.

[0035] In some specific embodiments, the principal component analysis is processed by computer software, in which a Mahalanobis distance analysis program is embedded, and preferably, the computer software is RStudio software.

[0036] In some specific embodiments, the principal component analysis is outputted in a form of visualized two-dimensional coordinate graph, and each distribution area is displayed in the coordinate graph.

[0037] In some specific embodiments, in step iii, the distribution area A, the distribution area B, and the distribution area C appear in the form of an ellipse.

[0038] Further, the present application provides a method for identifying the plant source of anthocyanins in food, which comprises the method according to any one of the above methods to identify whether the anthocyanins in the food are derived from bilberry or whether they are derived from any one of bilberry, blueberry, or cranberry.

[0039] In some specific embodiments, the food includes a nutritional supplement which is in liquid, semi-solid, or solid state at room temperature.

[0040] Effects of the application

[0041] Through the implementation of the above technical solutions, the present application can obtain the following technical effects:

[0042] 1) The present application first proposes a method based on LC-MS / MS anthocyanin fingerprint combined with principal component analysis (PCA) and Mahalanobis distance classification model, which is used to identify the characteristic composition of anthocyanins from three main Vaccinium plant raw materials (Vaccinium myrtillus, Vaccinium corymbosum, and Vaccinium macrocarpon);

[0043] 2) The model based on the above analytical method can accurately and efficiently verify whether the anthocyanins in the existing dietary supplements containing anthocyanins are derived from bilberry or not, or whether they are derived from any of bilberry, blueberry or cranberry or from other plant resources. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 : Average peak area percentages of 18 selected anthocyanins in the three target species (V. mytillus, V. corybosum and V. macrocarpon), bar represents standard deviation.

[0045] Figure 2 : Principal component analysis (PCA) results for distinguishing the three Vaccinium species: V. myrtillus (blue), V. macrocarpon (green) and V. corybosum (brown). PCA score plots show the models built from (A) 18 anthocyanins and (B) 18 anthocyanins and the ratio of cyanidin (Cy) to malvidin (Mv) derivatives, respectively.

[0046] Figure 3 : The proposed models were validated to correctly classify new samples by adding new BRMs from each target species. (A) V. myrtillus validation, (B) V. corybosum validation, and (C) V. macrocarpon validation.

[0047] Figure 4 : (A) Examples of classification of new, non-target samples outside the three target groups, red triangles represent new blackberry bio-regulators (BRMs) that are chemically different from the target groups. PCA shows that all non-target samples are different from the anthocyanin profiles of (B) V. mytillus, (C) V. corymbosum and (D) V. macrocarpon.

[0048] Figure 5 : Identification of consumer products using principal component analysis (PCA) based on 18 anthocyanins and the ratio of Cy / Mv derivatives: A) V. macrocarpon supplement; B) V. corybosum supplement; C) V. myrtillus supplement; D) S. nigra supplement. Model classification of (E) V. myrtillus supplement (arrow points to missing band in supplement) and (F) S. nigra supplement was confirmed by high-performance thin-layer chromatography (HPTLC). DETAILED DESCRIPTION

[0049] Hereinafter, the present application will be described in detail. The description of the technical features described below is based on representative embodiments, specific examples of the present application, but the present application is not limited to these embodiments, specific examples. Note that:

[0050] In the present specification, a numerical range represented by "numerical value A to numerical value B" means a range including the end point values A and B.

[0051] In the present specification, a numerical range represented by "above" or "below" means a range including the present number.

[0052] In the present specification, the meaning represented by "may" includes both the meaning of performing a certain process and the meaning of not performing a certain process.

[0053] In the present specification, "optionally" or "optional" means that a certain substance, component, step, condition, etc. is used or not used.

[0054] In the present specification, "room temperature" or "ambient temperature" means an indoor environmental temperature of "23±2°C".

[0055] In the present specification, "multiple" means a number of 2 or more.

[0056] In the present specification, the unit name used is the international standard unit name, and if not specifically stated, "%" used means a weight or mass percentage content.

[0057] In the present specification, "substantially" or "essentially" means that the standard deviation from a theoretical model, theoretical data, or target data is within a numerical range of 1%, preferably 0.8%, and more preferably 0.5%.

[0058] In the present specification, the terms "comprising" and / or "including" mean that a feature, step, operation, device, component, and / or combinations thereof are present.

[0059] In the present specification, "some specific / preferred embodiments", "other specific / preferred embodiments", "embodiments", and the like mean that the specific elements (e.g., features, structures, properties, and / or characteristics) described in relation to the embodiments are included in at least one embodiment described herein, and can be present in other embodiments or can not be present in other embodiments. In addition, it should be understood that the elements can be combined in various embodiments in any suitable manner.

[0060] The present application provides a method for identifying the source of plant ingredients based on a cyanine fingerprint combined with principal component analysis (PCA) and Mahalanobis distance classification model using liquid chromatography-mass spectrometry, which is used to identify three main Vaccinium plant raw materials (Vaccinium myrtillus, V. corymbosum, and V. macrocarpon) and identify adulterants that are different from the above three raw materials and can be used in nutritional supplements, thereby achieving efficient and accurate identification of the authenticity of dietary supplement ingredients.

[0061] Further, the method of the present application is based on the establishment of a standard model and the use and comparison (i.e., identification) of the model.

[0062] (Step of model establishment)

[0063] The establishment of the model of the present application is based on the principal component analysis and visual characterization of three standard samples of Vaccinium, blueberry, and cranberry.

[0064] Specifically, steps i-iii are included as follows:

[0065] Step i:

[0066] i. Vaccinium, blueberry, and cranberry are used as three standard samples, and the cyanine components contained in each standard sample are extracted.

[0067] There is no particular restriction on the source of Vaccinium, blueberry, and cranberry in principle, and they can be those of recognized quality, nutritional value, or commercial preference.

[0068] In some specific embodiments, for example, blueberry can be V. corymbosum L., cranberry can be V. macrocarpon Aiton, and Vaccinium can be V. myrtillus L.

[0069] There is no particular restriction on the extraction step in principle, for example, it can be alcohol extraction or water extraction, and in some preferred embodiments, an acidic alcohol solution, such as a methanol solution containing formic acid or an ethanol solution containing formic acid, can be used.

[0070] There is no particular restriction on the temperature of extraction, which can be performed at room temperature or at a temperature not exceeding 60°C.

[0071] In some specific embodiments, the standard can be crushed during extraction, preferably under ultrasonic / stirring conditions.

[0072] After a period of extraction, the sample extract solution containing the target components can be separated by dilution, centrifugation, filtration, etc. for subsequent detection.

[0073] Step ii:

[0074] ii. For each of the three standard sample extracts, the content of each of the selected anthocyanin components is analyzed by the detection system, and further, the content of each of the selected anthocyanin components is normalized based on the total content of the selected anthocyanin components to obtain the peak area ratio of each of the selected anthocyanin components relative to the total content of the selected anthocyanin components.

[0075] The detection system in the present application comprises a liquid chromatography-mass spectrometry system. The liquid chromatography is a high-performance thin-layer chromatography (HPTLC) system, and has a photodiode array detector (PDA) as a signal detector.

[0076] Further, for the liquid chromatography-mass spectrometry system, preferably, a multiple reaction monitoring (MRM) mode, in particular a positive ion mode, is adopted.

[0077] The above LC-MS / MS is used to detect the extracts of the three standard samples, and the content of each of the selected anthocyanins is analyzed.

[0078] The selected anthocyanin components at least include the anthocyanin components of cyanidin (Cy), delphinidin (Dp), petunidin (Pt), peonidin (Pn), pelargonidin (Pg), and malvidin (Mv).

[0079] In some preferred embodiments, the selected anthocyanins can at least include the following 18 anthocyanins: Dp-3-galactoside, Dp-3-glutamate, Dp-3-arabinoside, Cy-3-galactoside, Cy-3-glutamate, Pt-3-galactoside, Pt-3-glutamate, Pg-3-galactoside, Cy-3-arabinoside, Pg-3-glutamate, Pt-3-arabinoside, Pn-3-galactoside, Mv-3-galactoside, Pg-3-arabinoside, Pn-3-glutamate, Mv-3-glutamate, Pn-3-arabinoside, and Mv-3-arabinoside.

[0080] Step iii:

[0081] iii. The three groups of data obtained by the normalization process are subjected to principal component analysis to establish three non-overlapping distribution areas corresponding to the components of each group of data, wherein the principal component analysis uses anthocyanins and the ratio of cyanidin (Cy) to malvidin (Mv) as variables, and the distribution areas are established based on Hotelling's distance analysis with a confidence level of 95%, wherein the distribution area derived from the cranberry extract is designated as distribution area A, the distribution area derived from the blueberry extract is designated as distribution area B, and the distribution area derived from the raspberry extract is designated as distribution area C.

[0082] There is no particular limitation on the normalization method, which can be performed using software such as Excel, and the processed data can be further output or stored in a machine-readable data format in the principal component analysis. In some specific embodiments, the normalized data (peak area percentage) ratio can be exported or saved as a CSV file.

[0083] There is no particular limitation on the principal component analysis method, which can be performed using various software available in the art, for example, the above-mentioned CSV file can be analyzed in RStudio. In a separate CSV file, the samples are labeled as the target species (raspberry, blueberry, or cranberry), etc.

[0084] Principal component analysis (PCA) is used to visualize the anthocyanin pattern of a specific species, and the principal component analysis of the present application includes Mahalanobis distance analysis.

[0085] PCA is an unsupervised multivariate statistical technique and a data dimensionality reduction method. It simplifies multiple variables into smaller principal components (PCs), which describe how samples vary and correlate according to overall chemical characteristics. The resulting score plot brings samples with more similar characteristics together, while samples with more distinct characteristics are further apart. In this case, the model discovers the relationship between the percentage of each anthocyanin peak area, thereby classifying samples according to their overall anthocyanin characteristics.

[0086] By selecting appropriate variables in the PCA analysis, the distribution of the three substances can be obtained in a non-overlapping manner in a PC view (which can be a two-dimensional coordinate graph, such as a PC1-PC2 two-dimensional coordinate graph). Specifically, in the principal component analysis, anthocyanins (species / relative content) and the ratio of cyanidin (Cy) to malvidin (Mv) are used as variables, and the distribution area is established based on Hotelling's distance analysis with a confidence level of 95%, wherein the distribution area derived from the cranberry extract is set as distribution area A, the distribution area derived from the blueberry extract is set as distribution area B, and the distribution area derived from the cranberry extract is set as distribution area C. Moreover, the three distribution areas do not overlap with each other.

[0087] In some specific embodiments, the distribution areas are in the form of ellipses in the above-mentioned view.

[0088] It should be noted that in the establishment of the above-mentioned model, three specific species of cranberry, blueberry and cranberry are used as the basis to form a visual PCA graph. From the perspective of the universality of the model, different specific species of cranberry, blueberry and cranberry can be transformed to obtain multiple sets of PCA graphs, or these data or graphs can be presented in one PCA graph by data weighting or the like, so that distribution areas appear, and each distribution area can cover technical information of different species of, for example, cranberry, thereby facilitating subsequent detection.

[0089] (Discriminating step)

[0090] The discriminating step mainly uses the step of forming a model to obtain data comparable to the model data, and then compares the data with the model data. Specifically, the discriminating step includes:

[0091] Step i':

[0092] i'. Select the object to be discriminated and extract the anthocyanin components contained therein.

[0093] For such an object, for example, various foods to which plant raw materials are added, these raw materials (anthocyanins) can come from cranberry, blueberry or cranberry, or from other species of plants. At the same time, in addition to the anthocyanins in step ii above, these raw materials can also have other species of anthocyanins from plants, and these other species of anthocyanins can also be extracted according to the extraction process in step i above.

[0094] In some specific embodiments, the discriminated object can be a food, for example, a nutritional (dietary) supplement that is solid, semi-solid or liquid at room temperature.

[0095] Step ii':

[0096] ii'. For the extracted anthocyanin components, the content of selected anthocyanin components in the standard sample is analyzed by liquid chromatography-mass spectrometry, and further, the content of each of the selected anthocyanin components is normalized based on the total content of the selected anthocyanin components to obtain the ratio of each selected anthocyanin component relative to the total selected anthocyanin components. Preferably, the liquid chromatography-mass spectrometry and normalization in step ii' can be performed in the same manner as step ii.

[0097] The selected anthocyanin components are the same as, different from, or partially the same as the anthocyanin components in step ii.

[0098] Step iii':

[0099] iii'. The data obtained by the normalization in step iii' is processed in the same manner as step iii to obtain the distribution region D of the object to be identified.

[0100] The distribution region D can be an elliptical region, or a point value obtained by adjusting different parameters or variables in principal component analysis.

[0101] Further, the positional relationship between the distribution region D and the distribution regions A, B, and C is compared.

[0102] In some specific embodiments, if the distribution region D falls into any one of the distribution regions A, B, and C, it can be considered that the anthocyanin in the detected object is from the corresponding blueberry, bilberry, or cranberry.

[0103] In some other specific embodiments, if the distribution region D does not fall into any one of the distribution regions A, B, and C, it can be considered that the anthocyanin in the detected object is not from blueberry, bilberry, or cranberry.

[0104] Examples

[0105] The embodiments of the present application will be described in detail below with reference to examples, but those skilled in the art will understand that the following examples are only used to illustrate the present application and should not be regarded as limiting the scope of the present application. The specific conditions not specified in the examples are carried out according to the conventional conditions or the conditions recommended by the manufacturer. The reagents or instruments not specified by the manufacturer are all conventional products that can be obtained by purchase.

[0106] Materials and Methods Chemicals and Reagents

[0107] Anthocyanin-3-glucoside chloride and other anthocyanin reference standards were provided by ChromaDex.

[0108] Liquid chromatography-mass spectrometry grade methanol, acetonitrile, and Optima water were purchased from Fisher Scientific (St. Louis, MO, USA). Liquid chromatography-mass spectrometry grade formic acid and HPLC grade trifluoroacetic acid (TFA) were purchased from Thermo Scientific and Sigma Aldrich, respectively.

[0109] Sample collection

[0110] A total of 50 samples were included in this study. The samples used in the method development and model training phase included botanical reference materials (BRMs) purchased from certified vendors. The samples used in the method validation phase were commercial dietary supplements or ingredients purchased from local markets. These dietary supplements contained a single species of blueberry and were in various forms, such as concentrates, extracts, syrups, and juices. To prevent any potential bias in the analysis process, the brand information of the supplements was kept confidential.

[0111] Among the 50 samples, the following were included:

[0112] Target species group: Northern highbush blueberry (V. corymbosum L.) (n = 5) and American cranberry (V. macrocarpon (Aiton) (n = 13) and bilberry (V. myrtillus L.) (n = 11).

[0113] Non-target species group: Black rice (Oryza sativa L.) (n = 2), chokeberry (Aronia melanocarpa (Michx.) Elliott) (n = 3), European elderberry (Sambucus nigra L.) (n = 3), acai berry (Euterpe oleracea Mart.) (n = 5), black soybean (Glycine max L. Merr.) (n = 1), pomegranate (Punica granatum L.) (n = 1), blackberry (Rubus spp.) (n = 1), grape (Vitis vinifera L.) (n = 1), which are the most common adulterants reported in the literature.

[0114] Four dietary supplements (n = 4): containing a single species of blueberry or non-target species.

[0115] In the validation study, the true target ingredient sample information is shown in Table 1, the non-target ingredient sample information is shown in Table 2, the dietary supplement sample information is shown in Table 3, and the samples not listed in the above three tables were used as the training dataset to establish the classification model.

[0116] Table 1:

[0117] Table 1: Species Sample name Results of Mahalanobis distance analysis V. corymbosum Blueberry BRM3 V. corymbosum V. macrocarpon Cranberry BRM4 V. macrocarpon V. myrtillus Bilberry BRM6 V. myrtillus

[0118] Table 2:

[0119]

[0120] Table 3:

[0121]

[0122] Sample preparation

[0123] Fresh samples were mixed and homogenized by coffee grinder. According to the visual color of the sample, an appropriate amount of sample was weighed into a volumetric flask. Anthocyanins were extracted using acidified methanol (5% formic acid (v / v)) under ultrasonic conditions for 10 minutes. The extract was diluted to 20 mL with acidified water (5% formic acid (v / v)) and then mixed in the volumetric flask. A portion of the sample was moved to a centrifuge tube and centrifuged at 10,000 rpm for 4 minutes. The supernatant after centrifugation was diluted with acidified water (5% formic acid (v / v)) and filtered through a 0.22-micron PTFE filter membrane before being transferred to an HPLC vial for analysis.

[0124] Marker selection and method development

[0125] Six major anthocyanin aglycons (Cy, Dp, Pt, Pn, Pg, Mv) were selected, which contain galactose (-gal), glucose (-glu) and arabinose (-arab) glycosylation. Since it is very difficult to separate all 18 compounds according to retention time in the chromatographic column, an LC-MS / MS method was developed to identify and quantify the selected anthocyanins by mass-to-charge ratio (m / z) and fragmentation pattern.

[0126] The sample extract was analyzed using a Waters Acquity Xevo TQ MS (Waters Corp. Phenyl-hexyl chromatographic column (2.1 x 150 mm, 1.7 pm) of Acquity UPLC CSH) was used. The UPLC system included a binary pump, a cooled autosampler maintained at 20 °C, a 5 uL sample loop and a photodiode array (PDA) detector. The oven temperature was set at 45 °C, mobile phase A consisted of water, acetonitrile, formic acid and TFA in the ratio of 19:1:0.002:0.01; mobile phase B consisted of methanol, acetonitrile, water, formic acid and TFA in the ratio of 7:8:5:0.002:0.01. 50% methanol was used for needle washing. The gradient was set as follows:

[0127] Initial concentration of 2% mobile phase B, increased to 4% mobile phase B in 5 minutes;

[0128] From 4% to 20% mobile phase B in 5 to 8 minutes;

[0129] From 20% to 90% mobile phase B in 8 to 8.5 minutes;

[0130] 10% mobile phase B at 9.5 minutes;

[0131] 96% mobile phase B at 10 minutes;

[0132] 96% mobile phase B at 13 minutes.

[0133] This gradient ensures that anthocyanins with the same mass to charge ratio (m / z) and fragmentation pattern can be separated by retention time.

[0134] The mass spectrometer was run in positive ion mode. The capillary voltage was set at 2.0 kV, the cone gas flow at 150 L / h, and the source temperature at 150 °C. The desolvation gas flow was set at 1000 L / h, the desolvation temperature at 600 °C, and the nebulizer pressure at 7 bar. For each compound, the MRM transition with the highest signal to noise ratio was selected as the quantifier. The identification of anthocyanins was confirmed by reference to the standard. The instrument was controlled and data were acquired using Masslynx software. Data processing was performed using TargetLynx software.

[0135] A three-fold injection of 50 ppm of the reference standard of cyanidin-3-glucoside (C3g) was injected at the beginning of each run, and periodically throughout the batch, to check system suitability. The peak area and the relative standard deviation (RSD) of the retention time were calculated by the TargetLynx software. Only batches with RSDs of both peak area and retention time ≤ 2% were included in the analysis. This criterion ensures a full evaluation of system suitability and instrument performance.

[0136] High performance thin layer chromatography (HPTLC) for adulteration confirmation

[0137] HPTLC was performed on a single plate, with samples including the samples to be tested listed in Table 3, all prepared at a concentration of 5 mg / ml in methanol. The samples were applied to silica gel 60 HPTLC plates (Merck, fluorescent indicator F254) in 10 mm wide bands by CAMAG automatic TLC sampler. The bands were eluted in a CAMAG automatic developing chamber with butanol, formic acid and water (8:2:3) as solvent. The plates were visualized by CAMAG viewer 2.

[0138] Data processing and chemometrics

[0139] Peak areas for each anthocyanin were exported as a CSV file for analysis in RStudio (Posit, version 2024.09.0, Build 375, using R 4.3.2). Samples were labeled as the corresponding target species (cranberry, blueberry, or lingonberry), non-target species, or unknown species, and this information was recorded in a separate CSV file.

[0140] Peak area percentages were exported as a CSV file for analysis in RStudio (Posit, version 2024.09.0, Build 375, using R 4.3.2). Samples were labeled as the corresponding target species (cranberry, blueberry, or lingonberry), non-target species, or unknown species, and this information was recorded in a separate CSV file.

[0141] Principal component analysis (PCA) was performed in RStudio for data analysis. The data was first log-transformed, all ‘0’ cells were replaced with a small value (0.0001), and auto-scaled using the prep.autoscale function from the mdatools package. PCA based on singular value decomposition was done using the prcomp function from the stats package, and score plots were generated with ggplot2. Hotelling confidence levels were calculated for each species (including lingonberry, blueberry, cranberry, or non-target species) and represented with ellipses. Through a scree plot (not shown), principal component 1 (PC1) and principal component 2 (PC2) were confirmed to contain the highest variance and were suitable for sample comparison. The developed model used only reference materials to determine species grouping and utilized Hotelling 95% confidence intervals to establish confidence regions for each species. To assess the species identity of new samples, each sample was added individually to the dataset and labeled as “unknown”. Mahalanobis distance was used to measure the distance of the new sample to the cluster center of each target species. If the new sample fell within the Hotelling 95% confidence interval based on this distance, it was considered to belong to the cluster of that species. This effectively created a Mahalanobis distance classification model that used a decision boundary approach to classify new samples. The classification of new samples was subsequently determined based on which specific confidence ellipse the sample fell within.

[0142] Results

[0143] Anthocyanin analysis of main Vaccinium species Figure 1 ) revealed significant differences in anthocyanin profiles between the three target species, V. myrtillus, V. macrocarpon, and V. croybosum. Figure 1The average anthocyanin content of the three target species reference materials used in the study is shown. The predominant anthocyanin varies for each species, for example, Pn-3-gal accounts for an average of over 40% of the total peak area. V. macrocarpon samples contain Vp-3-gal, but only in trace amounts in the other two species Figure 1 The standard deviation for each anthocyanin is represented by the black bars, highlighting the anthocyanin variability within the same species reference material.

[0144] The anthocyanins were identified as Dp-3-galactoside, Dp-3-glutamate, Dp-3-arabinoside, Cy-3-galactoside, Cy-3-glutamate, Pt-3-galactoside, Pt-3-glutamate, Pg-3-galactoside, Cy-3-arabinoside, Pg-3-glutamate, Pt-3-arabinoside, Pn-3-galactoside, Mv-3-galactoside, Pg-3-arabinoside, Pn-3-glutamate, Mv-3-glutamate, Pn-3-arabinoside, Mv-3-arabinoside. The most abundant anthocyanin in the V. myrtillus USP reference standard was Mv-3-glu, accounting for 13.4% of the relative abundance of total anthocyanins, followed by Cy-3-arab, accounting for 11.8%. Other major anthocyanins included Dp-3-galactoside, Dp-3-glucoside, Dp-3-arabinoside, Cy-3-galactoside, Cy-3-glucoside, Pn-3-glucoside, Pt-3-glucoside, and Mv-3-galactoside, with relative contents ranging between 5-10%.

[0145] The anthocyanin profile within the V. myrtillus species was generally consistent, with standard deviations ranging from 0.02% to 1.60% for each anthocyanin. Overall, the fruit extract of V. myrtillus contained large amounts of Cy-, Dp-, and Mv-derivatives, while Pg-derivatives were almost non-existent, indicating a high degree of similarity in the anthocyanin profile Figure 1

[0146] In contrast to V. myrtillus, the anthocyanin profile of V. macrocarpon fruit was much simpler. The main anthocyanins identified in the extract of V. macrocarpon included Pn-3-galactoside, Pn-3-arabinoside, Cy-3-galactoside, and Cy-3-arabinoside. These four anthocyanins accounted for approximately 95% of the total anthocyanin content Figure 1 ​). Pn-3-glucoside, Mv-3-galactoside, Pg-3-arabinoside and Mv-3-arabinoside were detected at low levels in only a few samples. Pn-3-glu, Mv-3-gal, Pg-3-arab and Mv-3-arab were detected at low levels in only a few V. macrocarpon fruits. Similar to V. myrtillus, the anthocyanin profile within V. macrocarpon varieties showed a high degree of consistency with standard deviations (SD) ranging from 2.2% to 4.9% for the main anthocyanins.

[0147] As shown in Figure 1, Figure 1 the anthocyanin pattern of V. corybosum and V. myrtillus was similar, but the proportion of Mv-derivatives was consistently higher in V. corybosum than in V. myrtillus.

[0148] PCA shows differences between target species

[0149] After anthocyanin profiling by LC-MS / MS, principal component analysis (PCA) was employed to visualize the anthocyanin pattern of specific species. PCA is an unsupervised multivariate statistical technique that reduces multiple variables into smaller principal components (PCs) that describe how samples vary and correlate according to their overall chemical characteristics. The resulting score plot brings together samples with more similar characteristics, while samples with more distinct characteristics are further apart. Figure 2

[0150] In this case, the model found relationships between the percentage area of each anthocyanin peak, allowing the classification of samples according to their similarity in overall anthocyanin characteristics.

[0151] Figure 2 Figure 1 A shows that principal component 1 (PC1) and principal component 2 (PC2) explained 56% and 28% of the total variance, respectively, and were able to distinguish V. macrocarpon from the other two target species by means of the 18 target anthocyanins. Each ellipse represents the 95% confidence interval for the specified group. However, there was an overlap between the ellipses of V. myrtillus and V. corybosum, meaning that these two species could not be completely distinguished by means of the selected anthocyanins. The anthocyanins that contributed most to PC1 were Pn-3-arabinoside, Pn-3-galactoside, Dp-3-galactoside, DP-3-arabinoside and Pt-3-galactoside; the anthocyanins that contributed most to PC2 were Pn-3-glucoside, Cy-3-glucoside, Dp-3-glucoside, Mv-3-arabinoside and Pt-3-glucoside.

[0152] As mentioned previously, Figure 1 ​It was shown that V. myrtillus and V. corybosum differ in the overall level of anthocyanins and luteolin. To improve the separation, the ratio of cyanidin (Cy) to malvidin (Mv) was calculated as an additional variable, which allowed for a complete separation of the three species (see Figure 2 of B).

[0153] Values

[0154] It is noted that the addition of the Cy / Mv ratio as an additional variable did not change the explained variance of each principal component or the order of the variables in the principal components, indicating that the overall chemical variation explained by the principal components was still preserved. Therefore, this model was chosen as the basis for the subsequent steps.

[0155] Classification model validation study

[0156] To test the performance of the classification model, a method validation study was performed to check whether the model can accurately classify target and non-target species. Although principal component analysis (PCA) is an unsupervised chemometric method and does not have the ability to predict classifications by itself, we used it in combination with Mahalanobis distance to analyze new samples in the model built from the three target species reference materials. This approach allowed us to determine whether the new samples clustered within the expected cluster on the score plot, thus aiding in the classification of the samples.

[0157] In the validation study, we performed three tests. First, we validated each target species (not included in the initial model building) individually. Each new sample always clustered within the 95% confidence interval of the model for its own species, showing a 100% model accuracy ( Figure 3 , Table 1).

[0158] We further assessed the model’s ability to distinguish non-target samples (i.e., any sample not belonging to blueberry, cranberry, or lingonberry) from target samples. We introduced 17 non-target BRMs (Table 2) each, and all samples did not fall within the cluster of the three target species ( Figure 4 of A, Table 2). This indicates that the model was able to successfully identify samples not belonging to blueberry, cranberry, or lingonberry using the profiles of the 18 anthocyanins and the Cy / Mv ratio without additional information. Mahalanobis distance calculations confirmed that these non-target samples indeed lay outside the confidence interval of the target clusters. We included all non-target plant reference materials (BRMs) simultaneously and compared them to the training dataset of the three target species by principal component analysis (PCA) ( Figure 4 of B-D). Although some non-target samples had similarities to lingonberry. Due to the similarity of the anthocyanin profiles, Figure 4Figure 6D). The separation of non-target samples from each target species was demonstrated by Mahalanobis distance decision boundaries (visualized using confidence ellipses), further supporting the unique anthocyanin profiles observed in the authentication model. Figure 4 B-D of Figure 6 served as a supplementary visualization tool to qualitatively assess the ability of principal component analysis to differentiate non-target samples from each target species. Unlike Figure 4 A of Figure 6, which was optimized for a classification model of the three target species, Figure 4 B-D of Figure 6 provided a view of the non-target variation without restricting it to the principal component analysis space designed for all three target species. This step was used for visualization only and did not affect the final authentication model.

[0159] Evaluation of identification

[0160] To assess the applicability of the model to commercially available products, we tested four dietary nutritional supplements, each labeled to contain one of the target species, and one of which was also labeled to contain Sambucus nigra. Table 3 summarizes the results of the testing of these products. Each product was extracted and analyzed in the same manner as the model, and its anthocyanin profile was added to the classification model.

[0161] The clustering of the V. macrocarpon, V. corymbosum, and S. nigra supplements was consistent with their label claims Figure 5 A, B, D of Figure 6, Table 3). Of these, the V. macrocarpon product showed a high concentration of Pn-3-gla (46%), which is consistent with the BRM profile of V. macrocarpon. The product of V. corymbosum showed the highest concentration of Mv-3-glul (22%), which is consistent with the BRM profile of V. corymbosum (Supplemental Table S2). The product of S. nigra contained 97% Cy-3-glul and trace amounts of other anthocyanins, and upon addition to the principal component analysis model, it clustered significantly away from the target species Figure 5 D of Figure 6). This was confirmed by HPTLC analysis Figure 5 F of Figure 6), which showed a band pattern consistent with the in-house S. nigra BRM. The relative amounts of Cy-3-glu (19%), Pn-3-glu (15%), and Mv-3-glu (14%) were high in the V. myrtillus supplement, which is similar to the profile of a mixture of V. myrtillus and V. corymbosum. When this supplement was added to the classification model, it fell outside of any known cluster, suggesting that its label might be incorrect Figure 5C). HPTLC analysis further confirmed that the band pattern of this supplement did not match V. myrtillus (Fig. 1C). Figure 5

[0162] Conclusions

[0163] Among the Vaccinium species studied, blueberry (V. corybosum) contains multiple purpose-bred varieties that have been selected for size, firmness, and disease resistance, resulting in a diverse anthocyanin profile. Our study showed significant differences in anthocyanin profiles between different BRMs. The anthocyanin profiles of ChromaDex and AHP BRMs were similar, characterized by a high proportion of Cy-, Dp-, and Pt-derivatives, and a low content of glucose-glycosylated anthocyanins. In contrast, the NIST sample contained a high proportion of Mv- and Pn-derivatives, with approximately 33% of the anthocyanins being glucose-glycosylated. This observation is consistent with previous findings that identified and quantified anthocyanins in 19 blueberry samples and concluded that while the types of anthocyanins were similar across different varieties, the specific proportions of each anthocyanin varied by variety.

[0164] In contrast to V. corybosum, V. myrtillus is typically obtained from wild harvesting and less frequently cultivated. In our study, we analyzed eight BRMs (biological reference materials) of V. myrtillus from different suppliers, all of which showed a standard deviation of less than 2.0% for the identified anthocyanins, confirming the consistent anthocyanin profile observed in previous studies.

[0165] During the development process, the differences in anthocyanin profiles between BRMs of the same species must be considered. In many cases, the adulteration of cranberry can be detected by comparing the anthocyanin profile of the sample to a single reference material. However, as previously discussed, Vaccinium species show significant differences in the proportions of individual anthocyanins, and therefore a single reference material cannot represent the entire species. Additionally, the anthocyanin composition pattern of cranberry and V. myrtillus is similar to another species. Both plants contain anthocyanins from five different aglycons (Cy-, Dp-, Mv-, Pn-, Pt-), and each aglycon contains three sugar groups (galactose, glucose, and arabinose), making it difficult to distinguish between the two plants based on the anthocyanin profile alone without the aid of chemometrics.

[0166] ​In our study, individual anthocyanins were semi-quantitatively analyzed using peak area percentage. Selectively assessing the relative, rather than absolute, abundance of chosen labeled anthocyanins allows for a comprehensive evaluation of the entire anthocyanin spectrum, thus assessing the overall relative abundance of various anthocyanins in a sample while minimizing normalization and chromatographic data processing. Furthermore, since plant components appear at varying concentrations in different dietary supplements and raw materials, using relative abundance allows for the disregarding of their absolute concentrations, resulting in a uniform treatment across all samples. In summary, our semi-quantitative method provides a holistic perspective on the distribution and relative proportions of labeled compounds, contributing to more comprehensive pattern recognition analysis.

[0167] Principal component analysis (PCA) is a simple and easy-to-use chemometric tool that addresses interspecies variation by generating comprehensive chemical profiles through semi-quantitative methods. By reducing the dimensionality of datasets, PCA can reveal patterns and visualize relationships between different samples. In our study, PCA based on anthocyanin profiles highlighted the variation and similarities between different species. Furthermore, using a standardized set of anthocyanins for PCA analysis simplifies the data collection process by eliminating the need to obtain information on potential marker compounds from other non-target samples. To further simplify the analysis, we focused on the 18 most common anthocyanins found in *Vaccinium* species.

[0168] Preliminary PCA analysis using 18 selected anthocyanin markers revealed significant overlap between *V. corybosum* and *V. myrtillus*. The clustering of *V. myrtillus* stemmed from significant differences in anthocyanin content within the *V. corybosum* species. Figure 2 (A). Figure 1 Highlighting the significant differences in the accumulation of individual anthocyanins in *V. corybosum* across different samples, the percentage of malvidin derivatives remained consistently high (51.0% ± 11.8%). In contrast, cyanidin derivatives accounted for only a small fraction of the total anthocyanin content (7.5% ± 1.7%). On the other hand, the distribution of cyanidin (24.0% ± 5.9%) and malvidin derivatives (23.9% ± 4.9%) in bilberry extract was more balanced, with both accounting for approximately 25% of the total anthocyanin content (Table 1). Figure 2 Therefore, we hypothesize that the ratio between cyanidin and malvidin derivatives can reflect species-specific differences, thus providing an additional predictive indicator for distinguishing the two Vaccinium species in chemometric analysis.

[0169] By using the ratio of cyanidin (Cy) to mallow pigment (Mv) as the 19th marker in the principal component analysis model, the three blueberry groups were visually completely separated. Figure 2B). The introduction of this marker proved to be effective. Without changing the amount of variation explained by each principal component or the key variables contributing to each component, better differentiation between species was achieved. The effectiveness of the marker selection was further confirmed during the model validation process by adding new target BRMs to the model and successfully clustering them with the corresponding species each time ( Figure 3

[0170] To ensure consistency in classification, principal component analysis (PCA) was always performed using the training dataset of the three target species and an unknown sample. This approach ensures that the unknown sample undergoes the same transformation process as the training data, avoiding inconsistencies that can arise when validating samples are projected onto a pre-computed PCA model. Although each time an unknown sample is incorporated into the PCA calculation slightly changes the clustering structure, these effects are minimal because the training dataset from the three target species still dominates. This approach ensures that the validation sample is consistently transformed with the training dataset before Mahalanobis distance classification, maintaining the accuracy of the classification. Our classification method is based on Mahalanobis distance, a statistical method used in chemometric applications. Similar techniques were employed in an earlier study that combined Mahalanobis distance and residual variance analysis for the classification of near-infrared (NIR) spectra.25Although their study focused on pattern recognition in spectroscopy, our study is the first to utilize Mahalanobis distance for anthocyanin fingerprinting based on liquid chromatography-tandem mass spectrometry (LC-MS / MS). By combining principal component analysis (PCA) with Mahalanobis distance classification, we achieved accurate species identification of botanical ingredients and dietary supplements, directly addressing the true cGMP regulatory requirements of the dietary supplement industry. Both studies demonstrate that Mahalanobis distance is a powerful classification tool applicable to different analytical techniques and data types.

[0171] In addition to demonstrating that the classification model successfully differentiated the three Vaccinium species, it was also important to distinguish true Vaccinium extracts from non-target species. The AOAC SMPR provided a list containing 14 fruits and 9 non-fruit sources to differentiate Vaccinium plants (AOAC SMPR 2014.07, 2014). Considering the availability of reference materials and previously reported adulterants, the non-target group in this study included extracts of acai (Euterpe oleracea), black soybean (Glycine max), black currant (Aronia melanocarpa), pomegranate (Punica granatum), blackberry (Rubus spp.), elderberry (Sambucus nigra), grape (Vitis vinifera), and black rice (Oryza sativa L.).

[0172] ​When each non-target BRM was added to the model one at a time, they all fell outside of the target clusters Figure 4 A, Table 2), indicating that the anthocyanin profiles of these non-target BRMs were different from the three target species. When all non-target samples were added to the model at the same time, there was no overlap between the target and non-target samples Figure 4 B and D), indicating that our method was suitable for determining whether a new sample was not the target species.

[0173] The unique advantage of this study was the use of multiple BRMs to construct and validate the model. By selecting BRMs and validated internal samples, we were able to cover a wide range of sample source variability. In other words, we utilized a diverse sample library to expand the reference anthocyanin profile. Therefore, samples that might not fit the single BRM anthocyanin profile but were indeed the correct species were adequately represented.

[0174] It is worth noting that an ideal model should contain a large number of samples belonging to each target and non-target species. However, for most plant products, the number of BRMs available on the market is very small.

[0175] To evaluate whether this classification model could be used as a quality control method to identify adulterated dietary supplements, 18 anthocyanins from four commercially available supplements were analyzed (see Table 3). Based on the positions of these anthocyanins in the score plot, the V. corybosum and V. macrocarpon supplements were correctly labeled by the manufacturers Figure 5 A, B). Similarly, the non-target supplement, black chokeberry (Aronia arbutifolia), was separated from the three target groups Figure 5 D), which was confirmed by comparing the HPTLC profile of this supplement with the HPTLC profiles of Aronia melanocarpa BRMs Figure 5 F). However, the V. myrtillus supplement did not appear in the target cluster, indicating that it might be a mislabeled product. We confirmed this sample was likely not V. myrtillus by HPTLC analysis, as it did not have the characteristic band pattern of V. myrtillus BRMs in its chromatogram Figure 5 E).

[0176] It should be noted that although the technical solutions of the present application are described with specific examples, those skilled in the art can understand that the present application should not be limited thereto.

[0177] Having described various embodiments of the application, it is to be understood that the above description is meant to be illustrative only and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art, without departing from the scope and spirit of the described embodiments. The choice of words in this document is intended to best explain the principles of the embodiments, the practical application, or technical improvement over the existing technology, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for identifying plant raw materials, characterized in that, It is an identification method based on mass spectrometry and principal component analysis combined with Mahalanobis distance analysis, which includes: The steps involved in model building and identification. The steps for establishing the model include: i. Using blueberry, blueberry and cranberry as three standard samples, the anthocyanin components contained in each standard sample were extracted to obtain the three standard sample extracts. ii. For each of the three standard sample extracts, the content of a selected anthocyanin component in the standard sample is analyzed by a detection system including liquid chromatography-mass spectrometry. Further, the content of each of the selected anthocyanin components is normalized based on the total content of the selected anthocyanin components to obtain the mass spectrometry peak area ratio of each selected anthocyanin component relative to the total content of the selected anthocyanin components. The selected anthocyanin components include at least the anthocyanin components classified as cyanidin (Cy), delphinidin (Dp), petuniadin (Pt), paeoniflorin (Pn), pelargonidin (Pg), and mallow pigment (Mv). iii. Principal component analysis was performed on the three sets of data obtained after the normalization process. Three non-overlapping distribution regions were established based on the components corresponding to each set of data. In the principal component analysis, the ratios of anthocyanins and cyanidin (Cy) to malvidin (Mv) were used as variables. The distribution regions were established based on Mahalanobis distance analysis with a Hotelling confidence level of 95%. The distribution region derived from blueberry extract was designated as distribution region A, the distribution region derived from blueberry extract as distribution region B, and the distribution region derived from cranberry extract as distribution region C. The identification steps include: i'. Select the object to be identified and extract the anthocyanin components contained therein; ii'. For the extracted anthocyanin components, the content of selected anthocyanin components in the standard sample is analyzed using a detection system. Furthermore, based on the total content of selected anthocyanin components, the content of each selected anthocyanin component is normalized to obtain the ratio of each selected anthocyanin component to the total selected anthocyanin components. The selected anthocyanin component is the same as, different from, or partially the same as the anthocyanin component in step ii. iii'. The data obtained by the normalization process is processed in the same way as in step iii to obtain the distribution area D of the object to be identified. Furthermore, the positional relationship between the distribution region D and the distribution regions A, B, and C is compared.

2. The method according to claim 1, characterized in that, The extraction is carried out in an acidic alcohol solution.

3. The method according to claim 1 or 2, characterized in that, The detection system includes a high-performance thin-layer chromatography (HPTLC) system with a photodiode array detector (PDA).

4. The method according to any one of claims 1 to 3, characterized in that, The liquid chromatography-mass spectrometry (LC-MS) was performed in multiple reaction monitoring (MRM) mode, and the data was acquired and processed using software.

5. The method according to any one of claims 1 to 4, characterized in that, The anthocyanins in step ii include at least the following 18 anthocyanins: Dp-3-galactoside, Dp-3-glutamic acid, Dp-3-arabinoside, Cy-3-galactoside, Cy-3-glutamic acid, Pt-3-galactoside, Pt-3-glutamic acid, Pg-3-galactoside, Cy-3-arabinoside, Pg-3-glutamic acid, Pt-3-arabinoside, Pn-3-galactoside, Mv-3-galactoside, Pg-3-arabinoside, Pn-3-glutamic acid, Mv-3-glutamic acid, Pn-3-arabinoside, and Mv-3-arabinoside.

6. The method according to any one of claims 1 to 5, characterized in that, After the normalization process, the normalized data will be output or stored in the form of machine-readable data from the principal component analysis.

7. The method according to any one of claims 1 to 6, characterized in that, The principal component analysis is performed using computer software, which includes a Mahalanobis distance analysis program. Preferably, the computer software is RStudio software.

8. The method according to any one of claims 1 to 7, characterized in that, The principal component analysis is output as a visualized two-dimensional coordinate graph, in which each distribution region is displayed.

9. The method according to claim 8, characterized in that, In step iii, the distribution regions A, B, and C appear in an elliptical shape.

10. A method for identifying the plant source of anthocyanins in food, characterized in that, The method includes the method according to any one of claims 1 to 9, to identify whether the anthocyanins in the food are derived from blueberries or from any one of blueberries, blueberries or cranberries.

11. The method according to claim 10, characterized in that, The food products include nutritional supplements that are liquid, semi-solid, or solid at room temperature.