A method for protecting and developing and utilizing phellinus igniarius germplasm resources
By analyzing the harvesting cycle and characteristic information of Phellinus linteus germplasm resources, effective Phellinus linteus were screened, application value score outliers and matching degree were calculated, and the protection strategy was adjusted using a latent semantic model. This solved the problem of inaccurate application value scores in the protection of Phellinus linteus germplasm resources and improved the accuracy and utilization rate of the protection strategy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI AGRICULTURAL UNIVERSITY
- Filing Date
- 2025-05-30
- Publication Date
- 2026-04-24
AI Technical Summary
Existing methods fail to fully consider the information of Sanghuang itself and the information of the germplasm resources in the protection of Sanghuang germplasm resources, resulting in inaccurate application value score matrix, which affects the accuracy of protection strategies and the mature utilization rate.
By acquiring information on the harvesting cycle and characteristics of Phellinus linteus from various regions associated with germplasm resources, effective Phellinus linteus are screened, application value score outliers and matching degree values are calculated, latent semantic models are used to adjust protection strategies, and a predictive application value matrix is constructed to improve the accuracy of protection strategies.
This improved the mature utilization rate of Sanghuang germplasm resources, ensured the accuracy and effectiveness of protection strategies, avoided the impact of malicious marketing, and enhanced the satisfaction and user experience of germplasm resources.
Smart Images

Figure CN120632482B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of genetic data processing and information technology in bioinformatics, and in particular to a method for the protection and development of Phellinus linteus germplasm resources. Background Technology
[0002] With the development of big data, adaptive protection strategy information adjustment services can be provided for biological genetic data resources such as germplasm resources through big data. In particular, adaptive adjustment can be made according to the characteristics of germplasm resources to make the content more suitable for germplasm resources, thereby improving the demand satisfaction of germplasm resources.
[0003] Existing methods adaptively adjust the protection strategy information of germplasm resources, but only obtain the final predicted application value score matrix based on the harvesting situation of Sanghuang from germplasm resources to determine the protection strategy information to be adjusted to germplasm resources. This ignores the information of Sanghuang itself and the germplasm resources themselves, which leads to an inability to fully understand the underlying reasons for the behavior of germplasm resources. For example, different cycles of using protection strategies for germplasm resources result in significant differences in the harvesting situation of Sanghuang, making the application value scores of some germplasm resources for certain Sanghuang inaccurate. This makes the final predicted application value score matrix inaccurate, leading to inaccurate adjustment of the protection strategy information of germplasm resources and reducing the mature utilization rate of germplasm resources. Summary of the Invention
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] According to a first aspect of the present invention, the present invention claims protection for a method for the protection and development of Phellinus linteus germplasm resources, the method comprising the following steps:
[0006] Obtain the harvesting cycle and frequency of Sanghuang in each region associated with a preset number of germplasm resources, as well as the regional characteristic information of Sanghuang in each region; set the first harvest time point of each harvest of Sanghuang in each region associated with each germplasm resource as the harvest time point.
[0007] Based on the harvesting cycle, obtain the application value score of each germplasm resource associated with Sanghuang in each region; based on the matching status of the application value score of each germplasm resource with each other germplasm resource associated with Sanghuang harvested together, the number of strains of Sanghuang harvested together, and the difference in application value score of Sanghuang harvested together, obtain the outlier value of the application value score of each germplasm resource.
[0008] Any species of *Sanghuang* is designated as a valid *Sanghuang*. Based on the differences in characteristic information of the same strains between valid *Sanghuang* and *Sanghuang* from other regions, the matching degree value of valid *Sanghuang* with *Sanghuang* from other regions is obtained, and candidate *Sanghuang* for valid *Sanghuang* is screened. Based on the outlier value of the application value score of each germplasm resource, the matching degree value between valid *Sanghuang* and candidate *Sanghuang* from each region, and the application value score of each germplasm resource associated with candidate *Sanghuang* from each region, the confidence level of the application value score associated with valid *Sanghuang* for each germplasm resource is obtained.
[0009] The application value score confidence level is adjusted based on the associated application value score confidence level to obtain the adjusted application value score confidence level of each germplasm resource associated with effective Phellinus linteus.
[0010] The application value score confidence of each germplasm resource associated with the latent semantic model is adjusted to obtain the predicted application value matrix, and the protection strategy information of each germplasm resource is then adjusted accordingly.
[0011] Furthermore, the method for obtaining the application value score is as follows:
[0012] The harvesting cycle of Sanghuang in each region associated with each germplasm resource was recorded in a standardized manner, and an application value score was set for each germplasm resource associated with Sanghuang in each region.
[0013] Furthermore, the method for obtaining outliers in the application value score is as follows:
[0014] The Sanghuang harvested jointly from germplasm resources a and b is designated as the Sanghuang for testing.
[0015] The application value score of Sanghuang in each region is associated with the a-th germplasm resource and set as the status information of the first region.
[0016] The application value scores of Phellinus linteus detected in each region are associated with the b-th germplasm resource and set as the status information of the second region.
[0017] Obtain the correlation between the status information of the first region and the status information of the second region, and set it as the value difference score (correlation) between the a-th germplasm resource and the b-th germplasm resource;
[0018] Obtain the value difference score (correlation) between the a-th germplasm resource and each other germplasm resource, and set the other germplasm resources corresponding to the value difference score (correlation) that are greater than the preset value difference score (correlation) threshold as candidate germplasm resources of the a-th germplasm resource;
[0019] For any candidate germplasm resource of germplasm resource a, the first development value between germplasm resource a and candidate germplasm resource is obtained based on the value difference score (correlation) between germplasm resource a and candidate germplasm resource, the number of strain regions that detect Sanghuang, and the difference in application value scores of Sanghuang detection in each region associated with germplasm resource a and candidate germplasm resource.
[0020] The record of the sum of the first development value between the a-th germplasm resource and each candidate germplasm resource is set as the application value score outlier of the a-th germplasm resource.
[0021] Furthermore, the method for obtaining the matching degree value of effective Sanghuang and Sanghuang from other regions based on the characteristic information of the same strains between effective Sanghuang and Sanghuang from other regions, and screening out candidate Sanghuang for effective Sanghuang, is as follows:
[0022] For any feature information of Sanghuang, the record that is negatively correlated and standardized between Sanghuang in region k and Sanghuang in region g is set as the matching record of that feature information between Sanghuang in region k and Sanghuang in region g.
[0023] The mean of all matching records of species feature information between Phellinus linteus in region k and Phellinus linteus in region g is obtained and set as the matching degree value between Phellinus linteus in region k and Phellinus linteus in region g.
[0024] Obtain the matching degree value of Sanghuang in region k with Sanghuang in other regions, and sort the matching degree values of Sanghuang in region k with Sanghuang in other regions from largest to smallest to obtain the matching degree value result;
[0025] Set the first preset number of matching scores in the matching score results as valid matching scores;
[0026] For each valid matching degree value, other Sanghuang (Phellinus thunbergii) that are not in the k-th region are set as candidate Sanghuang for the k-th region.
[0027] Furthermore, the method for obtaining the application value score confidence level is as follows:
[0028] For any candidate Sanghuang that is effective, the first germplasm value of effective Sanghuang and candidate Sanghuang is obtained based on the matching degree value between effective Sanghuang and candidate Sanghuang, and the application value score of each germplasm resource associated with candidate Sanghuang.
[0029] The record of summing the first quality value of effective Phellinus linteus and the candidate Phellinus linteus from each region is set as the reliable candidate value of effective Phellinus linteus.
[0030] The records of outliers in the application value score of each germplasm resource that are negatively correlated and standardized are set as the reliable adjustment values for the corresponding germplasm resources.
[0031] The product of the reliable adjusted value of each germplasm resource and the reliable candidate value of the effective Phellinus linteus is set as the confidence score of the application value associated with the effective Phellinus linteus of the corresponding germplasm resource.
[0032] Furthermore, the method for adjusting the cost function of the latent semantic model associated with the confidence scores of each germplasm resource for Sanghuang in each region to obtain the predicted application value matrix is as follows:
[0033] Each germplasm resource is associated with the application value score of Sanghuang from each region as an element to construct an initial application value matrix; where each row of the initial application value matrix corresponds to a germplasm resource and each column corresponds to a Sanghuang strain.
[0034] Elements in the initial application value matrix that are greater than the third preset constant are set as valid elements. Based on the coordinates of the valid elements in the initial application value matrix, the confidence level of the adjusted application value score corresponding to each valid element is determined.
[0035] The cost function of the latent semantic model is adjusted based on the confidence score of the adjusted application value score corresponding to each valid element, and the adjusted cost function is obtained.
[0036] Based on the adjusted cost function, the potential index matrix of germplasm resources and the potential index matrix of Sanghuang are obtained, and a predictive application value matrix is constructed.
[0037] This invention relates to the field of agricultural information technology, and particularly to a method for the protection and development of *Phellinus linteus* germplasm resources. The method involves acquiring a preset number of germplasm resources associated with *Phellinus linteus* from various regions, including their harvesting cycles and frequencies, as well as the regional characteristics of *Phellinus linteus* from each region. Any *Phellinus linteus* species is designated as a valid *Phellinus linteus*. Based on the differences in characteristic information of the same strains between valid *Phellinus linteus* and *Phellinus linteus* from other regions, a matching degree value is obtained for valid *Phellinus linteus*. Candidate valid *Phellinus linteus* are then selected, and the adjusted application value score confidence level of each germplasm resource associated with valid *Phellinus linteus* is obtained. The cost function of the latent semantic model is adjusted based on the adjusted application value score confidence level of each germplasm resource associated with *Phellinus linteus* from various regions to obtain a predicted application value matrix. The protection strategy information associated with each germplasm resource is then adjusted. This invention can accurately adjust the protection strategy information for germplasm resources, improving the mature utilization rate of germplasm resources. Attached Figure Description
[0038] Figure 1 This is a flowchart illustrating a method for the protection and development of Phellinus linteus germplasm resources as claimed in an embodiment of the present invention. Detailed Implementation
[0039] The following will describe the embodiments of the present invention clearly and completely with reference to the accompanying drawings and related technical solutions. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0040] The terms "first," "second," and "third" used in this invention are for descriptive purposes only and should not be construed as indicating or implying associated importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this invention are only used to explain the associated positional relationships, motion states, etc., between components in a specific orientation (as shown in the figures). If the specific orientation changes, the directional indication also changes accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0041] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0042] The following description, in conjunction with the accompanying drawings, details the specific scheme of a method for the protection and development of Phellinus linteus germplasm resources provided by this invention.
[0043] The specific scenario of this invention is as follows: by analyzing the harvesting status of Sanghuang in each region associated with each germplasm resource, the initial application value matrix is obtained. Then, by using a latent semantic model and combining the parameters of different Sanghuang, the potential index vectors of germplasm resources and Sanghuang are obtained, and a predictive application value matrix is constructed to adjust the protection strategy information associated with each germplasm resource.
[0044] The purpose of this invention is to address the issue that, in reality, the cycles for using protection strategies differ among different germplasm resources. Simply inferring the affinity of germplasm resources for Sanghuang (a type of medicinal mushroom) based on their harvesting cycles is inaccurate. Furthermore, there may be instances of malicious marketing of some Sanghuang, making it impossible to accurately predict the affinity between germplasm resources and Sanghuang. This leads to the inability of latent semantic models to accurately predict the application value score of each germplasm resource associated with Sanghuang from different regions, hindering accurate adjustments to protection strategy information. Therefore, this invention analyzes information on the joint harvesting of Sanghuang among different germplasm resources to identify outliers in the application value score of each germplasm resource, avoiding significant errors caused by individual factors in associating application value scores. It further analyzes the attributes of various Sanghuang to obtain matching Sanghuang from each region. Based on the application value score status of the matched Sanghuang associated with germplasm resources, the reliability of the corresponding application value score for each region is determined. Finally, based on the application value score status of the matched Sanghuang associated with germplasm resources and the outliers in the application value score of germplasm resources, the confidence level of the application value score associated with Sanghuang is obtained. To avoid the negative impact of malicious promotion of Sanghuang (a type of medicinal mushroom), and considering that the application value score degrades over time, the confidence level of the application value score associated with the germplasm resource is further adjusted based on the time point corresponding to the harvesting of Sanghuang and the current time point. This yields the adjusted application value score confidence level for each germplasm resource associated with Sanghuang from various regions. Furthermore, the cost function of the latent semantic model is accurately adjusted based on the adjusted application value score confidence level, improving the accuracy of the latent semantic model in predicting the application value score. This, in turn, allows for accurate adjustment of the germplasm resource protection strategy information, effectively enhancing the satisfaction and usability of the germplasm resources. The latent semantic model is a well-known technology and will not be elaborated upon further.
[0045] Please see Figure 1 The diagram illustrates a flowchart of a method for the protection and development of Phellinus linteus germplasm resources according to an embodiment of the present invention. The method includes the following steps:
[0046] Step S1: Obtain the harvesting cycle and harvesting frequency of Sanghuang in each region associated with a preset number of germplasm resources, as well as the regional characteristic information of Sanghuang in each region; set the first harvesting time point of each Sanghuang in each region associated with each germplasm resource as the harvesting time point.
[0047] Specifically, information regarding the harvesting of *Sanghuang* (a type of medicinal mushroom) for germplasm resources under protection strategies, as well as information on *Sanghuang* from various regions, is recorded in big data. Each germplasm resource and each region's *Sanghuang* has a unique ID. Since a single type of *Sanghuang* is completely identical, the big data allows for the acquisition of the harvesting cycle and frequency of *Sanghuang* associated with each germplasm resource in various regions, as well as the regional characteristics of *Sanghuang* from each region. The harvesting cycle of each germplasm resource associated with *Sanghuang* from various regions is the sum of all harvesting cycles associated with that germplasm resource in those regions. The characteristic information of *Sanghuang* from each region includes various features such as price, applicable age, size, and weight. This embodiment of the invention uses 500 germplasm resources as an example for analysis, i.e., the preset quantity is set to 500. The implementer can set the preset quantity according to the actual situation; this is not limited here. The IDs of the 500 germplasm resources are all different. Each harvest of Sanghuang associated with each germplasm resource in each region corresponds to a specific time period. In order to determine the harvesting status of Sanghuang associated with each germplasm resource in each region, the first harvest time point of each harvest of Sanghuang associated with each germplasm resource in each region recorded in the big data is set as the harvesting time point.
[0048] Step S2: Based on the harvesting cycle, obtain the application value score of each germplasm resource associated with Sanghuang in each region; based on the matching status of the application value scores of each germplasm resource with each other's Sanghuang harvested together, the number of strains harvested together, and the differences in the application value scores of the harvested Sanghuang, obtain the outlier values of the application value score of each germplasm resource.
[0049] Specifically, the more suitable a germplasm resource is for associating with a certain type of *Sanghuang* (a type of medicinal mushroom), the longer the corresponding harvesting period. Therefore, by recording the harvesting period of *Sanghuang* from each region associated with each germplasm resource, an evaluation index for each germplasm resource associated with *Sanghuang* from each region is obtained. In this embodiment of the invention, the harvesting period of each germplasm resource associated with *Sanghuang* from each region is recorded in a standardized manner and set as an application value score for each germplasm resource associated with *Sanghuang* from each region. The higher the application value score, the more suitable the corresponding germplasm resource is for associating with the corresponding *Sanghuang*. When a germplasm resource associated with a certain type of *Sanghuang* is not harvested, the corresponding harvesting period is set to 0, and the standardized value mapped to 0 is also 0. Therefore, the application value score for that germplasm resource associated with that type of *Sanghuang* is also 0. Thus, the application value score for each germplasm resource associated with *Sanghuang* from each region is obtained. Considering the differences in the cycle and frequency of protection strategies used for different germplasm resources, standardizing the harvesting cycle of *Sanghuang* (a type of medicinal mushroom) associated with each germplasm resource across different regions can lead to excessively low application value scores for *Sanghuang* associated with germplasm resources using fewer protection strategies. Conversely, it can also result in inaccurate differentiation of application value scores for *Sanghuang* associated with germplasm resources using more protection strategies. To accurately adjust the protection strategy information associated with each germplasm resource, the application value score status of each germplasm resource is analyzed to determine its accuracy. It is known that highly correlated germplasm resources should have matching application value scores for *Sanghuang* harvested in common. This invention obtains the common *Sanghuang* harvesting relationships between each germplasm resource and every other germplasm resource. Based on the matching status of the application value scores of common *Sanghuang* harvesting relationships between each germplasm resource and every other germplasm resource, highly correlated candidate germplasm resources are determined for each germplasm resource. It is certain that there is a common *Sanghuang* harvesting relationship between each germplasm resource and its candidate germplasm resources. The greater the difference in application value score between each germplasm resource and its candidate germplasm resources in the co-harvesting of Phellinus linteus, the more likely there is an anomaly in the application value score of the corresponding germplasm resource. Therefore, in this embodiment of the invention, the application value score anomaly value of each germplasm resource is obtained based on the matching status of the application value score of each germplasm resource in the co-harvesting of Phellinus linteus with each other germplasm resource, the number of strains in the co-harvesting of Phellinus linteus, and the difference in application value score of the co-harvesting of Phellinus linteus.
[0050] Preferably, the method for obtaining outliers in the application value score is as follows: *Sanghuang* harvested jointly from germplasm resources a and b are designated as *Sanghuang* for testing; the application value scores of *Sanghuang* from each region associated with germplasm resource a are designated as the first region status information; the application value scores of *Sanghuang* from each region associated with germplasm resource b are designated as the second region status information; the correlation between the first and second region status information is obtained and designated as the value difference score (correlation) between germplasm resources a and b; the larger the value difference score (correlation), the more matched the first and second region status information are, indirectly indicating a greater correlation between germplasm resources a and b. The method for obtaining the correlation is a well-known technique and will not be elaborated further. The value difference score (correlation) between the a-th germplasm resource and each other germplasm resource is obtained. Other germplasm resources with value difference scores (correlation) greater than a preset threshold are designated as candidate germplasm resources for the a-th germplasm resource. This identifies germplasm resources with high correlation to the a-th germplasm resource, improving the efficiency of analyzing the application value score status associated with the a-th germplasm resource. For any candidate germplasm resource of the a-th germplasm resource, a higher value difference score (correlation) and a larger number of strains of *Sanghuang* detected indicate a better match between the a-th germplasm resource and the candidate germplasm resource. Conversely, a greater difference (lower correlation) in the application value scores of *Sanghuang* detected in each region associated with the a-th germplasm resource and the candidate germplasm resource suggests a higher likelihood of anomalies in the application value scores of *Sanghuang* detected in each region associated with the a-th germplasm resource. Therefore, based on the relationship between the a-th germplasm resource and the candidate germplasm resource... The value difference score (correlation) between sources, the number of strains of *Phellinus linteus* detected, and the difference in application value scores between the *a*-th germplasm resource and the candidate germplasm resource in each associated region are used to obtain the first development value between the *a*-th germplasm resource and the candidate germplasm resource. A higher first development value indicates a more likely anomaly in the application value score of the *a*-th germplasm resource associated with *Phellinus linteus*. The sum of the first development values between the *a*-th germplasm resource and each candidate germplasm resource is set as the outlier of the application value score of the *a*-th germplasm resource. A larger outlier indicates a less accurate application value score for the *a*-th germplasm resource associated with *Phellinus linteus*.
[0051] Taking the i-th germplasm resource as an example, this embodiment of the invention sets the preset value difference score (relevance) threshold to 0.6. The implementer can set the preset value difference score (relevance) threshold according to the actual situation; no limitation is made here. This allows for the acquisition of candidate germplasm resources for the i-th germplasm resource. The specific method for obtaining outliers in the application value score of the i-th germplasm resource is as follows:
[0052] (1) Obtain the first development value.
[0053] Taking the acquisition of the first development value between the i-th germplasm resource and its u-th candidate germplasm resource as an example, the formula for calculating the first development value between the i-th germplasm resource and its u-th candidate germplasm resource is as follows:
[0054] In the formula, P i,u The first development value between the i-th germplasm resource and its u-th candidate germplasm resource; n i,u γ represents the number of strain regions of *Phellinus linteus* detected between the i-th germplasm resource and its u-th candidate germplasm resource; i,u f is the value difference score (correlation) between the i-th germplasm resource and its u-th candidate germplasm resource; i,j The application value score of *Phellinus linteus* was detected for the i-th germplasm resource associated with the j-th region. The application value score of Phellinus linteus is detected by associating the u-th candidate germplasm resource of the i-th germplasm resource with the j-th region.
[0055] It should be noted that γ i,u The value must be greater than 0.6, meaning there must be a common harvesting method for *Sanghuang* between the i-th germplasm resource and its u-th candidate germplasm resource. i,u ≠0. n i,u The larger γ is i,u The larger the value, the higher the correlation between the i-th germplasm resource and its u-th candidate germplasm resource. The larger the value of P, the greater the difference in application value scores between the i-th germplasm resource and the u-th candidate germplasm resource harvested together, and the less normal the situation. This indicates that the application value score of the i-th germplasm resource associated with the pheasant is more likely to be abnormal. i,u The larger P is; therefore, i,u The larger the value, the more likely the application value score of the i-th germplasm resource associated with Phellinus linteus is to be inaccurate.
[0056] According to the method of obtaining the first development value between the i-th germplasm resource and its u-th candidate germplasm resource, the first development value between the i-th germplasm resource and each of its candidate germplasm resources is obtained.
[0057] (2) Obtain outliers in the application value score.
[0058] To more accurately analyze the abnormal state of the application value score of the i-th germplasm resource associated with Phellinus linteus, the record of the sum of the first development value between the i-th germplasm resource and each of its candidate germplasm resources is obtained and set as the abnormal value of the application value score of the i-th germplasm resource. The larger the abnormal value of the application value score of the i-th germplasm resource, the less accurate the application value score of the i-th germplasm resource associated with Phellinus linteus.
[0059] Based on the method for obtaining the application value score outliers of the i-th germplasm resource, obtain the application value score outliers of each germplasm resource.
[0060] Step S3: Set any one type of Sanghuang as effective Sanghuang. Based on the differences in characteristic information of the same strains between effective Sanghuang and Sanghuang from other regions, obtain the matching degree value between effective Sanghuang and Sanghuang from other regions, and screen out candidate Sanghuang for effective Sanghuang. Based on the outlier value of the application value score of each germplasm resource, the matching degree value between effective Sanghuang and candidate Sanghuang from each region, and the application value score of each germplasm resource associated with candidate Sanghuang from each region, obtain the confidence level of the application value score associated with effective Sanghuang for each germplasm resource.
[0061] Specifically, in practice, the extent of promotion of *Sanghuang* (a type of medicinal mushroom) affects the application value score of *Sanghuang* associated with germplasm resources, leading to an imbalance in the application value scores and making the application value score more biased towards a certain type of *Sanghuang* that has been maliciously promoted. To prevent maliciously promoted *Sanghuang* from affecting the application value score of *Sanghuang* associated with germplasm resources, thus causing inaccurate predictions of the application value score and inaccurate adjustments to the protection strategy information of associated germplasm resources, this embodiment of the invention analyzes the differences in characteristic information of the same strains between *Sanghuang* from different regions and *Sanghuang* from other regions. The smaller the difference in characteristic information of the same strains between a certain type of *Sanghuang* and another type of *Sanghuang*, the more compatible the two types of *Sanghuang* are. Therefore, based on the differences in characteristic information of the same strains between *Sanghuang* from different regions and other regions, a matching degree value is obtained between any two types of *Sanghuang*. The larger the matching degree value, the more compatible the two types of *Sanghuang* are. For clarity, this embodiment of the invention randomly selects one type of *Sanghuang* as the effective *Sanghuang*, obtains the matching degree value between the effective *Sanghuang* and *Sanghuang* from other regions, and filters out candidate *Sanghuang* that match the effective *Sanghuang* based on the magnitude of the matching degree value. The higher the application value score of the candidate Sanghuang associated with a germplasm resource, the more likely the effective Sanghuang is to be a popular Sanghuang, and the greater the possibility of malicious promotion of the effective Sanghuang. Therefore, the application value score of the effective Sanghuang associated with the germplasm resource is more likely to be inaccurate. As shown in step S2, the larger the outlier of the application value score of a germplasm resource, the less accurate the application value score of the Sanghuang associated with that germplasm resource. To more accurately analyze the reliability of the application value score of each effective Sanghuang associated with a germplasm resource, this embodiment of the invention obtains the confidence level of the application value score of each effective Sanghuang associated with a germplasm resource based on the outlier of the application value score of each germplasm resource, the matching degree value between the effective Sanghuang and the candidate Sanghuang in each region, and the application value score of each germplasm resource associated with the candidate Sanghuang in each region. The higher the confidence level of the application value score, the more accurate the application value score of the effective Sanghuang associated with the corresponding germplasm resource. The specific method for obtaining the confidence level of the application value score of each effective Sanghuang associated with a germplasm resource is as follows:
[0062] (1) Obtain the matching degree value.
[0063] Preferably, the method for obtaining the matching degree value is as follows: For any feature information of Sanghuang, the records that are negatively correlated and standardized between Sanghuang in region k and region g are set as matching records of that feature information between Sanghuang in region k and region g; the larger the number of matching records, the more matched the feature information is between Sanghuang in region k and region g. The mean of all matching records of feature information between Sanghuang in region k and region g is obtained and set as the matching degree value between Sanghuang in region k and region g.
[0064] Taking this as an example, to obtain the matching degree value between Phellinus linteus from region k and region g, the formula for calculating the matching degree value between Phellinus linteus from region k and region g is as follows:
[0065]
[0066] In the formula, S k,g denoted as , where is the matching degree value between Sanghuang in region k and region g; U is the total number of Sanghuang feature information; x k,u x represents the u-th feature information of Sanghuang in the k-th region; g,u Let be the uth characteristic information of Sanghuang in the g-th region; exp is an exponential function with the natural constant as the base; || is an absolute value function.
[0067] It should be noted that |x k,u -x g,u The smaller x is k,u With x g,u The more equal they are, the smaller the difference in the u-th feature information between Phellinus linteus in region k and region g, indirectly indicating a better match between Phellinus linteus in region k and region g, exp(-|x k,u -x g,u The larger |) is, The larger S is k,g The larger, therefore, S k,g The larger the value, the better the match between the *Sanghuang* in region k and region g. k,g The value range is from 0 to 1.
[0068] Based on the method for obtaining the matching degree value between Phellinus linteus in region k and Phellinus linteus in region g, the matching degree value of Phellinus linteus in region k and Phellinus linteus in other regions is obtained.
[0069] (2) Obtain candidate Sanghuang.
[0070] Taking Sanghuang from region k as an example, the method for obtaining candidate Sanghuang from region k is as follows: Arrange the matching degree values of Sanghuang from region k to Sanghuang from other regions in descending order to obtain matching degree value results; set the first preset number of matching degree values in the matching degree value results as valid matching degree values; in this embodiment, the first preset number is set to 10, but the implementer can set the size of the first preset number according to the actual situation, which is not limited here. Other Sanghuang from regions other than region k corresponding to each valid matching degree value are set as candidate Sanghuang from region k.
[0071] Based on the method for obtaining candidate Sanghuang in the k-th region, candidate Sanghuang in each region is obtained.
[0072] (3) Obtain the confidence score of the application value score.
[0073] Preferably, the method for obtaining the confidence level of the application value score is as follows: For any candidate Sanghuang that is effective, the first germplasm value of the effective Sanghuang and the candidate Sanghuang is obtained based on the matching degree value between the effective Sanghuang and the candidate Sanghuang, and the application value score of each germplasm resource associated with the candidate Sanghuang; when the first germplasm value is larger, the application value score of the germplasm resource associated with the candidate Sanghuang is lower, the candidate Sanghuang is more likely to be a low-quality Sanghuang, which indirectly indicates that the effective Sanghuang is also more likely to be a low-quality Sanghuang, the degree to which the effective Sanghuang is maliciously promoted is smaller, and the application value score corresponding to the effective Sanghuang is more accurate.
[0074] Low-quality *Sanghuang* (a type of medicinal mushroom) is less affected by heat. Therefore, the record of summing the first germplasm value of effective *Sanghuang* and candidate *Sanghuang* from each region is set as the reliable candidate value of effective *Sanghuang*. The larger the reliable candidate value, the lower the quality of the effective *Sanghuang*, the less likely it is to be maliciously marketed, and the more accurate the application value score corresponding to the effective *Sanghuang*. To accurately obtain the application value score accuracy of each germplasm resource associated with effective *Sanghuang*, the record of negatively correlated and standardized application value score outliers of each germplasm resource is set as the reliable adjustment value of the corresponding germplasm resource. The larger the reliable adjustment value, the smaller the application value score outliers of the corresponding germplasm resource, and the more accurate the application value score of the corresponding germplasm resource. Therefore, the product of the reliable adjustment value of each germplasm resource and the reliable candidate value of effective *Sanghuang* is set as the confidence level of the application value score associated with the corresponding germplasm resource. The larger the confidence level of the application value score, the more accurate the application value score of the corresponding effective *Sanghuang* associated with the germplasm resource.
[0075] As an example, taking the i-th germplasm resource and the Sanghuang from the j-th region as an example, the Sanghuang from the j-th region is considered a valid Sanghuang. The specific method for obtaining the application value score confidence score of the i-th germplasm resource associated with the Sanghuang from the j-th region is as follows:
[0076] (3-1) Obtain the first quality value.
[0077] Taking the candidate Sanghuang in region v of region j as an example, the formula for calculating the first quality value of Sanghuang in region j and its candidate Sanghuang in region v is as follows:
[0078]
[0079] In the formula, δ j,v The first germplasm value of Phellinus jumbo in region j and Phellinus jumbo in region v; S j,v denoted as the matching degree value of Phellinus j-th region and its candidate Phellinus j-th region; M represents the number of germplasm resources. The application value score of the candidate Sanghuang in the v region associated with the m-th germplasm resource in the j-th region; α is a first preset constant, which is greater than 0.
[0080] In this embodiment of the invention, α is set to 1 to avoid the denominator being 0. The implementer can set the size of α according to the actual situation, and there is no limitation here.
[0081] It should be noted that S j,v The larger, The more accurate, The smaller the value, the less suitable the candidate Sanghuang from region v is for associating the m-th germplasm resource with Sanghuang from region j, indirectly inferring that associating it with Sanghuang from region j is also unsuitable. The larger; when The larger the value, the more diverse the associated resources, indicating that the *Sanghuang* (a type of medicinal herb) in region j is unsuitable. The lower the quality of *Sanghuang* in region j, the more accurate the application value score corresponding to *Sanghuang* in region j. j,v The larger it is, the greater δ is. j,v The larger the value, the more accurate the application value score of Sanghuang in region j.
[0082] Based on the method for obtaining the first germplasm value of Phellinus jatamansi in region j and the first germplasm value of Phellinus jatamansi in region v, the first germplasm value of Phellinus jatamansi in region j and the first germplasm value of Phellinus jatamansi in each region is obtained.
[0083] (3-2) Obtain the confidence score of the application value score.
[0084] The formula for calculating the confidence score of the application value of the i-th germplasm resource associated with the j-th region of *Sanghuang* is as follows:
[0085]
[0086] In the formula, B i,j L represents the confidence score of the application value of the i-th germplasm resource associated with the j-th region of *Phellinus linteus*. i Y represents an outlier in the application value score of the i-th germplasm resource; j,w The first germplasm value of *Sanghuang* from region j and its w-th candidate *Sanghuang*; W is the number of candidate *Sanghuang* from region j, which is set to 10 in this embodiment; norm(-L) i ) represents the reliable adjustment value for the i-th germplasm resource; This is a reliable candidate value for Phellinus japonica in the j-th region.
[0087] It should be noted that L i The smaller the value, the more accurate the application value score of the i-th germplasm resource associated with Phellinus linteus. The reliable adjustment value norm(-L) i The larger ) is, the more B i,j The larger; Y j,w The larger the value, the lower the quality of the *Sanghuang* in region j, and the more accurate the corresponding application value score, making it a reliable candidate value. The larger B is i,j The larger, therefore, B i,jThe larger the value, the more accurate the application value score of the i-th germplasm resource associated with the j-th region of Sanghuang.
[0088] Based on the method of obtaining the confidence score of the application value of the i-th germplasm resource associated with the j-th region of Sanghuang, the confidence score of the application value of each germplasm resource associated with each region of Sanghuang is obtained.
[0089] Step S4: Adjust the confidence level of the associated application value score to obtain the adjusted confidence level of the application value score of each germplasm resource associated with effective Phellinus linteus.
[0090] Specifically, adjustments are made based on the harvesting frequency of effective Sanghuang associated with each germplasm resource, the difference between the first and last harvest times of effective Sanghuang for each germplasm resource, the fluctuation of the difference between two consecutive harvest times of effective Sanghuang for each germplasm resource, and the difference between the last harvest time of effective Sanghuang for each germplasm resource and the current time. The application value score of Sanghuang associated with each germplasm resource will degrade to a certain extent over time. The latent semantic model does not take into account the temporal characteristics of the application value score of Sanghuang associated with germplasm resources. Therefore, in order to more accurately adjust the protection strategy information associated with germplasm resources, the adjustment coefficient of the confidence score of the application value score of Sanghuang associated with each germplasm resource in each region is obtained based on the difference between the most recent harvest time of Sanghuang associated with each germplasm resource in each region and the current time. Meanwhile, considering that the harvesting status of Sanghuang in each region varies for each germplasm resource, and to avoid the application value score confidence level being adjusted too low over time, it is necessary to analyze the harvesting status of Sanghuang in each region associated with each germplasm resource, and further adjust the adjustment coefficient of the application value score confidence level of Sanghuang in each region associated with each germplasm resource.Therefore, this embodiment of the invention obtains the harvesting frequency of *Sanghuang* (a type of medicinal mushroom) associated with each germplasm resource in various regions and the harvesting time of each germplasm resource for harvesting *Sanghuang* in each region. The absolute value of the difference between the first and last harvesting times of a certain germplasm resource for a particular type of *Sanghuang* is set as the first time point difference value. The larger the first time point difference value, the larger the total harvesting cycle associated with that germplasm resource for that type of *Sanghuang* is likely to be, and the more suitable the association between the germplasm resource and that type of *Sanghuang* is. The adjustment coefficient for the confidence score of the application value score associated with that germplasm resource should be smaller. To avoid the situation where the actual harvesting frequency associated with that germplasm resource is very low, but the time interval between the first and last harvests of that type of *Sanghuang* is long, and thus mistakenly believing that the association between the germplasm resource and that type of *Sanghuang* is suitable, further analysis is performed on the harvesting frequency associated with that germplasm resource for that type of *Sanghuang*. The higher the harvesting frequency associated with that germplasm resource for that germplasm resource, the more accurately it can be concluded that the association between the germplasm resource and that type of *Sanghuang* is more suitable. The smaller the adjustment coefficient for the confidence score of the application value associated with the germplasm resource of this species of Sanghuang, the better. The absolute value of the difference between two consecutive harvesting times of this germplasm resource of this species of Sanghuang is set as the second time-point difference value. The smaller the fluctuation of the second time-point difference value, the more stable the harvesting frequency of this germplasm resource associated with this species of Sanghuang, and the smaller the adjustment coefficient for the confidence score of the application value associated with this germplasm resource of this species of Sanghuang should be. Therefore, the standard deviation of the second time-point difference value is obtained. The smaller the standard deviation of the second time-point difference value, the smaller the adjustment coefficient for the confidence score of the application value associated with this germplasm resource of this species of Sanghuang should be. Therefore, based on the first time-point difference value, the harvesting frequency of this germplasm resource associated with this species of Sanghuang, and the second time-point difference value, the adjustment coefficient for the confidence score of the application value associated with this germplasm resource of this species of Sanghuang is modified to obtain the modified adjustment coefficient. The confidence score of the application value associated with this germplasm resource of this species of Sanghuang is then accurately adjusted based on the modified adjustment coefficient. Therefore, based on the harvesting frequency of effective Sanghuang associated with each germplasm resource, the difference between the first and last harvesting times of effective Sanghuang for each germplasm resource, the fluctuation of the difference between two consecutive harvesting times of effective Sanghuang for each germplasm resource, and the difference between the last harvesting time of effective Sanghuang for each germplasm resource and the current time, the confidence level of the associated application value score is adjusted to obtain the adjusted application value score confidence level of effective Sanghuang associated with each germplasm resource.
[0091] As an example, taking the i-th germplasm resource and the Sanghuang from the j-th region in step S3 as an example, the Sanghuang from the j-th region is considered valid. The daily harvesting frequency e of the Sanghuang from the j-th region associated with the i-th germplasm resource is determined using big data. i,j And the harvesting time of each harvest of Sanghuang in the j-th region associated with the i-th germplasm resource, where the current time is t. nowThe most recent harvest time of Sanghuang associated with the i-th germplasm resource in the j-th region is... when The larger it is, the more it indicates With t now The greater the temporal difference, the less accurate the confidence score of the application value score of the i-th germplasm resource associated with the j-th region of *Sanghuang* at the current time point. Therefore, the degree of adjustment for the confidence score of the application value score of the i-th germplasm resource associated with the j-th region of *Sanghuang* should be. Thus, the embodiments of the present invention... The adjustment coefficient for the application value score confidence of the i-th germplasm resource associated with the j-th region of *Sanghuang* is obtained. However, in reality, the harvesting status of the i-th germplasm resource associated with the j-th region of *Sanghuang* may exhibit certain patterns. The more frequent and longer the harvesting cycle of the i-th germplasm resource associated with the j-th region of *Sanghuang*, the better. The smaller the value, the better, to ensure that the confidence score of the application value of the i-th germplasm resource associated with the j-th region of Sanghuang does not decrease too much over time. Therefore, the absolute value of the difference between the first and last harvest times of the i-th germplasm resource and the harvest time of the j-th region of Sanghuang is obtained. Set as the first time point difference value. Among them, For the i-th germplasm resource, the e-th i,j The harvesting time point for Sanghuang in the j-th area; Let be the harvest time point of the first harvest of Sanghuang from region j for the i-th germplasm resource; the larger the difference value at the first time point, the higher the harvest frequency e of the i-th germplasm resource associated with Sanghuang from region j. i,j A larger value indicates a longer total harvesting cycle for the i-th germplasm resource associated with the j-th region of *Sanghuang*, and a more suitable association between the i-th germplasm resource and the j-th region of *Sanghuang*. The absolute value of the difference between two consecutive harvesting times for the i-th germplasm resource in the j-th region of *Sanghuang* is the second time-point difference value. A smaller standard deviation of the second time-point difference value indicates a more stable harvesting frequency for the i-th germplasm resource associated with the j-th region of *Sanghuang*. The difference is further calculated using the first time-point difference value, the second time-point difference value, and e. i,j , related Adjustments are made to accurately obtain the confidence score of the adjusted application value score of *Sanghuang* from the j-th region associated with the i-th germplasm resource. The formula for calculating the confidence score of the adjusted application value score of *Sanghuang* from the j-th region associated with the i-th germplasm resource is as follows:
[0092]
[0093] In the formula, B′ i,j σ represents the adjusted application value score confidence level for the i-th germplasm resource associated with the j-th region of *Phellinus linteus*. i,j Let e be the standard deviation of the absolute value of the difference between two consecutive harvesting times of *Sanghuang* from region j for the i-th germplasm resource;i,j The harvesting frequency of Phellinus japonica in the i-th germplasm resource is associated with the i-th germplasm resource. For the first time of the i-th germplasm resource and the e-th germplasm resource i,j The difference in harvesting time points between the second harvest of Sanghuang in region j; For the i-th germplasm resource, the e-th i,j The harvesting time point for Sanghuang in the j-th area; Let Δt′ be the harvesting time point of the first harvest of Sanghuang from the i-th germplasm resource in the j-th region; i,j For the i-th germplasm resource, the e-th i,j The difference between the harvest time of the second harvest of Sanghuang in region j and the current time; t now B represents the current time point. i,j ω is the confidence score of the application value of the i-th germplasm resource associated with the j-th region of Sanghuang; ω is the second preset constant, which is greater than 0.
[0094] In this embodiment of the invention, ω is set to 1 to avoid the numerator and denominator being 0. The implementer can set the value of ω according to the actual situation, and there is no limitation here.
[0095] It should be noted that σ i,j The smaller the value, the more stable the harvesting frequency of *Sanghuang* in the j-th region associated with the i-th germplasm resource. i,j The larger, The larger the value, the longer the harvesting cycle of *Sanghuang* from region j associated with the i-th germplasm resource, and the more suitable the i-th germplasm resource is for *Sanghuang* from region j. The smaller the value, the stronger the correlation Δt′. i,j The smaller the degree of adjustment, The larger the value, the more associated with B. i,j The smaller the degree of adjustment, the better B′ i,j The larger, therefore, B′ i,j The larger the value, the more stable the application value score of the i-th germplasm resource associated with the j-th region of Sanghuang.
[0096] Based on the method for obtaining the adjusted application value score confidence of the i-th germplasm resource associated with the j-th region of Sanghuang, the adjusted application value score confidence of each germplasm resource associated with each region of Sanghuang is obtained.
[0097] Step S5: Adjust the cost function of the latent semantic model based on the confidence score of the adjusted application value of each germplasm resource associated with the Sanghuang of each region, obtain the predicted application value matrix, and adjust it based on the protection strategy information of each germplasm resource.
[0098] Specifically, the application value score of each germplasm resource associated with Sanghuang from each region is set as an element to construct an initial application value matrix. Each row of the initial application value matrix corresponds to one germplasm resource, and each column corresponds to one Sanghuang strain. Elements in the initial application value matrix that are greater than a third preset constant are set as valid elements. In this embodiment, the third preset constant is set to 0, meaning that only Sanghuang harvested from each germplasm resource is analyzed. The cost function E = Σ of the latent semantic model is obtained through the initial application value matrix. (x,y)∈K (f x,y -p x T q y ) 2 Where E is the cost function; K is the set of coordinates of the effective elements in the first application value matrix; f x,y p represents the application value score of the x-th germplasm resource associated with the y-th species of *Phellinus sylvestris*. x T Let q be the transpose of the potential index vector corresponding to the x-th germplasm resource; y Let be the latent indicator vector corresponding to the y-th Sanghuang. Based on the coordinates of the effective elements in the initial application value matrix, determine the confidence level of the adjusted application value score corresponding to each effective element; based on the confidence level of the adjusted application value score corresponding to each effective element, associate the cost function E = ∑ (x,y)∈K (f x,y -p x T q y ) 2 Adjustments are made to obtain the adjustment cost function; based on the adjustment cost function, the potential index matrix of germplasm resources and the potential index matrix of Sanghuang are obtained, and a predictive application value matrix is constructed.
[0099] This invention is now complete.
[0100] In summary, this invention obtains the harvesting cycle of *Sanghuang* associated with germplasm resources and determines the application value score of *Sanghuang* associated with germplasm resources; it obtains application value score outliers based on the application value score status of *Sanghuang* harvested jointly by germplasm resources; it obtains matching degree values based on the differences in characteristic information of the same strains between *Sanghuang* and other *Sanghuang*; it obtains application value score confidence based on application value score outliers, matching degree values, and application value scores of *Sanghuang* associated with germplasm resources; it adjusts the application value score confidence based on the harvesting status of *Sanghuang* associated with germplasm resources, obtains adjusted application value score confidence, adjusts the cost function of the associated latent semantic model, obtains a predicted application value matrix, and adjusts the protection strategy information of associated germplasm resources. This invention adaptively obtains the adjusted application value score confidence of *Sanghuang* associated with each germplasm resource in each region, thereby accurately obtaining the predicted application value score of *Sanghuang* associated with each germplasm resource in each region and accurately adjusting the protection strategy information associated with each germplasm resource.
[0101] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0102] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
[0103] The specific embodiments of the above-described related inventions have been described in detail, but they are only provided as examples. The present invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions made related to this invention are also within the scope of the present invention. Therefore, all equivalent transformations, modifications, and improvements made without departing from the spirit and principles of the present invention should be covered within the scope of the present invention.
Claims
1. A method for the protection and development of Phellinus linteus germplasm resources, characterized in that, The method includes the following steps: Obtain the harvesting cycle and frequency of Sanghuang in each region associated with a preset number of germplasm resources, as well as the regional characteristic information of Sanghuang in each region; set the first harvest time point of each harvest of Sanghuang in each region associated with each germplasm resource as the harvest time point. Based on the harvesting cycle, obtain the application value score of each germplasm resource associated with Sanghuang in each region; based on the matching status of the application value score of each germplasm resource with each other germplasm resource associated with Sanghuang harvested together, the number of strains of Sanghuang harvested together, and the difference in application value score of Sanghuang harvested together, obtain the outlier value of the application value score of each germplasm resource. Any species of *Sanghuang* is designated as a valid *Sanghuang*. Based on the differences in characteristic information of the same strains between valid *Sanghuang* and *Sanghuang* from other regions, the matching degree value of valid *Sanghuang* with *Sanghuang* from other regions is obtained, and candidate *Sanghuang* for valid *Sanghuang* is screened. Based on the outlier value of the application value score of each germplasm resource, the matching degree value between valid *Sanghuang* and candidate *Sanghuang* from each region, and the application value score of each germplasm resource associated with candidate *Sanghuang* from each region, the confidence level of the application value score associated with valid *Sanghuang* for each germplasm resource is obtained. The application value score confidence level is adjusted based on the associated application value score confidence level to obtain the adjusted application value score confidence level of each germplasm resource associated with effective Phellinus linteus. The application value score confidence of each germplasm resource associated with the latent semantic model is adjusted to obtain the predicted application value matrix, and the protection strategy information of each germplasm resource is also adjusted accordingly. The higher the confidence level of the application value score, the more accurate the application value score of the corresponding germplasm resource associated with effective Phellinus linteus. The method for obtaining outliers in the application value score is as follows: The Sanghuang harvested jointly from germplasm resources a and b is designated as the Sanghuang for testing. The application value score of Sanghuang in each region is associated with the a-th germplasm resource and set as the status information of the first region. The application value scores of Phellinus linteus detected in each region are associated with the b-th germplasm resource and set as the status information of the second region. Obtain the correlation between the status information of the first region and the status information of the second region, and set it as the value difference score between the a-th germplasm resource and the b-th germplasm resource; Obtain the value difference score between the a-th germplasm resource and each other germplasm resource, and set the other germplasm resources corresponding to the value difference score that is greater than the preset value difference score threshold as candidate germplasm resources of the a-th germplasm resource; For any candidate germplasm resource of germplasm resource a, the first development value between germplasm resource a and candidate germplasm resource is obtained based on the value difference score between germplasm resource a and candidate germplasm resource, the number of strain regions that detect Sanghuang, and the difference in application value scores of Sanghuang in each region associated with germplasm resource a and candidate germplasm resource. The record of the sum of the first development value between the a-th germplasm resource and each candidate germplasm resource is set as the application value score outlier of the a-th germplasm resource.
2. The method for the protection and development of Phellinus linteus germplasm resources according to claim 1, characterized in that, The method for obtaining the application value score is as follows: The harvesting cycle of Sanghuang in each region associated with each germplasm resource was recorded in a standardized manner, and an application value score was set for each germplasm resource associated with Sanghuang in each region.
3. The method for the protection and development of Phellinus linteus germplasm resources according to claim 1, characterized in that, The method for obtaining the matching degree value of effective Sanghuang and Sanghuang from other regions based on the characteristic information of the same strains between effective Sanghuang and other Sanghuang from other regions, and screening out candidate Sanghuang for effective Sanghuang, is as follows: For any feature information of Sanghuang, the record that is negatively correlated and standardized between Sanghuang in region k and Sanghuang in region g is set as the matching record of that feature information between Sanghuang in region k and Sanghuang in region g. The mean of all matching records of species feature information between Phellinus linteus in region k and Phellinus linteus in region g is obtained and set as the matching degree value between Phellinus linteus in region k and Phellinus linteus in region g. Obtain the matching degree value of Sanghuang in region k with Sanghuang in other regions, and sort the matching degree values of Sanghuang in region k with Sanghuang in other regions from largest to smallest to obtain the matching degree value result; Set the first preset number of matching scores in the matching score results as valid matching scores; For each valid matching degree value, other Sanghuang (Phellinus thunbergii) that are not in the k-th region are set as candidate Sanghuang for the k-th region.
4. The method for the protection and development of Phellinus linteus germplasm resources according to claim 1, characterized in that, The method for obtaining the application value score confidence level is as follows: For any candidate Sanghuang that is effective, the first germplasm value of effective Sanghuang and candidate Sanghuang is obtained based on the matching degree value between effective Sanghuang and candidate Sanghuang, and the application value score of each germplasm resource associated with candidate Sanghuang. The record of summing the first germplasm value of effective Phellinus linteus and the candidate Phellinus linteus from each region is set as the reliable candidate value of effective Phellinus linteus; the record of negatively correlated and standardized application value score outliers of each germplasm resource is set as the reliable adjusted value of the corresponding germplasm resource. The product of the reliable adjusted value of each germplasm resource and the reliable candidate value of the effective Phellinus linteus is set as the confidence score of the application value associated with the effective Phellinus linteus of the corresponding germplasm resource.
5. The method for the protection and development of Phellinus linteus germplasm resources according to claim 2, characterized in that, The method for adjusting the cost function of the latent semantic model based on the confidence score of the applied value score associated with each germplasm resource of Sanghuang in each region to obtain the predicted application value matrix is as follows: Each germplasm resource is associated with the application value score of Sanghuang from each region as an element to construct an initial application value matrix; where each row of the initial application value matrix corresponds to a germplasm resource and each column corresponds to a Sanghuang strain. Elements in the initial application value matrix that are greater than the third preset constant are set as valid elements. Based on the coordinates of the valid elements in the initial application value matrix, the confidence level of the adjusted application value score corresponding to each valid element is determined. The cost function of the latent semantic model is adjusted based on the confidence score of the adjusted application value score corresponding to each valid element, and the adjusted cost function is obtained. Based on the adjusted cost function, the potential index matrix of germplasm resources and the potential index matrix of Sanghuang are obtained, and a predictive application value matrix is constructed.
Citation Information
Patent Citations
Efficient comprehensive evaluation method and system for mulberry germplasm resources
CN115983679A
Smart screen information recommendation method based on big data
CN117994010A