A forest tree genetic breeding screening system based on big data
Through the big data forest genetic breeding screening system, the modular architecture and tensor flow adaptive screening algorithm are used to solve the problems of low screening efficiency and result deviation in the existing technology, and efficient and accurate breeding scheme generation and screening are achieved.
Patent Information
- Application Number
- CN202510607215.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-13
AI Technical Summary
When existing forest genetic breeding screening systems based on machine learning and deep learning are facing noise interference, labeling errors or sample imbalances, it is difficult to effectively capture the true genetic laws in high-dimensional sparse data, resulting in low screening efficiency and biased results. The deep learning model consumes high computing resources, is costly, and has poor model interpretability, making it difficult to process large-scale dynamic data.
A forest genetic breeding screening system based on big data is adopted, including information collection module, breeding scheme generation module, data processing module, adaptive screening module, evaluation and acquisition module and optimization module. Through noise cleaning, genetic relationship analysis, feature marking, dual attention mechanism screening and weight redistribution, combined with tensor flow adaptive screening algorithm, the optimal breeding scheme that meets user needs is generated.
It improves the screening accuracy and efficiency of breeding programs, can process larger data volume and real-time data, quickly integrate multi-region genetic information, and provide efficient and accurate breeding decision support.
Smart Images

Figure CN120124988B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a forest tree genetic breeding screening system based on big data. Background Art
[0002] For a forest tree genetic breeding screening system based on big data, the quality of its performance mainly depends on the accuracy and efficiency of screening. Currently, the mainstream screening schemes based on machine learning algorithms and deep learning models have exposed significant defects. In machine learning algorithms, their dependence on data quality is too high. Once there are noisy data or missing values, it is extremely easy to cause deviations in the results. While for deep learning models, they face a series of thorny problems. On the one hand, their training process consumes extremely large amounts of computing resources and requires the strong support of high-performance hardware, which not only results in high costs and greatly increases scientific research and production inputs, but also has poor model interpretability and shows black-box characteristics. This black-box characteristic seriously affects the credibility of key decisions. On the other hand, deep learning models also bring great difficulties to the debugging work. When the model has problems, due to the inability to intuitively understand its internal mechanism, the debugging efficiency is low, and it is difficult to quickly locate and solve the problems.
[0003] The core problems existing in the existing systems include: it is difficult to effectively capture the true genetic laws in high-dimensional sparse data when facing noise interference, mislabeling, or sample imbalance, etc.; at the same time, data distribution drift will lead to prediction failure, thus requiring frequent online learning, but this process may introduce short-term noise; in addition, rare scenarios will result in low confidence in the screening results due to insufficient data, making it difficult for the system to balance screening efficiency and result accuracy when processing large-scale dynamic data. These problems ultimately lead to extremely low screening efficiency and significant deviations in the screening results when screening breeding schemes, greatly restricting the actual application effectiveness of the system. Summary of the Invention
[0004] The present invention provides a forest tree genetic breeding screening system based on big data, and its main purpose is to solve the problem of low screening efficiency and screening quality when screening breeding schemes.
[0005] To achieve the above object, the present invention provides a forest tree genetic breeding screening system based on big data, and the system includes a status information collection module, a breeding scheme generation module, a data processing module, an adaptive screening module, an evaluation acquisition module, an optimization module, and an updated adaptive screening module, wherein:
[0006] The information collection module is used to obtain the genetic breeding data of forest trees;
[0007] The breeding scheme generation module is used to perform correlation analysis on the genetic breeding data to obtain the breeding scheme of the forest trees;
[0008] The data processing module is used to perform feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan;
[0009] The adaptive screening module is used to obtain the breeding requirements of the user, and based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm, screen the breeding plan to obtain the optimal plan corresponding to the breeding requirements;
[0010] The evaluation acquisition module is used to obtain the score of the user for the optimal plan;
[0011] The optimization module is used to perform weight redistribution on the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula;
[0012] The updated adaptive screening module is used to synchronize the updated adaptive screening formula to the TensorFlow adaptive screening algorithm to obtain an updated TensorFlow adaptive screening algorithm, and screen the breeding plan according to the new requirements obtained next time and the updated TensorFlow adaptive screening algorithm to obtain the target breeding plan for the new requirements.
[0013] In a preferred embodiment, when the information collection module executes to obtain the genetic breeding data of forest trees, it specifically is used for:
[0014] Collect phenotypic data, genotype data, environmental factor data, and historical breeding records related to forest tree genetic breeding.
[0015] In a preferred embodiment, when the breeding plan generation module executes to perform correlation analysis on the genetic breeding data to obtain the breeding plan of the forest trees, it specifically is used for:
[0016] Perform noise cleaning on the genetic breeding data to obtain the standard data of the genetic breeding data;
[0017] Perform genetic relationship analysis on the standard data to obtain the trait heritability evaluation map of the standard data;
[0018] Perform correlation analysis on the trait heritability evaluation map to obtain the breeding plan of the forest trees.
[0019] In a preferred embodiment, when the data processing module executes to perform feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan, it specifically is used for:
[0020] Classify the genetic characteristics of the forest trees to obtain the feature tags of the forest trees;
[0021] Extract the features of the breeding program based on the ETL tool to obtain the feature data of the breeding program;
[0022] Annotate the feature data based on the feature tags to obtain the dynamic annotation of the breeding program.
[0023] In a preferred embodiment, when the adaptive screening module performs adaptive screening on the breeding program based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm to obtain the optimal solution corresponding to the breeding requirements, it is specifically used for:
[0024] Perform initial screening on the dynamic annotation based on the adaptive screening formula and the breeding requirements to obtain the screening value of the dynamic annotation, where the adaptive screening formula is:
[0025]
[0026] In the formula, is the screening value of the dynamic annotation, is the quantity of the breeding requirements, is the quantity of the dynamic annotation, is the correlation between the breeding requirements and the dynamic annotation, is the maximum possible entropy of the joint distribution of the breeding requirements and the dynamic annotation, is the breeding requirement, is the dynamic annotation.
[0027] Screen the breeding program based on the screening value and the dynamic annotation to obtain the available breeding programs for the breeding requirements;
[0028] Calculate the feature value of the available breeding programs based on the dual attention mechanism, where the dual attention mechanism is:
[0029]
[0030] In the formula, is the feature value of the available breeding programs, is the Sigmoid activation function, is the spatial feature weight, is the channel weight, is all the features of the available breeding programs, is the significant feature of the available breeding programs, is the available breeding program;
[0031] Select the available breeding plan with the highest eigenvalue as the optimal plan corresponding to the breeding demand.
[0032] In a preferred embodiment, when the evaluation acquisition module executes to obtain the score of the user for the optimal plan, it is specifically configured to:
[0033] The user scores the optimal plan to obtain the score of the optimal plan, where the score of the optimal plan includes: the accuracy score, the coverage score, and the timeliness score of the optimal plan.
[0034] In a preferred embodiment, when the optimization module executes to re - allocate weights to the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula, it is specifically configured to:
[0035] Analyze the optimal plan based on the score and the reward function to obtain the screening score of the optimal plan, where the reward function is:
[0036]
[0037] In the formula, is the screening score, is the accuracy score of the optimal plan, is the coverage score of the optimal plan, is the timeliness score of the optimal plan;
[0038] Re - allocate weights to the adaptive screening formula based on the optimization formula and the screening score to obtain the weight matching degree of the adaptive screening formula, where the optimization formula is:
[0039]
[0040] In the formula, is the re - allocated weight matching degree, is the attenuation degree of weight allocation at time is the weight allocation value attenuation, is the credibility of weight allocation at time is the influence of weight allocation source, is the dispersion factor of weight allocation at time is the dispersion degree of weight allocation at time is the weight allocation at time is The weight assignment value at a moment, is the total number of weight assignment values, is the time factor, is the number of the weight assignment value, is the screening score, is the hyperbolic tangent function;
[0041] Based on the weight assignment value with the highest weight matching degree, re - assign weights to the adaptive screening formula to obtain an updated adaptive screening formula.
[0042] In a preferred embodiment, the weight assignment value at a moment has the following calculation formula:
[0043]
[0044] In the formula, is the weight assignment value at a moment, is the initial weight assignment value, is the weight assignment progress at a moment, is the optimal adjustment direction of the weight assignment value, is the hyperbolic tangent function, is the screening score at a moment, is the time factor.
[0045] In a preferred embodiment, the tensor - flow adaptive screening algorithm further includes: an adaptive threshold formula and an adversarial verification formula. Among them, the adaptive threshold formula is:
[0046]
[0047] In the formula, is the screening adaptive threshold at a moment, is the screening degree, is the historical screening score, is the number of breeding programs at a moment, is the influence of the number of breeding programs at a moment on the screening result, is the time factor.
[0048] In a preferred embodiment, the adversarial verification formula is:
[0049]
[0050] In the formula, is the actual value of the weight assignment value, is random noise, is the mathematical expectation, is the discriminator in the tensor flow adaptive screening algorithm, is the generator in the tensor flow adaptive screening algorithm, is the distribution probability of the actual value, is the distribution probability of the random noise, is the discriminator the probability of identifying the actual value, is based on the generator and the random noise the predicted value generated, is the discriminator the probability of identifying the predicted value, is for the discriminator the logarithm value of the probability that the discriminator fails to identify the predicted value, is for the discriminator the logarithm value of the probability that the discriminator identifies the predicted value, is to measure the ability of the discriminator in the tensor flow adaptive screening algorithm to distinguish the actual value and the predicted value of the weight assignment value.
[0051] The present invention generates a breeding plan by collecting and analyzing the genetic breeding data of forest trees, and then screens the breeding plan to obtain a plan that meets the user's requirements; during the screening process, by using the tensor flow adaptive screening algorithm, it is determined whether the data meets the requirements, and the data that does not meet the requirements is discarded in a timely manner. At the same time, importance assessment is carried out for different features, focusing on the features that have a significant impact on the breeding results, avoiding screening biases caused by ignoring important features or being interfered by irrelevant features, and then accurately identifying the key information closely related to breeding from many genetic features, greatly improving the accuracy of the screening results; in addition, a modular architecture design is adopted, and this design pattern is conducive to function expansion, enabling the present invention to process a larger amount of data, being able to quickly collect and integrate the genetic information of forest trees from different regions, and being able to efficiently process a large amount of real-time data in a short time, effectively improving the screening efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is the system architecture diagram of the forest tree genetic breeding screening system based on big data provided by an embodiment of the present invention;
[0053] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE INVENTION
[0054] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. "Plural" generally includes at least two.
[0056] Depending on the context, the words "if", "when" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0057] In addition, the sequence of steps in the following method embodiments is only an example, not strictly limited.
[0058] In fact, the server devices deployed by the big data-based forest genetic screening system may be composed of one or more devices. The above-mentioned big data-based forest genetic breeding screening system can be implemented as: a business instance, a virtual machine, or a hardware device. For example, the big data-based forest genetic breeding screening system can be implemented as a business instance deployed on one or more devices in a cloud node. Simply put, the big data-based forest genetic breeding screening system can be understood as a software deployed on a cloud node for providing the big data-based forest genetic screening system for each client. Or, the big data-based forest genetic breeding screening system can also be implemented as a virtual machine deployed on one or more devices in a cloud node. An application software for managing each client is installed in the virtual machine. Or, the big data-based forest genetic screening system can also be implemented as a server composed of many identical or different types of hardware devices, and one or more hardware devices are set to provide the big data-based forest genetic breeding screening system for each client.
[0059] In terms of implementation form, the big data-based forest tree genetic screening system and the user terminal adapt to each other. That is, if the big data-based forest tree genetic screening system is an application installed on the cloud service platform, then the user terminal is a client that establishes a communication connection with this application; or if the big data-based forest tree genetic breeding screening system is implemented as a website, then the user terminal is implemented as a web page; or if the big data-based forest tree genetic breeding screening system is implemented as a cloud service platform, then the user terminal is implemented as a small program in an instant messaging application.
[0060] As Figure 1 shown, it is the system architecture diagram of the big data-based forest tree genetic breeding screening system provided by an embodiment of the present invention.
[0061] The big data-based forest tree genetic screening system 100 of the present invention can be set in a cloud server. In terms of implementation form, it can be one or more service devices, or can be an application installed on the cloud (such as the server of a mobile service operator, a server cluster, etc.), or can also be developed into a website. According to the functions to be achieved, the big data-based forest tree genetic screening system 100 can include a status information collection module 101, a breeding plan generation module 102, a data processing module 103, an adaptive screening module 104, an evaluation acquisition module 105, an optimization module 106, and an updated adaptive screening module 107. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0062] In the embodiment of the present invention, in the big data-based forest tree genetic screening system, each of the above modules can be independently implemented and called with other modules. Here, the call can be understood as that a certain module can be connected to multiple modules of another type and provide corresponding services for the multiple modules it is connected to. For example, the sharing and evaluation module can call the same information collection module to obtain the information collected by this information collection module. Based on the above characteristics, in the big data-based forest tree genetic breeding screening system provided by the embodiment of the present invention, without modifying the program code, the applicable range of the big data-based forest tree genetic screening system architecture can be adjusted by adding modules and directly calling them, so as to achieve cluster-level horizontal expansion, so as to achieve the purpose of quickly and flexibly expanding the big data-based forest tree genetic breeding screening system. In practical applications, the above modules can be set in the same device or different devices, or can also be set in virtual devices, such as service instances in a cloud server.
[0063] Next, in combination with specific embodiments, the respective components and specific working processes of the big data-based forest tree genetic screening system will be described respectively:
[0064] The information collection module 101 is used to obtain the genetic breeding data of forest trees.
[0065] In an embodiment of the present invention, when the information collection module executes the operation of obtaining the genetic breeding data of forest trees, it specifically is used for:
[0066] Collect phenotypic data, genotype data, environmental factor data, and historical breeding records related to the genetic breeding of forest trees.
[0067] The breeding plan generation module 102 is used to perform a correlation analysis on the genetic breeding data to obtain the breeding plan of the forest trees.
[0068] In an embodiment of the present invention, when the breeding plan generation module executes the operation of performing a correlation analysis on the genetic breeding data to obtain the breeding plan of the forest trees, it specifically is used for:
[0069] Clean the noise points in the genetic breeding data to obtain the standard data of the genetic breeding data;
[0070] Analyze the genetic relationship of the standard data to obtain the trait heritability evaluation map of the standard data;
[0071] Perform a correlation analysis on the trait heritability evaluation map to obtain the breeding plan of the forest trees.
[0072] Specifically, for the collected genetic breeding data, identify and remove the possible noise points in the data set. In this process, through multiple means such as setting a reasonable data threshold range, using an outlier detection model, and manual review, comprehensively check for noise data, including but not limited to outliers caused by measurement errors, error values that occur during data transmission, etc. After cleaning, obtain the standard data of the genetic breeding data.
[0073] Specifically, based on the cleaned standard data, deeply analyze the genetic relationship inside the data. Through the excavation and analysis of genetic information such as gene loci and chromosome segments, combined with the biological inheritance theory, accurately evaluate the heritability of various traits, and present the evaluation results in a visual way to generate a detailed trait heritability evaluation map.
[0074] Generally speaking, the trait heritability evaluation map clearly shows the strength of heritability, inheritance patterns of different traits, and the potential associations between various traits, providing an intuitive basis for deeply understanding the genetic laws.
[0075] Specifically, a comprehensive correlation analysis is carried out on the genetic heritability evaluation map of traits. By comprehensively considering factors such as the heritability of different traits, the degree of correlation between them, and the compatibility with the target breeding characteristics, weight distribution and combination calculation are performed on various factors. Finally, according to the analysis results, a breeding plan for the forest trees is formulated.
[0076] Generally speaking, the breeding plan covers key contents such as seed selection strategies, hybrid combination design, and suggestions for cultivating environment regulation, so as to achieve the goals of efficient and precise forest tree genetic breeding.
[0077] The data processing module 103 is used to perform feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan.
[0078] In the embodiment of the present invention, when the data processing module performs feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan, it specifically is used for:
[0079] Classify the genetic characteristics of the forest trees to obtain the feature tags of the forest trees;
[0080] Based on the ETL tool, perform feature extraction on the breeding plan to obtain the feature data of the breeding plan;
[0081] Based on the feature tags, annotate the feature data to obtain the dynamic annotation of the breeding plan.
[0082] Specifically, using bioinformatics methods and classification algorithms, comprehensively sort out and classify the genetic characteristics of forest trees. Starting from multiple dimensions such as gene sequence characteristics and phenotypic characteristics, according to the pre-set classification criteria and feature attributes, accurately classify a large number of genetic characteristics. During the classification process, a unique and clearly directive identifier will be assigned to each class of genetic characteristics as the feature tags of the forest trees.
[0083] Generally speaking, the feature tags can concisely and accurately summarize the essential attributes of various genetic characteristics, providing a clear index and classification basis for subsequent data processing and analysis.
[0084] Specifically, through a powerful ETL tool, perform feature extraction operations on the breeding plan. First, accurately extract various data related to the plan from various data sources involved in the breeding plan, such as experimental records, literature materials, historical data, etc. Secondly, use the conversion rules and algorithms built into the ETL tool to perform a series of processing on the extracted data, such as format conversion, data cleaning, integration calculation, etc., to refine a data set that can accurately reflect the core elements and key characteristics of the breeding plan, and finally obtain the feature data of the breeding plan.
[0085] Generally speaking, the characteristic data highly condenses the key information of the breeding program, laying a solid foundation for subsequent annotation and analysis.
[0086] Specifically, with the previously generated forest tree characteristic tags as the reference basis, the characteristic data of the extracted breeding program is carefully annotated. Each data item in the characteristic data is matched and associated with the corresponding characteristic tag, and it is given the corresponding category tag and attribute description, thus completing the initial annotation of the characteristic data. On this basis, considering the dynamics of the breeding process and the real-time change characteristics of the data, a dynamic annotation mechanism is introduced.
[0087] In detail, by continuously monitoring and analyzing newly generated data, changes in environmental factors, and feedback information during the implementation of the breeding program, the annotated characteristic data is updated and corrected in real time, generating a dynamic annotation that can reflect the real-time state and dynamic changes of the breeding program.
[0088] Generally speaking, dynamic annotation can provide timely, accurate, and comprehensive information support for breeding decisions, helping to achieve efficient and precise forest tree genetic breeding practices.
[0089] The adaptive screening module 104 is used to obtain the breeding requirements of the user, and based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm, the breeding program is screened to obtain the optimal program corresponding to the breeding requirements.
[0090] In the embodiment of the present invention, when the adaptive screening module executes screening of the breeding program based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm to obtain the optimal program corresponding to the breeding requirements, it is specifically used for:
[0091] Based on the adaptive screening formula and the breeding requirements, an initial screening of the dynamic annotation is performed to obtain the screening value of the dynamic annotation, where the adaptive screening formula is:
[0092]
[0093] In the formula, is the screening value of the dynamic annotation, is the quantity of the breeding requirements, is the quantity of the dynamic annotation, is the correlation between the breeding requirements and the dynamic annotation, is the maximum possible entropy of the joint distribution of the breeding requirements and the dynamic annotation, is the breeding requirement, is the dynamic annotation.
[0094] Screen the breeding plan based on the screening value and the dynamic annotation to obtain an available breeding plan for the breeding requirement;
[0095] Calculate the available breeding plan based on the dual attention mechanism to obtain the eigenvalue of the available breeding plan, where the dual attention mechanism is:
[0096]
[0097] In the formula, is the eigenvalue of the available breeding plan, is the Sigmoid activation function, is the spatial feature weight, is the channel weight, is all the features of the available breeding plan, is the significant feature of the available breeding plan, is the available breeding plan;
[0098] Select the available breeding plan with the highest eigenvalue as the optimal plan corresponding to the breeding requirement.
[0099] Specifically, from a large number of dynamic annotations, accurately screen out the data highly relevant to the breeding requirement according to the breeding requirement, provide data support for breeding decisions, and at the same time enable breeders to judge which dynamic annotations are the most critical for achieving the breeding requirement based on the screening value, so as to reasonably arrange breeding resources and formulate more effective breeding strategies.
[0100] For example, among the numerous dynamic annotations reflecting the growth status, genetic characteristics, etc. of forest trees, quickly locate the data that meet specific breeding goals (such as pest resistance, fast growth, etc.) to improve the screening efficiency and accuracy.
[0101] Generally speaking, the larger the screening value, the more the dynamic annotation matches the breeding requirement, and the more worthy of attention and retention in the screening process.
[0102] Specifically, retain the breeding plan with a screening value greater than 0.6 as the available breeding plan for the breeding requirement.
[0103] Specifically, the adaptive screening formula systematically evaluates numerous potential available breeding plans, quickly screens out the most potential plans, and based on the calculated eigenvalues, the advantages and disadvantages of each breeding plan can be clearly judged, so as to reasonably allocate resources, preferentially select the plans with high eigenvalues for implementation, and improve the breeding success rate and resource utilization efficiency.
[0104] For example, in a large number of different combinations of forest tree breeding programs, by calculating the eigenvalue, accurately locate those programs that perform well in terms of spatial characteristics and channel characteristics, etc., and whose significant characteristics cooperate well with all characteristics, thereby improving the screening efficiency and accuracy.
[0105] Generally speaking, the eigenvalue is the core output result of the entire formula, representing the overall performance level of the breeding program after comprehensively considering various characteristics. The higher the eigenvalue, the more the breeding program meets the user's needs.
[0106] The evaluation acquisition module 105 is used to obtain the user's score for the optimal solution.
[0107] In the embodiment of the present invention, when the evaluation acquisition module executes to obtain the user's score for the optimal solution, it is specifically used for:
[0108] The user scores the optimal solution to obtain the score of the optimal solution, where the score of the optimal solution includes: the accuracy score, the coverage score, and the timeliness score of the optimal solution.
[0109] The optimization module 106 is used to re - distribute the weights of the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula.
[0110] In the embodiment of the present invention, when the optimization module executes to re - distribute the weights of the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula, it is specifically used for:
[0111] Analyze the optimal solution based on the score and the reward function to obtain the screening score of the optimal solution, where the reward function is:
[0112]
[0113] In the formula, is the screening score, is the accuracy score of the optimal solution, is the coverage score of the optimal solution, is the timeliness score of the optimal solution;
[0114] Re - distribute the weights of the adaptive screening formula based on the optimization formula and the screening score to obtain the weight matching degree of the adaptive screening formula, where the optimization formula is:
[0115]
[0116] In the formula, is the weight matching degree after redistribution, is the attenuation degree of weight distribution at moment, is the attenuation of the weight distribution value, is the credibility of weight distribution at moment, is the influence of the weight distribution source, is the dispersion factor of weight distribution at moment, is the dispersion degree of weight distribution at moment, is the weight distribution at moment, is the weight distribution value at moment, is the total number of weight distribution values, is the time factor, is the number of the said weight distribution value, is the said screening score, is the hyperbolic tangent function;
[0117] Based on the weight distribution value with the highest weight matching degree, the weight of the said adaptive screening formula is redistributed to obtain an updated adaptive screening formula.
[0118] The said weight distribution value at moment has the following calculation formula:
[0119]
[0120] In the formula, is the said weight distribution value at moment, is the initial weight distribution value, is the progress of weight distribution at moment, is the optimal adjustment direction of the said weight distribution value, is the hyperbolic tangent function, is the screening score at moment, is the time factor.
[0121] The said tensor flow adaptive screening algorithm further includes: an adaptive threshold formula and an adversarial verification formula, where the said adaptive threshold formula is:
[0122]
[0123] In the formula, is the screening adaptive threshold at moment, is the screening degree, is the historical screening score, is the number of the breeding programs at time is the influence of the number of the breeding programs at time on the screening result, is the time factor.
[0124] The adversarial verification formula is:
[0125]
[0126] In the formula, is the actual value of the weight assignment value, is the random noise, is the mathematical expectation, is the discriminator in the tensor flow adaptive screening algorithm, is the generator in the tensor flow adaptive screening algorithm, is the distribution probability of the actual value, is the distribution probability of the random noise, is the discriminator to identify the probability of the actual value, is according to the generator and the random noise to generate the predicted value, is the discriminator to identify the probability of the predicted value, is to the discriminator the value obtained by taking the logarithm of the probability that the predicted value cannot be identified, is to the discriminator the value obtained by taking the logarithm of the probability that the predicted value is identified, is to measure the discriminator in the tensor flow adaptive screening algorithm to distinguish the ability of the actual value and the predicted value of the weight assignment value.
[0127] Specifically, the screening score is obtained by specific weighted calculation of the accuracy score, the coverage score and the real-time score. This calculation method reflects the difference in the degree of emphasis on different dimension scores, emphasizes the importance of the coverage score, and at the same time makes a certain degree of reverse consideration of the real-time score.
[0128] In detail, the screening score provides a clear and quantitative basis for breeding decisions. Users can, according to the level of the screening score, clarify which programs are more worthy of investing resources to implement and which programs need to be further optimized and improved, so as to reasonably allocate human, material and financial resources and improve the overall success rate of the breeding project.
[0129] Generally speaking, as the amount of data increases and computing power improves, the parameter calculations in the formula will be more accurate, thus further enhancing the accuracy and reliability of the screening scores. Meanwhile, according to different breeding scenarios and requirements, the coefficients in the formula may be dynamically adjusted to better adapt to diverse actual situations and provide more efficient and intelligent decision-making support for forest tree genetic breeding.
[0130] Specifically, the optimization formula optimizes the weight matching based on the screening scores, ensuring that the screening process favors those solutions with better comprehensive performance in terms of accuracy, coverage, and real-time performance. By continuously adjusting the weights, the adaptive screening formula can better select solutions that meet the current breeding requirements, avoid deviations in screening results caused by unreasonable weights, and improve the quality and reliability of the screening results.
[0131] For example, in the initial stage of a breeding project, when the amount of data is relatively small, it may rely more on the credibility and source influence of weight allocation. At this time, parameters such as the credibility of weight allocation and the source influence of weight allocation in the formula will play a more crucial role in weight reallocation. As time progresses and the data gradually becomes richer, the dispersion and attenuation degree of weight allocation may become important factors affecting weight adjustment, and the formula can correspondingly change the action intensity of each parameter dynamically to achieve precise weight adjustment.
[0132] In detail, For the attenuation degree of weight allocation at time, it reflects the attenuation speed of the weight allocation value over time as time goes by.
[0133] For example, if the experience of a breeding project shows that the weights determined early are gradually less important in the later stage, this attenuation trend will be controlled.
[0134] In detail, For the attenuation of the weight allocation value, it clarifies the specific numerical value of the attenuation amount of the weight in the time dimension.
[0135] In detail, For the credibility of weight allocation at time, it can be understood as a measure of the reliability of the current weight allocation method. For example, the weight allocation based on a more authoritative data source or a more mature algorithm has a higher credibility.
[0136] In detail, For the source influence of weight allocation, it reflects the degree of influence of the source of weight allocation on the final weight matching degree. Different data sources or decision-making bases may contribute differently to the rationality of the weights.
[0137] In detail, For The weight distribution dispersion factor at the moment is used to adjust the influence of the discrete degree of weight distribution on the weight matching degree. If you want the weight distribution to be more concentrated or dispersed, you can adjust to achieve.
[0138] In detail, for The discrete degree of weight distribution at a certain moment measures the dispersion of weight distribution values in different dimensions or factors.
[0139] For example, if you want the weights of some key factors to be relatively concentrated, while the weights of other factors to be relatively dispersed, This will reflect the dispersed state.
[0140] In detail, for The weight distribution of time, for different numbers The weight distribution value is The specific allocation method of the moments is related to the screening score This means that different screening scores will affect the size of the weight assignment value.
[0141] In detail, for The weight distribution value at the moment is the specific weight value. The total number of values assigned to the weights, is the time factor, The number of values assigned to the weights, these parameters together determine the The weight distribution values of different numbers at different times.
[0142] In detail, It is the time factor, which emphasizes that the entire weight redistribution process changes dynamically over time, and the parameter values at different times are different, thus realizing dynamic adjustment of weights.
[0143] In detail, is a hyperbolic tangent function, which performs a nonlinear transformation on the weighted summation result to make the weight matching degree Fluctuating within a certain range enhances the adaptability and ability to distinguish different situations and avoids excessive extreme weight adjustments.
[0144] Generally speaking, with the continuous improvement of sensor technology and data acquisition systems, more accurate and high-frequency time series data can be obtained, which will make the calculation of parameters related to the time factor in the formula more accurate, thus achieving more refined dynamic weight adjustment. At the same time, combined with big data analysis and artificial intelligence technologies, this formula is expected to further optimize the setting and calculation method of its own parameters. For example, through machine learning algorithms, it can automatically learn the optimal weight allocation patterns under different breeding scenarios and dynamically adjust the coefficients and parameters in the formula to better adapt to the complex and changeable breeding environment. In addition, with the in-depth development of interdisciplinary research, this formula may be further integrated with knowledge in fields such as genetics and ecology, considering influencing factors of weight allocation from more dimensions, and providing more comprehensive and efficient screening support for forest tree genetic breeding.
[0145] Specifically, the adaptive threshold formula constructs a dynamically flexible screening threshold setting mechanism, which takes into account the time factor and incorporates key elements such as the screening degree , historical screening scores and the number of breeding programs at a given time , changing the limitation of traditional fixed thresholds that cannot adapt to the dynamic changes in data during the screening process.
[0146] Specifically, by continuously adjusting the screening adaptive threshold , the TensorFlow adaptive screening algorithm can accurately screen breeding programs at different stages and with different data scales, improving the scientificity and accuracy of the screening process and providing a more refined control means for screening out breeding programs that meet the requirements.
[0147] Specifically, the adversarial verification formula continuously adjusts the parameters of the discriminator and the generator to maximize the value of , thereby enhancing the discrimination ability of the discriminator and further optimizing the entire TensorFlow adaptive screening algorithm.
[0148] For example, when the discriminator has difficulty distinguishing between actual values and predicted values, the generator can be adjusted to generate more deceptive predicted values, while improving the discriminator to enhance its recognition ability. In this process of confrontation and optimization, the performance of the algorithm is improved.
[0149] Generally speaking, the adversarial verification formula simulates the discriminator and the generator The confrontation process between them, by comprehensively considering the actual value and the predicted value generated by random noise, breaks through the limitations of traditional single evaluation methods, makes the evaluation of algorithm performance more comprehensive and in-depth, helps to improve the accuracy and reliability of the TensorFlow adaptive screening algorithm when dealing with weight allocation values, and lays a foundation for screening out more practical breeding demand-compliant solutions.
[0150] The updated adaptive screening module 107 is used to synchronize the updated adaptive screening formula into the TensorFlow adaptive screening algorithm to obtain an updated TensorFlow adaptive screening algorithm, and screen the breeding plan according to the new requirements obtained next time and the updated TensorFlow adaptive screening algorithm to obtain the target breeding plan for the new requirements.
[0151] The present invention generates a breeding plan by collecting and analyzing the genetic breeding data of forest trees, and then screens the breeding plan to obtain a plan that meets the user's needs; during the screening process, the TensorFlow adaptive screening algorithm is used to determine whether the data meets the requirements, and the data that does not meet the requirements is promptly discarded. At the same time, importance evaluations are carried out for different characteristics, focusing on the characteristics that have a significant impact on breeding results, avoiding screening biases caused by ignoring important characteristics or being interfered by irrelevant characteristics, and then precisely identifying the key information closely related to breeding from many genetic characteristics, greatly improving the accuracy of the screening results; in addition, a modular architecture design is adopted, and this design pattern is conducive to function expansion, enabling the present invention to process a larger amount of data, quickly collect and integrate the forest tree genetic information from different regions, and efficiently process a large amount of real-time data in a short time, effectively improving the screening efficiency.
[0152] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0153] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.
[0154] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, and is a theory, method, technology and application system that perceives needs, acquires knowledge and uses knowledge to obtain the best results.
[0155] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Terms such as first, second, etc. are used to denote names and do not denote any particular order.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A forest tree genetic breeding screening system based on big data, characterized in that, The system includes an information collection module, a breeding plan generation module, a data processing module, an adaptive screening module, an evaluation acquisition module, an optimization module, and an updated adaptive screening module, where: The information collection module is used to obtain the genetic breeding data of forest trees; The breeding plan generation module is used to perform a correlation analysis on the genetic breeding data to obtain the breeding plan of the forest trees; The data processing module is used to perform feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan; The adaptive screening module is used to obtain the breeding requirements of the user, and based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm, where the adaptive screening formula is: ; Wherein, is the screening value of the dynamic annotation, is the quantity of the breeding requirement, is the quantity of the dynamic annotation, is the correlation between the breeding requirement and the dynamic annotation, is the maximum possible entropy of the joint distribution of the breeding requirement and the dynamic annotation, is the breeding requirement, is the dynamic annotation; Based on the screening value and the dynamic annotation, screen the breeding plan to obtain the available breeding plan that meets the breeding requirements; Based on the dual attention mechanism, calculate the available breeding plan to obtain the eigenvalue of the available breeding plan, where the dual attention mechanism is: ; In the formula, is the eigenvalue of the available breeding scheme, is the Sigmoid activation function, is the spatial feature weight, is the channel weight, are all the features of the available breeding scheme, are the significant features of the available breeding scheme, is the available breeding scheme; Select the available breeding plan with the highest eigenvalue as the optimal plan corresponding to the breeding requirements; The evaluation acquisition module is used to obtain the score of the user for the optimal plan; The optimization module is used to reallocate the weights of the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain the updated adaptive screening formula; The updated adaptive screening module is used to synchronize the updated adaptive screening formula to the TensorFlow adaptive screening algorithm to obtain the updated TensorFlow adaptive screening algorithm, and screen the breeding plan according to the new requirements obtained next time and the updated TensorFlow adaptive screening algorithm to obtain the target breeding plan that meets the new requirements.
2. The a big data-based screening system for forest tree genetic breeding according to claim 1, characterized in that When the information collection module executes the operation of obtaining the genetic breeding data of forest trees, it specifically is used for: Collect the phenotypic data, genotype data, environmental factor data, and historical breeding records related to the genetic breeding of forest trees.
3. The a big data-based screening system for forest tree genetic breeding according to claim 1, wherein When the breeding plan generation module executes the operation of performing a correlation analysis on the genetic breeding data to obtain the breeding plan of the forest trees, it specifically is used for: Clean the noise of the genetic breeding data to obtain the standard data of the genetic breeding data; Analyze the genetic relationship of the standard data to obtain the trait heritability evaluation map of the standard data; Perform a correlation analysis on the trait heritability evaluation map to obtain the breeding plan of the forest trees.
4. The a big data-based screening system for forest tree genetic breeding according to claim 1, wherein When the data processing module executes the operation of performing feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan, it specifically is used for: Classify the genetic characteristics of the forest trees to obtain the feature tags of the forest trees; Extract features of the breeding plan based on the ETL tool to obtain the feature data of the breeding plan; Annotate the feature data based on the feature tags to obtain the dynamic annotation of the breeding plan.
5. The a big data-based screening system for forest tree genetic breeding according to claim 1, characterized in that When the evaluation acquisition module executes the operation of obtaining the score of the user for the optimal plan, it specifically is used for: The user scores the optimal solution to obtain the score of the optimal solution, where the score of the optimal solution includes: the accuracy score, the coverage score, and the timeliness score of the optimal solution.
6. The a big data-based screening system for forest tree genetic breeding according to claim 1, characterized in that When the optimization module performs weight reallocation on the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula, it is specifically used for: Analyze the optimal solution based on the score and the reward function to obtain the screening score of the optimal solution, where the reward function is: ; In the formula, is the screening score, is the accuracy score of the optimal solution, is the coverage score of the optimal solution, is the real-time score of the optimal solution; Perform weight reallocation on the adaptive screening formula based on the optimization formula and the screening score to obtain the weight matching degree of the adaptive screening formula, where the optimization formula is: ; Wherein, is the weight matching degree after redistribution, is the attenuation degree of weight distribution at the is the attenuation of the weight distribution value, is the reliability of weight distribution at the is the influence of the weight distribution source, is the dispersion factor of weight distribution at the is the dispersion degree of weight distribution at the is the weight distribution at the is the weight distribution value at the is the total number of weight distribution values, is the time factor, is the number of the said weight distribution value, is the said screening score, is the hyperbolic tangent function; Perform weight reallocation on the adaptive screening formula based on the weight assignment value with the highest weight matching degree to obtain an updated adaptive screening formula.
7. The a big data-based screening system for forest tree genetic breeding according to claim 6, wherein The weight assignment value at a moment is calculated as follows: ; In the formula, is the weight assignment value at the initial weight assignment value, is the weight assignment progress at the optimal adjustment direction of the weight assignment value, is the hyperbolic tangent function, is the screening score at the time factor.
8. The a big data-based forest genetic breeding screening system according to claim 1, characterized in that The TensorFlow adaptive screening algorithm further includes: an adaptive threshold formula and an adversarial verification formula, where the adaptive threshold formula is: ; In the formula, is the screening adaptive threshold at the moment, is the screening degree, is the historical screening score, is the number of the breeding programs at the moment, is the influence of the number of the breeding programs at the moment on the screening result, is the time factor.
9. The a big data-based screening system for forest tree genetic breeding according to claim 8, characterized in that, The adversarial verification formula is: ; In the formula, is the actual value of the weight assignment value, is the random noise, is the mathematical expectation, is the discriminator in the tensor flow adaptive screening algorithm, is the generator in the tensor flow adaptive screening algorithm, is the distribution probability of the actual value, is the distribution probability of the random noise, is the discriminator identifies the probability of the actual value, is based on the generator and the random noise generates the predicted value, is the discriminator identifies the probability of the predicted value, is the logarithm value of the probability that the discriminator fails to identify the predicted value, is the logarithm value of the probability that the discriminator identifies the predicted value, is to measure the ability of the discriminator in the tensor flow adaptive screening algorithm to distinguish the actual value and the predicted value of the weight assignment value.
Citation Information
Patent Citations
Local breeding pig improvement system and method based on big data
CN117853258A
Breeding method for peanut with excellent character
CN118127225A