Forest tree genetic breeding screening system based on big data
By designing a forest genetic breeding screening system based on big data, using tensor flow adaptive screening algorithm and dual attention mechanism, the breeding scheme is screened and optimized, and the existing system is difficult to effectively capture genetic laws when facing noise interference and data distribution drift, and efficient and accurate breeding scheme screening is achieved.
Patent Information
- Application Number
- CN202510607215.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing forest genetic breeding screening system based on machine learning algorithms and deep learning models is difficult to effectively capture the true genetic law when facing noise interference, labeling errors or sample imbalances, and data distribution drift leads to prediction failure, resulting in inefficient screening efficiency and deviation in result.
A forest genetic breeding screening system based on big data was designed, using the status information collection module, breeding scheme generation module, data processing module, adaptive screening module, evaluation and acquisition module, optimization module and update adaptive screening module. Through the tensor flow adaptive screening algorithm and dual attention mechanism, the breeding scheme is screened and optimized to improve the accuracy and efficiency of screening results.
By using tensor flow adaptive screening algorithm and dual attention mechanism, genetic laws can be accurately captured in high-dimensional sparse data, improving the accuracy and efficiency of breeding scheme screening, reducing screening bias, and improving the practical application efficiency of the system.
Smart Images

Figure CN120124988A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a forest tree genetic breeding screening system based on big data. Background Art
[0002] For a forest tree genetic breeding screening system based on big data, the quality of its performance mainly depends on the accuracy and efficiency of screening. Currently, the mainstream screening schemes based on machine learning algorithms and deep learning models have exposed significant defects. In machine learning algorithms, their dependence on data quality is too high. Once there are noise data or missing values, it is extremely easy to cause deviations in the results. While deep learning models face a series of thorny problems. On the one hand, their training process consumes extremely large amounts of computing resources and requires strong support from high-performance hardware. This not only results in high costs and significantly increases scientific research and production inputs, but also the model has poor interpretability and exhibits black-box characteristics. This black-box characteristic seriously affects the credibility of key decisions. On the other hand, deep learning models also bring great difficulties to the debugging work. When the model has problems, due to the inability to intuitively understand its internal mechanism, the debugging efficiency is low, and it is difficult to quickly locate and solve the problems.
[0003] The core problems existing in the existing systems include: when facing noise interference, annotation errors, or sample imbalance, etc., it is difficult to effectively capture the true genetic laws in high-dimensional sparse data; at the same time, data distribution drift will lead to prediction failure, thus requiring frequent online learning, but this process may introduce short-term noise; in addition, rare scenarios will result in low confidence in the screening results due to insufficient data, making it difficult for the system to balance screening efficiency and result accuracy when processing large-scale dynamic data. These problems ultimately lead to extremely low screening efficiency and significant deviations in the screening results when screening breeding schemes, greatly restricting the actual application effectiveness of the system. Summary of the Invention
[0004] The present invention provides a forest tree genetic breeding screening system based on big data, and its main purpose is to solve the problem of low screening efficiency and screening quality when screening breeding schemes.
[0005] To achieve the above object, a forest tree genetic breeding screening system based on big data provided by the present invention includes a status information collection module, a breeding scheme generation module, a data processing module, an adaptive screening module, an evaluation acquisition module, an optimization module, and an updated adaptive screening module, wherein:
[0006] The information collection module is used to obtain the genetic breeding data of forest trees;
[0007] The breeding scheme generation module is used to perform correlation analysis on the genetic breeding data to obtain the breeding scheme of the forest trees;
[0008] The data processing module is used to perform feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan;
[0009] The adaptive screening module is used to obtain the breeding requirements of the user, and based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm, screen the breeding plan to obtain the optimal plan corresponding to the breeding requirements;
[0010] The evaluation acquisition module is used to obtain the score of the user for the optimal plan;
[0011] The optimization module is used to reassign the weights of the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula;
[0012] The updated adaptive screening module is used to synchronize the updated adaptive screening formula to the TensorFlow adaptive screening algorithm to obtain an updated TensorFlow adaptive screening algorithm, and screen the breeding plan according to the new requirements obtained next time and the updated TensorFlow adaptive screening algorithm to obtain the target breeding plan for the new requirements.
[0013] In a preferred embodiment, when the information collection module executes to obtain the genetic breeding data of forest trees, it is specifically used for:
[0014] Collect phenotypic data, genotypic data, environmental factor data, and historical breeding records related to forest tree genetic breeding.
[0015] In a preferred embodiment, when the breeding plan generation module executes to perform a correlation analysis on the genetic breeding data to obtain the breeding plan of the forest trees, it is specifically used for:
[0016] Clean the noise points of the genetic breeding data to obtain the standard data of the genetic breeding data;
[0017] Analyze the genetic relationship of the standard data to obtain the trait heritability evaluation map of the standard data;
[0018] Perform a correlation analysis on the trait heritability evaluation map to obtain the breeding plan of the forest trees.
[0019] In a preferred embodiment, when the data processing module executes to perform feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan, it is specifically used for:
[0020] Classify the genetic characteristics of the forest trees to obtain the feature labels of the forest trees;
[0021] Extract the features of the breeding plan based on the ETL tool to obtain the feature data of the breeding plan;
[0022] Annotate the feature data based on the feature tags to obtain the dynamic annotation of the breeding plan.
[0023] In a preferred embodiment, when the adaptive screening module performs adaptive screening on the breeding plan based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm to obtain the optimal plan corresponding to the breeding requirements, it is specifically used for:
[0024] Perform initial screening on the dynamic annotation based on the adaptive screening formula and the breeding requirements to obtain the screening value of the dynamic annotation, where the adaptive screening formula is:
[0025]
[0026] In the formula, is the screening value of the dynamic annotation, is the quantity of the breeding requirements, is the quantity of the dynamic annotation, is the correlation between the breeding requirements and the dynamic annotation, is the maximum possible entropy of the joint distribution of the breeding requirements and the dynamic annotation, is the breeding requirement, is the dynamic annotation.
[0027] Screen the breeding plan based on the screening value and the dynamic annotation to obtain the available breeding plans for the breeding requirements;
[0028] Calculate the feature values of the available breeding plans based on the dual attention mechanism, where the dual attention mechanism is:
[0029]
[0030] In the formula, is the feature value of the available breeding plan, is the Sigmoid activation function, is the spatial feature weight, is the channel weight, is all the features of the available breeding plan, is the significant feature of the available breeding plan, is the available breeding plan;
[0031] Select the available breeding plan with the highest eigenvalue as the optimal plan corresponding to the breeding requirement.
[0032] In a preferred embodiment, when the evaluation acquisition module executes to obtain the score of the user for the optimal plan, it is specifically configured to:
[0033] The user scores the optimal plan to obtain the score of the optimal plan, where the score of the optimal plan includes: the accuracy score, the coverage score, and the timeliness score of the optimal plan.
[0034] In a preferred embodiment, when the optimization module executes to re - allocate the weights of the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula, it is specifically configured to:
[0035] Analyze the optimal plan based on the score and the reward function to obtain the screening score of the optimal plan, where the reward function is:
[0036]
[0037] In the formula, is the screening score, is the accuracy score of the optimal plan, is the coverage score of the optimal plan, is the timeliness score of the optimal plan;
[0038] Re - allocate the weights of the adaptive screening formula based on the optimization formula and the screening score to obtain the weight matching degree of the adaptive screening formula, where the optimization formula is:
[0039]
[0040] In the formula, is the re - allocated weight matching degree, is the attenuation degree of the weight assignment at time is the weight assignment value attenuation, is the credibility of the weight assignment at time is the influence of the weight assignment source, is the weight assignment dispersion factor at time is the dispersion degree of the weight assignment at time is the weight assignment at time is The weight assignment value at a moment is the total number of weight assignment values is the time factor is the number of the weight assignment value is the screening score is the hyperbolic tangent function
[0041] Based on the weight assignment value with the highest weight matching degree, re - assign the weights to the adaptive screening formula to obtain the updated adaptive screening formula
[0042] In a preferred embodiment, the weight assignment value at a moment has the following calculation formula
[0043]
[0044] In the formula is the weight assignment value at a moment is the initial weight assignment value is the weight assignment progress at a moment is the optimal adjustment direction of the weight assignment value is the hyperbolic tangent function is the screening score at a moment is the time factor
[0045] In a preferred embodiment, the tensor - flow adaptive screening algorithm further includes: an adaptive threshold formula and an adversarial verification formula. Among them, the adaptive threshold formula is
[0046]
[0047] In the formula is the screening adaptive threshold at a moment is the screening degree is the historical screening score is the number of breeding programs at a moment is the influence of the number of breeding programs at a moment on the screening result is the time factor
[0048] In a preferred embodiment, the adversarial verification formula is
[0049]
[0050] In the formula is the actual value of the weight assignment value, is random noise, is the mathematical expectation, is the discriminator in the tensor flow adaptive screening algorithm, is the generator in the tensor flow adaptive screening algorithm, is the distribution probability of the actual value, is the distribution probability of the random noise, is the discriminator the probability of identifying the actual value, is based on the generator and the random noise the predicted value generated, is the discriminator the probability of identifying the predicted value, is for the discriminator the value of the logarithm of the probability that the discriminator fails to identify the predicted value, is for the discriminator the value of the logarithm of the probability that the discriminator identifies the predicted value, is to measure the ability of the discriminator in the tensor flow adaptive screening algorithm to distinguish between the actual value and the predicted value of the weight assignment value.
[0051] The present invention generates a breeding plan by collecting and analyzing the genetic breeding data of forest trees, and then screens the breeding plan to obtain a plan that meets the user's needs; during the screening process, by using the tensor flow adaptive screening algorithm, it is determined whether the data meets the requirements, and the data that does not meet the requirements is discarded in a timely manner. At the same time, importance assessment is carried out for different features, focusing on the features that have a significant impact on the breeding results, avoiding screening biases caused by ignoring important features or being interfered by irrelevant features, and then accurately identifying the key information closely related to breeding from many genetic features, greatly improving the accuracy of the screening results; in addition, a modular architecture design is adopted, and this design pattern is conducive to function expansion, enabling the present invention to process a larger amount of data, being able to quickly collect and integrate the genetic information of forest trees from different regions, and being able to efficiently process a large amount of real-time data in a short time, effectively improving the screening efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is the system architecture diagram of the forest tree genetic breeding screening system based on big data provided by an embodiment of the present invention;
[0053] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. "Plural" generally includes at least two.
[0056] Depending on the context, the words "if" or "when" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0057] In addition, the sequence of steps in the following method embodiments is only an example and is not strictly limited.
[0058] In fact, the server device deployed by the forest tree genetic screening system based on big data may be composed of one or more devices. The above-mentioned forest tree genetic breeding screening system based on big data can be implemented as: a service instance, a virtual machine, or a hardware device. For example, the forest tree genetic breeding screening system based on big data can be implemented as a service instance deployed on one or more devices in a cloud node. Briefly speaking, the forest tree genetic breeding screening system based on big data can be understood as a software deployed on a cloud node for providing the forest tree genetic screening system based on big data for each client. Or, the forest tree genetic breeding screening system based on big data can also be implemented as a virtual machine deployed on one or more devices in a cloud node. An application software for managing each client is installed in the virtual machine. Or, the forest tree genetic screening system based on big data can also be implemented as a server composed of many identical or different types of hardware devices, and one or more hardware devices are set to provide the forest tree genetic breeding screening system based on big data for each client.
[0059] In terms of implementation form, the big data-based forest tree genetic screening system and the user terminal adapt to each other. That is, if the big data-based forest tree genetic screening system is an application installed on the cloud service platform, then the user terminal is the client that establishes a communication connection with this application; or if the big data-based forest tree genetic breeding screening system is implemented as a website, then the user terminal is implemented as a web page; or if the big data-based forest tree genetic breeding screening system is implemented as a cloud service platform, then the user terminal is implemented as a small program in an instant messaging application.
[0060] As Figure 1 shown, it is the system architecture diagram of the big data-based forest tree genetic breeding screening system provided by an embodiment of the present invention.
[0061] The big data-based forest tree genetic screening system 100 of the present invention can be set in a cloud server. In terms of implementation form, it can be one or more service devices, or can be an application installed on the cloud (such as the server of a mobile service operator, a server cluster, etc.), or can also be developed as a website. According to the functions to be achieved, the big data-based forest tree genetic screening system 100 can include a status information collection module 101, a breeding plan generation module 102, a data processing module 103, an adaptive screening module 104, an evaluation acquisition module 105, an optimization module 106, and an updated adaptive screening module 107. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by the processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0062] In the embodiment of the present invention, in the big data-based forest tree genetic screening system, each of the above modules can be independently implemented and called by other modules. Here, the call can be understood as that a certain module can be connected to multiple modules of another type and provide corresponding services for the multiple modules it is connected to. For example, the sharing and evaluation module can call the same information collection module to obtain the information collected by this information collection module. Based on the above characteristics, in the big data-based forest tree genetic breeding screening system provided by the embodiment of the present invention, without modifying the program code, the applicable range of the big data-based forest tree genetic screening system architecture can be adjusted by adding modules and directly calling them, so as to achieve cluster-level horizontal expansion, so as to achieve the purpose of quickly and flexibly expanding the big data-based forest tree genetic breeding screening system. In practical applications, the above modules can be set in the same device or different devices, or can also be set in virtual devices, such as service instances in a cloud server.
[0063] Next, in combination with specific embodiments, the respective components and specific working processes of the big data-based forest tree genetic screening system will be described separately:
[0064] The information collection module 101 is used to obtain the genetic breeding data of forest trees.
[0065] In an embodiment of the present invention, when the information collection module executes obtaining the genetic breeding data of forest trees, it is specifically used for:
[0066] Collect phenotypic data, genotype data, environmental factor data, and historical breeding records related to forest tree genetic breeding.
[0067] The breeding plan generation module 102 is used to perform a correlation analysis on the genetic breeding data to obtain the breeding plan of the forest trees.
[0068] In an embodiment of the present invention, when the breeding plan generation module executes performing a correlation analysis on the genetic breeding data to obtain the breeding plan of the forest trees, it is specifically used for:
[0069] Clean the noise points in the genetic breeding data to obtain the standard data of the genetic breeding data;
[0070] Analyze the genetic relationship of the standard data to obtain the trait heritability evaluation map of the standard data;
[0071] Perform a correlation analysis on the trait heritability evaluation map to obtain the breeding plan of the forest trees.
[0072] Specifically, for the collected genetic breeding data, identify and remove the possible noise points in the dataset. In this process, through multiple means such as setting a reasonable data threshold range, using an outlier detection model, and manual review, comprehensively check the noise data, including but not limited to outliers caused by measurement errors, error values that occur during data transmission, etc. After cleaning, obtain the standard data of the genetic breeding data.
[0073] Specifically, based on the cleaned standard data, deeply analyze the genetic relationship inside the data. Through the mining and analysis of genetic information such as gene loci and chromosome segments, combined with biological genetic theory, accurately evaluate the heritability of various traits, and present the evaluation results in a visual way to generate a detailed trait heritability evaluation map.
[0074] Generally speaking, the trait heritability evaluation map clearly shows the strength of heritability, genetic patterns of different traits, and the potential associations between various traits, providing an intuitive basis for deeply understanding the genetic laws.
[0075] Specifically, a comprehensive correlation analysis is carried out on the trait heritability evaluation map, comprehensively considering factors such as the heritability of different traits, the degree of correlation between them, and the degree of fit with the target breeding characteristics. Weight distribution and combination calculation are performed on each factor. Finally, according to the analysis results, a breeding plan for the forest trees is formulated.
[0076] Generally speaking, the breeding plan covers key contents such as seed selection strategies, hybridization combination design, and suggestions for cultivating environment control, so as to achieve the goals of efficient and precise forest tree genetic breeding.
[0077] The data processing module 103 is used to perform feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan.
[0078] In the embodiment of the present invention, when the data processing module performs feature marking on the breeding plan to obtain the dynamic annotation of the breeding plan, it specifically is used for:
[0079] Classify the genetic characteristics of the forest trees to obtain the feature tags of the forest trees;
[0080] Based on the ETL tool, perform feature extraction on the breeding plan to obtain the feature data of the breeding plan;
[0081] Based on the feature tags, annotate the feature data to obtain the dynamic annotation of the breeding plan.
[0082] Specifically, using bioinformatics methods and classification algorithms, comprehensively sort out and classify the genetic characteristics of forest trees. Starting from multiple dimensions such as gene sequence characteristics and phenotypic characteristics, according to the preset classification criteria and feature attributes, accurately classify a large number of genetic characteristics. During the classification process, a unique and clearly directional identifier will be assigned to each class of genetic characteristics as the feature tag of the forest trees.
[0083] Generally speaking, the feature tags can concisely and accurately summarize the essential attributes of various genetic characteristics, providing a clear index and classification basis for subsequent data processing and analysis.
[0084] Specifically, through a powerful ETL tool, perform feature extraction operations on the breeding plan. First, accurately extract various data related to the plan from various data sources involved in the breeding plan, such as experimental records, literature materials, historical data, etc. Secondly, use the conversion rules and algorithms built into the ETL tool to perform a series of processing on the extracted data, such as format conversion, data cleaning, integration calculation, etc., to refine a data set that can accurately reflect the core elements and key characteristics of the breeding plan, and finally obtain the feature data of the breeding plan.
[0085] Generally speaking, the characteristic data highly condenses the key information of the breeding program, laying a solid foundation for subsequent annotation and analysis.
[0086] Specifically, with reference to the previously generated forest tree characteristic tags, a detailed annotation work is carried out on the extracted breeding program characteristic data. Each data item in the characteristic data is matched and associated with the corresponding characteristic tag, and a corresponding category tag and attribute description are assigned to it, thereby completing the initial annotation of the characteristic data. On this basis, considering the dynamic nature of the breeding process and the real-time change characteristics of the data, a dynamic annotation mechanism is introduced.
[0087] In detail, by continuously monitoring and analyzing newly generated data, changes in environmental factors, and feedback information during the implementation of the breeding program, the annotated characteristic data is updated and corrected in real time to generate a dynamic annotation that can reflect the real-time state and dynamic changes of the breeding program.
[0088] Generally speaking, dynamic annotation can provide timely, accurate, and comprehensive information support for breeding decisions, helping to achieve efficient and precise forest tree genetic breeding practices.
[0089] The adaptive screening module 104 is used to obtain the breeding requirements of the user, and based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm, the breeding program is screened to obtain the optimal program corresponding to the breeding requirements.
[0090] In the embodiment of the present invention, when the adaptive screening module executes screening the breeding program based on the breeding requirements, the dynamic annotation, and the adaptive screening formula in the TensorFlow adaptive screening algorithm to obtain the optimal program corresponding to the breeding requirements, it specifically is used for:
[0091] Based on the adaptive screening formula and the breeding requirements, an initial screening is performed on the dynamic annotation to obtain the screening value of the dynamic annotation, where the adaptive screening formula is:
[0092]
[0093] In the formula, is the screening value of the dynamic annotation, is the quantity of the breeding requirements, is the quantity of the dynamic annotation, is the correlation between the breeding requirements and the dynamic annotation, is the maximum possible entropy of the joint distribution of the breeding requirements and the dynamic annotation, is the breeding requirement, is the dynamic annotation.
[0094] Screen the breeding plan based on the screening value and the dynamic annotation to obtain an available breeding plan for the breeding requirement;
[0095] Calculate the available breeding plan based on the dual attention mechanism to obtain the eigenvalue of the available breeding plan, where the dual attention mechanism is:
[0096]
[0097] In the formula, is the eigenvalue of the available breeding plan, is the Sigmoid activation function, is the spatial feature weight, is the channel weight, is all the features of the available breeding plan, is the significant feature of the available breeding plan, is the available breeding plan;
[0098] Select the available breeding plan with the highest eigenvalue as the optimal plan corresponding to the breeding requirement.
[0099] Specifically, from a large number of dynamic annotations, accurately screen out the data highly relevant to the breeding requirement according to the breeding requirement, provide data support for breeding decision-making, and at the same time enable breeders to judge which dynamic annotations are the most critical for achieving the breeding requirement based on the screening value, so as to reasonably arrange breeding resources and formulate more effective breeding strategies.
[0100] For example, among many dynamic annotations reflecting the growth status, genetic characteristics, etc. of forest trees, quickly locate the data that meet specific breeding goals (such as pest resistance, fast growth, etc.) to improve the screening efficiency and accuracy.
[0101] Generally speaking, the larger the screening value, the more the dynamic annotation matches the breeding requirement, and the more worthy of attention and retention in the screening process.
[0102] Specifically, retain the breeding plans with a screening value greater than 0.6 as the available breeding plans for the breeding requirement.
[0103] Specifically, the adaptive screening formula systematically evaluates many potential available breeding plans, quickly screens out the most potential plans, and based on the calculated eigenvalues, the advantages and disadvantages of each breeding plan can be clearly judged, so as to reasonably allocate resources and preferentially select the plans with high eigenvalues for implementation, improving the breeding success rate and resource utilization efficiency.
[0104] For example, in a large number of different combinations of forest tree breeding programs, by calculating the eigenvalue, those programs that perform well in terms of spatial characteristics and channel characteristics, and whose significant characteristics match well with all characteristics, are accurately located, thus improving the screening efficiency and accuracy.
[0105] Generally speaking, the eigenvalue is the core output result of the entire formula, representing the overall performance level of the breeding program after comprehensively considering various characteristics. The higher the eigenvalue, the more the breeding program meets the user's needs.
[0106] The evaluation acquisition module 105 is used to obtain the score given by the user to the optimal solution.
[0107] In the embodiment of the present invention, when the evaluation acquisition module executes the operation of obtaining the score given by the user to the optimal solution, it specifically is used for:
[0108] The user gives a score to the optimal solution to obtain the score of the optimal solution, where the score of the optimal solution includes: the accuracy score, the coverage score, and the timeliness score of the optimal solution.
[0109] The optimization module 106 is used to reassign weights to the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula.
[0110] In the embodiment of the present invention, when the optimization module executes the operation of reassigning weights to the adaptive screening formula based on the score and the optimization formula in the TensorFlow adaptive screening algorithm to obtain an updated adaptive screening formula, it specifically is used for:
[0111] Analyze the optimal solution based on the score and the reward function to obtain the screening score of the optimal solution, where the reward function is:
[0112]
[0113] In the formula, is the screening score, is the accuracy score of the optimal solution, is the coverage score of the optimal solution, is the timeliness score of the optimal solution;
[0114] Reassign weights to the adaptive screening formula based on the optimization formula and the screening score to obtain the weight matching degree of the adaptive screening formula, where the optimization formula is:
[0115]
[0116] In the formula, is the weight matching degree after redistribution, is the attenuation degree of weight distribution at the moment, is the attenuation of the weight distribution value, is the credibility of weight distribution at the moment, is the influence of the weight distribution source, is the dispersion factor of weight distribution at the moment, is the dispersion degree of weight distribution at the moment, is the weight distribution at the moment, is the weight distribution value at the moment, is the total number of weight distribution values, is the time factor, is the number of the said weight distribution value, is the said screening score, is the hyperbolic tangent function;
[0117] Based on the weight distribution value with the highest weight matching degree, the weight of the said adaptive screening formula is redistributed to obtain an updated adaptive screening formula.
[0118] The said weight distribution value at the moment has the following calculation formula:
[0119]
[0120] In the formula, is the said weight distribution value at the moment, is the initial weight distribution value, is the progress of weight distribution at the moment, is the optimal adjustment direction of the said weight distribution value, is the hyperbolic tangent function, is the screening score at the moment, is the time factor.
[0121] The said tensor flow adaptive screening algorithm further includes: an adaptive threshold formula and an adversarial verification formula, where the adaptive threshold formula is:
[0122]
[0123] In the formula, is the screening adaptive threshold at the moment, is the screening degree, is the historical screening score, is the number of the breeding programs at time is the influence of the number of the breeding programs at time on the screening result, is the time factor.
[0124] The adversarial verification formula is:
[0125]
[0126] In the formula, is the actual value of the weight assignment value, is the random noise, is the mathematical expectation, is the discriminator in the tensor flow adaptive screening algorithm, is the generator in the tensor flow adaptive screening algorithm, is the distribution probability of the actual value, is the distribution probability of the random noise, is the discriminator to identify the probability of the actual value, is according to the generator and the random noise to generate the predicted value, is the discriminator to identify the probability of the predicted value, is to the discriminator the value of taking the logarithm of the probability that the predicted value cannot be identified, is to the discriminator the value of taking the logarithm of the probability that the predicted value is identified, is to measure the discriminator in the tensor flow adaptive screening algorithm the ability to distinguish the actual value and the predicted value of the weight assignment value.
[0127] Specifically, the screening score is obtained by specific weighted calculation of the accuracy score, the coverage score and the real-time score. This calculation method reflects the difference in the degree of emphasis on different dimension scores, emphasizes the importance of the coverage score, and at the same time makes a certain degree of reverse consideration of the real-time score.
[0128] In detail, the screening score provides a clear and quantitative basis for breeding decisions. Users can, according to the level of the screening score, clarify which programs are more worthy of investing resources for implementation and which programs need to be further optimized and improved, so as to reasonably allocate human, material and financial resources and improve the overall success rate of the breeding project.
[0129] Generally speaking, with the increase in data volume and the improvement of computing power, the parameter calculation in the formula will be more accurate, thus further improving the accuracy and reliability of the screening score. At the same time, according to different breeding scenarios and requirements, the coefficients in the formula may be dynamically adjusted to better adapt to diverse actual situations and provide more efficient and intelligent decision-making support for forest tree genetic breeding.
[0130] Specifically, the optimization formula optimizes the weight matching based on the screening score to ensure that the screening process is more inclined to those solutions with better comprehensive performance in terms of accuracy, coverage, and real-time performance. By continuously adjusting the weights, the adaptive screening formula can better screen out solutions that meet the current breeding requirements, avoid screening result deviations caused by unreasonable weights, and improve the quality and reliability of the screening results.
[0131] For example, in the initial stage of a breeding project, when the data volume is relatively small, it may be more dependent on the credibility and source influence of weight allocation. At this time, parameters such as the credibility of weight allocation and the source influence of weight allocation in the formula will play a more crucial role in weight reallocation. As time progresses and the data gradually becomes rich, the dispersion and attenuation degree of weight allocation may become important factors affecting weight adjustment, and the formula can dynamically change the action intensity of each parameter accordingly to achieve precise weight adjustment.
[0132] Specifically, is the attenuation degree of weight allocation at time, which reflects the attenuation speed of the weight allocation value over time as time goes by.
[0133] For example, if the experience of a breeding project shows that the weights determined early are gradually less important in the later stage, this attenuation trend will be controlled.
[0134] Specifically, is the attenuation of the weight allocation value, which clarifies the specific numerical value of the attenuation amount of the weight in the time dimension.
[0135] Specifically, is the credibility of weight allocation at time, which can be understood as a measure of the reliability of the current weight allocation method. For example, weight allocations based on more authoritative data sources or more mature algorithms have higher credibility.
[0136] Specifically, is the source influence of weight allocation, which reflects the influence degree of the source of weight allocation on the final weight matching degree. Different data sources or decision-making bases may contribute to the rationality of weights to different degrees.
[0137] Specifically, is The weight distribution dispersion factor at a moment is used to adjust the influence degree of the dispersion of weight distribution on the weight matching degree. If it is desired to make the weight distribution more concentrated or dispersed, it can be achieved by adjusting to realize.
[0138] Specifically, is the weight distribution dispersion degree at a moment, which measures the dispersion of weight distribution values in different dimensions or factors.
[0139] For example, if it is desired that the weights of some key factors are relatively concentrated while the weights of other factors are more dispersed, it will reflect this dispersion state.
[0140] Specifically, is the weight distribution at a moment, which is the specific distribution method of weight distribution values with different numbers at a moment. It is associated with the screening score at a moment, meaning that different screening scores will affect the size of this weight distribution value.
[0141] Specifically, is the weight distribution value at a moment, which is the specific weight value. is the total number of weight distribution values. is the time factor. is the number of the said weight distribution value. These parameters jointly determine the weight distribution value situation with different numbers at a moment.
[0142] Specifically, the time factor emphasizes that the entire weight redistribution process changes dynamically with time, and the parameter values at different moments are different, thus realizing the dynamic adjustment of weights.
[0143] Specifically, is the hyperbolic tangent function, which performs a non - linear transformation on the result of weighted summation, making the weight matching degree fluctuate within a certain range, enhancing the adaptability and discrimination ability for different situations and avoiding overly extreme weight adjustment.
[0144] Generally speaking, with the continuous improvement of sensor technology and data acquisition systems, more accurate and high-frequency time series data can be obtained, which will make the calculation of parameters related to the time factor in the formula more accurate, thus realizing a more refined dynamic adjustment of weights. At the same time, combined with big data analysis and artificial intelligence technologies, this formula is expected to further optimize the setting and calculation method of its own parameters. For example, through machine learning algorithms, it can automatically learn the optimal weight allocation patterns in different breeding scenarios and dynamically adjust the coefficients and parameters in the formula to better adapt to the complex and changeable breeding environment. In addition, with the in-depth development of interdisciplinary research, this formula may be further integrated with knowledge in fields such as genetics and ecology, considering influencing factors of weight allocation from more dimensions, and providing more comprehensive and efficient screening support for forest tree genetic breeding.
[0145] Specifically, the adaptive threshold formula constructs a dynamically flexible screening threshold setting mechanism, which takes the time factor into consideration and integrates key elements such as the screening degree , historical screening scores and the number of the breeding programs at the moment etc., changing the limitation that traditional fixed thresholds cannot adapt to the dynamic changes of data in the screening process.
[0146] Specifically, by continuously adjusting the screening adaptive threshold , the TensorFlow adaptive screening algorithm can accurately screen breeding programs at different stages and with different data scales, improving the scientificity and accuracy of the screening process and providing a more refined control means for screening out breeding programs that meet the requirements.
[0147] Specifically, the adversarial verification formula continuously adjusts the parameters of the discriminator and the generator to maximize the value of , thereby enhancing the discrimination ability of the discriminator and further optimizing the entire TensorFlow adaptive screening algorithm.
[0148] For example, when the discriminator has difficulty distinguishing between the actual value and the predicted value, more confusing predicted values can be generated by adjusting the generator , and at the same time, the discriminator is improved to enhance its recognition ability. In this process of confrontation and optimization, the performance of the algorithm is improved.
[0149] Generally speaking, the adversarial verification formula simulates the discriminator and the generator The confrontation process between them, by comprehensively considering the predicted value generated by the actual value and random noise, breaks through the limitations of traditional single evaluation methods, makes the evaluation of algorithm performance more comprehensive and in-depth, helps to improve the accuracy and reliability of the TensorFlow adaptive screening algorithm when dealing with weight assignment values, and lays a foundation for screening out more practical breeding requirement-compliant solutions.
[0150] The updated adaptive screening module 107 is used to synchronize the updated adaptive screening formula into the TensorFlow adaptive screening algorithm to obtain an updated TensorFlow adaptive screening algorithm, and screen the breeding plan according to the new requirements obtained next time and the updated TensorFlow adaptive screening algorithm to obtain the target breeding plan for the new requirements.
[0151] The present invention generates a breeding plan by collecting and analyzing forest genetic breeding data, and then screens the breeding plan to obtain a plan that meets the user's needs; during the screening process, by using the TensorFlow adaptive screening algorithm, it determines whether the data meets the requirements and promptly discards the data that does not meet the requirements. At the same time, it conducts importance evaluation for different features, focusing on the features that have a significant impact on breeding results, and avoiding screening biases caused by ignoring important features or being interfered by irrelevant features, thereby accurately identifying the key information closely related to breeding from many genetic features, greatly improving the accuracy of the screening results; in addition, it adopts a modular architecture design, which is conducive to function expansion, enables the present invention to process a larger amount of data, can quickly collect and integrate forest genetic information from different regions, and can efficiently process a large amount of real-time data in a short time, effectively improving the screening efficiency.
[0152] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0153] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
[0154] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive needs, acquire knowledge, and use knowledge to obtain the best results.
[0155] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or devices stated in the system claims may also be implemented by one unit or device through software or hardware. Terms such as first, second, etc. are used to denote names and do not denote any particular order.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A forest tree genetic breeding screening system based on big data, characterized in that: The system includes an information collection module, a breeding scheme generation module, a data processing module, an adaptive screening module, an evaluation acquisition module, an optimization module and an updated adaptive screening module, wherein: The information collection module is used to obtain genetic breeding data of forest trees; The breeding scheme generation module is used to perform correlation analysis on the genetic breeding data to obtain a breeding scheme for the forest tree; The data processing module is used to perform feature marking on the breeding scheme to obtain a dynamic labeling of the breeding scheme; The adaptive screening module is used to obtain the breeding needs of the user, screen the breeding schemes based on the breeding needs, the dynamic annotations and the adaptive screening formula in the tensor flow adaptive screening algorithm, and obtain the optimal scheme corresponding to the breeding needs; The evaluation acquisition module is used to obtain the user's rating of the optimal solution; The optimization module is used to redistribute the weight of the adaptive screening formula based on the score and the optimization formula in the tensor flow adaptive screening algorithm to obtain an updated adaptive screening formula; The updated adaptive screening module is used to synchronize the updated adaptive screening formula to the tensor flow adaptive screening algorithm to obtain the updated tensor flow adaptive screening algorithm, and screen the breeding scheme according to the new requirements obtained next time and the updated tensor flow adaptive screening algorithm to obtain the target breeding scheme of the new requirements.
2. A forest tree genetic breeding screening system based on big data as claimed in claim 1, characterized in that: When the information collection module is executed to obtain the genetic breeding data of the forest tree, it is specifically used to: Collect phenotypic data, genotypic data, environmental factor data and historical breeding records related to forest tree genetic breeding.
3. The forest tree genetic breeding screening system based on big data as claimed in claim 1, characterized in that: When the breeding scheme generation module performs a correlation analysis on the genetic breeding data to obtain the breeding scheme of the forest tree, it is specifically used to: Cleaning the genetic breeding data for noise to obtain standard data of the genetic breeding data; Performing genetic relationship analysis on the standard data to obtain a trait heritability evaluation map of the standard data; Association analysis is performed on the trait heritability evaluation map to obtain a breeding plan for the forest tree.
4. The forest tree genetic breeding screening system based on big data according to claim 1, characterized in that: When the data processing module performs feature marking on the breeding scheme to obtain a dynamic labeling of the breeding scheme, it is specifically used to: Classifying the genetic characteristics of the trees to obtain characteristic labels of the trees; Extracting features of the breeding scheme based on an ETL tool to obtain feature data of the breeding scheme; The characteristic data is annotated based on the characteristic tag to obtain a dynamic annotation of the breeding scheme.
5. The forest tree genetic breeding screening system based on big data as claimed in claim 1, characterized in that: When the adaptive screening module performs screening of the breeding scheme based on the breeding requirements, the dynamic labeling and the adaptive screening formula in the tensor flow adaptive screening algorithm to obtain the optimal scheme corresponding to the breeding requirements, it is specifically used to: The dynamic annotation is initially screened based on the adaptive screening formula and the breeding requirements to obtain the screening value of the dynamic annotation, wherein the adaptive screening formula is: In the formula, is the filter value of the dynamic annotation, is the number of breeding requirements, is the number of dynamic annotations, is the correlation between the breeding requirements and the dynamic marking, is the maximum possible entropy of the joint distribution of the breeding requirements and the dynamic annotation, For the breeding needs, marking the dynamic image; Screening the breeding scheme based on the screening value and the dynamic annotation to obtain an available breeding scheme for the breeding demand; The available breeding schemes are calculated based on a dual attention mechanism to obtain characteristic values of the available breeding schemes, wherein the dual attention mechanism is: In the formula, is the characteristic value of the available breeding scheme, is the Sigmoid activation function, is the spatial feature weight, is the channel weight, for all the characteristics of the available breeding schemes, is a distinctive feature of the breeding scheme that can be used, For the available breeding schemes; The available breeding scheme with the highest characteristic value is selected as the optimal scheme corresponding to the breeding requirement.
6. The forest tree genetic breeding screening system based on big data according to claim 1, characterized in that: When the evaluation acquisition module is executed to acquire the user's rating of the optimal solution, it is specifically used to: The user scores the optimal solution to obtain a score of the optimal solution, wherein the score of the optimal solution includes: an accuracy score, a coverage score, and a real-time score of the optimal solution.
7. The forest tree genetic breeding screening system based on big data according to claim 1, characterized in that: When the optimization module performs weight redistribution on the adaptive screening formula based on the score and the optimization formula in the tensor flow adaptive screening algorithm to obtain an updated adaptive screening formula, it is specifically used to: The optimal solution is analyzed based on the score and the reward function to obtain a screening score of the optimal solution, wherein the reward function is: In the formula, For the screening score, Score the accuracy of the optimal solution, Score the coverage of the optimal solution, Score the real-time degree of the optimal solution; The adaptive screening formula is weighted redistributed based on the optimization formula and the screening score to obtain the weight matching degree of the adaptive screening formula, wherein the optimization formula is: In the formula, is the weight matching degree after redistribution, for The degree of weight distribution attenuation at the moment, Assign value decay to the weights, for The weight distribution credibility of the moment, Assign source influences to the weights, for The weight distribution discreteness factor at the moment, for The discrete degree of weight distribution at each moment, for The weight distribution of time, for The weight distribution value at the moment, The total number of values assigned to the weights, is the time factor, a number to assign a value to said weight, is the screening score, is the hyperbolic tangent function; The adaptive screening formula is weighted redistributed based on the weight distribution value with the highest weight matching degree to obtain an updated adaptive screening formula.
8. The forest tree genetic breeding screening system based on big data according to claim 7, characterized in that: Said The weight distribution value at the moment The calculation formula is as follows: In the formula, for The weight distribution value at the time, Assign values to the initial weights, for The weight distribution progress at each moment, is the optimal adjustment direction of the weight distribution value, is the hyperbolic tangent function, for The screening score at the moment, is the time factor.
9. The forest tree genetic breeding screening system based on big data according to claim 1, characterized in that: The tensor flow adaptive screening algorithm also includes: an adaptive threshold formula and an adversarial verification formula, wherein the adaptive threshold formula is: In the formula, for The adaptive threshold for filtering at the moment, For the degree of screening, Filter scores for history, for The number of breeding programs at the time, for The influence of the number of breeding schemes at the time on the screening results, is the time factor.
10. A forest tree genetic breeding screening system based on big data as claimed in claim 9, characterized in that: The adversarial verification formula is: In the formula, is the actual value of the weight assignment value, is random noise, is the mathematical expectation, is the discriminator in the tensor flow adaptive screening algorithm, is the generator in the tensorflow adaptive filtering algorithm, is the distribution probability of the actual value, is the distribution probability of the random noise, For the discriminator The probability of identifying the actual value, According to the generator and the random noise The generated prediction values are For the discriminator The probability of identifying the predicted value, For the discriminator The logarithm of the probability of failing to identify the predicted value, For the discriminator The logarithm of the probability of identifying the predicted value is, To measure the discriminator in the tensor flow adaptive screening algorithm The ability to distinguish between actual and predicted values of weight assignments.
Citation Information
Patent Citations
Local breeding pig improvement system and method based on big data
CN117853258A
Breeding method for peanut with excellent character
CN118127225A
Inventory allocation optimization method and system based on genetic algorithm
CN119090408A
High-quality breeding group selection method, device and equipment and storage medium
CN119694406A
Goat selective breeding method and system based on genetic algorithm
CN119920307A