Sample data expansion method and device for whole genome selection, equipment and medium

By enhancing the signal of the actual genotype data of the target crop and expanding the sample data using linkage information, the problem of insufficient phenotypic data is solved, and the accuracy and reliability of crop trait prediction are improved.

CN121938451APending Publication Date: 2026-04-28SDIC SEED TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SDIC SEED TECHNOLOGY CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, the high cost of acquiring phenotypic data and its susceptibility to environmental interference result in insufficient data resources for genome-wide selection models, limiting their widespread application in the agricultural field.

Method used

By acquiring the actual genotype and phenotypic data of the target crop, the initial sample set is segmented using linkage feature information, and segments are replaced based on auxiliary gene sample data to generate mixed gene samples and expand the sample data.

Benefits of technology

It increases the scale and diversity of training samples, reduces overfitting during model training, and significantly improves the prediction efficiency and reliability of crop trait prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938451A_ABST
    Figure CN121938451A_ABST
Patent Text Reader

Abstract

The invention discloses a sample data expansion method and device for whole genome selection, equipment and a medium, and the sample data expansion method for whole genome selection comprises the steps: obtaining real genotype data of a target crop in a historical crop breeding process and phenotype data corresponding to the real genotype data, generating an initial sample set according to the real genotype data and the phenotype data; based on linkage feature information of the real genotype data, performing segment segmentation on initial gene sample data in the initial sample set, and performing segment replacement on segmented gene segments based on auxiliary gene sample data to generate a mixed gene sample; and training a pre-constructed crop prediction model based on the mixed gene sample and the initial sample set. According to the technical scheme, the prediction efficiency of crop character prediction and the reliability of the prediction result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a method, apparatus, device, and medium for expanding sample data through whole-genome selection. Background Technology

[0002] Under current technological conditions, acquiring phenotypic data is costly and easily affected by environmental and other natural factors; at the same time, datasets focused on constructing genome-wide selection models are relatively limited. This results in a severe shortage of data resources with complete genotypes and phenotypes that can be used for breeding training and prediction, thus restricting the training effectiveness and prediction accuracy of genome-wide selection models and limiting the widespread application of precision prediction models in agriculture. Therefore, there is an urgent need for a method to expand sample data to overcome the aforementioned technological bottlenecks. Summary of the Invention

[0003] This application provides a method, apparatus, device, and storage medium for expanding sample data of whole-genome selection, so as to improve the data dimensionality and diversity of gene samples, thereby improving the model performance of prediction models.

[0004] According to one aspect of this application, a method for augmenting sample data through whole-genome selection is provided, the method comprising:

[0005] Obtain the true genotype data of the target crop in the historical crop breeding process and the corresponding phenotypic data, and generate an initial sample set based on the true genotype data and the phenotypic data;

[0006] Based on the linkage feature information of the real genotype data, the initial gene sample data in the initial sample set is segmented, and the segmented gene segments are replaced based on the auxiliary gene sample data to generate a hybrid gene sample; wherein, the initial gene sample data and the auxiliary gene sample data belong to different sample data in the initial sample set; the hybrid gene sample is a new sample data that has been expanded.

[0007] The pre-constructed crop prediction model is trained based on the mixed gene sample and the initial sample set.

[0008] According to another aspect of this application, a sample data augmentation device for whole-genome selection is provided, the device comprising:

[0009] The data acquisition module is used to acquire the real genotype data of the target crop in the historical crop breeding process and the phenotypic data corresponding to the real genotype data, and generate an initial sample set based on the real genotype data and the phenotypic data.

[0010] The mixed sample module is used to segment the initial gene sample data in the initial sample set based on the linkage feature information of the real genotype data, and to replace the segmented gene segments based on the auxiliary gene sample data to generate a mixed gene sample; wherein the initial gene sample data and the auxiliary gene sample data belong to different sample data in the initial sample set; the mixed gene sample is a new sample data that has been expanded.

[0011] The training module is used to train a pre-built crop prediction model based on the mixed gene samples and the initial sample set.

[0012] According to another aspect of this application, an electronic device is provided, the electronic device comprising:

[0013] One or more processors;

[0014] Memory, used to store one or more programs;

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the whole-genome selection sample data augmentation methods provided in the embodiments of this application.

[0016] According to another aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements any of the whole-genome selection sample data augmentation methods provided in the embodiments of this application.

[0017] According to another aspect of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the whole-genome selection sample data augmentation methods provided in the embodiments of this application.

[0018] This application enhances the signal of historically collected real genotype data of target crops by introducing linkage information, thereby achieving efficient parsing of genotype data, generating new mixed gene samples, increasing the data scale and diversity of training samples, reducing overfitting problems in the model training process, improving the prediction accuracy of complex traits, overcoming the limitation of insufficient samples, comprehensively capturing nonlinear genetic signals, and significantly improving the prediction efficiency and reliability of crop trait prediction. Attached Figure Description

[0019] Figure 1 This is a flowchart of a method for augmenting sample data for whole-genome selection according to Embodiment 1 of this application;

[0020] Figure 2This is a flowchart of a sample data augmentation method for whole-genome selection according to Embodiment 2 of this application;

[0021] Figure 3 This is a schematic diagram of a sample data augmentation device for whole-genome selection according to Embodiment 3 of this application;

[0022] Figure 4 This is a schematic diagram of the electronic device used to implement the whole-genome selection sample data augmentation method of Embodiment 4 of this application. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0025] Furthermore, it should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of crop genetic data and other related data involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0026] Example 1

[0027] Figure 1 This is a flowchart of a whole-genome selection sample data amplification method according to Embodiment 1 of this application. This embodiment is applicable to the amplification of crop gene samples and can be executed by a whole-genome selection sample data amplification device. This whole-genome selection sample data amplification device can be implemented in hardware and / or software and can be configured in a computer device, such as a server.Figure 1 As shown, the method includes:

[0028] S110. Obtain the true genotype data of the target crop in the historical crop breeding process and the corresponding phenotypic data, and generate an initial sample set based on the true genotype data and phenotypic data.

[0029] The target crop can refer to the crop species for which trait prediction is to be performed.

[0030] In this embodiment of the invention, genotype data of the target crop can be collected during the crop breeding process. This data includes the genotype sequence information of the target crop, as well as phenotypic data such as plant height, ear length, and grain number. Based on the genotype and phenotypic data, gene samples are used to form an initial sample set. It should be noted that phenotypic data can serve as sample label data for genotype data, and the data format of the genotype data can be uniformly standardized to multi-site genotype coding.

[0031] Optionally, after obtaining the true genotype data of the target crop and the corresponding phenotypic data, the process also includes: comprehensive quality control of the collected gene samples, including detecting the missing rate, genotype errors, and anomalies of gene samples and loci, removing low-quality samples and abnormal loci, and ensuring data accuracy and completeness. Simultaneously, data cleaning is performed on outliers and erroneous records in the phenotypic data, removing duplicate samples to ensure data quality.

[0032] For example, missing values ​​in genotype data are uniformly labeled with the characters "NN" without complex interpolation or inference to ensure the authenticity of missing information and the objectivity of the analysis. Missing values ​​in phenotypic data are filled using mean imputation, that is, the sample mean of the phenotypic indicator is used to replace the missing value to ensure data integrity and reduce bias caused by missing values.

[0033] Optionally, after obtaining the true genotype data of the target crop and the corresponding phenotypic data, the method further includes: encoding and converting the data format of the true genotype data according to the combination of bases in the diallelic genes.

[0034] For example, for biseleural base combinations in genotype data, the following integer encoding method can be used for conversion: Homozygous reference alleles are used as the baseline, represented by the encoding value "0". For example, "AA" is a homozygous reference allele combination. Heterozygous genotypes containing one reference allele and one variant allele are encoded as "1", such as "AC" and "AG". Homozygous variant genotypes containing two identical variant alleles are encoded as "2", such as "TT". Finally, for missing, untested, or indeterminate genotypes (e.g., "NN"), a special encoding value "9" is assigned to distinguish them. Through this encoding, the original genotype string is converted into an integer numerical matrix, providing a unified and standardized data format for subsequent computational analysis.

[0035] S120. Based on the linkage feature information of real genotype data, the initial gene sample data in the initial sample set is segmented, and the segmented gene segments are replaced based on the auxiliary gene sample data to generate a mixed gene sample.

[0036] Among them, the initial gene sample data and the auxiliary gene sample data can refer to different sample data belonging to the initial sample set; the mixed gene sample can be new sample data that has been expanded.

[0037] Linkage characteristics can be used to measure the degree of association between two or more gene loci in the genome corresponding to gene sample data. Optionally, linkage characteristics can be represented by linkage disequilibrium indices.

[0038] It should be noted that gene sample data of a predetermined sample ratio can be extracted from the initial sample set in advance as auxiliary gene sample data. Optionally, the predetermined sample ratio can be adaptively set according to those skilled in the art.

[0039] S130. Based on the mixed gene samples and the initial sample set, the pre-constructed crop prediction model is trained.

[0040] Optionally, the model structure of the crop prediction model can be adapted to the needs of those skilled in the art. For example, the model structure of the crop prediction model can be determined based on the phenotypic characteristics (phenotypic data) of the target crop to be predicted, such as convolutional neural networks, deep embedding models, crop prediction converters, and lightweight gradient boosters.

[0041] Specifically, the gene sample data from the initial sample set and the expanded mixed gene samples can be used as model inputs to a pre-built crop prediction model to train the crop prediction model.

[0042] This application embodiment enhances the signal of historically collected real genotype data of target crops by introducing linkage information, thereby achieving efficient parsing of genotype data, generating new mixed gene samples, increasing the data scale and diversity of training samples, reducing overfitting problems in the model training process, improving the prediction accuracy of complex traits, overcoming the limitation of insufficient samples, comprehensively capturing nonlinear genetic signals, and significantly improving the prediction efficiency and reliability of crop trait prediction.

[0043] Example 2

[0044] Figure 2 This is a flowchart of a method for augmenting sample data for whole-genome selection according to Embodiment 2 of this application. Based on the technical solutions of the above embodiments, this embodiment further refines the process of "segmenting the initial gene sample data in the initial sample set based on the linkage feature information of the real genotype data, and replacing the segmented gene segments based on auxiliary gene sample data to generate a mixed gene sample." It should be noted that for parts not detailed in this embodiment, please refer to the relevant descriptions in other embodiments. Figure 2 As shown, the method includes:

[0045] S210. Obtain the true genotype data of the target crop in the historical crop breeding process and the corresponding phenotypic data, and generate an initial sample set based on the true genotype data and phenotypic data.

[0046] S220. Based on the linkage disequilibrium index, the initial gene sample data in the initial sample set is divided into linkage blocks to generate at least two target gene segments.

[0047] Among them, there are no overlapping segments between different gene segments.

[0048] Optionally, based on the linkage disequilibrium index, the initial gene sample data in the initial sample set is divided into linkage blocks to generate at least two target gene segments. This includes: randomly determining an initial gene sample data in the initial sample set, and dividing the initial gene sample data into blocks according to the linkage disequilibrium index of the initial gene sample data to generate at least two candidate gene blocks; randomly sampling from a symmetrical Beta distribution, and using the probability value corresponding to the sampling point as the gene mixing ratio; and randomly selecting a target number of gene blocks from the at least two candidate gene blocks as the target gene segments according to the gene mixing ratio.

[0049] For example, in this embodiment of the invention, a genetic linkage disequilibrium index in the initial gene sample data can be calculated first. Based on a preset index threshold, gene blocks (gene segments) with linkage disequilibrium indices greater than the preset threshold are identified from the initial gene sample data. Based on this, contiguous sites with strong linkage are divided into linkage blocks as candidate gene blocks, and the marker sites within each linkage block exhibit high genetic relevance. For example, the linkage disequilibrium index r... 2 It can be determined using the following formula:

[0050] ;

[0051] Wherein: the chain imbalance coefficient D is:

[0052] ;

[0053] in, Both allele A and allele B are present. It is the frequency of allele A. This refers to the frequency of allele B. It should be noted that allele A and allele B can refer to a pair of genes located at the same gene locus on homologous chromosomes.

[0054] Optionally, after generating at least two candidate gene blocks, the candidate gene blocks can be randomly shuffled. This further emphasizes the randomness of the mixture.

[0055] The target number can be determined based on the gene mixing ratio and the total number of candidate gene blocks.

[0056] Gene mixing ratios can be used to determine the overall proportion of replaced segments in a candidate gene block.

[0057] Optionally, a sample feature matrix corresponding to the initial gene sample data can be further constructed. This feature matrix can be used to generate subsequent mixed gene samples, and the preset index thresholds can be adaptively set according to those skilled in the art. It should be noted that the sample feature matrix may include the feature vector corresponding to each linkage block in the initial gene sample data.

[0058] Optionally, the Beta distribution can be represented by the following formula:

[0059]

[0060] in, The gene mixing coefficients determined by random sampling from the Beta distribution are obtained by controlling... The parameter controls the strength of the Beta distribution, thus affecting the sampling of λ. When α = 1.0, the Beta distribution is close to a uniform distribution, and the mixing coefficient λ is close to 0.5. When α is close to 0, the mixing coefficient λ tends to be 0 or 1, meaning the generated sample is closer to the original sample. Optionally, α ∈ (0.1, 1.0).

[0061] By adjusting the shape parameter α of the Beta distribution, the central tendency and distribution pattern of the sampling results can be flexibly changed, so that the distribution of replacement segments at the block scale is both random and can adapt to the mixed enhancement needs of different traits and datasets.

[0062] S230. Based on the segment information of the target gene segment, determine the replacement gene segment with the same segment information in the auxiliary gene sample data of the initial sample set, and replace the target gene segment with the replacement gene segment in the initial gene sample data to generate mixed gene sample data.

[0063] For example, multiple target gene segments can be represented by a set of target segments. Each target gene segment represents the starting position of the signal fragment to be replaced. and end position For each target gene segment, based on the start and end positions of the target gene segment, gene segments with the same start and end positions are identified in the auxiliary gene sample data. These gene segments are then swapped with the target gene segment. After completing the segment replacement operation for all target gene segments, the mixed gene sample data is output.

[0064] Optionally, the total length of at least two target gene regions can be determined using the following formula:

[0065] ;

[0066] in, Let k be the length of the target gene segment. K represents the gene mixing ratio, and K represents the total number of target gene segments.

[0067] It should be noted that the following is adopted: This is to adjust the nonlinear relationship of the actual replacement length, enhance the diversity of mixed gene sample data, and thus improve the generalization performance of the model.

[0068] S240. Based on the gene mixing ratio of the mixed gene sample data, perform a weighted summation on the sample label data corresponding to the initial gene sample data and the sample label data corresponding to the auxiliary gene sample data to determine the sample label data corresponding to the mixed gene sample data.

[0069] Specifically, the mixed gene sample data and the corresponding sample label data are combined to form a mixed gene sample, generating a new mixed gene sample set.

[0070] Optionally, in one specific implementation, the sample labels corresponding to the initial gene sample data and the sample labels corresponding to the auxiliary gene sample data are linearly weighted and summed according to the gene mixing ratio to obtain the sample labels corresponding to the mixed gene sample data. For example, the sample labels corresponding to the mixed gene sample data can be determined by the following formula:

[0071] ;

[0072] in, The sample labels corresponding to the mixed gene sample data. y represents the gene mixing ratio, and y represents the sample label corresponding to the initial gene sample data. These are the sample labels corresponding to the auxiliary gene sample data.

[0073] S250, based on mixed gene samples and an initial sample set, trains a pre-constructed crop prediction model.

[0074] Optionally, in this embodiment of the invention, the process of determining the mixed gene sample may further include: randomly selecting two sample data from an initial sample set as first gene sample data and second gene sample data; randomly sampling from a symmetrical Beta distribution and using the probability value corresponding to the sampling point as a first mixing probability value, and determining the mixed gene segment to be mixed in the first gene sample data according to the first mixing probability value; wherein, the segment length of the mixed gene segment is the mixing segment length; randomly determining a random gene segment with a length equal to the mixing segment length from the second gene sample data, replacing the random gene segment with the mixed gene segment, and using the first gene sample data after segment replacement as the mixed gene sample data; and performing a weighted summation of the sample label data corresponding to the first gene sample data and the sample label data corresponding to the second gene sample data according to the first mixing probability value to determine the sample label data corresponding to the mixed gene sample data.

[0075] The first mixing probability value can be used to determine the length of the mixed gene segment. Optionally, the length of the mixed segment can be determined by the following formula:

[0076] ;

[0077] in, Where L is the length of the mixed segment, and L is the total length of the first gene sample data. is the first mixing probability value, and c is the mixing parameter used to determine the gene mixing method, such as linear scaling mixing or area ratio uniform mixing.

[0078] Optionally, in this embodiment of the invention, the process of determining the mixed gene sample may further include: randomly selecting two sample data from the initial sample set as the third gene sample data and the fourth gene sample data; randomly sampling from a symmetrical Beta distribution and using the probability value corresponding to the sampling point as the second mixing probability value; and performing linear interpolation mixing on the third gene sample data, the fourth sample data, and the sample label data corresponding to both based on the second mixing probability value to generate the mixed gene sample.

[0079] The second mixing probability value can be used to characterize the mixing ratio of the linear interpolation. It should be noted that in this embodiment of the invention, the gene mixing ratio, the first mixing probability value, and the second mixing probability value are determined in the same way.

[0080] This application embodiment divides the genotype feature data of a sample into multiple regions based on linkage information present in the genotype data, and fills them with the corresponding regions of another sample to generate new training samples. This achieves chromosome segmentation and amplification, overcomes the limitations of single segment replacement in expressing local variations, and can simulate the diverse genetic variation patterns in linked segments with finer granularity, thereby improving the accuracy and stability of complex trait prediction models.

[0081] Example 3

[0082] Figure 3 This is a schematic diagram of a whole-genome selection sample data amplification device according to Embodiment 3 of this application. It is applicable to the amplification of crop gene samples. This whole-genome selection sample data amplification device can be implemented in hardware and / or software, and can be configured in a computer device, such as a server. Figure 3 As shown, the device includes:

[0083] The data acquisition module 310 is used to acquire the real genotype data of the target crop in the historical crop breeding process and the phenotypic data corresponding to the real genotype data, and generate an initial sample set based on the real genotype data and the phenotypic data.

[0084] The mixed sample module 320 is used to segment the initial gene sample data in the initial sample set based on the linkage feature information of the real genotype data, and to replace the segmented gene segments based on the auxiliary gene sample data to generate a mixed gene sample; wherein the initial gene sample data and the auxiliary gene sample data belong to different sample data in the initial sample set; the mixed gene sample is a new sample data that has been expanded.

[0085] Training module 330 is used to train a pre-constructed crop prediction model based on the mixed gene sample and the initial sample set.

[0086] This application embodiment enhances the signal of historically collected real genotype data of target crops by introducing linkage information, thereby achieving efficient parsing of genotype data, generating new mixed gene samples, increasing the data scale and diversity of training samples, reducing overfitting problems in the model training process, improving the prediction accuracy of complex traits, overcoming the limitation of insufficient samples, comprehensively capturing nonlinear genetic signals, and significantly improving the prediction efficiency and reliability of crop trait prediction.

[0087] Optionally, the mixed sample module 320 includes:

[0088] The segmentation unit is used to divide the initial gene sample data in the initial sample set into linked blocks according to the linkage disequilibrium index, generating at least two target gene segments; wherein, different gene segments do not overlap with each other.

[0089] The segment replacement unit is used to determine a replacement gene segment with the same segment information in the auxiliary gene sample data of the initial sample set according to the segment information of the target gene segment, and replace the target gene segment with the replacement gene segment in the initial gene sample data to generate mixed gene sample data.

[0090] The tag mixing unit is used to perform a weighted summation of the sample tag data corresponding to the initial gene sample data and the sample tag data corresponding to the auxiliary gene sample data based on the gene mixing ratio of the mixed gene sample data, so as to determine the sample tag data corresponding to the mixed gene sample data.

[0091] Optionally, the segment division units include:

[0092] Candidate segment subunits are used to randomly determine an initial gene sample data in the initial sample set, and to divide the initial gene sample data into blocks according to the linkage disequilibrium index of the initial gene sample data to generate at least two candidate gene blocks.

[0093] The mixing ratio subunit is used to randomly sample from a symmetrical Beta distribution and use the probability value corresponding to the sampling point as the gene mixing ratio;

[0094] The segment mixing subunit is used to randomly select a target number of gene blocks from the at least two candidate gene blocks as target gene segments according to the gene mixing ratio; wherein the target number is determined based on the gene mixing ratio and the total number of candidate gene blocks.

[0095] Optionally, the device may also include:

[0096] The first mixed sample module can be specifically used for:

[0097] Two sample data points are randomly selected from the initial sample set as the first gene sample data and the second gene sample data.

[0098] Random sampling is performed from a symmetrical Beta distribution, and the probability value corresponding to the sampling point is used as the first mixing probability value. The mixing gene segment to be mixed in the first gene sample data is determined according to the first mixing probability value; wherein, the segment length of the mixing gene segment is the mixing segment length.

[0099] Random gene segments of the same length as the mixed segment are randomly selected from the second gene sample data, and the random gene segments are replaced with the mixed gene segments. The first gene sample data after the segment replacement is used as the mixed gene sample data.

[0100] Based on the first mixed probability value, the sample label data corresponding to the first gene sample data and the sample label data corresponding to the second gene sample data are weighted and summed to determine the sample label data corresponding to the mixed gene sample data.

[0101] Optionally, the device may also include:

[0102] The second mixed sample module can be specifically used for:

[0103] Two sample data points are randomly selected from the initial sample set as the third gene sample data and the fourth gene sample data;

[0104] Randomly sample from a symmetric Beta distribution and use the probability value corresponding to the sampled point as the second mixture probability value;

[0105] Based on the second mixing probability value, the third gene sample data and the fourth sample data, as well as the sample label data corresponding to both, are linearly interpolated and mixed to generate mixed gene samples.

[0106] The whole-genome selection sample data augmentation device provided in this application embodiment can execute the whole-genome selection sample data augmentation method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing each whole-genome selection sample data augmentation method.

[0107] According to embodiments of this application, this application also provides an electronic device, a readable storage medium, and a computer program product.

[0108] Example 4

[0109] Figure 4 This is a schematic diagram of the structure of an electronic device 410 implementing the whole-genome selection sample data augmentation method of the embodiments of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0110] like Figure 4 As shown, the electronic device 410 includes at least one processor 411 and a memory, such as a read-only memory 412 or a random access memory 413, communicatively connected to the at least one processor 411. The memory stores computer programs executable by the at least one processor. The processor 411 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 412 or loaded from storage unit 418 into the random access memory 413. The random access memory 413 can also store various programs and data required for the operation of the electronic device 410. The processor 411, read-only memory 412, and random access memory 413 are interconnected via a bus 414. An input / output interface 415 is also connected to the bus 414.

[0111] Multiple components in electronic device 410 are connected to input / output interface 415, including: input unit 416, such as keyboard, mouse, etc.; output unit 417, such as various types of monitors, speakers, etc.; storage unit 418, such as disk, optical disk, etc.; and communication unit 419, such as network card, modem, wireless transceiver, etc. Communication unit 419 allows electronic device 410 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0112] Processor 411 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 411 performs the various methods and processes described above, such as whole-genome selection sample data augmentation methods.

[0113] In some embodiments, the genome-wide selection sample data augmentation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program can be loaded and / or mounted on electronic device 410 via read-only memory 412 and / or communication unit 419. When the computer program is loaded into random access memory 413 and executed by processor 411, one or more steps of the genome-wide selection sample data augmentation method described above can be performed. Alternatively, in other embodiments, processor 411 can be configured for the genome-wide selection sample data augmentation method by any other suitable means (e.g., by means of firmware).

[0114] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), payload programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0115] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable genome-wide selection sample data augmentation device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0116] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0118] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0119] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system to address the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.

[0120] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.

[0121] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for augmenting sample data through whole-genome selection, characterized in that, include: Obtain the true genotype data of the target crop in the historical crop breeding process and the corresponding phenotypic data, and generate an initial sample set based on the true genotype data and the phenotypic data; Based on the linkage feature information of the real genotype data, the initial gene sample data in the initial sample set is segmented, and the segmented gene segments are replaced based on the auxiliary gene sample data to generate a hybrid gene sample; wherein, the initial gene sample data and the auxiliary gene sample data belong to different sample data in the initial sample set; the hybrid gene sample is a new sample data that has been expanded. The pre-constructed crop prediction model is trained based on the mixed gene sample and the initial sample set.

2. The method according to claim 1, characterized in that, The linkage feature information based on the real genotype data is used to segment the initial gene sample data in the initial sample set, and the segmented gene segments are replaced based on auxiliary gene sample data to generate a mixed gene sample, including: Based on the linkage disequilibrium index, the initial gene sample data in the initial sample set are divided into linkage blocks to generate at least two target gene segments; wherein, different gene segments do not overlap with each other. Based on the segment information of the target gene segment, a replacement gene segment with the same segment information is determined in the auxiliary gene sample data of the initial sample set, and the target gene segment is replaced with the replacement gene segment in the initial gene sample data to generate mixed gene sample data; Based on the gene mixing ratio of the mixed gene sample data, the sample label data corresponding to the initial gene sample data and the sample label data corresponding to the auxiliary gene sample data are weighted and summed to determine the sample label data corresponding to the mixed gene sample data.

3. The method according to claim 2, characterized in that, The process involves dividing the initial gene sample data in the initial sample set into linkage blocks based on the linkage disequilibrium index, generating at least two target gene segments, including: An initial gene sample data is randomly selected from the initial sample set, and the initial gene sample data is divided into blocks according to the linkage disequilibrium index of the initial gene sample data to generate at least two candidate gene blocks. Random sampling is performed from a symmetrical Beta distribution, and the probability value corresponding to the sampling point is used as the gene mixing ratio; Based on the gene mixing ratio, a target number of gene blocks are randomly selected from the at least two candidate gene blocks as target gene segments; wherein, the target number is determined based on the gene mixing ratio and the total number of candidate gene blocks.

4. The method according to claim 1, characterized in that, The process of identifying mixed gene samples also includes: Two sample data points are randomly selected from the initial sample set as the first gene sample data and the second gene sample data. Random sampling is performed from a symmetrical Beta distribution, and the probability value corresponding to the sampling point is used as the first mixing probability value. The mixing gene segment to be mixed in the first gene sample data is determined according to the first mixing probability value; wherein, the segment length of the mixing gene segment is the mixing segment length. Random gene segments of the same length as the mixed segment are randomly selected from the second gene sample data, and the random gene segments are replaced with the mixed gene segments. The first gene sample data after the segment replacement is used as the mixed gene sample data. Based on the first mixed probability value, the sample label data corresponding to the first gene sample data and the sample label data corresponding to the second gene sample data are weighted and summed to determine the sample label data corresponding to the mixed gene sample data.

5. The method according to claim 1, characterized in that, The process of identifying mixed gene samples also includes: Two sample data points are randomly selected from the initial sample set as the third gene sample data and the fourth gene sample data; Randomly sample from a symmetric Beta distribution and use the probability value corresponding to the sampled point as the second mixture probability value; Based on the second mixing probability value, the third gene sample data and the fourth sample data, as well as the sample label data corresponding to both, are linearly interpolated and mixed to generate mixed gene samples.

6. The method according to claim 1, characterized in that, After obtaining the true genotype data of the target crop and the corresponding phenotypic data, the process also includes: The data format of the real genotype data is encoded and converted according to the combination of bases in the dialleles.

7. An apparatus for augmenting sample data through whole-genome selection, characterized in that, include: The data acquisition module is used to acquire the real genotype data of the target crop in the historical crop breeding process and the phenotypic data corresponding to the real genotype data, and generate an initial sample set based on the real genotype data and the phenotypic data. The mixed sample module is used to segment the initial gene sample data in the initial sample set based on the linkage feature information of the real genotype data, and to replace the segmented gene segments based on the auxiliary gene sample data to generate a mixed gene sample; wherein the initial gene sample data and the auxiliary gene sample data belong to different sample data in the initial sample set; the mixed gene sample is a new sample data that has been expanded. The training module is used to train a pre-built crop prediction model based on the mixed gene samples and the initial sample set.

8. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the whole genome selection sample data augmentation method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the whole-genome selection sample data augmentation method as described in any one of claims 1-6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the whole-genome selection sample data augmentation method according to any one of claims 1-6.