Chrysanthemum breeding experiment data aggregation analysis method and system based on somatic variation

By constructing a data aggregation and correlation mechanism for chrysanthemum breeding experiments, the problem of scattered data in chrysanthemum breeding experiments was solved, and unified data management and cross-stage analysis were realized, supporting the systematic management and decision-making of breeding experiment results.

CN121883200APending Publication Date: 2026-04-17CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2026-01-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In the chrysanthemum somatic cell mutation breeding experiment, the experimental data generated in each experimental stage lacked a unified standard, resulting in data dispersion and making it difficult to conduct effective correlation and systematic analysis across experimental stages at the sample level.

Method used

By introducing unified sample identifiers and experimental stage attributes, a data aggregation and association mechanism is constructed across experimental stages. Experimental data is collected and identified, data consistency and standardization processes are performed, and data models are established for analysis.

Benefits of technology

It has enabled the structured organization and unified management of chrysanthemum breeding experimental data, ensuring data consistency within the same analytical framework and supporting effective analysis and decision-making across stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883200A_ABST
    Figure CN121883200A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of plant tissue culture and breeding, in particular to a chrysanthemum breeding experimental data aggregation analysis method and system based on somatic variation. Experimental data generated in the experimental stages of explant treatment, callus induction, regeneration differentiation, rooting culture, variation character statistics and the like are collected in a unified mode; and a sample identifier and an experiment batch identifier are configured for the experiment data. Through data aggregation association, data processing and data analysis, association organization of cross-experimental-stage experimental data in a sample dimension is realized, culture condition parameters, regeneration differentiation indexes and variation character data are subjected to conjoint analysis, and an analysis result is output in a structured data form. Meanwhile, the invention further provides a system for implementing the method, and the system is used for completing acquisition, convergence association, processing, analysis and result output of experimental data and provides technical support for data management and analysis of somatic variation chrysanthemum breeding experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of plant tissue culture and breeding technology, and in particular to a method and system for pooling and analyzing experimental data of chrysanthemum breeding based on somatic cell variation. Background Technology

[0002] Somatic cell variation breeding of chrysanthemums is a breeding method that obtains new germplasm resources through genetic or phenotypic variations generated during somatic cell culture. It is widely used in areas such as flower color improvement, plant type optimization, and trait innovation. This breeding method typically involves multiple experimental stages, including explant treatment, callus induction, regeneration differentiation, rooting culture, and statistical analysis of variant traits. During the breeding process, a large amount of experimental data related to culture conditions, regeneration differentiation processes, and phenotypic traits is continuously generated. Chrysanthemum breeding experiments based on somatic cell variation are characterized by long experimental cycles, multiple stages, and complex data types, making it one of the important technical approaches for the selection and breeding of new chrysanthemum varieties.

[0003] In existing chrysanthemum somatic cell mutation breeding experiments, the experimental data generated at each experimental stage are usually recorded and saved separately by the experimenters. There is a lack of unified standards for the data recording methods, field structures and identification rules at different stages. It is difficult to effectively correlate the data generated by the same experimental object at different experimental stages in terms of sample dimensions. As a result, the experimental data are scattered and it is difficult to form a complete data chain across experimental stages. This is not conducive to the systematic analysis and management of the process and results of somatic cell mutation breeding experiments. Summary of the Invention

[0004] To overcome the above shortcomings, this invention provides a method and system for pooling and analyzing experimental data of chrysanthemum breeding based on somatic cell variation. It aims to improve the problems in existing somatic cell variation breeding experiments where experimental data generated at different experimental stages lack unified correlation and data of the same experimental object at different experimental stages are difficult to correspond.

[0005] In a first aspect, the present invention provides the following technical solution: a method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation, comprising: S1. Collect experimental data generated at different stages during the chrysanthemum somatic cell mutation breeding experiment, and configure data identification information for the collected experimental data to identify the experimental objects and experimental processes; S2. The experimental data from different experimental stages are aggregated and processed according to a unified data structure. Based on the data identification information, a data association structure is established to characterize the correspondence between the same experimental object in multiple experimental stages, thereby obtaining a breeding experimental data set across experimental stages. S3. Perform data consistency processing and standardization processing on the breeding experiment data set to obtain standardized experimental data for analysis; S4. Based on the standardized experimental data, the correlation between culture condition parameters, regeneration and differentiation indicators and variation trait data is analyzed and processed to obtain the analysis results characterizing the relationship between breeding experimental conditions and somatic cell variation characteristics. S5. Based on the analysis results, output data-driven decision-making information for chrysanthemum somatic cell mutation breeding.

[0006] Preferably, in step S1, the step of collecting experimental data generated in multiple experimental stages includes: Experimental data were acquired according to a unified data collection rule at different experimental stages; Experimental data obtained at different experimental stages are bound and stored with corresponding sample identifiers and experimental batch identifiers; Record the corresponding experimental stage attribute information for the collected experimental data.

[0007] Preferably, in step S1, the step of configuring data identification information for experimental data includes: Generate unique sample identifiers for the same experimental subject at different experimental stages; Associate the sample identifier with the experimental batch information; Write the corresponding sample identifier and experimental batch identifier into the experimental data collected at different experimental stages.

[0008] Preferably, in step S2, the step of pooling the experimental data from the multiple experimental phases includes: Field mapping is performed on experimental data from different experimental stages according to the preset data field structure; The experimental data with completed field mappings are merged and stored according to the sample identifier; Based on the merged experimental data, establish data records that characterize the correspondence between the same experimental object in different experimental stages.

[0009] Preferably, in step S3, the steps of performing data consistency processing and standardization processing on the breeding experiment dataset include: Experimental data collected at different experimental stages were processed using a unified unit according to the corresponding experimental index type; Perform dimensional normalization on numerical experimental data; Based on the experimental stage, data records that do not match the sample labels are identified and removed.

[0010] Preferably, in step S3, the standardized experimental data includes: Experimental data used to characterize differences in culture conditions, experimental data used to characterize the regeneration and differentiation process, and experimental data used to characterize phenotypic variation; The experimental data are associated and stored with the same sample identifier in a cross-experimental data association structure.

[0011] Preferably, in step S4, the step of analyzing and processing the correlation between the experimental data includes: Correlation analysis was performed based on culture condition parameters and regeneration and differentiation indicators; Perform multidimensional feature analysis based on variable trait data; Establish a data model to characterize the relationship between culture condition parameters and variant traits.

[0012] Preferably, in step S4, the analysis and processing step further includes: Clustering of regenerated plants was performed based on data of variable traits; The regenerated plants were categorized based on the clustering results; Assign corresponding category identifiers to different categories of regenerated plants.

[0013] Preferably, in step S5, the step of outputting data-driven decision information includes: The analysis results are transformed into structured data records corresponding to the sample identifiers; The structured data records are summarized according to experimental batches; Output the summarized structured data results.

[0014] Secondly, this invention provides the following technical solution: a chrysanthemum breeding experimental data aggregation and analysis system based on somatic cell variation, comprising: The data acquisition module is configured to collect experimental data generated during multiple experimental stages in the chrysanthemum somatic cell mutation breeding experiment, and to configure data identification information for the collected experimental data to uniquely identify the experimental object and the experimental process. The data aggregation and association module is configured to aggregate experimental data generated in multiple experimental stages according to a unified data structure, and establish a data association structure based on the data identification information to characterize the correspondence between the same experimental object in different experimental stages, so as to obtain a cross-stage breeding experimental data set. The data processing module is configured to perform data consistency processing and standardization processing on the breeding experiment data set to obtain standardized experimental data for analysis. The data analysis module is configured to analyze and process the correlation between experimental data based on the standardized experimental data to obtain analysis results; The results output module is configured to output data-driven decision-making information for chrysanthemum somatic cell variation breeding based on the analysis results.

[0015] The present invention has the following beneficial effects: 1. In this invention, in order to address the problem of scattered and difficult-to-associate data in multiple experimental stages of somatic cell mutation breeding experiments, a data aggregation and association mechanism across experimental stages is constructed at the system level by introducing unified sample identifiers and experimental stage attributes. This enables experimental data generated by the same experimental object in different experimental stages to form a stable correspondence in the sample dimension, thereby realizing the structured organization and unified management of somatic cell mutation breeding experimental data.

[0016] 2. In this invention, by performing consistency verification, unit unification, and dimension normalization on experimental data collected at different experimental stages in the data processing module, experimental data from different sources and with different index types maintain numerical and structural consistency under the same analytical framework. This avoids data incomparability issues caused by differences in experimental stages or recording methods, and provides a reliable data foundation for subsequent correlation analysis of experimental data.

[0017] 3. In this invention, by jointly analyzing culture condition parameters, regeneration differentiation indicators and phenotypic variation data, and establishing corresponding data models and category identification mechanisms in the system, the regenerated plants are classified and managed in terms of variation traits. This enables the analysis results to be output in the form of structured data and summarized by experimental batch, facilitating the organization, comparison and subsequent application of somatic cell variation breeding experimental results. Attached Figure Description

[0018] Figure 1 This is a flowchart of the method and system for pooling and analyzing experimental data of chrysanthemum breeding based on somatic cell variation proposed in this invention. Figure 2 This is a system architecture diagram of the chrysanthemum breeding experimental data aggregation and analysis method and system based on somatic cell variation proposed in this invention. Detailed Implementation

[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Example 1: In the first embodiment of the present invention, the present invention provides a method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation, such as... Figure 1 As shown, it includes the following steps: S1. Collect experimental data generated at different stages during the chrysanthemum somatic cell mutation breeding experiment, and configure data identification information for the collected experimental data to identify the experimental objects and experimental processes; Furthermore, in step S1, the steps for collecting experimental data generated in multiple experimental phases include: Experimental data were acquired according to a unified data collection rule at different experimental stages; Experimental data obtained at different experimental stages are bound and stored with corresponding sample identifiers and experimental batch identifiers; Record the corresponding experimental stage attribute information for the collected experimental data.

[0021] Specifically, during implementation, unified data acquisition rules are predefined and executed at each experimental stage. These unified rules include at least: a set of data record fields, field data types, data acquisition timestamp recording methods, and data storage formats. Data acquired at each experimental stage is generated in structured record format, with each record containing at least: a data acquisition timestamp, experimental stage attribute fields, experimental parameter fields, and experimental result fields. Data acquisition can be done manually or via equipment. During equipment acquisition, the outputs of environmental sensors and digital measurement devices are converted into data records consistent with the field set before being written to the data.

[0022] For each structured data record, sample identifier and experimental batch identifier fields are simultaneously written to the storage medium, thus binding the data to the experimental object and experimental batch. The bound storage adopts a record-level binding method, meaning each data record contains a corresponding sample identifier and experimental batch identifier, and data records from different experimental stages use the same identifier field naming and data type constraints. In terms of storage structure, a batch index is created based on the experimental batch, and a sample index is created under the batch index based on the sample identifier, allowing records of the same sample from multiple experimental stages to be aggregated and read based on the sample identifier.

[0023] When generating each data record, experimental stage attribute information is written. This information consists of predefined stage enumeration values ​​or stage codes, indicating the experimental stage category to which the data record belongs. The experimental stage attribute information, along with the timestamp, sample identifier, and experimental batch identifier, constitutes the retrieval criteria for the data record, enabling subsequent steps to filter, aggregate, and validate the data by stage.

[0024] By implementing this step, data from each experimental stage can meet unified field rules during collection, and record-level binding of sample identifiers and experimental batch identifiers can be completed. At the same time, experimental stage attribute information is written, thereby forming a structured data record set that can be retrieved and aggregated by batch, sample, and stage.

[0025] Furthermore, in step S1, the step of configuring data identification information for experimental data includes: Generate unique sample identifiers for the same experimental subject at different experimental stages; Associate sample identifiers with experimental batch information; Write the corresponding sample identifier and experimental batch identifier into the experimental data collected at different experimental stages.

[0026] Specifically, before the somatic cell mutation breeding experiment begins, a unique sample identifier is generated for each experimental subject. This sample identifier remains unique and non-repeating throughout the entire experimental period, used to identify the data source for the same experimental subject in different experimental stages, including explant treatment, callus induction, regeneration differentiation, rooting culture, and statistical analysis of variant traits. The sample identifier is generated using a standardized coding method, which includes at least an experimental subject number field and an experimental round field. Field combinations ensure that the sample identifier does not duplicate across different experimental subjects. The generated sample identifier is recorded when the experimental subject is established and is continuously used in subsequent experimental stages.

[0027] For each round of somatic cell mutation breeding experiments, a corresponding experimental batch identifier is generated to distinguish experimental batches conducted at different times or under different experimental conditions. An association is established between the experimental batch identifier and the sample identifier to identify the experimental batch to which a sample belongs. At the storage level, the sample identifier and the experimental batch identifier are recorded as associated fields. These associated fields are mapped in the sample information table, enabling unified management of data generated by the same sample at different experimental stages across both the experimental batch and sample dimensions.

[0028] When collecting experimental data at different experimental stages, the corresponding sample identifier and experimental batch identifier are written into each experimental data record. The identifier information is embedded in the experimental data record as a field and stored together with the experimental parameter field and the experimental result field. By writing the sample identifier and experimental batch identifier during the data collection stage, the data generated at different experimental stages is identified and bound at the time of storage, avoiding problems such as inconsistent data sources or sample confusion in subsequent processing.

[0029] By implementing this step, the experimental data generated during the somatic cell mutation breeding experiment are configured and written with sample identification and experimental batch identification during the collection stage. This enables data generated by the same experimental object in different experimental stages to be associated and distinguished based on a unified identification, providing a stable data identification foundation for subsequent cross-experimental data aggregation and analysis.

[0030] S2. Collect and process experimental data from different experimental stages according to a unified data structure, and establish a data association structure based on data identification information to characterize the correspondence between the same experimental object in multiple experimental stages, thereby obtaining a breeding experimental data set across experimental stages. Furthermore, in step S2, the step of pooling experimental data from multiple experimental phases includes: Field mapping is performed on experimental data from different experimental stages according to the preset data field structure; The experimental data with completed field mappings are merged and stored according to the sample identifier; Based on the merged experimental data, establish data records that characterize the correspondence between the same experimental object in different experimental stages.

[0031] Specifically, during implementation, a unified data field structure is predefined to describe the experimental data at each stage of the somatic cell mutation breeding experiment. This unified data field structure includes at least a sample identifier field, an experimental batch field, an experimental stage field, an experimental parameter field, and an experimental result field. For experimental data collected at different experimental stages, field mapping is performed on the original data fields based on the correspondence between their original record fields and the unified data field structure. During field mapping, semantically consistent but differently named data fields from different experimental stages are converted to unified field names and formatted according to the unified field data type requirements, thus ensuring consistency of experimental data at the field level across different experimental stages.

[0032] After completing the field mapping, experimental data from different experimental stages are merged and stored according to the sample identifier. During the merging process, the sample identifier is used as the primary index to aggregate multiple experimental data records belonging to the same sample identifier into the same logical data set. During the merging process, data records from different experimental stages are distinguished by the experimental stage field, and data records within the same experimental stage are arranged in chronological order of collection time, thus forming a data storage structure organized by sample identifier and distinguished by experimental stage.

[0033] After merging and storing the data, a cross-experimental data correspondence record is established based on the merged experimental data. This correspondence record characterizes the data association between the same experimental object and different experimental stages. The record includes at least the sample identifier, the experimental stage identifier, and the data record index information corresponding to each experimental stage. This record allows for locating the data position of the same sample in each experimental stage. The established data correspondence record serves as the input basis for subsequent data consistency processing, standardization processing, and association analysis steps.

[0034] By implementing this step, experimental data generated in different experimental stages can be merged and stored according to a unified data structure after field mapping is completed, and cross-experimental data correspondence records can be established based on sample identifiers, thereby forming a cross-stage experimental data set that can be used for subsequent processing and analysis.

[0035] S3. Perform data consistency processing and standardization on the breeding experiment data set to obtain standardized experimental data for analysis; Furthermore, in step S3, the steps of performing data consistency processing and standardization processing on the breeding experiment dataset include: Experimental data collected at different experimental stages were processed using a unified unit according to the corresponding experimental index type; Perform dimensional normalization on numerical experimental data; Based on the experimental stage, data records that do not match the sample labels are identified and removed.

[0036] Specifically, by performing data consistency processing and standardization on the cross-experimental breeding experiment data set formed in the above steps, it is ensured that data from different experimental stages and different indicator types can be compared and correlated within the same analytical framework.

[0037] For data from different experimental stages within the breeding experiment dataset, the data is first categorized according to the type of experimental indicators. These indicators include at least those related to culture conditions, indicators of regeneration and differentiation processes, and statistical indicators of phenotypic variation. For data records with inconsistent units within the same indicator type, unit unification is performed according to pre-defined unit conversion rules. These rules are stored as configuration parameters in the system; the conversion process only reverts the data units without altering the physical meaning of the data. After unit unification, data on the same type of experimental indicator collected from different experimental stages are numerically comparable.

[0038] After completing the unified processing by the unit, dimensional normalization was performed on the numerical experimental data in the breeding experiment dataset. Normalization is used to transform numerical experimental data with different value ranges into a unified numerical interval to facilitate subsequent correlation analysis and model processing.

[0039] In this embodiment, dimensional normalization adopts linear normalization, and the normalization formula is as follows: ; in: This represents the original experimental data value before normalization. This represents the minimum data value under the same type of experimental index. This represents the maximum data value under the same experimental index type; This represents the experimental data value after normalization.

[0040] The minimum and maximum data values ​​are obtained by statistical analysis of data from the same experimental batch or within the same analytical range, and the normalization results are written into the standardized experimental dataset.

[0041] After unifying units and normalizing dimensions, a consistency check was performed on the breeding experiment dataset using experimental stage attributes. During the check, the experimental stage attributes in each data record were matched against the experimental stage information corresponding to the sample identifier. If an inconsistency was detected between the experimental stage attribute of a data record and its sample identifier in the corresponding experimental stage, the data record was marked as a mismatch and removed from the standardized experimental dataset used for analysis. This method avoids data confusion at the sample level across different experimental stages, ensuring the correct correspondence between samples and experimental stages in the dataset.

[0042] Through this step, the cross-experimental breeding experiment data set has completed unit unification, dimension normalization, and data consistency verification based on experimental stage attributes before entering the analysis and processing stage. This results in a standardized experimental data set with consistent structure, comparable values, and accurate sample correspondence, providing a reliable data foundation for subsequent experimental data correlation analysis.

[0043] Furthermore, in step S3, the standardized experimental data includes: Experimental data used to characterize differences in culture conditions, experimental data used to characterize the regeneration and differentiation process, and experimental data used to characterize phenotypic variation; Experimental data are linked and stored using the same sample identifier in a data association structure across experimental phases.

[0044] Specifically, after completing the data unit unification, dimension normalization, and consistency verification processes described above, a standardized experimental data set is formed for subsequent analysis and processing. The standardized experimental data set consists of multiple types of experimental data and is stored uniformly within a cross-experimental data association structure. The standardized experimental data includes the following three types of experimental data: The first category is experimental data used to characterize differences in culture conditions. This type of experimental data is used to describe experimental parameter information under different culture conditions during somatic cell mutation breeding experiments. The experimental data comes from culture condition-related records collected at each experimental stage and is written into a standardized experimental data set after unit unification and dimension normalization.

[0045] The second category is experimental data used to characterize the regeneration and differentiation process. This type of experimental data is used to describe the regeneration and differentiation status of experimental subjects in stages such as callus induction, adventitious shoot regeneration, and rooting culture. The experimental data includes numerical or statistical data that reflect the progress of the regeneration process, and is associated with the corresponding sample labels after standardization.

[0046] The third category is experimental data used to characterize phenotypic variation. This type of experimental data is used to describe the differences in morphological traits among regenerated plants. The experimental data comes from the collection records during the statistical stage of phenotypic variation and is written into the standardized experimental data set as phenotypic feature data after standardization.

[0047] The three types of experimental data mentioned above are distinguished by different data fields in the standardized experimental dataset, and the data types and storage formats are kept consistent at the field level.

[0048] In the storage of standardized experimental data, the three types of experimental data mentioned above are stored according to a cross-experimental stage data association structure. Standardized experimental data generated in different experimental stages are all stored using the same sample identifier as the association key, enabling a correspondence to be established at the sample level for data generated by the same experimental object in different experimental stages. In the cross-experimental stage data association structure, each sample identifier corresponds to a set of standardized experimental data records. This set of data records includes culture condition data, regeneration and differentiation process data, and phenotypic variation data corresponding to that sample in different experimental stages. Through this association storage method, standardized experimental data maintain a consistent correspondence across both the sample and experimental stage dimensions.

[0049] Through the implementation of the above steps, the standardized experimental dataset clearly includes three types of experimental data: differences in culture conditions, regeneration and differentiation processes, and phenotypic variations. These data are then linked and stored in a cross-experimental data association structure using the same sample identifier, thus providing a complete data composition and a stable data organization form for subsequent experimental data association analysis based on sample and experimental stage dimensions.

[0050] S4. Based on standardized experimental data, the correlation between culture condition parameters, regeneration and differentiation indicators, and variation trait data is analyzed and processed to obtain the analysis results characterizing the relationship between breeding experimental conditions and somatic cell variation characteristics. Furthermore, in step S4, the steps for analyzing and processing the correlation between experimental data include: Correlation analysis was performed based on culture condition parameters and regeneration and differentiation indicators; Perform multidimensional feature analysis based on variable trait data; Establish a data model to characterize the relationship between culture condition parameters and variant traits.

[0051] Specifically, culture condition parameter data and corresponding regeneration and differentiation index data are extracted from the standardized experimental dataset according to sample identifiers, and a one-to-one correspondence is established at the sample level. Based on this, correlation analysis is performed on the correspondence between changes in culture condition parameter values ​​and changes in regeneration and differentiation indexes to obtain analytical results characterizing the degree of correlation between culture condition parameters and regeneration and differentiation indexes. These analytical results are recorded as intermediate analytical data.

[0052] After completing the correlation analysis between culture condition parameters and regeneration and differentiation indicators, the variant trait data corresponding to the sample identifiers were extracted from the standardized experimental dataset. The variant trait data consisted of multiple phenotypic feature dimensions. By performing feature parsing processing on the multidimensional phenotypic feature data, feature components used to characterize the main variant features were extracted, and the feature components were associated and stored with the corresponding sample identifiers.

[0053] After completing the multidimensional feature analysis, a data model is established based on the culture condition parameter data and the corresponding variant trait components to characterize the relationship between culture condition parameters and variant traits. The data model describes the mapping relationship between culture condition parameters and variant trait characteristics, and its basic expression is as follows: ; in: This represents the set of input features consisting of cultivation condition parameters; This represents the set of output features composed of variable trait characteristic components; This represents a mapping model based on standardized experimental data.

[0054] The model structure and model parameter information are recorded as analysis results for use in subsequent data queries and decision information generation.

[0055] Through this step, the correlation analysis between culture condition parameters and regeneration and differentiation indicators was completed, the multidimensional feature analysis of variant trait data was performed, and a data model for characterizing the relationship between culture condition parameters and variant traits was established, providing basic model support for further analysis and processing of subsequent experimental data.

[0056] Furthermore, in step S4, the analysis and processing steps further include: Clustering of regenerated plants was performed based on data of variable traits; The regenerated plants were categorized based on the clustering results; Assign corresponding category identifiers to different categories of regenerated plants.

[0057] Specifically, the variant trait data corresponding to each regenerated plant was extracted from the standardized experimental dataset, and a variant trait feature dataset was constructed using the sample identifier as an index. The variant trait data consisted of multiple phenotypic feature dimensions, with different dimensions used to describe the differences in morphological traits among the regenerated plants. Based on the variant trait feature dataset, clustering was performed on the regenerated plants. During the clustering process, the regenerated plants were grouped according to the similarity of variant trait features among different plants, ensuring that regenerated plants within the same group had similar data distribution characteristics in terms of variant trait features. The results of the clustering process were recorded in the form of sample identifiers and corresponding cluster group numbers.

[0058] After clustering, the regenerated plants are categorized based on the clustering results. Each cluster corresponds to a category of regenerated plants, and regenerated plants within the same category exhibit similar variability traits. The categorization results are stored as a correspondence between sample identifiers and category numbers, and are associated with the corresponding variability data to characterize the classification of different regenerated plants along the variability dimension.

[0059] For each category of regenerated plant, a corresponding category identifier is assigned. These category identifiers are generated using an coded format to distinguish between different categories of regenerated plants. During data storage, the category identifiers are associated with their corresponding sample identifiers and written into standardized experimental data records, ensuring that each regenerated plant possesses both sample and category identifier information at the data level. The category identifiers serve as crucial index fields for subsequent data statistics, filtering, and decision-making information output.

[0060] Through this step, clustering, classification, and category labeling of regenerated plants were completed based on the data of variable traits. This enabled the regenerated plants to form a clear classification structure in terms of variable traits, providing structured support for subsequent data output based on the category dimension and the generation of breeding decision information.

[0061] S5. Output data-driven decision-making information for chrysanthemum somatic cell mutation breeding based on the analysis results; Furthermore, in step S5, the step of outputting data-driven decision information includes: The analysis results are transformed into structured data records corresponding to the sample identifiers; Structured data records are summarized according to experimental batches; Output the summarized structured data results.

[0062] Specifically, after completing the correlation analysis and clustering of the experimental data, the analysis results undergo structured transformation. During this transformation, the analysis results are broken down into predefined data fields, forming structured data records. Each structured data record uses a sample identifier as an index field, associating the analysis result for each regenerated plant with its sample identifier. The structured data record includes at least a sample identifier field, an analysis result field, and a category identifier field. These fields are recorded according to a unified data format, ensuring a clear correspondence between the analysis results and the sample dimension.

[0063] After generating the structured data records, they are aggregated according to experimental batches. During the aggregation process, the experimental batch identifier is used as the aggregation dimension to group structured data records belonging to the same experimental batch. The structured data records within an experimental batch maintain the correspondence between sample identifiers and analysis results during aggregation, while structured data records from different experimental batches are distinguished by their experimental batch identifiers. The aggregated data results form independent datasets on a batch-by-batch basis.

[0064] After completing the batch-based aggregation process, the resulting structured data is output. The output data maintains a structured format, including batch identifiers, sample identifiers, and corresponding analysis results. The output method for the structured data results is determined based on the specific application scenario; it can be output to local storage media or an external data interface. The data field structure and content remain unchanged during the output process, ensuring the integrity and consistency of the data results.

[0065] By implementing this step, the analysis results obtained during the experimental data analysis and processing can be transformed into structured data records corresponding to sample identifiers, and the data can be summarized and output according to experimental batches, so that the analysis results are finally output in the form of structured data, forming a complete data processing closed loop.

[0066] Example 2: In a chrysanthemum somatic cell mutation breeding experiment, the breeding process involves multiple experimental stages, including explant treatment, callus induction, regeneration differentiation, rooting culture, and statistical analysis of variant traits. The same experimental subject generates multiple types of experimental data at different stages, with differences in collection time, data type, and recording method. In actual experiments, the sources of experimental data from different stages are scattered, the field structures are inconsistent, and there is a lack of unified identification and association mechanisms across experimental stages. This makes it difficult to effectively aggregate and analyze culture condition parameters, regeneration differentiation indicators, and phenotypic variation data at the sample level. To solve these problems, the chrysanthemum breeding experiment data aggregation and analysis system based on somatic cell mutation, as provided in this invention, is used. Its structure is as follows: Figure 2 As shown. The specific implementation process of this system is as follows: The data acquisition module is used to collect experimental data generated at different experimental stages during somatic cell mutation breeding experiments. The collected data includes experimental parameter data under different culture conditions, experimental index data during regeneration and differentiation, and statistical data on variant traits. During implementation, the data acquisition module acquires experimental data from different experimental stages according to unified data acquisition rules and converts the data into a structured data format during the acquisition phase. Simultaneously, sample identifiers and experimental batch identifiers are written to the experimental data during data acquisition, ensuring that experimental data generated from the same experimental subject at different experimental stages have a unified data identification basis.

[0067] The data aggregation and association module is used to aggregate and associate experimental data from different experimental stages. This module maps experimental data from different stages according to a preset data field structure and merges and stores the mapped data based on sample identifiers. During the aggregation process, the module also retains the experimental stage attribute information, enabling the aggregated experimental data to distinguish the data sources from different experimental stages. It also establishes a data correspondence between the same experimental object and different experimental stages at the sample level, thus forming a cross-experimental stage data association structure.

[0068] The data processing module performs consistency and standardization processing on the aggregated experimental data. This module standardizes the units of the experimental data according to the type of experimental indicator and performs dimensional normalization on numerical experimental data. Simultaneously, the data processing module performs consistency checks on the experimental data based on experimental stage attributes and sample identifiers, removing data records whose experimental stage attributes do not match the sample identifiers, thus forming a standardized set of experimental data for subsequent analysis and processing.

[0069] The data analysis module performs correlation analysis on standardized experimental data. This module analyzes specific combinations of culture condition parameters, regeneration differentiation indicators, and phenotypic variation data generated in somatic cell mutation breeding experiments. During implementation, the data analysis module performs correlation analysis based on culture condition parameters and regeneration differentiation indicators, and performs multidimensional feature analysis based on the variant trait data. Based on this, a data model is established to characterize the relationship between culture condition parameters and variant traits. Then, the data analysis module clusters the regenerated plants based on the variant trait data, classifies the regenerated plants according to the clustering results, and assigns corresponding category labels to different categories of regenerated plants.

[0070] The results output module is used to structure the analysis results output by the data analysis module. This module converts the analysis results into structured data records corresponding to sample identifiers and summarizes these records according to experimental batches. The summarized structured data results are organized using sample identifiers and experimental batches as index dimensions and output in structured data format for subsequent analysis and management of somatic cell mutation breeding experimental data.

[0071] Finally, it should be noted that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A data aggregation and analysis method for chrysanthemum breeding experiments based on somatic cell variation, characterized in that, include: S1. Collect experimental data generated at different stages during the chrysanthemum somatic cell mutation breeding experiment, and configure data identification information for the collected experimental data to identify the experimental objects and experimental processes; S2. The experimental data from different experimental stages are aggregated and processed according to a unified data structure. Based on the data identification information, a data association structure is established to characterize the correspondence between the same experimental object in multiple experimental stages, thereby obtaining a breeding experimental data set across experimental stages. S3. Perform data consistency processing and standardization processing on the breeding experiment data set to obtain standardized experimental data for analysis; S4. Based on the standardized experimental data, the correlation between culture condition parameters, regeneration and differentiation indicators and variation trait data is analyzed and processed to obtain the analysis results characterizing the relationship between breeding experimental conditions and somatic cell variation characteristics. S5. Based on the analysis results, output data-driven decision-making information for chrysanthemum somatic cell mutation breeding.

2. The method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation according to claim 1, characterized in that, In step S1, the step of collecting experimental data generated in multiple experimental stages includes: Experimental data were acquired according to a unified data collection rule at different experimental stages; Experimental data obtained at different experimental stages are bound and stored with corresponding sample identifiers and experimental batch identifiers; Record the corresponding experimental stage attribute information for the collected experimental data.

3. The method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation according to claim 1, characterized in that, In step S1, the step of configuring data identification information for the experimental data includes: Generate unique sample identifiers for the same experimental subject at different experimental stages; Associate the sample identifier with the experimental batch information; Write the corresponding sample identifier and experimental batch identifier into the experimental data collected at different experimental stages.

4. The method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation according to claim 1, characterized in that, In step S2, the step of pooling the experimental data from the multiple experimental phases includes: Field mapping is performed on experimental data from different experimental stages according to the preset data field structure; The experimental data with completed field mappings are merged and stored according to the sample identifier; Based on the merged experimental data, establish data records that characterize the correspondence between the same experimental object in different experimental stages.

5. The method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation according to claim 1, characterized in that, In step S3, the steps of performing data consistency processing and standardization processing on the breeding experiment dataset include: Experimental data collected at different experimental stages were processed using a unified unit according to the corresponding experimental index type; Perform dimensional normalization on numerical experimental data; Based on the experimental stage, data records that do not match the sample labels are identified and removed.

6. The method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation according to claim 1, characterized in that, In step S3, the standardized experimental data includes: Experimental data used to characterize differences in culture conditions, experimental data used to characterize the regeneration and differentiation process, and experimental data used to characterize phenotypic variation; The experimental data are associated and stored with the same sample identifier in a cross-experimental data association structure.

7. The method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation according to claim 1, characterized in that, In step S4, the steps for analyzing and processing the correlation between the experimental data include: Correlation analysis was performed based on culture condition parameters and regeneration and differentiation indicators; Perform multidimensional feature analysis based on variable trait data; Establish a data model to characterize the relationship between culture condition parameters and variant traits.

8. The method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation according to claim 1, characterized in that, In step S4, the analysis and processing step further includes: Clustering of regenerated plants was performed based on data of variable traits; The regenerated plants were categorized based on the clustering results; Assign corresponding category identifiers to different categories of regenerated plants.

9. The method for pooling and analyzing chrysanthemum breeding experimental data based on somatic cell variation according to claim 1, characterized in that, In step S5, the step of outputting data-driven decision information includes: The analysis results are transformed into structured data records corresponding to the sample identifiers; The structured data records are summarized according to experimental batches; Output the summarized structured data results.

10. A data aggregation and analysis system for chrysanthemum breeding experiments based on somatic cell variation, characterized in that, include: The data acquisition module is configured to collect experimental data generated during multiple experimental stages in the chrysanthemum somatic cell mutation breeding experiment, and to configure data identification information for the collected experimental data to uniquely identify the experimental object and the experimental process; The data aggregation and association module is configured to aggregate experimental data generated in multiple experimental stages according to a unified data structure, and establish a data association structure based on the data identification information to characterize the correspondence between the same experimental object in different experimental stages, so as to obtain a cross-stage breeding experimental data set. The data processing module is configured to perform data consistency processing and standardization processing on the breeding experiment data set to obtain standardized experimental data for analysis. The data analysis module is configured to analyze and process the correlation between experimental data based on the standardized experimental data to obtain analysis results; The results output module is configured to output data-driven decision-making information for chrysanthemum somatic cell variation breeding based on the analysis results.