A data integration method and device for accurate identification of crop phenotypes

By building a cloud platform and interconnecting it with data acquisition equipment, and utilizing data acquisition models and blockchain evidence storage technology, we have achieved efficient fusion of multi-source heterogeneous crop phenotypic data. This has solved the problems of data resource waste and inaccurate identification results in traditional methods, and improved breeding efficiency and identification accuracy.

CN121093181BActive Publication Date: 2026-08-04BEIJING PAIDE WEIYE TECH DEV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING PAIDE WEIYE TECH DEV
Filing Date
2025-08-08
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Traditional data processing methods struggle to effectively integrate crop phenotypic data from multiple sources, of different types, at different scales, and with heterogeneity, leading to wasted data resources and inaccurate identification results.

Method used

The cloud platform is interconnected with the data acquisition equipment. Tasks are assigned through the data acquisition model to collect and integrate data. Blockchain is used to ensure data integrity and multimodal data fusion is performed.

Benefits of technology

This improved the accuracy and reliability of crop phenotypic identification, increased breeding efficiency, and ensured the utilization rate of data resources and the reliability of identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121093181B_ABST
    Figure CN121093181B_ABST
Patent Text Reader

Abstract

The application provides a data integration method and device for crop phenotype accurate identification, the method comprises the following steps: constructing a cloud platform for crop phenotype accurate identification, and interconnecting with a data acquisition device through a network; constructing a corresponding data acquisition model for the data acquisition device, obtaining a data acquisition task through the cloud platform and distributing it to the data acquisition device, collecting crop phenotype data and environmental data, collecting and processing crop phenotype data in different data formats, associating structured phenotype data to be integrated with environmental data, and then associating and aligning unstructured phenotype data to be integrated, obtaining multi-modal fusion data of each crop variety. The application integrates multi-source and multi-modal data of crop phenotype, which helps to improve the accuracy and reliability of crop phenotype accurate identification, and has important significance for improving crop breeding efficiency and accelerating the process of new variety breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural intelligent information processing technology, and in particular to a data integration method and apparatus for precise identification of crop phenotypes. Background Technology

[0002] Precise crop phenotyping refers to the high-precision and systematic identification and analysis of crop morphology, growth, quality, and other traits using modern biotechnology, information technology, and engineering technology. With the continuous development of engineered breeding technology, precise crop phenotyping has become one of the important means for crop genetic improvement, selection of high-quality varieties, and research on stress resistance. The accuracy and efficiency of precise crop phenotyping data directly affect the efficiency of breeding and promotion of new crop varieties. With the widespread application of multi-source data acquisition methods such as sensors, drones, and satellite remote sensing, the ways to obtain crop phenotypic and corresponding environmental information have become diversified. Crop growth observation data exhibits characteristics such as multiple sources, multiple types, multiple scales, and heterogeneity. These data differ significantly in terms of time, space, and application, making it difficult for a single data source to comprehensively and accurately characterize the phenotypic features of crops. Different types of data are often distributed across different data sources. How to effectively integrate these data to obtain multi-level and multi-dimensional comprehensive crop characteristics has become a key challenge in precise crop phenotyping technology.

[0003] Traditional data processing methods struggle to meet the demands of integrating massive and diverse crop breeding data, resulting in a waste of data resources and impacting the accuracy and reliability of phenotypic identification results. They also fail to effectively integrate data from multiple sources, of different types, at multiple scales, and exhibiting heterogeneity. By integrating multi-source heterogeneous data, we can overcome data barriers, avoid biases in crop variety identification caused by independent and scattered data, and construct a comprehensive and accurate crop characteristic profile, providing strong data support for precise crop phenotypic identification. Summary of the Invention

[0004] This invention provides a data integration method and apparatus for precise identification of crop phenotypes. By integrating multi-source and multimodal data of crop phenotypes, the accuracy and reliability of precise identification of crop phenotypes can be improved, which is of great significance for improving crop breeding efficiency and accelerating the process of new variety selection.

[0005] This invention provides a data integration method for precise identification of crop phenotypes, including: A cloud platform for precise identification of crop phenotypes is constructed. The cloud platform is interconnected with data acquisition devices via a network. The data acquisition devices include multi-source phenotyping devices for collecting crop phenotyping data and environmental monitoring devices for collecting environmental data. A corresponding data acquisition model is constructed for the data acquisition device. Data acquisition tasks for crop phenotypic data and environmental data are obtained through the cloud platform, and the data acquisition tasks are assigned to the data acquisition device so that the data acquisition device can acquire crop phenotypic data and environmental data according to the acquisition instances in the data acquisition model. Field aggregation processing is performed on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, and the unstructured crop phenotypic data is aggregated through the cloud platform to obtain unstructured phenotypic data to be integrated. The structured phenotypic data to be integrated is integrated with the environmental data collected by the environmental monitoring equipment. The resulting phenotypic-environmental integrated data is then correlated and aligned with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety.

[0006] In some embodiments, after constructing the crop phenotypic precision identification cloud platform, the method further includes: registering the data acquisition device in the cloud platform, wherein the registration process of the data acquisition device includes: After a communication link is established between the data acquisition device and the cloud platform, the data acquisition device is triggered to send a registration request message to the cloud platform and performs two-way authentication with the cloud platform through certificate exchange. The cloud platform parses the registration request message, generates a device access token for the data acquisition device after the registration request is verified, and pushes the configuration parameters for collecting crop phenotypic data and environmental data to the data acquisition device. The cloud platform stores the device registration information obtained after parsing the registration request message, and the data acquisition device stores the device access token and the configuration parameters.

[0007] In some embodiments, constructing a corresponding data acquisition model for the data acquisition device includes: Build a unified interface that adapts to the data acquisition device; Create a data acquisition instance, which includes a data acquisition device identifier, acquisition parameters, and data storage method. The acquisition parameters include the basic interface information of the unified interface, the source data parameters of the data acquisition device, the target data parameters of the crop phenotypic data, and the data transformation rules of the crop phenotypic data. The data acquisition instances are grouped according to the device attributes of the data acquisition equipment and the data acquisition requirements. For each group, set group configuration parameters, and set the collection instances within the group to inherit the group configuration parameters. Then, deploy the collection instances that inherit the group configuration parameters in the unified interface to obtain the data collection model.

[0008] In some embodiments, prior to the field aggregation processing of the structured and semi-structured crop phenotypic data, the method further includes: The on-chain notarization of the collected crop phenotypic data and environmental data is performed according to the data acquisition model. The on-chain notarization process of the crop phenotypic data and environmental data includes: Construct a consortium blockchain network composed of data acquisition instances from the data acquisition model; Generate content identifiers for the crop phenotypic data and the environmental data, and extract key information from the crop phenotypic data and the environmental data. The key information includes: raw data summary, timestamp, data acquisition device identifier, crop variety identifier, and test site identifier. The key information is encapsulated into a blockchain storage operation record with a data signature according to the consortium blockchain network, and the content identifiers of the crop phenotypic data and the environmental data are bound to the original data digest. The data signature is used to verify the identity and legitimacy of the data operator. The blockchain evidence storage operation record information is broadcast to multiple participating nodes in the consortium blockchain network, so that the participating nodes can verify the blockchain evidence storage operation record information according to a preset consensus mechanism and reach a consensus confirmation. The blockchain evidence storage operation record information is packaged into the consortium blockchain network, and on-chain evidence storage information of the crop phenotypic data and the environmental data is generated.

[0009] In some embodiments, field aggregation processing is performed on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, including: The semi-structured crop phenotypic data is converted into target structured data; Determine the target data fields for the phenotypic data to be integrated, and configure the source data fields for each target acquisition instance. The target acquisition instance is the acquisition instance corresponding to each type of multi-source phenotypic device in the data acquisition model. For each target acquisition instance, configure the mapping relationship between the target data field and the source data field; The mapping relationship is parsed to obtain the mapping formula. The actual values ​​of the source data fields in the structured crop phenotypic data and the target structured data are substituted into the mapping formula, and the actual values ​​of the target data fields are calculated to obtain the structured phenotypic data to be integrated.

[0010] In some embodiments, the step of performing data aggregation processing on the unstructured crop phenotypic data through the cloud platform to obtain unstructured phenotypic data to be integrated includes: The unstructured crop phenotypic data is encapsulated into data packets in the data acquisition device, and the data packets are transmitted to the edge node of the cloud platform; The data packets are parsed through the edge nodes, the parsed unstructured data is preprocessed, and the preprocessed data file is associated with the metadata of the unstructured crop phenotypic data to obtain a data collection record file; A multi-level index based on metadata is constructed for the data collection record file, representing unstructured tabular data to be integrated.

[0011] In some embodiments, after performing field aggregation processing on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, the method further includes: For the structured phenotypic data to be integrated, construct postfix expressions for complex traits; The trait variables in the postfix expression and the corresponding trait data of the complex traits in the structured trait data are stored in a stack for complex trait calculation to obtain the complex trait index value of the structured trait data.

[0012] In some embodiments, constructing a postfix expression for a complex trait from the structured phenotypic data to be integrated includes: Based on the structured trait data included in the structured phenotypic data to be integrated, determine the computational expression for calculating the complex traits; Traverse each element of the operational expression from left to right, and perform the following processing on each element: a. If the element is an operand, add the element to the output queue; b. If the element is a left parenthesis operator, then push the element onto the operator stack; c. If the element is a right parenthesis operator, then pop the operators from the operator stack one by one until a left parenthesis operator is encountered, and then store the popped operators into the output queue one by one. d. If the element is the target operator, then perform the following steps: d1. When the operator stack is empty, or the top element of the operator stack is an operator with a left parenthesis, or the operation priority of the target operator is greater than that of the top element, the target operator is pushed onto the operator stack. d2. When the priority of the target operator is less than or equal to the top element of the stack, pop the top element of the stack and add it to the output queue; d3. Repeat step d2 until the conditions for executing step d1 are met; After traversing the last element of the operation expression, if there are still stack elements in the operator stack, the stack elements are popped one by one and added to the output queue. The elements in the output queue are dequeued one by one to obtain a postfix expression of complex state.

[0013] In some embodiments, the process of integrating the structured phenotypic data to be integrated with the environmental data collected by the environmental monitoring equipment includes: According to the preset integration level, the trait data in the structured phenotypic data to be integrated are merged to obtain the corresponding trait merged value. The trait merged value of each crop variety is then associated with the corresponding environmental data collected by the environmental monitoring equipment to obtain phenotypic-environment integrated data. The integration hierarchy includes at least one of the following: single-plant hierarchy, plot hierarchy, experimental hierarchy, single-year multi-point experimental hierarchy, and multi-year multi-point experimental hierarchy; The environmental data includes meteorological and soil data when crop varieties are planted at the test sites.

[0014] In some embodiments, the step of associating and aligning the phenotypic-environment integrated data obtained through integration processing with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety includes: Obtain genotype data for each crop variety; The phenotypic-environment integrated data of each crop variety is associated and integrated with the genotype data to obtain the genotype-phenotype-environment integrated data of each crop variety. Fields were extracted from the genotype-phenotype-environment integrated data to obtain field information, which included crop variety number, experiment number, collection location, and collection time. Metadata is extracted from the unstructured phenotypic data to be integrated to obtain metadata information, which includes crop variety number, experimental plot number, collection location and collection time. The metadata information is associated and matched with the field information to obtain multimodal fusion data for each crop variety. The multimodal fusion data is used for accurate identification of crop phenotypes on the cloud platform.

[0015] The present invention also provides a data integration device for precise identification of crop phenotypes, comprising: A building module is used to build a cloud platform for precise identification of crop phenotypes. The cloud platform is interconnected with data acquisition devices via a network. The data acquisition devices include multi-source phenotyping devices for collecting crop phenotyping data and environmental monitoring devices for collecting environmental data. The allocation module is used to construct a corresponding data acquisition model for the data acquisition device, obtain data acquisition tasks of crop phenotypic data and environmental data through the cloud platform, and allocate the data acquisition tasks to the data acquisition device so that the data acquisition device can acquire crop phenotypic data and environmental data according to the acquisition instances in the data acquisition model. The aggregation module is used to perform field aggregation processing on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, and to perform data aggregation processing on the unstructured crop phenotypic data through the cloud platform to obtain unstructured phenotypic data to be integrated. The integration module is used to integrate the structured phenotypic data to be integrated with the environmental data collected by the environmental monitoring equipment, and to associate and align the phenotypic-environmental integrated data obtained by the integration process with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the data integration method for precise identification of crop phenotypes as described above.

[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data integration method for precise identification of crop phenotypes as described above.

[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the data integration method for precise identification of crop phenotypes as described above.

[0019] The data integration method and apparatus for precise crop phenotypic identification provided by this invention firstly connects data acquisition devices to a cloud platform at the data acquisition level. The cloud platform precisely assigns data acquisition tasks to the data acquisition devices, and the data acquisition models constructed by the devices are used to collect crop phenotypic data and environmental data, laying the foundation for subsequent data aggregation and integration. Secondly, at the data integration level, the crop phenotypic data collected from multiple phenotyping devices are aggregated according to their data structures. Thirdly, at the crop variety level, the structured phenotypic data to be integrated is first integrated with environmental data, and then further correlated and aligned with unstructured data to achieve multimodal data fusion of heterogeneous crop phenotypic data. This multi-source heterogeneous data integration process can meet the fusion needs of massive and diverse multi-source heterogeneous data, improve the utilization rate of data resources, and ensure the accuracy and reliability of phenotypic identification results. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the data integration method for precise identification of crop phenotypes provided by the present invention.

[0022] Figure 2 This is a schematic diagram of the data integration device for precise identification of crop phenotypes provided by the present invention.

[0023] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0025] The data integration method and apparatus for precise crop phenotypic identification of the present invention are described below with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating the data integration method for precise crop phenotypic identification provided by the present invention, as shown below. Figure 1As shown, the method includes the following steps 101 to 104.

[0026] Step 101: Construct a cloud platform for precise identification of crop phenotypes. The cloud platform and data acquisition equipment are interconnected via a network.

[0027] Data acquisition equipment includes multi-source phenotyping devices for collecting crop phenotypic data and environmental monitoring equipment for collecting environmental data. Specifically, data acquisition equipment includes various phenotyping devices and software tools, such as high-throughput plant phenotyping platforms, portable phenotyping acquisition devices, field data processing software, phenotyping imaging analysis systems, various sensors, and meteorological observation instruments.

[0028] A cloud platform is a human-computer interactive platform software system. Data processed by the cloud platform is used for precise identification of crop phenotypes. The cloud platform is used to manage various phenotype devices and to collect, parse, process, fuse, and store precise identification data of crop phenotypes from different sources, with different structures, and at different scales. Its main functions include: device access management, analytical model management, data acquisition, data parsing, data processing, data storage, data query, statistical analysis, information visualization, operation monitoring, and decision support.

[0029] Network interconnection between the cloud platform and data acquisition devices refers to data transmission between various data acquisition devices and the cloud platform via the internet or wireless LAN. The communication media or technologies used for data transmission include: Wi-Fi, cellular networks (4G / 5G), and network cables.

[0030] In some embodiments, in order to bind the data acquisition device to the cloud platform and facilitate the control of crop phenotypic data and environmental data acquisition and the management of multi-source phenotypic devices through the cloud platform, after building the crop phenotypic precision identification cloud platform, it is necessary to register the data acquisition device on the cloud platform, which is explained in detail below.

[0031] Generally, data acquisition devices are automatically registered on the cloud platform. First, the data acquisition device automatically scans the preset network configuration (such as Wi-Fi SSID and password) and uses the MQTT over TLS protocol to establish a communication link with the cloud platform.

[0032] After establishing a communication link between the data acquisition device and the cloud platform, the data acquisition device sends a registration request message to the cloud platform and performs two-way authentication with the cloud platform through certificate exchange. Here, the registration request message uses a private key to digitally sign the message digest, ensuring message integrity and credible origin. The registration request message specifically includes: the device ID number (e.g., IMEI number) of the multi-source phenotypic device, the device certificate, and metadata (e.g., device model, firmware version). The device certificate contains a public key, a private key, and a digital signature.

[0033] The data acquisition device establishes an encrypted channel with the cloud platform via the TLS 1.3 protocol, exchanges certificates, and completes mutual authentication. Specifically, the cloud platform verifies the device's certificate chain and validity, while the phenotypic device verifies the cloud platform's domain name certificate and root certificate to prevent third-party attacks.

[0034] Next, the cloud platform parses the registration request message. After successful registration verification, a device access token is generated for the data acquisition device, and configuration parameters for collecting crop phenotypic data are pushed to the data acquisition device. When parsing the registration request from the data acquisition device, the cloud platform verifies the uniqueness of the device ID, the validity of the certificate, and the validity of the signature. Upon successful verification, a device access token (such as a JSON Web Token, JWT) is generated, bound to the phenotypic device ID and permission policy, and the data access scope is restricted. Then, the cloud platform pushes configuration parameters encapsulated in JSON format (such as data acquisition frequency and transmission protocol) to the data acquisition device based on the device type. The data acquisition device directly parses the configuration parameters and activates the service.

[0035] Finally, the device registration information obtained after parsing the registration request message is stored in the cloud platform, and the device access token and configuration parameters are stored in the data acquisition device. The cloud platform stores the device registration information (including device ID, device access token, and configuration version) in a distributed database and updates the device status of the data acquisition device to "online". The data acquisition device stores the device access token and configuration parameters and establishes a heartbeat mechanism (e.g., sending a heartbeat every 60 seconds) to maintain a long-term connection with the cloud platform.

[0036] An anomaly handling and fault tolerance mechanism is also established between the data acquisition device and the cloud platform. If the data acquisition device fails to register (such as due to network interruption or certificate verification failure), the multi-source phenotypic device will retry according to the exponential backoff algorithm (initial interval of 1 second, maximum 10 retries). The cloud platform provides a manual intervention interface to support batch import of device certificates to deal with extreme scenarios.

[0037] In addition, in some embodiments, data acquisition devices can be manually registered on the cloud platform. For trusted data acquisition devices or those without automatic registration functions, basic information and parameter configuration information of the data acquisition devices can be manually entered or imported through the registration interface provided by the cloud platform.

[0038] In this embodiment of the invention, after constructing a cloud platform for precise identification of crop phenotypes, data acquisition devices are registered in the cloud platform to achieve configuration connection between the cloud platform and the data acquisition devices. Furthermore, through two-way authentication of certificates, the security of the transmission of collected crop phenotype data and environmental data is ensured.

[0039] Step 102: Construct a corresponding data acquisition model for the data acquisition device, obtain crop phenotypic data and environmental data acquisition tasks through the cloud platform, and assign the data acquisition tasks to the data acquisition device so that the data acquisition device can collect crop phenotypic data and environmental data according to the acquisition instances in the data acquisition model.

[0040] To collect crop phenotypic and environmental data, a corresponding data acquisition model is constructed for the data acquisition devices. It is important to note that the data acquisition devices include multiple devices, and the data acquisition model needs to be compatible with multiple devices simultaneously. The construction process of the data acquisition model is explained in detail below.

[0041] First, a unified interface is constructed to adapt to data acquisition devices. For phenotypic devices from different manufacturers and using different communication protocols, a highly abstract interface mechanism is built to achieve data acquisition device adaptation. Specifically, various data acquisition devices are abstracted into an abstract interface with unified commands or attributes. This unified interface is completely independent of the specific type, manufacturer, and actual deployment location of the data acquisition device. Simultaneously, this unified interface, through serialization and reflection mechanisms, allows crop phenotypic data and environmental data collected by various data acquisition devices to be consistently accessed and exchanged.

[0042] Next, we need to create collection instances. A collection instance is a specific and operable implementation of crop phenotypic data collection, created according to the rules and requirements of the data collection model framework and the actual application scenario. In other words, it is the concrete implementation of the data collection model in a specific environment, containing detailed information such as data collection device identification (e.g., ID), collection parameters, and data storage methods. When collecting crop phenotypic data and environmental data, each data collection device collects data according to the collection instance format.

[0043] The data collection instance includes the data collection device identifier, collection parameters, and data storage method. The collection parameters include the basic information of the unified interface, the source data parameters of the data collection device, the target data parameters of crop phenotypic data and environmental data, and data transformation rules.

[0044] Basic interface information describes the data acquisition device, including: device ID, device status, MAC address, device serial number, device IP address and port number, device type, device manufacturer, device model, target system IP address and port number, acquisition mode (push / pull), communication protocol, security authentication information (username, password, token, certificate, device key), and exception handling rules.

[0045] The source data parameters of the data acquisition devices are used to describe the specific parameters of the source data collected from each data acquisition device, including: source data (data item name, data type, data unit, data precision), data format, encoding method, source database name, source data table name, data file format, data generation time, data status, data tag, encryption algorithm and key.

[0046] The target data parameters for crop phenotypic data and environmental data describe the storage and representation of the collected crop phenotypic data and environmental data after processing. These parameters include: target data (data item name, data type, data unit, data precision), data format, encoding method, target database name, target data table name, data file format, file path, and file naming rules.

[0047] The data transformation rules for crop phenotypic data and environmental data describe the standards, methods, and steps that should be followed when converting source data collected by data acquisition equipment into target data, including: field mapping rules, data format conversion rules, and data type conversion rules.

[0048] The data format can be any one or a combination of structured, semi-structured, and unstructured data. Structured data consists of several fields and their values, and is represented, stored, and transmitted in the form of JSON, XML, CSV, YAML, key-value pairs, or two-dimensional tables. Unstructured data is represented, stored, and transmitted in the form of data files, including formats such as RGB images, fluorescence images, hyperspectral data, multispectral data, audio, video, and point clouds. The data type can be numeric, text, date, etc.

[0049] Next, the data acquisition instances are grouped according to the device attributes of the data acquisition equipment and the data acquisition requirements. Here, to achieve the function of defining model parameters once and accessing them in batches, this embodiment of the invention implements a grouping strategy for the created acquisition instances. Specifically, based on the device attributes of the data acquisition equipment and the specific data acquisition requirements for crop phenotypic data and environmental data, acquisition instances meeting the same parameter configuration conditions are grouped into the same group, while acquisition instances with different parameter configuration conditions are grouped into different groups. Device attributes can be the hardware configuration, functional characteristics, or application scenarios of the data acquisition equipment. For example, high-throughput plant phenotyping platforms and portable phenotyping equipment both belong to application scenarios and can be grouped into one group, while temperature sensors and image sensors both belong to sensors and can be grouped into one group, and field data processing software and phenotypic imaging analysis systems both belong to software configurations and can be grouped into one group. Data acquisition requirements can be data format and acquisition strategy. For example, those acquiring structured data belong to the same data format and can be grouped into one group, while those using the same data acquisition strategy can also be grouped into one group.

[0050] Next, group configuration parameters are set for each group, and the collection instances within the group are configured to inherit these group configuration parameters to obtain the data collection model. Specifically, a centralized model parameter definition operation is performed for each parameter configuration group. Each collection instance within the group automatically inherits these group configuration parameters, thereby achieving batch parameter deployment of collection instances through a single parameter configuration, eliminating the need to configure parameters for each collection instance within the group individually.

[0051] Preferably, the data collection instances can also be grouped based on the data collection device type, manufacturer, model, region, and user. This can be set according to actual needs to obtain as few groups as possible.

[0052] After setting group configuration parameters for each group, the collection instances that inherit the group configuration parameters are deployed on a unified interface to obtain the data collection model. Here, the collection instances with parameter configuration for each group are imported into the constructed unified interface, thereby obtaining the constructed data collection model.

[0053] This invention provides a unified configuration interface for multiple data acquisition devices. It then designs data instances for a model and implements a grouping strategy to configure parameters for these data instances. This allows for batch deployment of model instances through a single parameter configuration, significantly improving configuration efficiency and reducing operational complexity. Finally, the configured data instances are imported into a unified interface, enabling multiple data acquisition devices to use a unified data acquisition model to collect multi-source, heterogeneous crop phenotypic data, reducing data acquisition difficulty and meeting diverse data acquisition needs.

[0054] Furthermore, data acquisition tasks for crop phenotypic data and environmental data are obtained through the cloud platform, and these tasks are assigned to data acquisition devices so that the devices can collect crop phenotypic data and environmental data based on the acquisition instances in the data acquisition model.

[0055] Here, data acquisition tasks are generated based on preset data acquisition schemes, which are manually set according to specific business needs. A data acquisition scheme includes: scheme number, scheme name, scheme description, acquisition target, acquisition scope, start and end time, acquisition frequency, acquisition mode, data items to be acquired, data format, and data type. These can be manually set one by one in the cloud platform, which then automatically generates corresponding data acquisition tasks for these data acquisition schemes.

[0056] For example, if we want to collect data on two traits (plant height and ear position) from four maize planting resource performance identification trials (TR001 and TR002) in a certain region (Region A) in 2024, we can formulate the data collection plan shown in Table 1 below: Table 1:

[0057] Data acquisition tasks include: task identifier (ID), associated data acquisition scheme, device ID of the data acquisition equipment, task status, acquisition range, start time, triggering conditions, data format, and data type. The acquisition range can be determined after filtering based on one or more conditions such as administrative region, experimental site, ecological region, crop variety, planting plot, greenhouse, plot, single plant, and crop trait. Data types include: trait data, meteorological data, soil environmental data, physiological parameter data, image data, video data, spectral data, and point cloud data. Task status includes not executed or executed. The start time indicates when the data acquisition equipment begins collecting crop phenotypic data. Triggering conditions indicate the conditions under which the data acquisition equipment begins collecting crop phenotypic and environmental data, such as timed triggering. Data formats include structured data, semi-structured data, and unstructured data.

[0058] Therefore, based on the data acquisition scheme in Table 1 above, the corresponding data acquisition tasks can be automatically generated in the cloud platform, as shown in Table 2 below: Table 2:

[0059] Next, the cloud platform will allocate the generated data acquisition tasks to various data acquisition devices to collect crop phenotypic data. Specifically, all newly generated data acquisition tasks are first grouped according to the device ID number of the data acquisition device, resulting in data acquisition tasks generated by each data acquisition device. Then, the newly generated data acquisition tasks of each data acquisition device are added to the acquisition task list of the corresponding data acquisition device. The elements in the acquisition task list of each data acquisition device are sorted according to the start time of the data acquisition tasks. Finally, using the scheduled task trigger configured on the cloud platform, the data acquisition tasks to be executed at the current time are obtained, and data acquisition instructions are sent to the relevant data acquisition devices. After receiving the data acquisition instructions, the data acquisition devices can select the corresponding data acquisition task from the acquisition task list, and then collect crop phenotypic data and environmental data according to the acquisition instance in the data acquisition model.

[0060] The task execution interval of the scheduled task trigger can be defined according to business needs, typically set to 1 second or 1 minute. Data acquisition commands include two types: active push and passive pull. This means the data acquisition device actively pushes data to the cloud platform, while the cloud platform directly pulls data from the data acquisition device.

[0061] Step 103: Perform field aggregation processing on structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, and perform data aggregation processing on unstructured crop phenotypic data through the cloud platform to obtain unstructured phenotypic data to be integrated.

[0062] After collecting crop phenotypic and environmental data in step 102, since the data is collected from various data acquisition devices with different data structures and sources, data aggregation processing is required to achieve multi-source data fusion and obtain the corresponding structured and unstructured phenotypic data to be integrated.

[0063] However, in some embodiments, the collected crop phenotypic data needs to be stored on-chain according to the data collection model before performing data aggregation processing.

[0064] To ensure the integrity, immutability, and traceability of the collected data, this embodiment of the invention performs on-chain notarization of the metadata fingerprints and key feature values ​​of the collected crop phenotypic data and environmental data before performing data aggregation processing. The immutable distributed ledger characteristics of blockchain are used to perform data storage. On-chain notarization is completed on a cloud platform. The on-chain notarization process of crop phenotypic data and environmental data is described in detail below.

[0065] First, a consortium blockchain network consisting of data acquisition instances from the data acquisition model is constructed, using the Hyperledger Fabric framework. Then, content identifiers (CIDs) are generated for crop phenotypic data and environmental data. These CIDs are unique. Furthermore, key information from the crop phenotypic and environmental data is extracted, including: raw data digest, timestamp, data acquisition device identifier (ID), crop variety identifier (ID), and test site identifier (ID). The raw data digest is a fixed-length hash value calculated using a message digest algorithm on the crop phenotypic and environmental data from the data acquisition device. Message digest algorithms include SM3, MD5, and the SHA series.

[0066] Furthermore, the key information is encapsulated into a blockchain-based evidence storage operation record with a data signature by the consortium blockchain network. The encapsulation method can employ a lightweight smart contract based on the consortium blockchain architecture. The data signature is used to verify the identity and legitimacy of the data operator, and is specifically generated by authorized nodes or devices in the consortium blockchain network using their private keys.

[0067] In addition, the content identifiers of crop phenotypic data and environmental data are bound to the original data summary. The binding can still be implemented using lightweight smart contracts to ensure that off-chain data cannot be tampered with.

[0068] Next, the blockchain evidence storage operation record information is broadcast to multiple participating nodes in the consortium blockchain network, so that the participating nodes can verify the blockchain evidence storage operation record information according to the preset consensus mechanism and reach a consensus confirmation.

[0069] After consensus confirmation, the blockchain evidence storage operation record information is packaged into the consortium blockchain network and linked to the preceding blocks in the consortium blockchain through an encryption algorithm, forming an irreversible chain structure. This achieves permanent storage and tamper-proof protection of the data evidence storage information. Once the blockchain evidence storage operation record information is successfully written to the consortium blockchain, on-chain evidence storage information for crop phenotypic data and environmental data is generated. This on-chain evidence storage information can be automatically generated by the cloud platform system and specifically includes information such as the hash value of the evidence storage operation, block height, and evidence storage time. The on-chain evidence storage information is returned to the data operator or relevant authorized access entity, and an on-chain and off-chain phenotypic data mapping relationship is established for subsequent querying and verification.

[0070] This invention utilizes the immutable distributed ledger characteristics of blockchain to store the content identifiers and key information of crop phenotypic data collected by data acquisition devices on the blockchain, thereby realizing a trusted guarantee system for crop phenotypic data collection and ensuring the integrity, immutability, and traceability of the collected crop phenotypic and environmental data.

[0071] Next, data aggregation processing is performed. This mainly targets the collected crop phenotypic data. Different data aggregation processing strategies are formulated based on the different data formats of the crop phenotypic data. In other words, different aggregation processing strategies are selected to perform aggregation processing based on the first type of structured data, semi-structured data, and unstructured data. The details are explained below.

[0072] On the one hand, field aggregation processing is performed on structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated.

[0073] Specifically, when crop phenotypic data is in the form of first-order structured data and semi-structured data, the semi-structured crop phenotypic data is converted into target structured data. The conversion method can use a parser, such as a JSON parser, XML parser, or CSV parser.

[0074] Next, determine the target data fields for the phenotypic data to be integrated, and then configure the source data fields for each target acquisition instance. Here, the target acquisition instance is the acquisition instance corresponding to each type of multi-source phenotypic device set in the data acquisition model.

[0075] Here, the target data fields to be retained or generated after data aggregation and processing need to be determined in advance. These target data fields can be divided into two categories: identifier fields and trait fields. Identifier fields are relatively fixed and can be predefined by experts, including one or more of the following: field plot number, crop variety ID, test site ID, collection date, field row number, and equipment number. Trait fields can be dynamically added or removed by the user, such as: plant height, ear height, number of leaves, leaf color, number of male spike branches, silking date of female ear, pollen shedding date, and lodging rate. Then, based on equipment attributes such as equipment type, manufacturer, model, region, and user, the multi-source phenotyping equipment is classified, thereby obtaining the collection instance corresponding to each multi-source phenotyping equipment. Next, the source data fields are configured for the collection instances. Specifically, based on the type of multi-source phenotyping equipment, collection instances of each type of multi-source phenotyping equipment can be retrieved from the data collection model, and the corresponding data fields can be configured.

[0076] Furthermore, for each target acquisition instance, a mapping relationship between the target data field and the source data field is configured. The mapping formula supports custom functions and complex arithmetic expressions with multiple levels of nested parentheses. Parentheses appear in pairs in the mapping formula and can be nested an unlimited number of times. Specifically, the mapping relationship between the source data field and the target data field can be one-to-one (1:1) or many-to-one (N:1). When the mapping relationship is N:1, it means that the value of the target data field is obtained by processing the values ​​of multiple source data fields.

[0077] For example, for a multi-source phenotyping device at a certain test site, the mapping relationship shown in Table 1 below can be established. Here, AVG and MAX are user-defined mean and maximum functions, used to calculate the average and maximum values ​​of a set of data, respectively. The mapping formulas between the source data fields and the target data fields are shown in Table 3 below: Table 3:

[0078] Finally, the mapping relationship is parsed to obtain the mapping formula. The actual values ​​of the source data fields in the structured crop phenotypic data and the target structured data are substituted into the mapping formula, and the actual values ​​of the target data fields are calculated to obtain the structured phenotypic data to be integrated.

[0079] By parsing the mapping relationship to form the corresponding mapping formula, which can be a specific data calculation formula, and then substituting the actual values ​​of the source data fields in the structured crop phenotypic data and the target structured data into the mapping formula, the actual values ​​of the corresponding target data fields can be obtained. This achieves the data integration of the target data fields and obtains the structured phenotypic data to be integrated.

[0080] For example, using the mapping relationships in Table 3 above, the value of "tjbh" in the source data field can be mapped to "field plot number" in the target data field, the value of "pzbh" in the source data field can be mapped to "variety number" in the target data field, and the value of "zhugao" in the source data field can be mapped to "plant height" in the target data field. For target data fields with a mapping relationship of N:1, such as stem rot (stemRot) and leaf blight (leaf Blight), the mapping formula (including the mean function and the maximization function) can be obtained by parsing the mapping relationship, and then calculated based on the source data fields sr1, sr2, sr3 and lb1, lb2, lb3.

[0081] In this embodiment of the invention, when crop phenotypic data consists of first structured data and semi-structured data, the corresponding data fields are aggregated and processed. The semi-structured data is uniformly converted into target structured data. Then, the multiple source data fields of the structured crop phenotypic data are integrated into a unified target data field by calculating the mapping formula of the target data fields. This reduces the number of data fields and performs data simplification processing on the data. Thus, multi-source data integration is achieved for crop phenotypic data collected from multiple types of multi-source phenotypic devices.

[0082] On the other hand, unstructured crop phenotypic data is aggregated and processed through a cloud platform to obtain unstructured phenotypic data to be integrated.

[0083] Specifically, when crop phenotypic data is unstructured, a cloud platform is needed to preprocess, integrate, and store the unstructured data. The following describes the data collection and processing process for unstructured crop phenotypic data.

[0084] First, unstructured data is encapsulated into data packets in the data acquisition device, and then the data packets are transmitted to the edge node of the cloud platform.

[0085] Before being packaged into data packets, the data undergoes initial preprocessing at the data acquisition device. Unstructured crop phenotypic data comes in numerous and varied forms, including text, image, audio, video, and other specific format data files. Different types of unstructured data require different preprocessing steps. The cloud platform has a data preprocessing module that first validates the acquired crop phenotypic data to ensure its integrity and format correctness. For example, it checks whether the resolution of image data meets requirements and whether the encoding format of text data is consistent. Next, it performs format conversion, such as converting image data acquired from different devices to JPEG or PNG format and converting text data to UTF-8 encoding. Simultaneously, it extracts metadata from the multi-source phenotypic devices, such as device ID, acquisition time, and geographical location, and packages this information along with the crop phenotypic data into a standard data packet. Finally, the data packet is transmitted to the cloud platform's edge nodes via a secure channel (such as HTTPS), providing standardized input for subsequent processing.

[0086] Next, the data packets are parsed through edge nodes, the parsed unstructured data is preprocessed, and the preprocessed data is associated with the metadata of the unstructured data to obtain the data aggregation record.

[0087] Specifically, several edge nodes can be set up on the cloud platform to receive data packets from different multi-source phenotypic devices and parse the data packets.

[0088] During the parsing process, data is automatically routed to different processing pipelines based on its data type (text, image, audio, video, etc.) for further preprocessing. For example, text data is sent to the natural language processing pipeline to extract key information and perform preliminary classification, while image data enters the computer vision processing pipeline for preprocessing operations such as denoising and enhancement. Simultaneously, the processed data files are associated with the metadata of unstructured crop phenotypic data (i.e., image data, text data, etc. before preprocessing on the cloud platform) to form a data aggregation record file, which is then stored in a temporary storage area (such as the in-memory database Redis) for further processing and analysis.

[0089] Finally, a multi-level index of metadata is constructed for the data aggregation record files to store them. Preferably, different types of storage systems are selected based on data type and access patterns. High-performance distributed file systems are used for frequently accessed data, while low-cost cloud storage services are used for cold data. For frequently accessed hot data (such as recently acquired crop images), a high-performance distributed file system (such as Hadoop HDFS) is selected for storage to ensure fast access and analysis. For infrequently accessed cold data (such as historical audio data), lower-cost cloud storage services are selected to reduce storage costs.

[0090] Furthermore, to facilitate querying and retrieval, a multi-level index based on metadata is constructed for the stored data aggregation records. Frequently queried data is cached using the elastic computing resources of the cloud platform to improve query response speed. Preferably, the device ID, collection time, test point ID, variety ID, and trait ID of the multi-source phenotypic devices are selected as the core index dimensions, employing a hash-tree hybrid encoding mechanism. Specifically, the device ID is a hash partitioned index, physically isolated by device to reduce cross-device query interference; the collection time is a time-series R-tree index, supporting segmentation by year-month-day three-level timestamps. For example, for time-series data (such as crop growth data collected in chronological order), a partitioned index is built by device ID and time range to quickly locate data within a specific time period. For text data, a keyword inverted index is constructed to support full-text search. In addition, the elastic computing resources of the cloud platform are used to cache frequently queried data (e.g., using Redis caching), storing hot data in memory to reduce disk I / O operations and significantly improve query response speed.

[0091] In this embodiment of the invention, unstructured data in crop phenotypic data is preprocessed and stored on a cloud platform, which improves data quality and reduces data storage size. In addition, by constructing data aggregation record files and multi-level indexes, the query response speed can be effectively improved.

[0092] Therefore, for structured and semi-structured crop phenotypic data, the actual values ​​of the target data fields are obtained by integrating the target data fields. For unstructured crop phenotypic data, data encapsulation, preprocessing, and index storage are used to form a data collection record file. This achieves the collection and processing of crop phenotypic data, effectively improves data quality, and lays the foundation for further integration of heterogeneous data.

[0093] In some embodiments, after performing field aggregation processing on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, the present invention embodiments further perform complexity trait calculations on the structured trait data included in the structured phenotypic data to be integrated.

[0094] The structured phenotypic data to be integrated includes structured trait data, where the actual values ​​in the target data fields involve various crop traits, such as plant height, ear height, number of leaves, leaf color, number of male spike branches, silking date of female ear, pollen shedding date, and lodging rate. Therefore, this trait data also needs to be aggregated, requiring the execution of corresponding trait calculations. The calculation process for the trait data is explained in detail below.

[0095] First, for the structured trait data included in the structured phenotypic data to be integrated, postfix expressions for complex traits are constructed. The calculation of complex traits refers to calculating the value of the complex trait based on the values ​​of several other trait indicators, using basic operational rules such as addition, subtraction, multiplication, and division. In this process, the operational expressions support the use of multiple levels of nested parentheses to precisely define the order and logical relationship of different operations, ensuring the accuracy and reliability of the calculation results.

[0096] For example: index values ​​of complex traits of a certain crop variety It is based on the trait values ​​of the other 10 traits. The calculated formula can be expressed as follows: .

[0097] Preferably, to facilitate the parsing and program implementation of the complex state index calculation formula and improve computational efficiency and speed, the calculation formula can be converted into a postfix expression. For example, the above calculation formula, after being converted into a postfix expression, can be represented as: After the above transformation, the multi-layered nested parentheses in the original calculation formula can be eliminated. That is, the computer does not need to deal with complex parenthesis nesting and priority rules, and can directly scan in sequence to perform the calculation.

[0098] The following section details the implementation process of constructing postfix expressions with complex states.

[0099] First, based on the structured phenotypic data to be integrated, the computational expressions for calculating complex traits are determined, for example... .

[0100] Next, iterate through each element of the expression from left to right and perform the following processing on each element: a) If the element is an operand, add it to the output queue. An operator stack S1 and an output queue Q1 are created to store these elements, which include operators and variables. Operators can be numbers or variables, while operators are computational operators such as addition, subtraction, multiplication, and division. The operator stack S1 stores operators, and the output queue Q1 stores the elements of the transformed postfix expression. If the current element being iterated is an operand (number or variable), it is directly added to the output queue Q1.

[0101] b. If the element is a left parenthesis operator, push the element onto the operator stack. The operator is stored in the operator stack S1. If the current element being traversed is a left parenthesis "(", push it onto the operator stack S1.

[0102] c. If the element is a right parenthesis operator, then pop the operators from the operator stack one by one until a left parenthesis operator is encountered, and store the popped operators into the output queue. Here, if the current element being traversed is a right parenthesis ")", then pop the operators from the operator stack S1 one by one and add them to the output queue Q1, until a left parenthesis "(" is encountered, at which point the left parenthesis will also be popped from the operator stack, but will not be output to queue Q1.

[0103] d. If the element is the target operator, then perform the following steps: d1. When the operator stack is empty, or the top element of the operator stack is an operator with a left parenthesis, or the target operator has a higher precedence than the top element of the stack, push the target operator onto the operator stack.

[0104] d2. When the priority of the target operator is greater than that of the top element of the stack, pop the top element of the stack and add it to the output queue.

[0105] Specifically, if the current element being iterated is an operator (such as "+", "-", "*", " / ", "^", etc.), then the following cases are handled: First, if the operator stack S1 is empty, or the top element of operator stack S1 is a left parenthesis "(", or the current operator has a higher priority than the operator at the top of the stack, then push the current operator onto operator stack S1. Second, if the current operator has a priority equal to or lower than the operator at the top of the stack, pop the operator from the stack and add it to output queue Q1. The second case generally needs to be repeated until the first case is satisfied, at which point the current operator can be pushed onto operator stack S1.

[0106] After traversing the last element of the expression, if there are still stack elements in the operator stack, pop the stack elements one by one and add them to the output queue; dequeue the elements in the output queue one by one to obtain the postfix expression of complex form.

[0107] Finally, the remaining operators are processed. After traversing the last element of the expression, the formula parsing is complete. If there are still operators in the operator stack S1, all stack elements are popped and added to the output queue Q1. Finally, all elements in the output queue Q1 are dequeued in sequence, and the dequeued elements are concatenated according to their dequeue order to obtain the complex postfix expression.

[0108] In this embodiment of the invention, by using two different data structures, operator stack S1 and output queue Q1, the "last in first out" and "first in first out" access strategies are executed respectively, converting the computational expression of the complexity state into a postfix expression. This facilitates the parsing and program implementation of the complexity state index calculation formula, thereby improving computational efficiency and speed.

[0109] Next, the trait variables in the postfix expression and the corresponding trait data of complex traits in the structured trait data are stored in a stack for mathematical operations to obtain the complex trait index values ​​of the structured trait data.

[0110] Specifically, mathematical operations can be performed through the following steps: (1) Establish a variable mapping table (dictionary or hash table structure) to store the mapping relationship between variable names and their corresponding attribute values ​​in postfix expressions. For example: {"t1":50, "t2":10}.

[0111] (2) Traverse each element in the postfix expression. If the current element is a variable, obtain the trait value corresponding to the variable in the variable mapping table and replace the element with the trait value corresponding to the variable.

[0112] (3) Create an empty stack S2 to store operands.

[0113] (4) Iterate through each element of the postfix expression from left to right again, and perform the following operations depending on the case: If the current element is a number, push it directly onto stack S2; If the current element is an operator, first check if there are at least two operands in stack S2 (if not, an error is reported or the calculation is terminated because the expression is invalid); then, pop the top operand from stack S2 and denot it as "operand 2"; next, pop the top operand from stack S2 and denot it as "operand 1"; according to the current operator, perform the corresponding operation (such as addition, subtraction, multiplication, division, etc.) on "operand 1" and "operand 2", obtain the operation result, and push the operation result onto stack S2.

[0114] (5) Repeat the above traversal and processing steps until all elements of the postfix expression have been processed. At this point, the only element remaining in stack S2 is the calculation result of the postfix expression, which is also the complexity index value of the structured phenotypic data.

[0115] In this embodiment of the invention, after the structured and semi-structured crop phenotypic data are processed by field aggregation, the complex trait calculation is performed on the structured phenotypic data to be integrated. This can effectively simplify the expression of the structured trait data and facilitate crop phenotypic identification.

[0116] Step 104: Integrate the structured phenotypic data to be integrated with the environmental data collected by the environmental monitoring equipment, and then align the phenotypic-environmental integrated data obtained from the integration with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety.

[0117] Here, after performing complexity calculations on the structured phenotypic data to be integrated, it is then integrated with environmental data collected by environmental monitoring equipment. The specific process of integration is explained below.

[0118] First, according to the preset integration level, the trait data in the structured phenotypic data to be integrated are merged to obtain the corresponding trait merged value. The integration level includes at least one of the following: single plant level, plot level, experimental level, single-year multi-point experimental level, and multi-year multi-point experimental level.

[0119] Merging data involves using different data processing strategies for different traits. These strategies include taking the latest value, average value, maximum value, minimum value, cumulative value, statistical value, and interval value. Crop traits include plant height, ear height, number of leaves, leaf color, number of male ear branches, silking date of female ear, pollen shedding date, and lodging rate. In merging data, the corresponding trait values ​​are directly merged. Since the indicators for the complex traits have already been calculated, the trait data for the corresponding traits can be directly merged here.

[0120] At the individual plant level, multiple observation records for each trait of each individual plant in the planting experiment are merged and processed into a single value to obtain the value of each trait after treatment. For example, for the traits "plant height" and "ear height" of a certain maize variety, the average value of the corresponding multiple trait observation records is taken as the value of "plant height" and "ear height" of the individual plant after treatment for that variety.

[0121] At the plot level, multiple observation records of a certain trait in a specific plot of the planting experiment are merged into a single value to obtain the plot-processed trait value. For example, in each plot, the average of multiple observation records corresponding to maize plant height is taken as the plot-processed trait value.

[0122] For each experimental level, the trait values ​​of each crop variety after treatment in multiple plots are combined into a single value, resulting in the trait value for that experimental level. For example, if a crop variety V has three plots (A, B, and C) in a planting experiment, when integrating the "plant height" trait at the experimental level, the "plant height" trait values ​​after treatment in the three plots (A, B, and C) can be combined. For instance, the average "plant height" value of the three plots (A, B, and C) can be taken as the "plant height" value for the experimental level of crop variety V.

[0123] For single-year multi-site trials, the trait values ​​of a crop variety after variety treatment at multiple test sites in a specific test year are combined into a single value to obtain the trait value of that crop variety after multi-site trial treatment in that year. For example, the yield values ​​of a crop variety after plot treatment at all test sites in 2024 are combined to obtain the yield value of that crop variety after "single-year multi-site" treatment in 2024.

[0124] For multi-year, multi-location trials, the trait observations of a crop variety after treatment in plots at different trial sites across multiple trial years are integrated and calculated. These values ​​are then combined and transformed into a single trait value that comprehensively reflects the characteristics of the multi-year, multi-location trials, thereby obtaining the comprehensive performance value of the corresponding crop variety under multi-year, multi-location trial conditions. For example, in 2007 and 2008, the trait values ​​of a maize variety after treatment in each plot at each trial site for these two years are combined (e.g., averaged). This can reflect the comprehensive trait performance of maize in different years and at different trial sites.

[0125] After merging the above trait data, we can obtain the trait values ​​of each variety at different levels, such as single plant, plot, trial, single-year multi-point trial, and multi-year multi-point trial.

[0126] Furthermore, the combined phenotypic values ​​of each crop variety are correlated with corresponding environmental data collected by environmental monitoring equipment to obtain integrated phenotypic-environmental data. Here, it is considered that crop phenotype is the result of the combined effects of genes and environment, and environmental factors (such as temperature, light, and soil conditions) have a significant impact on crop phenotype. For example, the same wheat variety may show significant differences in traits such as plant height and tiller number under drought and humid conditions. Integrating phenotypic data with environmental data can eliminate environmental interference, more accurately reflect the genetic characteristics of the crop itself, and improve the accuracy of crop phenotypic identification. Environmental data includes meteorological and soil data from the crop varieties during planting trials at the experimental sites. Secondly, it can identify which genes are expressed under specific environments and how the environment affects gene expression, thus helping breeders screen for crop varieties that perform well under specific conditions. For example, in breeding in arid regions, combining crop drought-resistance phenotypic data with local precipitation, soil moisture, and other environmental data can quickly locate drought-resistance gene resources, accelerate the breeding process, and cultivate superior varieties adapted to specific environments.

[0127] The following describes the process of integrating trait data (shape merged values) with corresponding environmental data. The environmental data used here includes meteorological and soil data.

[0128] (1) Obtain daily meteorological data for each test site, as well as soil data for the area where each test site is located. Meteorological data include factors such as maximum temperature, minimum temperature, average temperature, average wind speed, wind angle, maximum wind speed, surface air pressure, sunshine duration, wind force level, relative humidity, and precipitation. These can be obtained from professional meteorological websites, third-party databases, or collected from meteorological observation records at the test sites.

[0129] Soil data includes factors such as pH value, soil organic matter, total nitrogen, available phosphorus, and available potassium, which can be obtained from data sources such as the soil testing and fertilizer recommendation data management platform.

[0130] Optionally, when a certain meteorological or soil factor index value is missing at a certain test site, a spatial interpolation method can be used to complete it. Spatial interpolation methods include any one or more of the following: Kriging interpolation, inverse distance weighted interpolation, natural neighbor interpolation, and nearest neighbor interpolation.

[0131] (2) Using the test site ID as the association identifier, the meteorological data and soil data of each test site are associated and aligned to obtain the environmental data of each test site.

[0132] (3) Calculate the growth period of each variety in different experimental years and at different experimental sites, and further calculate the key performance indicators of each crop variety in different experimental years and at different experimental sites based on the growth period of the varieties. Specifically, the calculation method of crop variety growth period includes: grouping the trait data (combined trait values) by experimental year, crop variety ID, experimental site ID, and trait ID, selecting the trait data with the trait names "sowing period" and "maturity period", and calculating the growth period of each variety in different experimental years and at different experimental sites based on the trait data with the trait names "sowing period" and "maturity period".

[0133] It should be noted that the varietal growth period refers to the number of natural days from sowing to maturity. Key performance indicators include: average daily growth, leaf area index growth rate, plant height growth rate, photosynthetic productivity, and dry matter accumulation rate.

[0134] (4) Based on the test year and crop variety ID, query the phenotypic identification test that each crop variety participated in each year and obtain the test site ID; then, based on the test site ID, associate the phenotypic data and key performance indicators of the crop variety with the environmental data of the corresponding test site to obtain phenotypic-environment integrated data.

[0135] In this embodiment of the invention, when integrating structured data according to crop varieties, on the one hand, the trait data of each crop variety are merged according to each entire level, thereby obtaining the trait values ​​of each crop variety at different levels such as single plant, plot, trial, single-year multi-location trial, and multi-year multi-location trial, providing a data foundation for crop phenotypic identification. On the other hand, the key performance indicators for corresponding phenotypic identification are calculated based on the merged trait values ​​of each crop variety, and then correlated with the corresponding environmental data to more accurately reflect the genetic characteristics of the crop itself, improve the accuracy of crop phenotypic identification, and help breeders screen out crop varieties that perform well in specific environments.

[0136] Therefore, this embodiment of the invention integrates the structured data in the phenotypic data to be integrated. On the one hand, it enables the calculation of complex traits in crop varieties, effectively simplifying the expression of trait data for crops with multiple traits. On the other hand, according to the crop variety dimension, it further integrates the trait data by merging and associating environmental data. This not only simplifies the data expression of crop phenotypes but also fully utilizes the data complementarity between different data, enhancing the reliability and robustness of crop phenotypic identification, which is conducive to improving the efficiency of crop breeding and accelerating the process of new crop variety selection.

[0137] Finally, in this embodiment of the invention, the phenotypic-environment integrated data obtained through integration processing is further correlated and aligned with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety.

[0138] Specifically, the embodiments of the present invention aim to realize a data chain of "variety-experiment-collection location-collection time-trait" to complete the fusion of multi-source heterogeneous crop phenotypic data and form a complete dataset containing multiple data modalities, which will be described in detail below.

[0139] First, genotype data for each crop variety is obtained. Genotype data refers to quality-controlled single nucleotide polymorphism (SNP) molecular marker data. Quality control of genotype data mainly includes: outlier removal, sample filtering, SNP deletion site filtering, minimum allele frequency (MAF) filtering, Hardy-Weinberg equilibrium (HWE) test, missing value imputation, and SNP numerical encoding.

[0140] Preferably, molecular markers with a minimum allele frequency less than 0.05 or a molecular marker deletion rate greater than 0.1 can be considered outliers and deleted. Missing value imputation can be performed using the average value of each molecular marker, filling the missing value location. SNP numerical coding encodes different genotypes as {0, 1, 2} or {-1, 0, 1} according to certain rules. Optionally, whole-genome selection models can be used to estimate missing values, detect and correct outliers, and estimate breeding values ​​for crop varieties from the collected phenotypic data, helping breeders better obtain information on the phenotypic characteristics and additive genetic effects of each variety. Whole-genome selection models include GBLUP, rrBLUP, Bayes-type models, and crop phenotypic prediction models based on machine learning or deep learning.

[0141] Genotype data is only related to crop varieties. Therefore, when achieving association alignment, the phenotypic-environment integrated data of each crop variety is first associated and integrated with the genotype data to obtain the genotype-phenotype-environment integrated data for each crop variety. Here, a field for genotype data can be set up and added to the field of phenotypic-environment integrated data, so that the phenotypic-environment integrated data of each crop variety carries the field of genotype data.

[0142] Next, the genotype-phenotype-environment integrated data is further correlated and aligned with the unstructured phenotypic data to be integrated. Here, key information is extracted from both the genotype-phenotype-environment integrated data and the unstructured phenotypic data to be integrated, and this key information is used to achieve data correlation and alignment. Specifically, the genotype-phenotype-environment integrated data obtained through the integration process is processed to extract field information, including crop variety number, experiment number, collection location, and collection time. The phenotypic-environment integrated data is still structured data, generally stored in a two-dimensional table, and has various predefined fields. Therefore, direct field extraction yields the crop variety, experiment number, collection location, and collection time information.

[0143] Metadata extraction is performed on the unstructured phenotypic data to be integrated, yielding metadata information including crop variety ID, experimental plot ID, collection location, and collection time. This unstructured phenotypic data is stored in the form of data aggregation record files, requiring preprocessing after extraction. For text data, this includes data cleaning, word segmentation, part-of-speech tagging, and named entity recognition; for image, audio, and video data, computer vision and speech recognition technologies are used for basic correction, noise reduction and enhancement, and target region segmentation. Further metadata extraction is then performed.

[0144] Optionally, for text data, natural language processing tools are used to identify words and phrases in the text related to variety, experiment, collection location, and collection time. For example, collection time can be extracted by matching date and time strings in a specific format using regular expressions, and the name of the crop variety and collection location can be identified using named entity recognition technology.

[0145] For image data, computer vision algorithms are used to identify text information in the image (such as variety name, test name, and test location on the test label) or to determine the test scene to which the image belongs (such as crop images at different test stages) through image classification technology.

[0146] For audio and video data, speech recognition technology is used to convert audio content into text, and then key information is extracted from the text; for video data, key frames can be extracted first, and then image analysis and speech recognition can be performed on the key frames to obtain the required information in a comprehensive manner.

[0147] Finally, the metadata information is correlated and matched with the field information to obtain multimodal fusion data for each crop variety. Here, a composite spatiotemporal index system with "variety-experiment-location-time-trait" as the dimensions is constructed. One or more elements from "variety, experiment, collection location, collection time, trait" are used as correlation alignment identifiers to correlate and align structured genotype-phenotype-environment integrated data with unstructured phenotypic data to be integrated. Correlation alignment can be performed using graph databases or relational databases. Core fields include crop variety ID, experiment ID, collection location, and collection time. In addition, the correlation file path after correlation alignment needs to be determined (e.g., / data / P001 / T2025-1 / XXlocation / 20250515 / image.jpg). For time-series data (such as video and sensor logs), dynamic time warping or attention mechanisms are used for synchronous time alignment. The CLIP (Contrastive Language-Image Pre-training) model is used to map image features and text descriptions to a shared semantic space, and semantic alignment is performed through cosine similarity matching.

[0148] After the association and alignment are completed, the data is divided according to the crop variety ID, resulting in multimodal fusion data for each crop variety. Ultimately, the multimodal fusion data includes: genotype-phenotype-environment integrated data, text data, image data, and video data. This multimodal fusion data, processed by the cloud platform, is used for subsequent precise identification of crop phenotypes.

[0149] Therefore, this embodiment of the invention utilizes a two-stage process of data acquisition and data integration to integrate multi-source heterogeneous crop phenotypic data collected from various multi-source phenotyping devices. This effectively addresses the need for fusion of massive and diverse multi-source heterogeneous data, improves data resource utilization, and ensures the accuracy and reliability of phenotypic identification results. By fully leveraging the data complementarity between data with different structures, the reliability and robustness of crop phenotypic identification are enhanced. Furthermore, incorporating crop variety genotype data into the data integration process is of great significance for improving crop variety breeding efficiency and accelerating the process of new variety selection.

[0150] The data integration device for precise identification of crop phenotypes provided by the present invention is described below. The data integration device for precise identification of crop phenotypes described below and the data integration method for precise identification of crop phenotypes described above can be referred to in correspondence with each other.

[0151] like Figure 2 As shown, the data integration device for precise identification of crop phenotypes specifically includes: a construction module 201, an allocation module 202, a collection module 203, and an integration module 204.

[0152] Specifically, the construction module 201 is used to build a cloud platform for precise crop phenotypic identification. The cloud platform is interconnected with data acquisition devices via a network. The data acquisition devices include multi-source phenotyping devices for collecting crop phenotypic data and environmental monitoring devices for collecting environmental data. The allocation module 202 is used to construct corresponding data acquisition models for the data acquisition devices, obtain data acquisition tasks for the crop phenotypic data and environmental data through the cloud platform, and allocate the data acquisition tasks to the data acquisition devices, so that the data acquisition devices can collect crop phenotypic data based on the acquisition instances in the data acquisition model. The data includes environmental data; a collection module 203 is used to perform field collection processing on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, and to perform data collection processing on the unstructured crop phenotypic data through the cloud platform to obtain unstructured phenotypic data to be integrated; an integration module 204 is used to integrate the structured phenotypic data to be integrated with the environmental data collected by the environmental monitoring equipment, and to associate and align the phenotypic-environmental integrated data obtained by the integration processing with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety.

[0153] It should be noted that the beneficial effects of the data integration device for precise crop phenotypic identification described here correspond to those of the data integration method for precise crop phenotypic identification mentioned above. Therefore, the beneficial effects of the data integration device for precise crop phenotypic identification will not be elaborated here.

[0154] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, communications interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute a data integration method for precise crop phenotypic identification. This method includes: constructing a cloud platform for precise crop phenotypic identification, wherein the cloud platform is interconnected with data acquisition devices via a network, and the data acquisition devices include multi-source phenotypic devices for collecting crop phenotypic data and environmental monitoring devices for collecting environmental data; constructing a corresponding data acquisition model for the data acquisition devices; acquiring data acquisition tasks for the crop phenotypic data and the environmental data through the cloud platform; and allocating the data acquisition tasks to the data acquisition devices so that the data acquisition devices can... The crop phenotypic data and environmental data are collected according to the collection instances in the data acquisition model; the structured and semi-structured crop phenotypic data are processed by field aggregation to obtain structured phenotypic data to be integrated; the unstructured crop phenotypic data is processed by data aggregation through the cloud platform to obtain unstructured phenotypic data to be integrated; the structured phenotypic data to be integrated is integrated with the environmental data collected by the environmental monitoring equipment, and the phenotypic-environmental integrated data obtained by the integration process is correlated and aligned with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety.

[0155] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0156] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data integration method for precise identification of crop phenotypes provided by the above methods. This method includes: constructing a cloud platform for precise identification of crop phenotypes, wherein the cloud platform is interconnected with data acquisition devices via a network, and the data acquisition devices include multi-source phenotyping devices for acquiring crop phenotyping data and environmental monitoring devices for acquiring environmental data; constructing a corresponding data acquisition model for the data acquisition devices; and acquiring the data acquisition tasks of the crop phenotyping data and the environmental data through the cloud platform. The data acquisition task is assigned to the data acquisition device, which then collects crop phenotypic data and environmental data based on the acquisition instances in the data acquisition model. Field aggregation processing is performed on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated. Unstructured crop phenotypic data is then aggregated through the cloud platform to obtain unstructured phenotypic data to be integrated. The structured phenotypic data to be integrated is then integrated with the environmental data collected by the environmental monitoring device. The resulting phenotypic-environmental integrated data is then correlated and aligned with the unstructured phenotypic data to be integrated, yielding multimodal fusion data for each crop variety.

[0157] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a data integration method for precise crop phenotypic identification provided by the methods described above. This method includes: constructing a cloud platform for precise crop phenotypic identification, the cloud platform being interconnected with data acquisition devices via a network, the data acquisition devices including multi-source phenotyping devices for acquiring crop phenotypic data and environmental monitoring devices for acquiring environmental data; constructing a corresponding data acquisition model for the data acquisition devices; acquiring data acquisition tasks for the crop phenotypic data and the environmental data through the cloud platform; and allocating the data acquisition tasks to the cloud platform. A data acquisition device is used to collect crop phenotypic data and environmental data according to the acquisition instances in the data acquisition model; field aggregation processing is performed on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated; and data aggregation processing is performed on the unstructured crop phenotypic data through the cloud platform to obtain unstructured phenotypic data to be integrated; the structured phenotypic data to be integrated is integrated with the environmental data collected by the environmental monitoring device, and the phenotypic-environmental integrated data obtained by the integration process is associated and aligned with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety.

[0158] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Depending on actual needs, some or all of the modules can be selected to achieve the purpose of this embodiment. Those skilled in the art can understand and implement this without any creative effort.

[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data integration method for precise identification of crop phenotypes, characterized in that, include: A cloud platform for precise identification of crop phenotypes is constructed. The cloud platform is interconnected with data acquisition devices via a network. The data acquisition devices include multi-source phenotyping devices for collecting crop phenotyping data and environmental monitoring devices for collecting environmental data. A corresponding data acquisition model is constructed for the data acquisition device. Data acquisition tasks for crop phenotypic data and environmental data are obtained through the cloud platform, and the data acquisition tasks are assigned to the data acquisition device so that the data acquisition device can acquire crop phenotypic data and environmental data according to the acquisition instances in the data acquisition model. Field aggregation processing is performed on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, and the unstructured crop phenotypic data is aggregated through the cloud platform to obtain unstructured phenotypic data to be integrated. The structured phenotypic data to be integrated is integrated with the environmental data collected by the environmental monitoring equipment, and the phenotypic-environmental integrated data obtained by the integration process is correlated and aligned with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety. The process of integrating the structured phenotypic data to be integrated with the environmental data collected by the environmental monitoring equipment includes: According to the preset integration level, the trait data in the structured phenotypic data to be integrated are merged to obtain the corresponding trait merged value. The trait merged value of each crop variety is then associated with the corresponding environmental data collected by the environmental monitoring equipment to obtain phenotypic-environment integrated data. The integration hierarchy includes at least one of the following: single-plant hierarchy, plot hierarchy, experimental hierarchy, single-year multi-point experimental hierarchy, and multi-year multi-point experimental hierarchy; The environmental data includes meteorological and soil data when crop varieties are planted at the test site. The process involves associating and aligning the phenotypic-environment integrated data obtained through integration with the unstructured phenotypic data to be integrated, resulting in multimodal fusion data for each crop variety, including: Obtain genotype data for each crop variety; the genotype data refers to single nucleotide polymorphism molecular marker data that has undergone quality control. The phenotypic-environment integrated data of each crop variety is associated and integrated with the genotype data to obtain the genotype-phenotype-environment integrated data of each crop variety. Fields were extracted from the genotype-phenotype-environment integrated data to obtain field information, which included crop variety number, experiment number, collection location, and collection time. Metadata is extracted from the unstructured phenotypic data to be integrated to obtain metadata information, which includes crop variety number, experimental plot number, collection location and collection time. The metadata information is associated and matched with the field information to obtain multimodal fusion data for each crop variety. The multimodal fusion data is used for accurate identification of crop phenotypes on the cloud platform.

2. The data integration method for precise identification of crop phenotypes according to claim 1, characterized in that, After constructing the cloud platform for precise crop phenotypic identification, the method further includes: registering the data acquisition device in the cloud platform, wherein the registration process of the data acquisition device includes: After a communication link is established between the data acquisition device and the cloud platform, the data acquisition device is triggered to send a registration request message to the cloud platform and performs two-way authentication with the cloud platform through certificate exchange. The cloud platform parses the registration request message, generates a device access token for the data acquisition device after the registration request is verified, and pushes the configuration parameters for collecting crop phenotypic data and environmental data to the data acquisition device. The cloud platform stores the device registration information obtained after parsing the registration request message, and the data acquisition device stores the device access token and the configuration parameters.

3. The data integration method for precise identification of crop phenotypes according to claim 1, characterized in that, The construction of a corresponding data acquisition model for the data acquisition device includes: Build a unified interface that adapts to the data acquisition device; Create a data acquisition instance, which includes a data acquisition device identifier, acquisition parameters, and data storage method. The acquisition parameters include the basic interface information of the unified interface, the source data parameters of the data acquisition device, the target data parameters of the crop phenotypic data and the environmental data, and data transformation rules. The data acquisition instances are grouped according to the device attributes of the data acquisition equipment and the data acquisition requirements. For each group, set group configuration parameters, and set the collection instances within the group to inherit the group configuration parameters. Then, deploy the collection instances that inherit the group configuration parameters in the unified interface to obtain the data collection model.

4. The data integration method for precise identification of crop phenotypes according to claim 1, characterized in that, Prior to the field aggregation processing of the structured and semi-structured crop phenotypic data, the method further includes: The on-chain notarization of the collected crop phenotypic data and environmental data is performed according to the data acquisition model. The on-chain notarization process of the crop phenotypic data and environmental data includes: Construct a consortium blockchain network composed of data acquisition instances from the data acquisition model; Generate content identifiers for the crop phenotypic data and the environmental data, and extract key information from the crop phenotypic data and the environmental data. The key information includes: raw data summary, timestamp, data acquisition device identifier, crop variety identifier, and test site identifier. The key information is encapsulated into a blockchain storage operation record with a data signature according to the consortium blockchain network, and the content identifiers of the crop phenotypic data and the environmental data are bound to the original data digest. The data signature is used to verify the identity and legitimacy of the data operator. The blockchain evidence storage operation record information is broadcast to multiple participating nodes in the consortium blockchain network, so that the participating nodes can verify the blockchain evidence storage operation record information according to a preset consensus mechanism and reach a consensus confirmation. The blockchain evidence storage operation record information is packaged into the consortium blockchain network, and on-chain evidence storage information of the crop phenotypic data and the environmental data is generated.

5. The data integration method for precise identification of crop phenotypes according to claim 1, characterized in that, Field aggregation processing is performed on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, including: The semi-structured crop phenotypic data is converted into target structured data; Determine the target data fields for the phenotypic data to be integrated, and configure the source data fields for each target acquisition instance. The target acquisition instance is the acquisition instance corresponding to each type of multi-source phenotypic device in the data acquisition model. For each target acquisition instance, configure the mapping relationship between the target data field and the source data field; The mapping relationship is parsed to obtain the mapping formula. The actual values ​​of the source data fields in the structured crop phenotypic data and the target structured data are substituted into the mapping formula, and the actual values ​​of the target data fields are calculated to obtain the structured phenotypic data to be integrated.

6. The data integration method for precise identification of crop phenotypes according to claim 1, characterized in that, The process of aggregating and processing the unstructured crop phenotypic data through the cloud platform to obtain unstructured phenotypic data to be integrated includes: The unstructured crop phenotypic data is encapsulated into data packets in the data acquisition device, and the data packets are transmitted to the edge node of the cloud platform; The data packets are parsed through the edge nodes, the parsed unstructured data is preprocessed, and the preprocessed data file is associated with the metadata of the unstructured crop phenotypic data to obtain a data collection record file; A multi-level index based on metadata is constructed for the data collection record file, representing unstructured tabular data to be integrated.

7. The data integration method for precise identification of crop phenotypes according to claim 1, characterized in that, After performing field aggregation processing on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, the method further includes: For the structured phenotypic data to be integrated, construct postfix expressions for complex traits; The trait variables in the postfix expression and the corresponding trait data of the complex traits in the structured trait data are stored in a stack for complex trait calculation to obtain the complex trait index value of the structured trait data.

8. The data integration method for precise identification of crop phenotypes according to claim 7, characterized in that, The construction of postfix expressions for complex traits, based on the structured phenotypic data to be integrated, includes: Based on the structured trait data included in the structured phenotypic data to be integrated, determine the computational expression for calculating the complex traits; Traverse each element of the operational expression from left to right, and perform the following processing on each element: a. If the element is an operand, add the element to the output queue; b. If the element is a left parenthesis operator, then push the element onto the operator stack; c. If the element is a right parenthesis operator, then pop the operators from the operator stack one by one until a left parenthesis operator is encountered, and then store the popped operators into the output queue one by one. d. If the element is the target operator, then perform the following steps: d1. When the operator stack is empty, or the top element of the operator stack is an operator with a left parenthesis, or the operation priority of the target operator is greater than that of the top element, the target operator is pushed onto the operator stack. d2. When the priority of the target operator is less than or equal to the top element of the stack, pop the top element of the stack and add it to the output queue; d3. Repeat step d2 until the conditions for executing step d1 are met; After traversing the last element of the operation expression, if there are still stack elements in the operator stack, the stack elements are popped one by one and added to the output queue. The elements in the output queue are dequeued one by one to obtain a postfix expression of complex state.

9. A data integration device for precise identification of crop phenotypes, characterized in that, include: A building module is used to build a cloud platform for precise identification of crop phenotypes. The cloud platform is interconnected with data acquisition devices via a network. The data acquisition devices include multi-source phenotyping devices for collecting crop phenotyping data and environmental monitoring devices for collecting environmental data. The allocation module is used to construct a corresponding data acquisition model for the data acquisition device, obtain data acquisition tasks of crop phenotypic data and environmental data through the cloud platform, and allocate the data acquisition tasks to the data acquisition device so that the data acquisition device can acquire crop phenotypic data and environmental data according to the acquisition instances in the data acquisition model. The aggregation module is used to perform field aggregation processing on the structured and semi-structured crop phenotypic data to obtain structured phenotypic data to be integrated, and to perform data aggregation processing on the unstructured crop phenotypic data through the cloud platform to obtain unstructured phenotypic data to be integrated. The integration module is used to integrate the structured phenotypic data to be integrated with the environmental data collected by the environmental monitoring equipment, and to associate and align the phenotypic-environmental integrated data obtained by the integration process with the unstructured phenotypic data to be integrated to obtain multimodal fusion data for each crop variety. The process of integrating the structured phenotypic data to be integrated with the environmental data collected by the environmental monitoring equipment includes: According to the preset integration level, the trait data in the structured phenotypic data to be integrated are merged to obtain the corresponding trait merged value. The trait merged value of each crop variety is then associated with the corresponding environmental data collected by the environmental monitoring equipment to obtain phenotypic-environment integrated data. The integration hierarchy includes at least one of the following: single-plant hierarchy, plot hierarchy, experimental hierarchy, single-year multi-point experimental hierarchy, and multi-year multi-point experimental hierarchy; The environmental data includes meteorological and soil data when crop varieties are planted at the test site. The process involves associating and aligning the phenotypic-environment integrated data obtained through integration with the unstructured phenotypic data to be integrated, resulting in multimodal fusion data for each crop variety, including: Obtain genotype data for each crop variety; the genotype data refers to single nucleotide polymorphism molecular marker data that has undergone quality control. The phenotypic-environment integrated data of each crop variety is associated and integrated with the genotype data to obtain the genotype-phenotype-environment integrated data of each crop variety. Fields were extracted from the genotype-phenotype-environment integrated data to obtain field information, which included crop variety number, experiment number, collection location, and collection time. Metadata is extracted from the unstructured phenotypic data to be integrated to obtain metadata information, which includes crop variety number, experimental plot number, collection location and collection time. The metadata information is associated and matched with the field information to obtain multimodal fusion data for each crop variety. The multimodal fusion data is used for accurate identification of crop phenotypes on the cloud platform.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data integration method for precise identification of crop phenotypes as described in any one of claims 1 to 8.