A method and apparatus for processing multi-modal data

CN122673618APending Publication Date: 2026-09-01CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611019747.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0005]但是方式(1)的对齐方式,使用日历时间无法反映不同地域、不同季节、不同栽培方式下的农业物候差异,同一日历日期对应的作物发育阶段可能不同,会导致地域、季节、品种差异等造成的时序错位

Benefits of technology

[0020] This scheme uses phenological phase sequences as the basis for dividing phase grids and crop phenological phase sequences as resampling anchors. This allows multimodal data across regions, seasons, and crops to be compared under a unified phase, ensuring that data at the same phenological phase grid point corresponds to the same physiological developmental state of the crop. This eliminates temporal misalignments caused by regional, seasonal, and varietal differences, improving the portability of downstream world models when migrating across regions, seasons, or domains. Furthermore, it provides a cross-source timestamp fault-tolerant alignment strategy, setting a configurable tolerance window for legitimate timestamp offsets between sources. Samples within the tolerance are considered to be at the same time point, while those exceeding the tolerance are entered into an anomaly queue instead of being discarded entirely, thus improving the sample utilization rate of multi-source data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122673618A_ABST
    Figure CN122673618A_ABST
Patent Text Reader

Abstract

This application discloses a method and apparatus for processing multimodal data, relating to the technical fields of smart agriculture, controlled environment agriculture, and multimodal machine learning data engineering. It uses phenological phase sequences as the basis for dividing phase grids and crop phenological phase sequences as resampling anchors, enabling cross-regional, cross-seasonal, and cross-crop multimodal data to be compared under a unified phase. This ensures that data under the same phenological phase grid point corresponds to the same physiological developmental state of the crop, eliminating temporal misalignment caused by regional, seasonal, and varietal differences, and improving the portability of downstream world models when migrating across regions, seasons, or domains. Furthermore, it provides a cross-source timestamp fault-tolerant alignment strategy, setting a configurable tolerance window for legitimate timestamp offsets between sources. Samples within the tolerance are considered to be at the same time point, while those exceeding the tolerance are entered into an anomaly queue instead of being discarded entirely, improving the sample utilization rate of multi-source data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of smart agriculture, controlled environment agriculture and multimodal machine learning data engineering, and more specifically, to a method and apparatus for processing multimodal data. Background Technology

[0002] Agricultural multimodal data refers to various types of data with completely different formats and properties collected through different technical means during agricultural production, management, and scientific research.

[0003] The processing of agricultural multimodal data involves agricultural multimodal data alignment. Agricultural multimodal data alignment integrates information from multiple sources such as vision, text, and voice, enabling machines to understand complex agricultural scenarios more comprehensively and accurately. This provides core support for precise decision-making and intelligent services, ultimately driving agriculture towards a more efficient and intelligent direction.

[0004] Existing agricultural multimodal data alignment methods typically include the following: Method (1) Using calendar time as the resampling anchor point, aggregate each data source to a unified time grid by day, week or fixed time window; Method (2) Cross-source timestamp alignment often adopts a fixed main source strict matching strategy, and non-main source sample timestamps that are inconsistent with the main source are discarded.

[0005] However, the alignment method of method (1) cannot reflect the differences in agricultural phenology under different regions, seasons, and cultivation methods by using calendar time. The crop development stage corresponding to the same calendar date may be different, which will lead to time sequence misalignment caused by regional, seasonal, and variety differences. The fixed main source strict matching strategy of method (2) lacks fault tolerance for legitimate sampling clock offsets between sources. A large number of valid samples are discarded meaninglessly, which reduces the sample utilization rate of multi-source data. Summary of the Invention

[0006] In view of this, this application discloses a method and apparatus for processing multimodal data, which aims to improve the mobility of downstream world models when they migrate across regions, seasons or domains, and to improve the sample utilization rate of multi-source data.

[0007] To achieve the above objectives, the disclosed technical solution is as follows:

[0008] The first aspect of this application discloses a method for processing multimodal data, the method comprising:

[0009] A phenological phase sequence is defined for each data acquisition subject; wherein, the phenological phase sequence is the temporal position information corresponding to the multiple phenological stages, which are divided into multiple phenological stages according to the progress of phenological development.

[0010] According to the phenological phase sequence, the acquired multi-source time series data are resampled to a unified phase grid to obtain multi-source data from each data source after resampling at the unified phase grid.

[0011] The phase resampled multi-source data is cross-source timestamp fault-tolerant alignment is performed to obtain aligned multi-source data.

[0012] If missing values ​​are detected in the aligned multi-source data, the multi-source data with missing values ​​will be filled in hierarchically.

[0013] The multi-source data after hierarchical completion is verified by using the intermodal consistency verification rule set, and a misalignment report is generated and output.

[0014] A second aspect of this application discloses a multimodal data processing apparatus, the apparatus comprising:

[0015] A calibration unit is used to calibrate the phenological phase sequence for each data acquisition subject; wherein, the phenological phase sequence is the temporal position information corresponding to the multiple phenological stages, which are divided into multiple phenological stages according to the progress of phenological development.

[0016] The resampling unit is used to resample the acquired multi-source time-series data to a unified phase grid according to the phenological phase sequence, so as to obtain multi-source data resampled from each data source on the unified phase grid.

[0017] The fault-tolerant alignment unit is used to perform cross-source timestamp fault-tolerant alignment on the phase resampled multi-source data to obtain aligned multi-source data.

[0018] The hierarchical completion unit is used to perform hierarchical completion on the multi-source data with missing values ​​if missing values ​​are detected in the aligned multi-source data.

[0019] The verification output unit is used to verify the multi-source data after hierarchical completion through the intermodal consistency verification rule set, generate a misalignment report and output it.

[0020] This scheme uses phenological phase sequences as the basis for dividing phase grids and crop phenological phase sequences as resampling anchors. This allows multimodal data across regions, seasons, and crops to be compared under a unified phase, ensuring that data at the same phenological phase grid point corresponds to the same physiological developmental state of the crop. This eliminates temporal misalignments caused by regional, seasonal, and varietal differences, improving the portability of downstream world models when migrating across regions, seasons, or domains. Furthermore, it provides a cross-source timestamp fault-tolerant alignment strategy, setting a configurable tolerance window for legitimate timestamp offsets between sources. Samples within the tolerance are considered to be at the same time point, while those exceeding the tolerance are entered into an anomaly queue instead of being discarded entirely, thus improving the sample utilization rate of multi-source data. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating a method for processing multimodal data disclosed in an embodiment of this application;

[0023] Figure 2 This is a schematic diagram comparing calendar alignment and phenological phase alignment disclosed in an embodiment of this application;

[0024] Figure 3 This is a heatmap of cross-source tolerance window sensitivity disclosed in an embodiment of this application;

[0025] Figure 4 This is a schematic diagram of the supplementary traceability mark and downstream usage structure disclosed in the embodiments of this application;

[0026] Figure 5 This is a schematic diagram illustrating the completion confidence distribution disclosed in the embodiments of this application;

[0027] Figure 6 This is a schematic diagram of the intermodal consistency verification and handling process disclosed in the embodiments of this application;

[0028] Figure 7 This is a flowchart illustrating another method for processing multimodal data disclosed in an embodiment of this application;

[0029] Figure 8 This is a schematic diagram of the structure of a multimodal data processing device disclosed in an embodiment of this application;

[0030] Figure 9 This is a schematic diagram of the structure of the electronic device disclosed in the embodiments of this application. Detailed Implementation

[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0033] As can be seen from the background technology, existing agricultural multimodal data alignment methods typically include the following: Method (1) uses calendar time as the resampling anchor point, aggregating various data sources to a unified time grid by day, week, or fixed time window; Method (2) cross-source timestamp alignment often adopts a fixed main source strict matching strategy, and non-main source sample timestamps that are inconsistent with the main source are discarded. The alignment method of Method (1) will lead to temporal misalignment caused by regional, seasonal, and variety differences. The alignment method of Method (2) reduces the sample utilization rate of multi-source data.

[0034] To address the aforementioned issues, this application discloses a method and apparatus for processing multimodal data. It uses phenological phase sequences as the basis for dividing phase grids and crop phenological phase sequences as resampling anchors, enabling cross-regional, cross-seasonal, and cross-crop multimodal data to be compared under a unified phase. This ensures that data at the same phenological phase grid point corresponds to the same physiological developmental state of the crop, eliminating temporal misalignments caused by regional, seasonal, and varietal differences, and improving the portability of downstream world models when migrating across regions, seasons, or domains. Furthermore, it provides a cross-source timestamp fault-tolerant alignment strategy, setting a configurable tolerance window for legitimate timestamp offsets between sources. Samples within the tolerance are considered to be at the same time point, while those exceeding the tolerance are entered into an anomaly queue instead of being discarded entirely, improving the sample utilization rate of multi-source data. Specific implementation methods are described in detail through the following embodiments.

[0035] It should be noted that the multimodal data processing method and apparatus provided in this application relate to the technical fields of smart agriculture, controlled environment agriculture, and multimodal machine learning data engineering. Specifically, it involves resampling, alignment, and missing data completion of multi-source heterogeneous data (environmental sensing, imagery, physiological measurements, and management events) anchored by phenological phase for training agricultural world models. The above is merely an example and does not limit the application field of the multimodal data processing method and apparatus provided in this application.

[0036] refer to Figure 1 The image shows a method for processing multimodal data disclosed in an embodiment of this application. This method mainly includes the following steps:

[0037] S101: Calibrate the phenological phase sequence for each data acquisition subject; wherein, the phenological phase sequence is the temporal position information corresponding to the multiple phenological stages, which are divided into multiple phenological stages according to the progress of phenological development of the data acquisition subject.

[0038] It should be noted that a phenological phase sequence Φ(t) is calibrated for the entire data cycle of each data acquisition subject. A phenological phase refers to the developmental progress of a crop or animal according to its growth stages, such as emergence, vegetative growth, flowering, fruiting, and senescence. This sequence serves as an anchor point for multi-source data resampling, replacing calendar time. The phenological phase sequence is generated based on growth period records, image-driven classifiers, or a fusion of both.

[0039] The phenological phase sequence can be calibrated for the entire data cycle of each data acquisition subject using methods such as reproductive period record-driven, image-driven, and dual-source fusion. The phenological phase sequence is represented by Φ(t).

[0040] The process driven by reproductive period records is shown in A1-A2.

[0041] A1: Obtain the key phenological event timestamps of the data collection subjects; among which, the data collection subjects include at least greenhouse compartments, fields and individual animals; the key phenological event timestamps include at least the sowing date, seedling emergence date, flowering date and first harvest date.

[0042] A2: Between two adjacent key phenological event timestamps, interpolation is performed using crop phenological development temperature thresholds to obtain continuous phase variables or discrete phases, thus completing the process of calibrating the phenological phase sequence for each data collection subject.

[0043] The temperature threshold for crop phenological development is obtained by accumulating the base temperature and the effective accumulated temperature (GDD).

[0044] GDD is used to interpolate continuous phase variables between key phenological events.

[0045] Key phenological event timestamps provided by the cultivation entity, such as sowing date, emergence date, flowering date, and first harvest date, can be used to interpolate between two adjacent key phenological event timestamps by using crop phenological development temperature thresholds obtained from the base temperature and GDD, to obtain continuous phase variables (φ∈[0,1]) or discrete phases (φ∈{emergence,vegetative,flowering,fruiting,senescence}).

[0046] Among them, emergence is the seedling emergence date; vegetative is the vegetative growth period; flowering is the flowering date; fruiting is the fruiting period; and senescence is the senescence period.

[0047] By interpolating between phenological events using the cumulative base temperature and GDD (Gross Difference in Temperature), the advantage of this method is that it transforms discrete observation nodes into continuous physiological development processes, thereby improving prediction accuracy and management timeliness.

[0048] The image-driven process is shown in B1-B2.

[0049] B1: Obtain the phenological characteristics of crop canopy images of the data collection subject; wherein, the phenological characteristics of crop canopy images include at least the greenness index, coverage, and fruit visibility.

[0050] The greenness index is a quantitative indicator that represents the degree of greenness or growth status of vegetation.

[0051] Coverage rate refers to the ratio of the vertical projection area of ​​plants (including leaves, stems and branches) in a certain area to the total area of ​​that area, usually expressed as a percentage.

[0052] Fruit visibility refers to the degree of clarity or detectability of a fruit by a sensor or human eye against a complex background (such as foliage obstruction).

[0053] B2: Input the phenological features of crop canopy images into a pre-trained classifier, and output the phenological phase labels corresponding to the crop canopy images through the pre-trained classifier to complete the process of labeling the phenological phase sequence for each data collection subject.

[0054] When the growth period of the data acquisition subject is unavailable, i.e. when the growth period record driver is unavailable, the phenological characteristics such as greenness index, coverage, and fruit visibility of the crop canopy image of the data acquisition subject are used, and the phenological phase label corresponding to the crop canopy image is output by the pre-trained classifier.

[0055] When both the reproductive period recording driver and the image driver are available, the reproductive period recording driver is used as the main core and the image driver is used as the check. If the conflict value between the reproductive period recording driver and the image driver exceeds the preset threshold, it will enter the anomaly queue, as shown in C1-C4.

[0056] C1: Obtain the key phenological event timestamps of the data collection subject, and obtain the phenological characteristics of the crop canopy images of the data collection subject.

[0057] In C1, key phenological event timestamps such as sowing date, emergence date, flowering date, and first harvest date of the data collection subject are obtained, as well as phenological characteristics such as greenness index, coverage, and fruit visibility of crop canopy images of the data collection subject.

[0058] C2: The phenological phase sequence marked with the timestamps of key phenological events is used as the main sequence, and the phenological phase sequence marked with the phenological features of crop canopy images is used as the verification sequence.

[0059] C3: Calculate the conflict value between the main sequence and the check sequence.

[0060] The master sequence can be obtained through manual recording, environmental sensor sequences, or management logs.

[0061] The verification sequence can be obtained through canopy images, accumulated temperature models, expert rules, or other modalities.

[0062] The conflict value is used to measure the degree of inconsistency between the main sequence and the check sequence in terms of phenological phases or key events.

[0063] In C3, the main sequence and the check sequence are mapped to the same phenological phase grid or the same time grid point. Then, the phase difference, key event time difference, and label inconsistency term are calculated to determine the conflict value between the main sequence and the check sequence. The specific calculation method is shown in formula (1).

[0064] (1);

[0065] in, Indicates conflicting values; The phenological phase of the main sequence at grid point j is represented; N represents the number of grid points used to compare phenological phases, where N is a positive integer and is usually equal to the number of effective grid points after the main sequence and the check sequence are mapped to a unified phase grid; M represents the number of events used to compare the time difference of critical events, where M is a positive integer and is usually equal to the number of critical phenological events that both events share. This indicates the phenological phase of the verification sequence at the j-th grid point; This indicates the time of the e-th critical event in the main sequence; These represent the times of the e-th critical event in the verification sequence; This indicates the normalized scale of the allowable time deviation for the event; The phenological phase label or category label corresponding to the main sequence is usually determined by the main source data such as manual records, management logs, or timestamps of key phenological events. The phenological phase label or category label corresponding to the verification sequence is usually inferred from canopy images, accumulated temperature models, expert rules, or other modal data; Indicates an indicator function, and The value is 1 if the two are inconsistent, and 0 if they are consistent. , , Indicates the weight.

[0066] If the phase is a discrete label, such as seedling stage, vegetative growth stage, flowering stage, fruit setting stage, and maturity stage, each stage can be mapped to an ordered number first, and then the normalized stage difference can be calculated. When the conflict value exceeds a preset threshold, the corresponding sample is written into the misalignment report.

[0067] C4: If the conflict value is greater than the preset threshold, the corresponding sample will be stored in the exception queue.

[0068] The preset threshold is set according to the actual situation, and this application does not impose specific limitations.

[0069] An exception queue is a temporary collection of samples that exceed the tolerance but are not discarded directly.

[0070] This solution eliminates temporal misalignment caused by regional, seasonal, and varietal differences by calibrating the phenological phase sequence for each data collection subject, thereby improving the portability of the downstream world model when it migrates across regions, seasons, or domains.

[0071] S102: According to the phenological phase sequence, the acquired multi-source time series data are resampled to a unified phase grid to obtain multi-source data after resampling of each data source on the unified phase grid.

[0072] Multi-source time-series data refers to modality / multi-modal data collected through different sensing methods. Multi-source time-series data includes at least one or more of the following: environmental sensing time-series data, image data, physiological measurement data, and administrative event data. Multi-source time-series data is represented by X(t). Environmental sensing time-series data uses aggregated statistics, image and physiological sources use nearest neighbor aggregation, and event sources use count / intensity aggregation.

[0073] The unified phase lattice is represented by Φ_k, where k=1,...,K.

[0074] Environmental sensing time-series data include, but are not limited to, temperature, humidity, carbon dioxide (CO2), photosynthetically active radiation (PAR), and saturated vapor pressure difference (VPD).

[0075] PAR is one of the key channels for environmental sensing and has a causal relationship with the supplemental lighting setting.

[0076] VPD is derived from temperature and relative humidity and is often used as an indicator of plant stress.

[0077] Image data can be collected using equipment such as drones, remote sensing, and fixed cameras in greenhouses.

[0078] Physiological measurement data include, but are not limited to, leaf area index, plant height, and harvesting records.

[0079] Management event data includes, but is not limited to, irrigation, fertilization, pesticide application, setpoint changes, event-triggered data, and data with irregular intervals. Event-triggered data refers to data that is not generated according to a fixed sampling period, but rather recorded only when a certain type of agricultural management event or system state change occurs. Examples include irrigation start / end, fertilization, pesticide application, supplemental lighting activation, ventilation setpoint changes, CO2 setpoint changes, alarm triggers, and pest and disease treatment records. The timestamps of this type of data reflect the time of event occurrence, not a fixed sampling frequency. Irregular intervals refer to the time interval between two adjacent records from the same data source being inconsistent. For example, a management log might be recorded at 08:00, 08:07, 09:30, and 14:15, with intervals of 7 minutes, 83 minutes, and 285 minutes, respectively. This type of data cannot be directly input into the model at fixed time steps and requires fault-tolerant alignment or event window aggregation.

[0080] The specific process of S102, which resamples the acquired multi-source time-series data to a unified phase grid, is shown in D1-D5.

[0081] D1: Obtain the phenological phase intervals corresponding to each phase grid point in the unified phase grid.

[0082] D2: For environmental sensing time series data, aggregate and statistically analyze the environmental sensing time series data that fall within the phenological phase interval to obtain phase resampling environmental sensing data on a unified phase grid.

[0083] In D2, for environmental sensor time-series data from high-sampling-rate sources, weighted averages, maximum values, and minimum values ​​are aggregated and statistically analyzed within each phenological phase interval to obtain phase-resampled environmental sensor data at a unified phase grid. The aggregation method can be customized.

[0084] D3: For image data, perform nearest neighbor search on image data falling within the phenological phase interval to obtain image resampling data at a unified phase grid.

[0085] In D3, for image data from low sampling rate sources, the nearest neighbor search is performed on image data falling within the phenological phase interval. That is, the nearest neighbor sample is taken within the phenological phase interval to obtain image resampled data at a unified phase grid. If there are no image data samples within the phenological phase interval, the missing data completion stage is entered, that is, the hierarchical completion operation in S104 is performed.

[0086] D4: For physiological measurement data, perform nearest neighbor search on physiological measurement data that fall within the phenological phase interval to obtain physiological resampling data at a unified phase grid.

[0087] In D4, for physiological measurement data from low sampling rate sources, the physiological measurement data falling within the phenological phase interval are searched for nearest neighbors. That is, the nearest neighbor sample is taken within the phenological phase interval to obtain physiological resampled data at a unified phase grid. If there are no samples with physiological resampled data within the phenological phase interval, the missing data completion stage is entered, that is, the hierarchical completion operation in S104 below is performed.

[0088] D5: For managed event data, aggregate the event counts and event intensity that fall within the phenological phase interval to obtain event resampling data at a unified phase grid.

[0089] The key effects of this scheme's phenological phase sequence calibration are as follows: Figure 2 As shown.

[0090] Figure 2 The distribution of the same crop at different latitudes is shown to be separated under calendar alignment and aligned under phenological phase alignment.

[0091] Figure 2 In the diagram, (a) indicates misaligned data. The horizontal axis represents calendar time, and the vertical axis represents crop state or physiological state indicators. Solid lines represent low latitude samples, and dashed lines represent high latitude samples.

[0092] Figure 2 In the diagram, (b) indicates aligned data; the horizontal axis represents the phenological phase; the vertical axis represents the crop state or physiological state index; the solid line represents low latitude samples; and the dashed line represents high latitude samples.

[0093] Originally, crop states at different latitudes were not comparable in terms of calendar time, but after resampling to the same phenological phase, the states of the two can be directly aligned.

[0094] The phase resampling module also outputs a phase_index table, which records the original time interval, GDD range, distance to key phenological events, and data coverage for each phase grid point. The phase_index table is used for subsequent cross-regional, cross-seasonal, or cross-crop model comparisons to avoid ignoring situations where different data sources, although under the same phase number, have significantly different actual coverage time windows.

[0095] Using phenological phase sequences as resampling anchors makes data that were originally incomparable under a unified phase comparable. Downstream world models show significantly better performance stability than calendar-aligned baselines in tasks such as cross-domain migration, cross-scenario evaluation, and cross-seasonal prediction.

[0096] Using phenological phase sequences as resampling anchors is the core of this approach, distinguishing it from existing calendar alignment methods. Phenological phase anchoring is not simply replacing the x-axis; it involves a complete set of engineering implementations, including crop-specific GDD interpolation, image phenological verification, and configurable phase granularity. This allows for the comparison of multimodal data across regions, seasons, and crops under a unified phase, ensuring that data at the same phenological phase grid point corresponds to the same physiological developmental state of the crop. This eliminates temporal misalignments caused by regional, seasonal, and varietal differences, improving the portability of downstream world models when migrating across regions, seasons, or domains.

[0097] S103: Perform cross-source timestamp fault-tolerant alignment on the phase resampled multi-source data to obtain aligned multi-source data.

[0098] In S103, cross-source timestamp fault-tolerant alignment is performed on the multi-source data that have completed phase resampling to align the multi-source data that have completed phase resampling to the same timestamp. Cross-source timestamp fault-tolerant alignment includes main source selection, configuring tolerance window, merging within tolerance, and storing samples that exceed the tolerance in an exception queue.

[0099] Compared to the fixed timestamp alignment strategy, the cross-source timestamp fault-tolerant alignment strategy significantly improves the sample retention rate, especially in scenarios with low source timestamp quality, thereby improving data utilization.

[0100] The specific process of selecting the primary source, configuring the tolerance window, merging within the tolerance, and storing samples exceeding the tolerance into the exception queue is shown in E1-E5.

[0101] E1: Select the data with the highest sampling rate and stable timestamp from the multi-source data of phase resampling as the primary source, and use the remaining data as non-primary sources.

[0102] In E1, the primary source is selected from the multi-source data with the highest sampling rate and the highest timestamp stability score (data whose timestamps meet the stability criteria, usually environmental sensor data) from the phase resampling multi-source data. The primary source is not only determined by the highest sampling rate, but is also affected by missing rate, duplication rate, out-of-order rate, drift, etc.

[0103] The master source is the source selected as the time reference in cross-source alignment, typically the environmental sensor source with the highest sampling rate and the most stable timestamps. The master source is not a manually designated fixed data source, but rather selected through a timestamp stability score. The system can calculate metrics such as sampling interval stability, missing rate, duplication rate, out-of-order rate, clock drift, and coverage for each candidate data source, and select the data source with the highest overall score as the master source.

[0104] Environmental sensor data is usually sampled at a high frequency and has continuous timestamps, so it is often chosen as the primary source. However, this scheme does not limit the primary source to environmental sensor data.

[0105] For example, parameters such as the median or mean of the sampling interval, the coefficient of variation of the sampling interval, the standard deviation of timestamp jitter, the missing rate, the proportion of duplicate timestamps, the proportion of out-of-order timestamps, the clock drift trend, and the coverage duration ratio can be used to determine whether the data with stable timestamps are the main source.

[0106] E2: Obtain the master source timestamp corresponding to the master source and configure the tolerance window.

[0107] The tolerance window duration is based on the time grid determined by the main source sampling rate, and is configured in conjunction with the sampling rate of the non-main source data to be aligned, timestamp accuracy, upload latency, and historical offset distribution.

[0108] The primary source sampling rate is used to determine the primary time grid for uniform alignment. For example, if the primary source is sampled every 5 minutes, then the primary time grid is usually 5 minutes or an integer multiple thereof.

[0109] Main source timestamp passed express.

[0110] The tolerance window is the allowed time range of offset on either side of the primary source timestamp. The tolerance window is determined by... This indicates a relationship between the main source sampling rate, tolerance window, and misalignment rate, as detailed below. Figure 3 As shown. Figure 3 A heatmap of tolerance window sensitivity is shown.

[0111] Figure 3 The horizontal axis represents the tolerance window step size; the vertical axis represents the main source sampling interval; and the color and cell value represent the misalignment rate.

[0112] E3: For each sample in each non-main source, obtain the non-main source sample timestamp of that sample.

[0113] Among them, the timestamps of non-primary source samples are obtained through... express.

[0114] E4: When the timestamp of a non-primary source sample does not fall within the tolerance window, the samples that exceed the tolerance window and whose timestamps do not fall within the tolerance window are stored in the exception queue and reviewed.

[0115] It should be noted that samples exceeding the tolerance are not directly discarded but are placed in an anomaly queue, where they can be manually or semi-automatically reviewed using the accompanying anomaly review tool. Through a cross-source timestamp fault-tolerant alignment strategy, a configurable tolerance window is set for legitimate timestamp offsets between sources. Samples within the tolerance are considered to be at the same time point, while those exceeding the tolerance are placed in the anomaly queue instead of being discarded entirely. This cross-source timestamp fault-tolerant alignment strategy differs from the univariate discarding of fixed timestamp alignment. This strategy preserves the opportunity for review of boundary samples, improving the sample utilization rate of multi-source data.

[0116] E5: When the timestamp of a non-primary source sample falls within the tolerance window, it is determined that the timestamp of the non-primary source sample and the timestamp of the primary source are the same time point, and the offset between the timestamp of the non-primary source sample and the timestamp of the primary source is recorded in the alignment log to obtain the aligned multi-source data.

[0117] In E5, the offset between the timestamp of the non-primary source sample and the timestamp of the primary source sample is calculated by subtracting the timestamp of the non-primary source sample from the timestamp of the primary source sample. The expression for the offset is shown in formula (2).

[0118] (2);

[0119] in, This represents the offset between the timestamp of the non-primary source sample and the timestamp of the primary source sample. For non-primary source sample timestamps; The primary timestamp.

[0120] After alignment is completed, the output includes the utilization rate of each source sample, the average and maximum offset of each source, the length of the exception queue and the distribution of exception causes, and generates an alignment log (alignment_log).

[0121] The utilization rate of each source sample is calculated by dividing the number of samples successfully aligned at each non-primary source by the total number of samples at that source. The utilization rate of each source sample can be used to reflect the data quality of each source and the rationality of the tolerance window settings.

[0122] The average and maximum offsets of each source are used to reflect the degree of deviation and dispersion of the timestamp system between non-primary sources and the primary source.

[0123] The distribution of abnormal causes refers to the statistical analysis of the number of samples that exceed the tolerance and are transferred to the abnormal queue, as well as the classification of causes.

[0124] The alignment log is used to record the matching relationship between non-primary source samples and primary source samples one by one. The alignment log includes, but is not limited to, fields such as source_id, source_timestamp, primary_timestamp, offset, tolerance_window, match_status, and review_status.

[0125] If a source's offset approaches the tolerance boundary for an extended period, the system generates a clock drift alarm instead of simply merging the sample into the main source's time point. This solution proactively tracks the trend of each offset, and when it detects that the offset is consistently close to the tolerance window boundary (e.g., the offset remains within the 80%~100% range of the tolerance window), it generates a clock drift alarm in advance, notifying the user that the sensor's clock may have accumulated deviation and recommending calibration or maintenance. Users can proactively intervene before samples begin to be discarded to avoid data loss.

[0126] S104: If missing values ​​are detected in the aligned multi-source data, perform hierarchical imputation on the multi-source data with missing values.

[0127] The hierarchical completion includes completing multi-source data in the first missing state (short-term missing) by linear interpolation or physical consistency interpolation, completing multi-source data in the second missing state (medium-term missing) by regression based on homologous neighbor samples or multi-source completion based on cross-source association models, and labeling multi-source data in the third missing state (long-term missing).

[0128] First missing state: Use linear interpolation or physically consistent interpolation (such as temperature weighted by cumulative radiation over adjacent time periods).

[0129] Second missing state: use regression imputation based on homologous neighbor samples, or multi-source imputation based on cross-source association models (e.g., missing humidity is inferred from temperature and dew point).

[0130] The third missing state: marked as unfillable, corresponding to the output tensor entry and keeping the non-NaN marked in the mask tensor.

[0131] The mask annotation tensor is used to identify which locations are valid observations and which are missing or cannot be filled in. The mask has the same shape as the data tensor, and downstream models process the data according to the mask.

[0132] The specific process of hierarchical completion of multi-source data with missing values ​​is shown in F1-F4.

[0133] F1: If missing values ​​are detected in the aligned multi-source data, determine the duration of the missing values.

[0134] F2: When the duration of the missing data is less than the first threshold, the aligned multi-source data is determined to be in the first missing state, and linear interpolation or physical consistency interpolation is used to complete the multi-source data in the first missing state.

[0135] For short-term missing multi-source data, linear interpolation or physical consistency interpolation can be used to complete the short-term missing multi-source data. This can maintain the continuity of the data while avoiding the discarding of the entire data sample due to a small number of missing data, thus improving the sample utilization rate of multi-source data.

[0136] F3: When the duration of the missing data is greater than or equal to the first threshold and less than the second threshold, the aligned multi-source data is determined to be in the second missing state. The multi-source data in the second missing state is filled by using regression imputation based on homologous neighboring samples or multi-source imputation based on cross-source association model.

[0137] It should be noted that the first threshold is less than the second threshold. The first and second thresholds are set according to the actual situation, and this application does not impose specific limitations.

[0138] For multi-source data with missing mid-time, regression imputation based on homologous neighboring samples or multi-source imputation based on cross-source association models are used to complete the multi-source data with missing mid-time. This fully utilizes homologous temporal correlation or heterologous physical correlation (such as multi-source correlation information such as temperature, humidity and dew point) to infer missing values ​​and improves the accuracy of imputation by using multi-source correlation information.

[0139] F4: When the duration of the missing data is greater than or equal to the second threshold, the aligned multi-source data is determined to be in the third missing state, and the multi-source data in the third missing state is marked as uncompleted.

[0140] For multi-source data in the third missing state, mark the multi-source data in the third missing state as uncompleted, output the target entry corresponding to the marked multi-source data, and mark the target entry in the mask tensor as missing state at the corresponding position.

[0141] If missing values ​​are detected in the aligned multi-source data, after hierarchical imputation of the multi-source data with missing values, a structured imputation provenance is added to each imputed sample. The composition of the imputation provenance and the downstream model reading process are as follows: Figure 4 As shown. Figure 4 A schematic diagram of the completion traceability mark and downstream usage structure is shown.

[0142] The imputation provenance is the core mandatory field of this scheme, containing the imputation method, confidence level, and nearest observation distance. Imputation provenance flags include, but are not limited to, the `imputation_method`, `imputation_confidence`, `nearest_observation_distance_phase_units`, `source_used_for_imputation`, and `imputation_timestamp` fields. These flags are mandatory fields for the aligned output, and the downstream world model uses them for uncertainty modeling and loss weighting during training. This imputation provenance flag is not a post-hoc annotation but mandatory metadata on par with the data tensor, which is key to distinguishing this scheme from general machine learning imputation methods.

[0143] Figure 4 The process involves the following: imputation method identifier, confidence; review status, nearest observation distance, source used list, imputation timestamp, data tensor, mask tensor, phase index table, source index table, loss function weighting, aleatoric variance, and traceability after leave-one-out cross-validation (LOCO).

[0144] Among them, `imputation_method`: the completion method identifier, including fields such as `linear_interp`, `physics_consistent`, `cross_source_regression`, and `unfilled`;

[0145] imputation_confidence: Imputation confidence level, with values ​​ranging from [0, 1]; a diagram illustrating the imputation confidence level is shown below. Figure 5 As shown; Figure 5 The distribution of imputation_confidence under different modalities and the weighted interval of downstream loss are shown; Figure 5The dashed lines in the text represent low-weight thresholds. Figure 5 The horizontal axis represents multi-source data, including environmental sensor data, image data, physiological data, etc.; the vertical axis represents the imputation confidence.

[0146] The downstream world model reads this imputation provenance flag during training or inference and performs the following operations on the completed samples:

[0147] In loss calculation, weights are reduced by imputation_confidence;

[0148] In uncertainty modeling, an additional aleatoric variance is introduced.

[0149] During the evaluation, distinguish between observed sample indicators and supplementary sample indicators to avoid the supplementary quality being masked.

[0150] It's important to note that the imputation provenance flag is stored synchronously with the data tensor and must not exist solely as an external documentation file. This means that the imputation provenance flag must maintain a spatial correspondence with the data tensor in the form of a tensor or structured table. This ensures that training, validation, Leave-One-Cohort-Out (LOCO) evaluation, and counterfactual auditing retain the source information of the imputation even after arbitrary partitioning, and will not be lost due to data splitting, migration, or format conversion. LOCO is used to evaluate generalization ability across regions, seasons, and teams.

[0151] The imputation provenance flag enables the downstream world model to distinguish between observed samples and imputed samples, obtaining reliable inputs in multiple stages such as evaluation indicators, uncertainty modeling, and loss weighting, thus achieving reliable uncertainty modeling in the downstream model.

[0152] S105: Verify the multi-source data after hierarchical completion using the intermodal consistency verification rule set, generate a misalignment report and output it.

[0153] The inter-modal consistency verification rule set is an extensible, domain-driven, and visual set of readable rules describing the physical or logical relationships that modalities should satisfy. The inter-modal consistency verification rule set includes at least five categories of rules: irrigation-soil moisture, heating settings-air temperature, ventilation settings-temperature and humidity, supplemental lighting-PAR, and CO2 settings-CO2 concentration. These rules are not implicit losses or pre-trained objectives, but rather explicit, readable, auditable, and directly verifiable judgment criteria by agronomists.

[0154] The specific process of S105 is shown in G1-G4.

[0155] G1: Verify the multi-source data after hierarchical completion through the inter-modal consistency verification rule set; the inter-modal consistency verification rule set includes multiple rules, each of which includes at least the rule type, triggering condition and expected response.

[0156] The inter-modal consistency verification rule set is an extensible domain rule set, and the specific inter-modal consistency verification rule set is shown in Table 1. The inter-modal consistency verification and handling process is as follows: Figure 6 As shown.

[0157] Table 1

[0158] Irrigation - Soil Moisture Irrigation events ≥ threshold Soil moisture gradient > 0 Heating settings - air temperature t_heat_sp increased Temperature rises Ventilation settings—temperature and humidity t_vent_sp down Temperature drop + humidity drop Supplemental lighting — PAR lamp_sp is on PAR followed the rise <![CDATA[CO2 setting — CO2 concentration]]> co2_sp increased <![CDATA[CO2 follows and rises]]> Phenological phases—image greenness Entering vegetarian Greenness index above the threshold

[0159] Figure 6 The process for verifying and handling intermodal consistency was demonstrated. Figure 6 The system provides six rule types: irrigation-soil moisture, heating-air temperature, ventilation-temperature and humidity, supplemental lighting-PAR, CO2-concentration, and phenological phase-image greenness. Each rule type represents a complete event-driven validation unit.

[0160] For each of the above rules, perform verification according to the following dimensions:

[0161] Trigger events (identify whether management events or state changes have occurred);

[0162] Detection window (sets the time or phase range to be monitored after a triggered event occurs);

[0163] Desired direction (defining the expected trend of change in response channel data within the detection window);

[0164] Actual response (reading the actual changes in the response channel within the detection window);

[0165] Deviation magnitude (quantifying the degree to which the actual response deviates from the expected direction);

[0166] The system generates a verification conclusion for each rule based on the deviation magnitude. The verification conclusion is divided into three levels: failure, pass, and warning.

[0167] Based on the validation results, the samples were subjected to differential processing, as follows:

[0168] Pass samples are directly retained for training; warning samples are trained with reduced weight; fail samples are divided into two branches according to specific rules and deviation magnitude: either excluded from control simulation or transferred to the exception queue for manual review.

[0169] G2: For each rule, determine whether the multi-source data after hierarchical completion meets the triggering conditions of that rule.

[0170] For example, the heating setting of multi-source data after hierarchical completion—temperature t_heat_sp is increased to determine that the multi-source data meets the triggering conditions of the intermodal consistency verification rule set.

[0171] G3: If satisfied, then check whether the response channel data corresponding to the multi-source data after hierarchical completion meets the expected response of this rule.

[0172] For example, the heating setting of the multi-source data after hierarchical completion—the temperature t_heat_sp is increased, which determines that the multi-source data meets the triggering conditions of the intermodal consistency verification rule set, verifies the rise of the temperature data corresponding to the multi-source data after hierarchical completion, and determines that the temperature data corresponding to the multi-source data after hierarchical completion meets the expected response of the rule.

[0173] G4: For samples that meet the triggering conditions but whose response channel data does not meet the expected response rules, generate and output the misalignment report corresponding to the sample; the misalignment report shall include at least the misalignment timestamp and phase, the violated rule, the data source involved, the deviation magnitude, and the recommended review action.

[0174] For each sample that violates a consistency rule, a misalignment report is output. This report includes, but is not limited to, misaligned timestamps and phases, the violated rule, the data source involved, the magnitude of the deviation, and recommended review actions. The misalignment report is presented in a visual matrix format to facilitate quick identification of data quality issues. The consistency rule describes readable rules governing the physical or logical relationships that modalities should satisfy, such as soil moisture should increase after an irrigation event.

[0175] Consistency verification results are categorized into three levels: pass, warning, and fail. Samples receiving a warning can continue training but with reduced weights; samples receiving a fail are placed in an exception queue and are excluded from the control simulation training set by default. This process prevents timestamp misalignment in a single modality from contaminating action-state relationship learning.

[0176] The intermodal consistency verification rule set and misalignment report in this solution enable data engineers to quickly locate misalignment events, misalignment sources, and misalignment periods, thus realizing the visualization of data quality issues.

[0177] The output format of this solution is compatible with downstream world model terminology and reference architecture, agricultural action timestamp (AAT) specifications such as agricultural multimodal training data exchange format, and can serve as a de facto reference implementation for such standards.

[0178] The following describes specific embodiments and... Figure 7 The technical solution of this application will be further described. Figure 7 This is a schematic diagram of the overall process of this solution. Figure 7 The presentation showcases a five-stage process and main data flow, including phenological phase calibration, phase-based resampling, cross-source timestamp fault-tolerant alignment, missing data completion and addition of imputation provenance flags, and intermodal consistency verification.

[0179] Figure 7 Output a multimodal tensor quantization dataset with unified phenological phase anchoring, including an imputation provenance flag, a consistency verification report, and cross-source timestamp offset logs.

[0180] Figure 7 The presentation showcases a five-stage process and main data flow, including phenological phase calibration, phase-based resampling, cross-source tolerance alignment, missing data completion and data source / source traceability metadata (provenance), and inter-modal consistency verification. Provenance records how a piece of data was generated (observation, completion method, confidence level, etc.).

[0181] Figure 7 In the middle, the raw input layer contains four types of raw multi-source data: management events (such as irrigation, fertilization, pesticide application, and setpoint changes), environmental sensing (such as temperature, humidity, CO2, PAR, and VPD), physiological measurements (such as leaf area index, plant height, and harvesting records), and image data (such as image data collected by drones, remote sensing, and fixed cameras).

[0182] Phenological phase calibration and resampling stage: Raw multi-source data enters the phenological phase calibration module, which uses key phenological events combined with effective accumulated temperature (GDD) interpolation to generate a basic phenological phase sequence. Simultaneously, image phase verification is used for validation, and both are used together to output a validated phenological phase sequence. Subsequently, the system resamples the four types of raw data according to this phenological phase sequence, outputting a phase index table that records information such as the original time interval, GDD range, and distance to key phenological events corresponding to each phase grid point.

[0183] Cross-source timestamp fault-tolerant alignment stage: After phase resampling, the data enters the cross-source alignment module. The environmental sensor with the highest sampling rate and most stable timestamps is used as the primary source. Tolerance window matching is performed on each non-primary source. Samples whose timestamps fall within the tolerance window are merged into samples of the same time point, while samples exceeding the tolerance are transferred to an exception queue. After alignment is complete, an alignment log is output, recording information such as the matching relationship between non-primary source samples and primary source samples, timestamp offset, tolerance window size, and matching status.

[0184] Missing value imputation stage: After cross-source alignment, the system performs hierarchical imputation on missing values ​​still existing in the data tensor—short-term missing values ​​are imputed using linear interpolation or physically consistent interpolation, medium-term missing values ​​are imputed using homologous regression or multi-source imputation using cross-source association models, and long-term missing values ​​are marked as unimprovable and indicated in the mask tensor. Information such as the imputation method, confidence level, and phase distance to the nearest observation for each imputed sample is written into the proof table, making each imputed value traceable.

[0185] Core Outputs: The core products of the above stages converge into two tensors of the same shape: a data tensor, which stores the actual values ​​of each mode at each phase grid point; and a mask tensor, which marks whether the corresponding position in the data tensor is a valid observation, a completed value, or a long-term missing value that cannot be completed. The two sets of tensors work together to enable the downstream world model to clearly identify the data state of each value.

[0186] Consistency verification and final output stage: Combined data packets enter the intermodal consistency verification module, and multiple rules such as irrigation-soil moisture, heating-air temperature, ventilation-temperature and humidity, supplemental lighting-PAR, CO2-concentration, phase-image greenness are applied for verification. A consistency verification report is output, and each sample is marked as pass, warning or fail.

[0187] Ultimately, all the above components (data tensor, mask tensor, phase index, alignment log, proofance table, consistency report) are encapsulated into a unified world model dataset (the training dataset for downstream world models (such as downstream agricultural world models)) for use in training, validation, LOCO evaluation, and counterfactual auditing of downstream world models.

[0188] In a preferred embodiment, the output dataset is not a single numerical matrix, but a composite data packet consisting of data_tensor, mask_tensor, phase_index, source_index, action_event_table, imputation_provenance_table, alignment_log, and consistency_report. When the downstream world model reads this data packet, it obtains not only the values ​​of each modality, but also the source, confidence level, phase position, and whether each value was generated by completion.

[0189] The core protection of this solution lies in incorporating both "phenological phase anchoring" and "complete source tracing" into the data construction process simultaneously. Simply resampling by calendar or only performing missing data completion cannot yield the auditable agricultural world model training data described in this solution.

[0190] Example 1: Implementation process on a series of publicly available greenhouse crop datasets:

[0191] Step 1: Obtain data from multiple consecutive Open Greenhouse Challenge competitions, covering various crops such as cucumbers, cherry tomatoes, lettuce, and dwarf tomatoes;

[0192] Step 2: Perform phenological phase calibration on data from multiple public greenhouse challenge competitions. Key events are anchored using the publicly available cultivation calendar for each competition, and crop-specific GDD cumulative interpolation is used between events.

[0193] Step 3: Perform phase resampling on the greenhouse challenge data after phenological phase calibration, that is, resample all sources to a unified phase grid.

[0194] Step 4: Perform cross-source timestamp fault-tolerant alignment on the phase-resampled greenhouse challenge data to obtain aligned greenhouse challenge data; environmental sensing (e.g., 5-minute sampling) uses a smaller tolerance; management events use a larger tolerance. For sources with low timestamp accuracy, samples exceeding the tolerance are placed in an anomaly queue;

[0195] Step 5: If missing values ​​are detected in the aligned greenhouse challenge data, perform missing value completion. Forward imputation is used to complete the missing channels (such as partially supplementing the illumination setting channels). Each sample is marked with an imputationprovenance flag (imputation method identifier, imputation confidence, and distance to neighboring observation samples). The downstream world model can use this to reduce the weight of the imputed samples or to add additional variance in uncertainty modeling.

[0196] Step 6: Perform consistency verification on the completed greenhouse challenge data using the intermodal consistency verification rule set. Typical management events within the intermodal consistency verification rule set are checked using rules such as irrigation-soil moisture, heating-air temperature, and supplemental lighting-PAR.

[0197] Results: After the above operations, data from multiple periods were successfully aligned to a unified phenological phase grid. The cross-period comparability index proved that the physical consistency of the cross-period data distribution after phenological alignment was significantly better than that of the calendar alignment baseline.

[0198] Example 2: Implementation process on field crop datasets:

[0199] Step 1: Obtain multi-year, cross-state, and cross-hybrid field crop (e.g., corn) data;

[0200] Step 2: Perform phenological phase calibration on multi-year, cross-state, and cross-hybrid field crop data. Utilize key growth period records (e.g., VE / V6 / V12 / VT / R1 / R3 / R6, etc.), and inter-event interpolation is performed using GDD.

[0201] Step 3: Resample multi-year, cross-state, and cross-hybrid field crop data according to phase after phenological phase calibration. Meteorological, soil moisture, and satellite imagery data are resampled to a unified growth period phase grid.

[0202] Step 4: Perform cross-source timestamp fault-tolerant alignment on the multi-year, cross-state, and cross-hybrid field crop data after phase resampling. The tolerance for meteorological and soil moisture data is relatively small, while the tolerance for satellite imagery and administrative events is relatively large (field event granularity is coarser).

[0203] Step 5: If missing values ​​are detected in the aligned multi-year, cross-state, and cross-hybrid field crop data, imputation is performed. For image data missing due to cloud cover, regression analysis is performed using images of the same variety in neighboring areas at the same growth stage, with an imputation provenance flag added.

[0204] Step 6: Perform consistency verification on the completed multi-year, cross-state, and cross-hybrid field crop data using the inter-modal consistency verification rule set. Verify the irrigation event-soil moisture gradient rule; skip the rules for samples in arid regions without irrigation events.

[0205] Results: Cross-state and cross-year field crop data were successfully aligned to a unified phase grid, and cross-state meteorological distributions were significantly closer under the same phase than under the same calendar date.

[0206] Example 3: Application on animal multimodal behavior datasets:

[0207] Step 1: Acquire three sources of data from animals (such as dairy cows): vision, neck ring sensor, and rumination monitoring.

[0208] Step 2: Replace the phenological phase with the lactation phase, use the calving date as the anchor point to divide the phase according to the number of lactation weeks, and calibrate the three-source data of animal (e.g. dairy cow) vision, neck ring sensor, and rumination monitoring according to the phase division by the number of lactation weeks;

[0209] Steps 3 through 6: Same as the method described above, but with different phase semantics. The consistency rules are replaced with animal-specific rules, such as the consistency between feeding events and delayed rumination initiation, and the consistency between neck ring acceleration and visual behavioral tags.

[0210] Results: Animal multimodal data were aligned on a unified lactation phase grid, demonstrating the cross-domain applicability of this method (including both plants and animals). The above phenological phase sequences used the growth period phase for crops (tomatoes, cucumbers, lettuce, corn, wheat, etc.) and the lactation, fattening, and reproduction phases for animals (cattle, pigs, poultry).

[0211] It should be noted that the technical solutions described in this application can be implemented in various ways based on different technical considerations, as follows:

[0212] Improvement 1: Phenological phase calibration can be automated by combining multimodal image-driven models, reducing the dependence on the quality of cultivation records;

[0213] Improvement 2: The cross-source tolerance window can be data-driven and adaptive (based on statistics of historical offset distribution), reducing manual configuration;

[0214] Improvement 3: The imputation provenance flag can be expanded to a probability distribution instead of a point estimate, allowing for better integration with downstream Bayesian models;

[0215] Improvement 4: The set of rules for intermodal consistency verification can be defined by users using a domain-specific language (DSL), further reducing the barrier to entry.

[0216] Alternative solution:

[0217] Alternative Option 1: The phenological phase anchor point can be replaced with a continuous variable, the cumulative GDD value (instead of a discrete phase). This option retains continuity but loses the interpretability of discrete phases.

[0218] Alternative Solution 2: Cross-source timestamp fault-tolerant alignment is replaced with soft alignment based on Dynamic Time Warping (DTW). This solution is more robust to large offsets but has a higher computational cost.

[0219] Alternative Solution 3: Replace missing completion with conditional distribution sampling based on a generative model, outputting the sample distribution of each completed sample. This solution is more favorable for modeling uncertainty but requires training the completion model.

[0220] Alternative 4: The intermodal consistency verification rule set is implicitly learned by a neural network (based on a self-supervised task) instead of explicit rules. This approach automates rule discovery but loses auditability.

[0221] The mainstream implementation of this solution employs: discrete phase + tolerance window alignment + hierarchical completion + explicit rule verification. However, the scope of protection covers substantially equivalent implementations of the aforementioned alternative solutions.

[0222] The aforementioned improvements and alternatives can all be combined with the core method described in this application without departing from the scope of protection of this application. Those skilled in the art can choose an appropriate implementation method according to the specific application scenario.

[0223] The relevant technical parameters of this solution are described in Table 2.

[0224]

[0225] In this embodiment, phenological phase sequences are used as the basis for dividing unified phase grids, and crop phenological phase sequences are used as resampling anchors. This allows multimodal data across regions, seasons, and crops to be compared under a unified phase, ensuring that data under the same phenological phase grid corresponds to the same physiological development state of the crop. This eliminates temporal misalignment caused by regional, seasonal, and varietal differences, improving the portability of downstream world models when migrating across regions, seasons, or domains. Furthermore, a cross-source timestamp fault-tolerant alignment strategy is provided, setting a configurable tolerance window for legitimate timestamp offsets between sources. Samples within the tolerance are considered to be at the same time point, while those exceeding the tolerance are entered into an anomaly queue instead of being discarded entirely, improving the sample utilization rate of multi-source data.

[0226] Based on the above embodiments Figure 1 This application discloses a method for processing multimodal data, and also provides a corresponding apparatus for processing multimodal data, such as... Figure 8 As shown, the multimodal data processing device includes:

[0227] The calibration unit 801 is used to calibrate the phenological phase sequence for each data acquisition subject; wherein, the phenological phase sequence is the temporal position information corresponding to the multiple phenological stages, which are divided into multiple phenological stages according to the progress of phenological development.

[0228] The resampling unit 802 is used to resample the acquired multi-source time series data to a unified phase grid according to the phenological phase sequence, so as to obtain multi-source data resampled from each data source on the unified phase grid.

[0229] The fault-tolerant alignment unit 803 is used to perform cross-source timestamp fault-tolerant alignment on the phase resampled multi-source data to obtain aligned multi-source data.

[0230] The hierarchical completion unit 804 is used to perform hierarchical completion on the multi-source data with missing values ​​if missing values ​​are detected in the aligned multi-source data.

[0231] The verification output unit 805 is used to verify the multi-source data after hierarchical completion through the intermodal consistency verification rule set, generate a misalignment report and output it.

[0232] Furthermore, the calibration unit 801 includes:

[0233] The first acquisition module is used to acquire the key phenological event timestamps of the data collection subjects; wherein, the data collection subjects include at least greenhouse compartments, fields and individual animals; the key phenological event timestamps include at least the sowing date, seedling emergence date, flowering date and first harvest date;

[0234] The first calibration module is used to interpolate between two adjacent key phenological event timestamps using crop phenological development temperature thresholds to obtain continuous phase variables or discrete phases, thereby completing the process of calibrating the phenological phase sequence for each data acquisition subject.

[0235] Furthermore, the calibration unit 801 includes:

[0236] The second acquisition module is used to acquire the phenological characteristics of crop canopy images of the data acquisition subject; wherein, the phenological characteristics of crop canopy images include at least greenness index, coverage and fruit visibility;

[0237] The second calibration module is used to input the phenological features of crop canopy images into a pre-trained classifier, and output the phenological phase labels corresponding to the crop canopy images through the pre-trained classifier, so as to complete the process of calibrating the phenological phase sequence for each data acquisition subject.

[0238] Furthermore, the multi-source time-series data includes at least environmental sensing time-series data, image data, physiological measurement data, and management event data. The resampling unit 802 includes:

[0239] The third acquisition module is used to acquire the phenological phase intervals corresponding to each phase grid point in the unified phase grid.

[0240] The aggregation and statistics module is used to aggregate and statistically analyze environmental sensing time-series data that fall within the phenological phase interval, and obtain phase resampled environmental sensing data on a unified phase grid.

[0241] The first search module is used to perform nearest neighbor search on image data that falls within the phenological phase interval to obtain image resampled data at the unified phase grid point.

[0242] The second search module is used to perform nearest neighbor search on physiological measurement data that fall within the phenological phase interval to obtain physiological resampled data at a unified phase grid.

[0243] The aggregation module is used to aggregate the number of events and the intensity of events that fall within the phenological phase interval for managed event data, and to obtain event resampling data on a unified phase grid.

[0244] Furthermore, the fault-tolerant alignment unit 803 includes:

[0245] The fourth acquisition module is used to acquire the data with the highest sampling rate and stable timestamp from the multi-source data of phase resampling as the main source, and the remaining data as non-main sources.

[0246] The configuration module is used to obtain the timestamp of the main source corresponding to the main source and configure the corresponding tolerance window. The time length of the tolerance window is based on the time grid determined by the sampling rate of the main source, and is configured in combination with the sampling rate of the non-main source data to be aligned, the timestamp accuracy, the upload latency, and the historical offset distribution.

[0247] The fifth acquisition module is used to obtain the non-main source sample timestamp for each sample in each non-main source.

[0248] The verification module is used to store the samples that exceed the tolerance window when the timestamp of a non-primary source sample does not fall within the tolerance window into the exception queue and verify them.

[0249] The recording module is used to determine that the timestamp of the non-primary source sample and the timestamp of the primary source sample are the same time point when the timestamp of the non-primary source sample falls within the tolerance window, and to record the offset between the timestamp of the non-primary source sample and the timestamp of the primary source sample in the alignment log to obtain the aligned multi-source data.

[0250] Furthermore, the hierarchical completion unit 804 includes:

[0251] The first determination module is used to determine the duration of missing values ​​based on the missing values ​​if missing values ​​are detected in the aligned multi-source data.

[0252] The first determination module is used to determine that the aligned multi-source data is in a first missing state when the duration of the missing data is less than a first threshold, and to complete the multi-source data in the first missing state using linear interpolation or physical consistency interpolation.

[0253] The second determination module is used to determine that the aligned multi-source data is in a second missing state when the duration of the missing data is greater than or equal to the first threshold and less than the second threshold, and to complete the multi-source data in the second missing state by using regression imputation based on homologous neighbor samples or multi-source imputation based on cross-source association model.

[0254] The third determination module is used to determine that the aligned multi-source data is in a third missing state when the duration of the missing data is greater than or equal to the second threshold, and to mark the multi-source data in the third missing state as uncompleteable.

[0255] Furthermore, the verification output unit 805 includes:

[0256] The verification module is used to verify the hierarchical and completed multi-source data through the inter-modal consistency verification rule set. The inter-modal consistency verification rule set includes multiple rules, and each rule includes at least the rule type, triggering condition, and expected response.

[0257] The first judgment module is used to determine whether the multi-source data after hierarchical completion meets the triggering conditions of the rule for each rule.

[0258] The second judgment module is used to check whether the response channel data corresponding to the hierarchical and completed multi-source data meets the expected response of the rule if the hierarchical and completed multi-source data meets the triggering conditions of the rule.

[0259] The output generation module is used to generate and output a misalignment report for a sample that meets the triggering conditions but whose response channel data does not meet the rules. The misalignment report includes at least the misalignment timestamp and phase, the violated rules, the data source involved, the deviation magnitude, and the recommended review action.

[0260] Furthermore, the multimodal data processing device also includes:

[0261] An add module is used to add a completion source marker for each completed sample; wherein the completion source marker consists of at least the completion method identifier, completion confidence, phase distance to the nearest observed sample, list of sources used, and timestamp generated by the completion.

[0262] Furthermore, the multimodal data processing device also includes:

[0263] The sixth acquisition module is used to acquire the key phenological event timestamps of the data acquisition subject, as well as the phenological characteristics of the crop canopy images of the data acquisition subject;

[0264] The second determining module is used to take the phenological phase sequence marked with the timestamps of key phenological events as the main sequence and the phenological phase sequence marked with the phenological features of crop canopy images as the verification sequence.

[0265] The calculation module is used to calculate the conflict value between the main sequence and the check sequence;

[0266] The storage module is used to store the corresponding sample in the anomaly queue if the conflict value is greater than a preset threshold.

[0267] The beneficial effects of some embodiments of this device can be referred to the beneficial effects of the corresponding method embodiments, and will not be repeated here.

[0268] This application embodiment also provides a storage medium, the storage medium including stored instructions, wherein, when the instructions are executed, the device where the storage medium is located is controlled to perform the multimodal data processing method described above.

[0269] This application also provides an electronic device, the structural schematic diagram of which is shown below. Figure 9 As shown, it specifically includes a memory 901 and one or more instructions 902, wherein one or more instructions 902 are stored in the memory 901 and configured to be executed by one or more processors 903 to perform the above-mentioned multimodal data processing method.

[0270] The steps in the methods of the various embodiments of this application can be adjusted, combined, and deleted according to actual needs. It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0271] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0272] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for processing multimodal data, characterized in that, The method includes: A phenological phase sequence is defined for each data acquisition subject; wherein, the phenological phase sequence is the temporal position information corresponding to the multiple phenological stages, which are divided into multiple phenological stages according to the progress of phenological development. According to the phenological phase sequence, the acquired multi-source time series data are resampled to a unified phase grid to obtain multi-source data from each data source after resampling at the unified phase grid. The phase resampled multi-source data is cross-source timestamp fault-tolerant alignment is performed to obtain aligned multi-source data. If missing values ​​are detected in the aligned multi-source data, the multi-source data with missing values ​​will be filled in hierarchically. The multi-source data after hierarchical completion is verified by using the intermodal consistency verification rule set, and a misalignment report is generated and output.

2. The method according to claim 1, characterized in that, The calibration of the phenological phase sequence for each data acquisition subject includes: Obtain key phenological event timestamps for the data collection subjects; wherein, the data collection subjects include at least greenhouse compartments, fields, and individual animals; the key phenological event timestamps include at least the sowing date, seedling emergence date, flowering date, and first harvest date; Between two adjacent key phenological event timestamps, interpolation is performed using crop phenological development temperature thresholds to obtain continuous phase variables or discrete phases, thus completing the process of calibrating the phenological phase sequence for each data collection subject.

3. The method according to claim 1, characterized in that, The calibration of the phenological phase sequence for each data acquisition subject includes: Acquire phenological characteristics of crop canopy images of the data collection subject; wherein, the phenological characteristics of the crop canopy images include at least greenness index, coverage, and fruit visibility; The phenological features of the crop canopy images are input into a pre-trained classifier, which then outputs the phenological phase labels corresponding to the crop canopy images to complete the process of labeling the phenological phase sequence for each data acquisition subject.

4. The method according to claim 1, characterized in that, The multi-source time-series data includes at least environmental sensor time-series data, image data, physiological measurement data, and management event data. The multi-source time-series data is resampled to a unified phase grid according to the phenological phase sequence to obtain multi-source data resampled from each data source at the unified phase grid, including: Obtain the phenological phase intervals corresponding to each phase grid point in the unified phase grid; For the environmental sensing time series data, the environmental sensing time series data falling within the phenological phase interval are aggregated and statistically analyzed to obtain phase resampled environmental sensing data at the unified phase grid point; For the image data, a nearest neighbor search is performed on the image data falling within the phenological phase interval to obtain the image resampled data at the unified phase grid point; For the physiological measurement data, the physiological measurement data falling within the phenological phase interval are searched for nearest neighbors to obtain the physiological resampled data at the unified phase grid point; For the management event data, the event counts and event intensities that fall within the phenological phase interval are aggregated to obtain event resampling data at the unified phase grid.

5. The method according to claim 1, characterized in that, The step of performing cross-source timestamp fault-tolerant alignment on the phase-resampled multi-source data to obtain aligned multi-source data includes: From the multi-source data of phase resampling, the data with the highest sampling rate and the timestamp that meets the stability condition is selected as the main source, and the remaining data is selected as the non-main source. Obtain the timestamp of the main source corresponding to the main source and configure the tolerance window; the time length of the tolerance window is based on the time grid determined by the sampling rate of the main source, and is configured in combination with the sampling rate of the non-main source data to be aligned, the timestamp accuracy, the upload latency, and the historical offset distribution. For each sample in each non-primary source, obtain the timestamp of the non-primary source sample for that sample; When the timestamp of the non-primary source sample does not fall within the tolerance window, the sample that exceeds the tolerance corresponding to the timestamp of the non-primary source sample not falling within the tolerance window is stored in the exception queue and reviewed. When the timestamp of the non-primary source sample falls within the tolerance window, it is determined that the timestamp of the non-primary source sample and the timestamp of the primary source are at the same time point, and the offset between the timestamp of the non-primary source sample and the timestamp of the primary source is recorded in the alignment log to obtain the aligned multi-source data.

6. The method according to claim 1, characterized in that, If missing values ​​are detected in the aligned multi-source data, hierarchical completion is performed on the multi-source data with missing values, including: If missing values ​​are detected in the aligned multi-source data, the duration of the missing values ​​is determined based on the missing values. When the duration of the missing data is less than the first threshold, the aligned multi-source data is determined to be in the first missing state, and linear interpolation or physical consistency interpolation is used to complete the multi-source data in the first missing state. When the duration of the missing data is greater than or equal to the first threshold and less than the second threshold, the aligned multi-source data is determined to be in the second missing state. Regression imputation based on homologous neighboring samples or multi-source imputation based on cross-source association model are used to complete the multi-source data in the second missing state. When the duration of the missing data is greater than or equal to the second threshold, the aligned multi-source data is determined to be in a third missing state, and the multi-source data in the third missing state is marked as uncompleted.

7. The method according to claim 1, characterized in that, The process of verifying the multi-source data after hierarchical completion using a set of intermodal consistency verification rules, generating and outputting a misalignment report includes: The multi-source data after hierarchical completion is verified using a set of inter-modal consistency verification rules. The set of inter-modal consistency verification rules includes multiple rules, and each rule includes at least a rule type, triggering conditions, and expected response. For each rule, determine whether the multi-source data after hierarchical completion meets the triggering conditions of that rule; If satisfied, then check whether the response channel data corresponding to the multi-source data after hierarchical completion meets the expected response of the rule; For a sample that meets the triggering conditions but whose response channel data does not meet the expected response rules, a misalignment report corresponding to the sample is generated and output; wherein, the misalignment report includes at least the misalignment timestamp and phase, the violated rule, the data source involved, the deviation magnitude, and the suggested review action.

8. The method according to claim 1, characterized in that, After detecting missing values ​​in the aligned multi-source data and performing hierarchical completion on the multi-source data with missing values, the method further includes: Add a completion traceability marker to each completed sample; The completion traceability flag consists of at least the completion method identifier, completion confidence level, phase distance to the nearest observed sample, list of sources used, and timestamp generated by the completion.

9. The method according to claim 1, characterized in that, Also includes: Obtain key phenological event timestamps from the data collection subject, and obtain phenological characteristics of crop canopy images from the data collection subject; The phenological phase sequence calibrated by the timestamps of the key phenological events is used as the main sequence, and the phenological phase sequence calibrated by the phenological features of the crop canopy image is used as the verification sequence. Calculate the conflict value between the main sequence and the check sequence; If the conflict value is greater than a preset threshold, the corresponding sample will be stored in the anomaly queue.

10. A multimodal data processing apparatus, characterized in that, The device includes: A calibration unit is used to calibrate the phenological phase sequence for each data acquisition subject; wherein, the phenological phase sequence is the temporal position information corresponding to the multiple phenological stages, which are divided into multiple phenological stages according to the progress of phenological development. The resampling unit is used to resample the acquired multi-source time-series data to a unified phase grid according to the phenological phase sequence, so as to obtain multi-source data resampled from each data source on the unified phase grid. The fault-tolerant alignment unit is used to perform cross-source timestamp fault-tolerant alignment on the phase resampled multi-source data to obtain aligned multi-source data. The hierarchical completion unit is used to perform hierarchical completion on the multi-source data with missing values ​​if missing values ​​are detected in the aligned multi-source data. The verification output unit is used to verify the multi-source data after hierarchical completion through the intermodal consistency verification rule set, generate a misalignment report and output it.