A big data-based silicon fertilizer treatment and hamimelon quality correlation analysis system

CN122656471APending Publication Date: 2026-08-28哈密瓜鲜果农业科技发展有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611116754.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

现实中,此类多维度完整对齐的样本量难以达到统计建模的基本要求

Benefits of technology

1、通过构建基于成对偏好比较的群体偏好梯度场,避免了现有技术试图精确拟合绝对感官评分所面临的噪声干扰和尺度漂移难题。本发明不要求消费者给出评分等级,而是从自然语言评价和历史购买记录中自动提取“A优于B”的相对偏好关系,以此训练偏好比较网络。该网络输出的是理化空间中群体偏好的梯度方向,而非绝对分值。这一设计使得系统对消费者个体差异、评价噪声和市场偏好漂移具有天然的鲁棒性,优化决策只需依赖相对排序信息即可有效进行。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122656471A_ABST
    Figure CN122656471A_ABST
Patent Text Reader

Abstract

The application provides a silicon fertilizer treatment quality correlation analysis system for Hami melon based on big data, and belongs to the field of intelligent agriculture and big data application.The system comprises a quality prediction module, a group preference perception module and an optimization decision module.The quality prediction module is used for predicting physicochemical quality indexes of Hami melon based on growth environment data and silicon fertilizer application schemes of the Hami melon.The group preference perception module is used for acquiring feedback data of consumers on the Hami melon, and constructing a preference gradient field reflecting the group consumer sensory preference direction based on the feedback data.The optimization decision module is used for searching an optimal silicon fertilizer application scheme based on a mapping relationship between the physicochemical quality indexes output by the quality prediction module and the candidate silicon fertilizer schemes, and taking the preference gradient field as an optimization target.The group preference gradient field based on pair-wise preference comparison is constructed, so that the noise interference and scale drift problems faced by the prior art in an attempt to accurately fit the absolute sensory score are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart agriculture and big data applications, and specifically relates to a big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality. Background Technology

[0002] Silicon has been proven to have multiple positive effects on the quality formation of cucurbitaceous crops such as cantaloupe. Appropriate silicon application can improve leaf photosynthetic efficiency, enhance plant disease resistance, and significantly increase the content of soluble solids, vitamin C, flesh firmness, and characteristic aroma compounds in fruits by regulating the activity of key enzymes such as sucrose phosphate synthase and sucrose synthase. However, existing technologies face several unavoidable technical obstacles in translating the correlation between silicon fertilizer treatment and cantaloupe quality into an engineering-deployable decision-making system.

[0003] The first obstacle is the extreme scarcity of high-quality correlation data. Silicon fertilizer, unlike macronutrients such as nitrogen, phosphorus, and potassium, is applied in small quantities in agricultural production, farmers' practices are often arbitrary, and agricultural records are generally lacking. Establishing a reliable correlation model requires complete, aligned samples across multiple years, plots, and varieties, encompassing "soil available silicon content—silicon fertilizer type, application time and dosage—environmental time-series data throughout the entire growth period—post-harvest multidimensional quality detection values." In reality, the sample size for such complete multidimensional alignment is insufficient to meet the basic requirements of statistical modeling. Furthermore, in-situ continuous monitoring of soil available silicon content still lacks low-cost sensors, relying on laboratory chemical extraction methods, resulting in data feedback that lags significantly behind production decision-making needs. Accurate labeling of multidimensional fruit quality relies on destructive sampling, gas chromatography-mass spectrometry, texture analyzers, and well-trained sensory evaluation teams; single-sample testing is costly, and large-scale acquisition is economically infeasible.

[0004] The second obstacle is the ambiguity and time-varying nature of the concept of quality. "Quality" itself is not a scalar quantity that can be defined by a single physicochemical indicator. Consumers' sensory experience of cantaloupe involves multiple dimensions, such as sweetness, acidity, crispness, degree of mushy texture, and aroma intensity, with complex interactions and compensatory relationships between these dimensions. The same combination of physicochemical indicators may elicit drastically different preference evaluations from different consumer groups, in different seasons, and in different market regions. Existing technologies either simplify quality to a single soluble solids content, losing a significant amount of sensory information; or attempt to integrate multiple indicators with fixed weights, failing to reflect the dynamic shifts in consumer preferences. Attempting to obtain accurate scores through a human sensory evaluation panel faces extremely high costs and the difficulty of continuous calibration of evaluators, making it unsustainable for commercial promotion.

[0005] The third obstacle is the confusion between correlation and causation. The final quality of Hami melons is the result of a complex interaction between genotype, environmental conditions, and cultivation management practices. Statistical correlations between silicon fertilizer application and quality indicators in observational data are often mixed with interference from factors such as irrigation levels, nitrogen fertilizer management, and farmers' technical literacy. Recommendation models built directly based on such confounding correlations show a sharp decline in predictive ability when extended from one production area to another or from one variety to another, exhibiting serious overfitting problems. Current technology lacks effective means to decouple the net contribution of silicon fertilizer from observational data.

[0006] The fourth obstacle is the economic feasibility of system integration. A complete decision-making system encompassing sensor networks, edge computing devices, cloud platform analytics engines, and continuously updated and maintained AI models incurs high research and development and deployment costs. While silicon fertilizer raw materials are inexpensive, if the system cost far exceeds the quality premium that silicon fertilizer input can bring, a sustainable business model will be difficult to establish. Existing smart agriculture platforms are mostly geared towards bulk grain crops or high-value greenhouse vegetables; dedicated decision-making tools for cantaloupes, focusing specifically on silicon fertilizer as a micronutrient, are still lacking.

[0007] In summary, the existing technology lacks a solution that can solve data sparsity at a reasonable cost, robustly extract quality optimization directions from highly noisy consumer feedback, unbiasedly estimate the causal effects of silicon fertilizer, and achieve cross-variety and cross-production area generalization of silicon fertilizer treatment-hami melon quality correlation analysis and decision optimization. Summary of the Invention

[0008] In view of the above situation, the main objective of this invention is to propose a big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality, in order to solve the aforementioned technical problems.

[0009] This invention proposes a big data-based system for analyzing the correlation between silicon fertilizer treatment and Hami melon quality, the system comprising: The quality prediction module is used for: Based on the growth environment data of Hami melon and the silicon fertilizer application plan, the physicochemical quality indicators of Hami melon are predicted. The group preference perception module is used for: Obtain consumer feedback data on cantaloupes and construct a preference gradient field that reflects the sensory preference direction of the group of consumers based on the feedback data; The optimization decision-making module is used for: Using the preference gradient field as the optimization objective, the optimal silicon fertilizer application scheme is searched based on the mapping relationship between the physicochemical quality indicators output by the quality prediction module and the candidate silicon fertilizer schemes.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By constructing a group preference gradient field based on pairwise preference comparison, this invention avoids the noise interference and scale drift problems encountered by existing technologies that attempt to accurately fit absolute sensory ratings. This invention does not require consumers to provide rating levels; instead, it automatically extracts the relative preference relationship of "A is better than B" from natural language evaluations and historical purchase records to train the preference comparison network. The network outputs the gradient direction of group preferences in physicochemical space, rather than absolute scores. This design makes the system inherently robust to individual consumer differences, evaluation noise, and market preference drift; optimization decisions can be effectively made solely based on relative ranking information.

[0011] 2. By modeling time-varying preference density fields, the system can perceive and predict trend changes in consumer preferences. Using neural network differential equations, snapshots of the preference field observed at discrete time points are modeled as samples of a continuous dynamic system. This allows for the prediction of preference status at harvest and market launch at the beginning of the planting season, guiding adjustments to silicon fertilizer programs to align with market trends. This capability is not present in existing static quality evaluation systems.

[0012] 3. By introducing causal effect estimation based on instrumental variables, unbiased decoupling of the net effect of silicon fertilizer treatment on population preference was achieved. Using exogenous shocks unrelated to farmers' management levels, such as silicon fertilizer promotion events and accidental logistical disruptions, as instrumental variables, a deep instrumental variable network separated the pure causal effect curve of silicon fertilizer from the observation data, overcoming the defect of traditional association analysis where the mixture of environmental and management measures leads to model failure across production areas. The obtained causal effect curve also serves as the prior mean function for Bayesian optimization, significantly accelerating the convergence speed of the optimization search.

[0013] 4. By constructing a mechanism digital twin submodule using isotope tracing and a physical information neural network, the quality prediction module is endowed with the ability to maintain reasonable predictive behavior even in sparse training data regions. The introduction of physical constraint loss and causal path constraint loss embeds prior knowledge such as known carbon mass conservation, enzyme reaction kinetics, and causal pathway directions into the network training, ensuring that the model does not produce outputs that violate agronomic common sense when extrapolating to unseen environmental scenarios, thus improving the reliability of the prediction results.

[0014] 5. By utilizing silicon state proxy fingerprinting and domain adversarial migration technology, the high-precision model of the core base station was adapted to a wide range of target fields at low cost. Only one UAV hyperspectral flight and basic soil testing are required in the target field. The super twin can then be migrated and adapted to the variety and microclimate conditions of the target field through the domain adversarial network, significantly reducing the deployment and replication costs of the system. This makes it economically possible to support commercial applications in multiple production areas with a single precision experiment.

[0015] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by means of embodiments of the invention. Attached Figure Description

[0016] Figure 1 This invention presents the architecture of a big data-based silicon fertilizer treatment correlation analysis system for Hami melon quality. Figure 2 This is a diagram of the internal architecture of the group preference perception module of the present invention; Figure 3 This is a diagram of the internal architecture of the quality prediction module of the present invention; Figure 4 This is a diagram of the internal architecture of the optimization decision module of the present invention. Detailed Implementation

[0017] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0018] These and other aspects of the embodiments of the present invention will become clear from the following description and accompanying drawings. In these descriptions and drawings, some specific embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention; however, it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0019] Example 1: A big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality; Please see Figure 1 The purpose of this embodiment is to explain in detail the overall architecture of the system of the present invention and the basic composition and cooperation relationship of the three core modules.

[0020] This embodiment proposes a big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality. The system's architecture comprises three core modules: a quality prediction module, a consumer preference perception module, and an optimization decision-making module. Logically, these three modules form a closed loop from field data input to consumer preference perception and then to optimization decision output, continuously evolving based on consumer feedback in each planting season.

[0021] The quality prediction module is used to predict the physicochemical quality indicators of Hami melons based on their growth environment data and silicon fertilizer application schemes. Growth environment data includes, but is not limited to: atmospheric temperature and humidity sequence during the growth period, photosynthetically active radiation, soil temperature and moisture content, soil pH and electrical conductivity, and soil available nutrient content. Silicon fertilizer application schemes include, but are not limited to: the type of silicon fertilizer (potassium silicate, sodium silicate, nano-silica, steel slag silicon fertilizer, etc.), application method (root drip irrigation, foliar spraying), application period, and concentration or dosage for each application. Physicochemical quality indicators include, but are not limited to: soluble solids content, titratable acidity, flesh firmness, peel thickness, and concentration of major volatile aroma compounds. The output of the quality prediction module is a multi-dimensional physicochemical quality vector, which serves as the input for subsequent group preference perception and optimization decision-making modules.

[0022] The core innovation of this invention lies in its group preference perception module, which acquires consumer feedback data on cantaloupes and constructs a preference gradient field reflecting the sensory preferences of the group of consumers based on this feedback data. Feedback data refers to unstructured or semi-structured evaluations voluntarily submitted by consumers after purchasing and consuming cantaloupes, through methods such as scanning QR codes. This includes, but is not limited to, natural language text descriptions, voiceprint recordings, and optional tracing of historical purchase records. The group preference perception module does not attempt to convert this feedback into an absolute "preference score," but rather learns a preference gradient field from it. The preference gradient field is a vector field defined on the physicochemical quality space. At any physicochemical quality point, the gradient direction indicates the direction of physicochemical quality improvement that the group of consumers perceive as "tastier." In other words, the module answers the question: "If we want consumers to like the next batch of cantaloupes more likely, in which direction should we adjust the physicochemical quality?" rather than "What score would consumers give this cantaloupe?"

[0023] To quantify this concept, the physicochemical quality space is defined as an n-dimensional vector space. , where n equals the number of physicochemical quality indicators measured. Define the preference score function. Any physicochemical quality vector Mapped to a real-valued preference score, Let x represent the nth component of the physicochemical quality vector. For two physicochemical quality vectors... and , If the group of consumers tends to prefer A over B, then the preference gradient field G(x) is defined as follows: ; in Let G(X) represent the partial derivative of the preference score with respect to the j-th physicochemical quality index. Its sign indicates the direction of the influence of the increase or decrease of this index on the preference score, and its magnitude indicates the intensity of the influence. G(X) constitutes a vector field in the physicochemical quality space, with its field lines pointing in the steepest direction of preference enhancement.

[0024] It is important to note that the aforementioned preference gradient field G(x) indicates the steepest unconstrained direction of preference enhancement in the physicochemical quality space. In actual cantaloupe agronomic scenarios, there are physiological coupling constraints between physicochemical quality indicators (e.g., soluble solids accumulation is often accompanied by a decrease in flesh firmness), and the mapping defined by the quality prediction module... The image space is only A low-dimensional feasible manifold Therefore, G(x) may point to the manifold. Beyond these infeasible directions, quality improvement in these directions cannot be directly achieved through any silicon fertilizer solution. In this system, G(x) does not serve as a linear search path for the optimizer along its direction, but rather provides prior guidance on preference improvement trends for Bayesian optimization—the optimizer operates within the feasible manifold. The upsampling search only adopts directional components that both satisfy agronomic constraints and improve preference scores. This design ensures that the optimization process always takes place within the physically realizable silicon fertilizer solution space.

[0025] The optimization decision module, using the preference gradient field as the optimization objective, searches for the optimal silicon fertilizer application scheme based on the mapping relationship between the physicochemical quality indicators output by the quality prediction module and candidate silicon fertilizer schemes. The optimization decision module receives the candidate scheme-quality correspondence of the current field after mapping by the quality prediction module, queries the current preference gradient field provided by the group preference perception module, and calculates the relative advantage score of each candidate scheme in the preference field. Through multiple rounds of iterative search, it outputs a set of silicon fertilizer application schemes that maximize the group preference advantage, which is directly sent to the field fertigation equipment for execution or recommended to growers for reference.

[0026] Let the space of candidate silicon fertilizer schemes be T, and the quality prediction module defines a mapping function. , where E represents the growth environment parameter space. For a given fixed environment ,plan The corresponding predicted physicochemical quality vector is The goal of the optimization decision-making module is to find the optimal solution. , so that: ; Constrained by agronomic feasibility boundary conditions ,in Let be the j-th constraint function, and m be the number of constraints. This optimization framework establishes a direct mathematical link between consumer subjective preferences and controllable field operations.

[0027] The following simplified single-season operation process illustrates the collaboration among the three components. Assume the target field has sandy loam soil with low organic matter content. Before planting, the grower inputs basic field environmental data into the quality prediction module to obtain the predicted physicochemical quality vector of the default silicon fertilizer scheme under the current conditions. The group preference perception module provides the current preference gradient field in... The steepest gradient direction nearby indicates that "increasing sugar content and ester aroma concentration" is the main direction for improving consumer preference. Based on this, the optimization decision module generates several candidate silicon fertilizer adjustment schemes, calls the quality prediction module to obtain the predicted physicochemical points corresponding to each scheme, calculates its score in the preference field, and outputs the optimal scheme after searching: applying nano-silicon to the roots during the early fruit expansion stage and supplementing silicon on the leaves during the sugar accumulation stage. This scheme is implemented. After fruit harvest, consumer feedback flows back to the group preference perception module, updating the preference gradient field and providing updated target directions for optimization in the next season.

[0028] Example 2: Training a pairwise preference mining and preference comparison network; Please see Figure 2 This embodiment details the specific implementation of the pairwise preference mining unit and the preference comparison network training unit in the group preference perception module, and how to automatically construct pairwise preference relationships using natural language evaluation and consumer history records.

[0029] The pairwise preference mining unit is responsible for extracting structured preference comparison information from the raw data of consumer feedback. Each cantaloupe on the market has a label with a unique traceability code, which consumers scan to access the evaluation interface. Instead of a traditional star rating scale, the interface features an open-ended text input box with the prompt, "Please describe the taste of this melon, or compare it with other melons you've eaten before." Additionally, there's a non-mandatory "record audio" button at the bottom of the page, allowing consumers to upload the sound of their first bite. This design reduces the evaluation burden on consumers, while the acquired natural language text inherently contains rich comparative semantics.

[0030] Once a review text is submitted, the system calls a comparative semantic parser based on a pre-trained language model for processing. This parser is based on a BERT model fine-tuned for the agricultural field, with the final classification head retrained for comparative sentence structures. The training corpus consists of thousands of pre-collected real consumer reviews, manually labeled with three categories: "comparison object," "comparison dimension," and "comparison direction." Comparison dimensions include sweetness, crispness, melt-in-your-mouth texture, aroma, and peel thickness; comparison directions are categorized as "better than," "worse than," and "equal to." For example, the review "This batch is much sweeter than last month's batch" is parsed as: Comparison object = this batch vs. last month's batch, Comparison dimension = sweetness, Comparison direction = better than.

[0031] When the comparison object is identified to involve a consumer's own historical purchase records, the system automatically traces the consumer's past barcode scanning records. Assuming the traceback reveals that the consumer purchased batch A last month, its corresponding post-harvest physicochemical quality vector is... =(TSS=13.5%, Hardness=4.1kg / cm) 2 , methyl acetate = 0.8 μg / g,...), while the physicochemical quality vector of batch B corresponding to this evaluation is =(TSS=15.2%, Hardness=3.7kg / cm) 2 (Benzyl acetate = 1.2 μg / g, ...). The system automatically constructs a structured triplet (Melon A, Melon B, Preference Tag = Melon B), indicating that batch B is superior to batch A. For evaluations mentioning "tastier than ordinary melons on the market" without explicitly pointing to historical records, the system can retrieve the average physicochemical quality vector of the same variety from major wholesale markets nationwide during the same period as a virtual anchor point. This forms a triple (meat market, meat B, preference label = meat B) with the current meat. Through this mechanism, a large amount of originally useless free text is automatically transformed into preference pairs that can be used as training data for supervised learning.

[0032] The preference comparison network training unit trains a preference comparison network on the aforementioned triplet data. The core architecture of this network is a Siamese network. Specifically, the network consists of two feature extraction branches sharing all weights. Each branch is a three-layer fully connected network with the structure [input dimension 10] → 128 → 64 → 32. Each layer is followed by batch normalization and ReLU activation, finally outputting a 32-dimensional embedding vector. This is for the input physical and chemical quality vectors of melon A and melon B. and Calculate their embeddings respectively and Then calculate the difference vector. ,Will The data is fed into a single-output sigmoid layer, which outputs p (A is better than B). The loss function is the standard binary cross-entropy loss, with the preference labels in the triples used as the supervision signal.

[0033] Define the embedding function of the preference comparison network as follows: , where d is the dimension of the embedding space, mapping the physicochemical quality vector to a d-dimensional embedding space. For any pair of samples The probability of preference is defined as follows: ; in For the Sigmoid function, The weight vector of the difference vector. This is a bias term. This indicates a "preference" for the network. The training objective of the network is to minimize the cross-entropy loss. ; In the formula For expectation operation, D is the pairwise preference data distribution, y∈{0,1} is the binary label, and y=1 indicates y=0 means After training convergence, the network defines a mapping from the physicochemical space to preference probabilities. For any physicochemical point x, a preference score s(x) can be assigned by calculating the difference between its embedding and the embedding of a fixed reference point. The absolute value of s(x) has no independent meaning, but the difference between the scores of the two points directly reflects their relative superiority or inferiority in the group's preferences. More importantly, the gradient vector of s(x) with respect to x... s(x) represents the local preference gradient direction at that point, directly indicating how the physicochemical quality should be fine-tuned to maximize consumer preference. This gradient direction vector field is the digital representation of the preference gradient field.

[0034] Based on the aforementioned embedding function, the preference score function can be defined as the physicochemical vector x in the preference field space relative to a fixed reference point. Mapping difference: ; in This refers to an arbitrarily selected reference point for physicochemical properties. (The reference point item...) Since is a constant and does not affect the gradient calculation, it can be omitted for ease of expression without affecting the gradient direction, resulting in a simplified form of the preference score function: ; Based on the above embedding function, the preference scoring function can be defined as follows: The specific form of the preference gradient field is: Based on the simplified form above, the preference gradient field can be concisely represented as: ; This is the simplified form of the preference gradient field. If we use the complete form s(x), since the gradient of the reference point term is zero, the specific form of the preference gradient field is completely consistent with the simplified form: ; in For embedding functions The Jacobian matrix at x, its elements This represents the sensitivity of the i-th embedding dimension to the j-th physicochemical index. From the above derivation, it can be seen that the preference gradient field G(x) and the reference point... The selection of the reference point is completely irrelevant; it is an objective and uniquely determined vector field, which ensures that the optimization objective of this system is not affected by the subjective setting of the reference point.

[0035] Example 3: Counteracting individual biases and decoupling and anomaly detection; Please refer to it again. Figure 2 This embodiment details how to perform adversarial decoupling of individual consumer biases in bite voiceprint feedback, and how to perform gating rejection of abnormal feedback.

[0036] When consumers choose to upload recordings of their bite, they receive an audio clip typically 0.5 to 2 seconds long. This clip, besides reflecting the physical characteristics of cantaloupe flesh breaking under tooth bite, also incorporates numerous individual differences unrelated to the cantaloupe's quality. These factors include: the consumer's bite force and speed habits, oral cavity structure, the distance and angle between the recording device and the bite point, the device's frequency response characteristics, and environmental background noise. If these voiceprints are used directly to train sensory models without processing, individual noise will severely contaminate the group preference signal.

[0037] This invention introduces an individual bias decoupling unit to address this problem. For each active anonymous consumer in the system, a learnable personal factor embedding vector is assigned. The initial value of the embedding vector is randomly generated, and during training, it is gradually learned to represent the compressed representation of the consumer's personal bite style as more feedback data is contributed.

[0038] The acoustic encoder employs a one-dimensional residual convolutional network, with the following structure: The input is a raw waveform sampled at 22.05 kHz. First, it passes through a 1×7 convolutional layer and a max-pooling layer, followed by three residual blocks. Each residual block contains two 1×3 convolutional layers, batch normalization, and ReLU activation. Skip connections are identity mappings. Finally, it passes through global average pooling and a fully connected layer, outputting a 64-dimensional physical feature vector. .Should Only physical information related to the fracture of the flesh of fruits should be encoded, such as the time interval between fracture events, the rate of rise of the peak sound pressure level of a single fracture, and the centroid of the spectrum.

[0039] In order to achieve the separation of individual information, After output, it is concatenated with the corresponding personal factor u and fed into a user identity classifier. . Given a two-layer fully connected network, output the predicted identity of the consumer.

[0040] It should be noted that the training sample composition and annotation system of the above adversarial training framework are as follows. Each training sample is a set of cantaloupe bite sound print segments, accompanied by two labels: the first label is a quality label, composed of objective physical quantities such as the flesh crispness value measured by a texture analyzer, used to supervise the quality prediction task, which employs the mean squared error loss function. Optimization was carried out, including The value is the brittleness measured by a texture analyzer. The first label is used to predict the fragility value for the model; the second label is the user identity label, which is the unique identifier of the anonymous consumer who recorded the voiceprint. k ∈{1,2,…, K This is used to supervise a user identity classification task, which employs the cross-entropy loss function. Optimization was carried out, including To verify the identity of real users, Users predicted by the classifier k The probability. The two tasks share the physical characteristics of the acoustic encoder output. As input, but using their own independent subsequent network branches. During the training phase, the same consumer can contribute multiple voiceprint samples, each sample sharing its personal factor embedding vector u.

[0041] Within the aforementioned adversarial training framework, the parameter update rules for each network module are naturally determined by the objective function and the gradient reversal layer, and are completed synchronously within a single backpropagation iteration, eliminating the need for phased alternating training. Specifically: the acoustic encoder... Ψ The parameters of the classifier receive double gradients from the quality prediction loss and the user classification loss after gradient reversal, and the update direction is to minimize the quality prediction error while maximizing the user classification error; the user classifier The parameters of the network only receive gradients from the user classification loss, and are updated in the direction of minimizing the user classification error to maintain its ability to distinguish user identities. The personal factor embedding vector u serves as a learnable parameter of the network, updated along with the parameters of the user classifier, and gradually learns a compressed representation of each consumer's personal bite style during training. The weight constraints of the multi-task loss are determined by the gradient inversion coefficients. λ Unified regulation, whenλ As the value gradually increases from 0 to 0.1, the optimization objective of the acoustic encoder smoothly transitions from "completely ignoring user classification" to "maximizing user classification error while minimizing quality prediction error," ultimately reaching λ. λ When stable, an equilibrium state is achieved where individual biases are decoupled.

[0042] During training, during backpropagation to A gradient reversal layer is inserted along the path. The gradient reversal layer is an identity mapping during forward propagation, and multiplies the gradient by a negative coefficient -λ during backpropagation (λ gradually increases from 0 to 0.1 with the number of training steps).

[0043] Therefore, the parameter update direction of the acoustic encoder is to simultaneously minimize the loss for quality prediction tasks (such as fragility regression) but maximize the loss for user classification. The ultimate result of this adversarial training is... The data no longer contains information that can distinguish user identities; all individual differences are compressed into the personal factor u. Experiments show that after adversarial decoupling, different consumers eating the same melon... The vector cosine similarity improved from 0.62 to 0.91, proving that individual biases were effectively separated.

[0044] Define the acoustic encoder as a function Let A be the original audio waveform space and U be the personal factor space. Let the original waveform be... Individual factors are physical characteristics Define the user classifier. , where K is the total number of users. The objective function for adversarial training is: ; in Let k be the audio-user pairing data distribution, and k be the user identity label. Through the gradient inversion layer, the acoustic encoder ψ is updated in the direction of maximizing the user classification error, thereby... User-identifiable information is stripped from the data.

[0045] The anomaly detection subunit is used for quality gating of new voiceprint data entering the system. After the system has accumulated a certain amount of voiceprint data, the anomaly detection subunit... The vector distribution is fitted using a Gaussian mixture model, and the 95% confidence ellipse of the principal components is taken as the normal distribution region. When a new bite recording enters, it is first extracted by an acoustic encoder. Calculate its Mahalanobis distance. If this distance exceeds a preset threshold (corresponding to χ²),... 2 If the 99th percentile of the distribution is used, then the recording is determined to be an anomalous sample.

[0046] Define the parameters of the Gaussian mixture model for the normal sample distribution as follows: ,in The number of components in the mixture. The mixing weight for the r-th component, and These are their mean vector and covariance matrix, respectively. For the newly extracted physical features... The outlier score is defined as the minimum of the weighted Mahalanobis distance: ; when When τ is determined to be an anomaly and excluded, the sample is considered abnormal and excluded, where τ is based on The threshold is set at the 99th percentile of the distribution. Typical anomalies include: consumers intentionally bumping their teeth instead of biting normally, strong transient noise in the recording environment, and the recording equipment being obstructed. Anomalies are temporarily isolated and not used in the training of the preference comparison network or the updating of the preference gradient field. Isolated samples are returned to the consumer with a friendly message, such as "The recording quality is not ideal, thank you for your participation," to avoid harming the user's sense of engagement. Through this gating mechanism, the system maintains the quality of the acoustic data used to construct the preference field, ensuring that the signal-to-noise ratio remains within an acceptable range.

[0047] To further clarify the feasibility of this invention, the model operation logic of the offline training phase and the online inference phase is clearly distinguished here.

[0048] Offline training phase: Acoustic encoder Ψ Personal factor embedding vector u, user classifier The acoustic encoder, along with the quality prediction branch, participates in the training. The core objective of the training is to leverage the adversarial signals provided by the user classifier and the gradient inversion layer to force the acoustic encoder to... Ψ Output physical characteristics The system does not contain user-identifiable information, thus stripping all individual differences into the personal factor embedding vector u. After training, a fully adversarially trained acoustic encoder is obtained. .

[0049] Online inference phase: The inference model deployed online only contains the trained acoustic encoder. User classifier Neither the personal factor embedding vector nor the individual factor embedding vector participates in online inference. For any new consumer's uploaded occlusal voiceprint, it is directly processed through... Extracting physical features .because The ability to "ignore user identity" has already been forced to learn during the training phase, and the extracted data at this point... It only reflects the physical cracking characteristics of cantaloupe flesh and has nothing to do with consumers' personal biting habits, thus achieving zero-cost real-time removal of individual biases from new consumers.

[0050] For the problem of initializing personal factors for new consumers, this invention provides two optional lightweight processing methods. Method 1: For the first feedback from a new consumer, initialize their personal factors as the mean vector of the personal factors of all users in the training set. This approach already provides good decoupling in most scenarios. Approach Two: If the system needs to provide more refined personalized services for the new consumer, a lightweight incremental learning process can be run asynchronously in the background to fine-tune the personal factor embedding using the consumer's first few voiceprint data points, while the acoustic encoder... The parameters remain frozen. Regardless of the method used, the acoustic encoder's general decoupling capability for new consumers is not affected.

[0051] Example 4: Modeling of time-varying preference density fields; Please refer to it again. Figure 2 This embodiment details the time-varying preference field construction unit in the group preference perception module, and how to use neural differential equations to capture the trend drift of consumer preferences.

[0052] Consumer preferences for Hami melon quality are not static. Factors such as the rise of healthy eating habits, the influence of food bloggers on social media, and seasonal consumer psychology all cause the standard for "delicious" to drift slowly over time. For example, consumers may prefer a sweet and crisp texture in summer, while in winter they may prefer a richer, more creamy texture; the popularity of "soft and glutinous" as a buzzword at a particular time can also influence consumers' sensory preferences. Traditional methods, using static sensory models, cannot capture these shifts, causing optimization efforts to gradually deviate from actual market demands.

[0053] This invention employs a time-varying preference field construction unit and models the continuous-time evolution of the preference field based on neural differential equations. The core idea is to treat the sequence of preference functions observed at discrete time points as discrete samples of a continuous dynamic system, learn the vector field of this dynamic system, and thus be able to predict the preference state at any future time.

[0054] The specific implementation is as follows: Using a week as the time unit, aggregate the paired preference triples collected each week, and train a separate static preference comparison network for that week. This yields a series of snapshots of network parameters arranged weekly. The output layer bias and weight matrices of these networks are flattened to serve as a low-dimensional representation vector of the preference field state at that moment. Suppose there exists a potential continuous-time state vector h(t) whose evolution follows an ordinary differential equation. Where g is a three-layer fully connected network with parameter θ. The initial value of h(t) From the first observation Obtained through an encoding network.

[0055] Using the Neural ODE framework, h(t) is transformed from... Integrating to any time t yields h(t), which is then passed through a decoding network to output the predicted preference field state at that time. (t). By minimizing all observed times With reality The mean squared error between the two parameters is used to train the network g and the encoder / decoder. Training uses the adjoint sensitivity method to calculate the gradient of the loss with respect to the parameter θ, avoiding the high memory consumption of backpropagation through the ODE solver.

[0056] Define the latent state vector of the time-varying preference field as: ,in For time variables, Let be the dimension of the latent state. Its continuous-time evolution is described by the God ordinary differential equation: ; in For a fully connected neural network, For its trainable parameters, The number of parameters. Given an initial state. Where ρ is the encoding network, To encode the network parameters, the state at any time t can be obtained through integration: ; Where τ is the integral dummy variable, the decoding network Mapping the latent state back to the preference field parameter space: ,in These are the decoder parameters. The training objective is to minimize: ; Where N is the number of observation time points. In time A snapshot of the preference field parameters during actual aggregate training.

[0057] The trained model possesses two important capabilities. First, it can interpolate the preference field state at any historical moment to analyze the historical trajectory of preference changes. Second, it can extrapolate the preference field state at future moments. For example, at the beginning of the planting season (April), the system can predict the preference field state three months later, at harvest time (July), based on the preference field evolution pattern of the past 24 months. This predicted July preference field will be used as the objective function for the optimization decision-making this season. Once July arrives and consumer feedback for that month is collected, the residual between the predicted value and the actual observation will be used to further fine-tune the parameters of the NeuralODE model, continuously improving the predictive capability within the closed loop.

[0058] Taking a real-world scenario as an example, the system observed over two consecutive planting seasons that, in the preference gradient field, the gradient strength pointing towards "high hardness" was 0.8 (normalized value) in spring, decreasing to 0.5 in autumn, while the gradient strength pointing towards "high ester aroma concentration" increased from 0.4 to 0.7. The time-varying preference field construction unit captured this rotational trend and modeled it in the vector field g. Before the third planting season, the system predicted that the preference field at harvest time (three months later) would further weaken the hardness weight and strengthen the aroma weight, and passed this prediction to the optimization decision module. Based on this, the module added a silicon-potassium fertilization measure during the fruit expansion period, which is beneficial for the accumulation of aroma precursors, to the silicon fertilizer program in advance.

[0059] Example 5: Construction of the Mechanism Digital Twin Submodule; This embodiment details the construction process and internal structure of the mechanism digital twin submodule.

[0060] Please see Figure 3 The mechanistic digital twin submodule within the quality prediction module forms the core of the entire system's physical mechanism. Its task is to simulate the dynamic evolution of the multi-dimensional physicochemical quality within the Hami melon fruit, given a complete growth cycle environmental time-series data and a silicon fertilizer application scheme, and ultimately output the physicochemical quality vector at maturity. Unlike purely data-driven "black box" models, this submodule fully utilizes mechanistic knowledge acquired through isotope tracing and multi-omics technologies, embedding it into the neural network as physical information constraints, thereby achieving higher generalization ability and agronomic interpretability.

[0061] Construction took place within a strictly controlled artificial climate chamber core base station. The climate chamber allows for independent control of temperature (set accuracy ±0.5°C), relative humidity (±3%), and light intensity (0-1200 μmol / m²). 2 / s (adjustable red-blue light ratio) and CO2 concentration. Cultivation was carried out using substrate bag culture, with nutrient solution precisely dispensed to each plant through independent pipelines. Several main cultivated varieties were tested, including "Xizhoumi 25" and "Golden Years".

[0062] The core method for mechanistic data acquisition was stable isotope ^30Si pulse labeling technology. Potassium silicate labeled with ^30Si at an abundance greater than 95% was supplied via 24-hour replacement during key time windows in the cantaloupe growth cycle. Specific labeling windows were set as follows: once during the vine extension stage, once on day 5, day 15, and day 25 after female flower opening. Four destructive sampling time points were set after each labeling: days 1, 3, 7, and 14. Three biological replicates were obtained for each sampling, and each plant was separated into five organ groups: root, stem, functional leaves, pericarp, and pulp.

[0063] The following tests were performed on each sample: ① Laser ablation-inductively coupled plasma mass spectrometry (ICP-MS / MS) to obtain the two-dimensional spatial distribution of ^30Si on tissue sections of various organs, with a spatial resolution of 10 μm, clearly identifying the deposition patterns of silicon in microregions such as cell walls and vascular bundles. ② Broadly targeted metabolomics analysis (UPLC-MS / MS) to detect the relative content of more than 800 metabolites, including sugars, organic acids, amino acids, and volatile aroma precursors. ③ Transcriptome sequencing (Illumina NovaSeq 6000 platform) to determine the gene expression levels of enzymes in key sugar metabolism and aroma synthesis pathways. ④ Key enzyme activity assays, including sucrose phosphate synthase, sucrose synthase, phenylalanine ammonia-lyase, and lipoxygenase.

[0064] By integrating the above data, the core causal pathway of silicon regulation in cantaloupe quality was reconstructed. The following are some typical findings: After foliar application of silicon fertilizer, ^30Si was first deposited in the epidermal cell walls and phloem companion cells of the leaves. Silicon deposition significantly upregulated the expression of the CmeSPS1 gene (encoding sucrose phosphate synthase), with the upregulation magnitude positively correlated with the ^30Si concentration in the companion cells. Increased SPS enzyme activity promoted sucrose accumulation in the pulp. Simultaneously, silicon treatment also upregulated PAL enzyme activity in the phenylpropanoid metabolic pathway, increasing the supply of phenylalanine, which in turn increased the accumulation of characteristic aroma substances such as benzoyl acetate through the downstream aroma synthesis bypass pathway.

[0065] The above pathway relationships are abstracted into a directional causal graph, with nodes including "leaf silicon concentration", "root silicon concentration", "CmeSPS1 expression level", "SPS enzyme activity", "fruit pulp sucrose concentration", "PAL expression level", "aroma precursor concentration", etc., and edges representing the activation or inhibition relationships verified by experiments, with edge weights determined by regression coefficients.

[0066] This causal graph and reaction kinetics knowledge are embedded into a generative neural network. The network structure is a conditional generative adversarial network. The generator G takes as input a sequence of environmental conditions during the growing season (7 dimensions: daily maximum temperature, daily minimum temperature, daily average humidity, daily total radiation, soil moisture content, soil pH, CO2 concentration), a sequence of silicon fertilizer application schemes (3 dimensions: root silicon concentration, leaf silicon concentration, and spraying time encoding), and the initial fruit state. The output is the internal quality state of the fruit at each future time step (10 dimensions, including sugar, acid, firmness, and main aroma substances). The discriminator D is used to distinguish the generated quality evolution trajectory from the actual observed trajectory. The loss function adds two key constraints to the standard GAN loss: Physical constraint loss The forced generator predicts the total sugar content of the fruit according to the carbon mass conservation law, that is, the increase in fruit sugar content equals the input of photosynthetic products from the leaves minus the respiration consumption. The input function is determined by environmental variables and the leaf area index, while the respiration consumption is determined by the temperature-sensitive Q. 10 Function description: Forces the SPS enzyme reaction rate to follow the Michaelis-Menten equation, with its Vmax modulated by both silicon concentration and temperature.

[0067] Causal path constraint loss In the generator's hidden layers, the activation patterns of neurons corresponding to nodes in the known causal graph are forced to satisfy the dependency directions in the causal graph. For example, if the edge "leaf silicon concentration → CmeSPS1 expression level" exists in the causal graph, a regularization term is applied to the generator during training, so that when the leaf silicon concentration in the input increases, the activation value of the hidden layer node corresponding to the CmeSPS1 expression level also increases, and the time delay is consistent with experimental observations.

[0068] The generator for the defined mechanism digital twin submodule is as follows: ,in For the spatial time series of the environment, For the time series space of the silicon fertilizer scheme, Let this be the initial fruit state space. For the quality evolution trajectory space. The discriminator is... The physical constraint loss is defined as: ; In the formula, the first term represents the carbon mass conservation constraint. Let t be the total sugar content of the fruit. The input function for photosynthetic products depends on environmental variables. Leaf area index , The respiration consumption function depends on temperature. The second item is enzyme kinetic constraints. The SPS reaction rate predicted by the generator. The maximum reaction rate is determined by the amount of silicon deposited. and temperature Co-modulation, It is the Michaelis constant. The concentration is the substrate concentration.

[0069] The causal path constraint loss is defined as: ; Where C is the set of known causal edges, and (i,j) represents the directed edge from node i to node j in the causal graph. and These are the activation values ​​of the corresponding nodes in the generator's hidden layer. This is the causal effect strength coefficient estimated from experimental data.

[0070] The total loss function is: ; in Generate adversarial loss for the standard. and The hyperparameters of the constraint terms are all positive real numbers to balance the constraints.

[0071] Training data from approximately 50,000 spatiotemporal points were accumulated within the core base station. After training, the PINN-GAN can generalize to environments and silicon fertilizer combinations not covered by the training data, generating corresponding quality evolution trajectories. After validation set testing, the prediction coefficient of determination for TSS reached 0.89, and the prediction coefficient of determination for hardness reached 0.84. Furthermore, due to physical constraints, its prediction curves extrapolated to extreme environments still maintain a reasonable shape, without the oscillations or outputs that violate agronomic common sense often seen in pure black-box models. The coefficient of determination measures the goodness of fit of the regression model to the observed data. It indicates the proportion of variation in the dependent variable that can be explained by the variation in the independent variable.

[0072] Example 6: Field migration adapter submodule and silicon state proxy fingerprint; Please refer to it again. Figure 3 This embodiment details how to lightweightly adapt the mechanistic twin constructed from the core base station to a wide-area target field.

[0073] While the core base station's mechanistic digital twin submodule boasts high accuracy and complete mechanistic information, its construction cost is exorbitant, and the initial training data is based on a controlled environment in an artificial climate chamber and specific varieties. Requiring the reconstruction of a base station for each production area and each variety is economically infeasible. Therefore, this invention designs a field migration adaptation submodule to "clone" the super twin to any target field at extremely low cost.

[0074] The core bridge for migration adaptation is the silicon state proxy fingerprint. During the core base station experiment, before determining the silicon content of all collected leaf samples using LA-ICP-MS, reflectance spectra were first collected in a darkroom using a hyperspectral camera of the same model as the one used by the field drone. The spectral range was 450-2500 nm, with a spectral resolution of 4 nm in the visible and near-infrared range and 10 nm in the short-wave infrared range. After preprocessing the spectrum of each leaf using standard normal variables, paired data were generated with its corresponding silicon speciation content (the ratio of total silicon, hydrated silicon, and polymeric silicon).

[0075] The proxy fingerprint extraction model employs a one-dimensional convolutional neural network. The network input is a 2100-dimensional spectral reflectance vector, processed through three consecutive convolutional blocks: each block contains a one-dimensional convolutional layer with a kernel size of 7, batch normalization, LeakyReLU activation, and max pooling with a kernel size of 2. The number of channels in the three blocks are 32, 64, and 128, respectively. Finally, global average pooling and a fully connected layer output the predicted silicon morphology content. A channel attention module is added after the last convolutional block, generating weights for each channel through global average pooling and two fully connected layers. This attention mechanism allows the network to automatically focus on spectral regions sensitive to silicon deposition. After training, the spectral region localization of the extracted attention weights shows that the highest weight bands are concentrated around 370-390 nm (UV fluorescence excitation region, related to silicon-phenol complexes), around 1450 nm (first overtone of hydroxyl stretching vibration, related to bound water in silicon-cellulose complexes), and around 1930 nm (water molecule combination frequency). This set of automatically discovered sensitive narrowband reflectance ratios is known as the "silicon state proxy fingerprint".

[0076] For the target field, the adaptation process only requires one drone flight and data collection, as follows: In the early stages of Hami melon vine growth, a drone equipped with a hyperspectral imager was used to fly 30m above the target field with a horizontal and vertical overlap of 75% to acquire hyperspectral images of the entire field.

[0077] After radiometric correction, geometric correction and orthophoto stitching, the proxy fingerprint extraction model is used to process the data pixel by pixel to generate a spatial distribution map of the silicon state proxy fingerprint of the whole field. The spatial resolution is about 5 cm / pixel, and the differences in silicon state of individual plant canopies can be observed.

[0078] Simultaneously, surface soil samples were collected from the target field to determine pH, EC, and available silicon content; field weather station data and future weather forecasts for the most recent week were obtained; and management information such as variety and planting date was recorded.

[0079] Define the proxy fingerprint extraction function as follows ,in This represents the number of spectral bands. This represents the number of silicon morphology categories. For the input spectral vector... Output the predicted silicon morphology content vector This function is parameterized by the convolution and attention mechanisms of the proxy fingerprint extraction model.

[0080] The domain adversarial migration calibration network then performs an adaptation. The network consists of three components: The system consists of a shared feature extractor F, a quality predictor P, and a domain discriminator Ddom. The shared feature extractor takes the environmental sequence of the target field and the silicon state surrogate fingerprint as input, and outputs high-level features. The quality predictor takes the output of the shared feature extractor as input and predicts the final physicochemical quality vector. Its skeleton is the generator part of the core base station mechanism twin (parameters are frozen), but the last three layers are replaced with trainable adaptation layers. The domain discriminator takes the features output by the shared feature extractor as input and attempts to determine whether the features originate from the source domain (core base station) or the target domain (target field).

[0081] During training, a gradient reversal layer is inserted between the shared feature extractor and the domain discriminator. The total loss function is: ,in For quality prediction error, For domain classification error, The balancing coefficient is gradually increased from 0 to 0.05. Through this adversarial training, the shared feature extractor is forced to extract domain-invariant features between the source and target domains—features that reflect the essential silicon-quality response of plants and are unaffected by differences between the artificial climate chamber and the field environment. Domain-specific biases (such as fluctuations in field light intensity and the effect of wind speed on the spectrum) are absorbed by the domain discriminator. Training typically converges within 100 epochs, requiring only the initial data from the first season of the target field for adaptation.

[0082] Assume the data distribution of the source domain (core base station) is as follows: The target domain (target field) data distribution is as follows The optimization objective of domain adversarial migration calibration is: ; in For the joint input containing the environment sequence and silicon state proxy fingerprint, For the true physicochemical quality vector, For feature extractor, For quality predictors, This is the domain discriminator. The first two terms are the supervised loss for quality prediction, and the last term is the domain adversarial loss.

[0083] After adaptation, the quality predictor can serve as a dedicated quality prediction engine for the target field. When the complete environmental data of the target field's growth period and candidate silicon fertilizer schemes are input, the quality predictor outputs the physicochemical quality vector prediction value corresponding to the scheme under the field conditions. Field verification in multiple production areas shows that after domain adversarial adaptation, the root mean square error of the TSS prediction is reduced by approximately 45% compared to before adaptation, achieving prediction accuracy similar to that within the core base station, while the cost is only the cost of a single drone flight and soil testing.

[0084] Example 7: Estimation of causal effects of instrumental variables; Please see Figure 4 This embodiment details how to use exogenous shocks in farmers' historical behavior data as instrumental variables to estimate the pure causal effect of silicon fertilizer on group preferences.

[0085] In real-world agricultural big data, the observed correlation between silicon fertilizer application and final quality contains significant confounding biases. For example, meticulously managed farmers often apply silicon fertilizer, optimize irrigation, and precisely control temperature simultaneously, resulting in high quality that is the combined effect of multiple measures and difficult to attribute solely to silicon fertilizer itself. Directly using this contaminated correlation to guide optimization could overestimate the marginal benefits of silicon fertilizer and even misdirect resources that should be allocated to irrigation optimization towards silicon fertilizer. Therefore, causal decoupling must be performed before optimization decisions are made.

[0086] This invention employs a causal inference method based on instrumental variables. Instrumental variables must meet three conditions: ① they must be related to the treatment variable (here, silicon fertilizer application rate); ② they must influence the outcome variable (here, the group preference score) only through the treatment variable; and ③ they must be independent of confounding factors. In the agricultural context, a natural class of instrumental variables was identified—accidental exogenous events related to silicon fertilizer accessibility.

[0087] The specific data mining process involved obtaining complete records from over 1,000 melon farmers in the cooperative production area over the past three years. This included daily agricultural operation logs, detailed agricultural input purchase records (including variety, quantity, price, purchase date, supplier name and address), as well as physicochemical testing of harvested fruits and vegetables and summaries of consumer feedback for corresponding batches. Through data mining, the following two typical instrumental variables were identified: Instrumental variable Z1: A certain brand of silicon fertilizer implemented a limited-time price reduction of 30% from June 15th to July 15th, 2022. Farmers who purchased silicon fertilizer during the promotion period had a significantly higher average application rate of silicon fertilizer in the current season than farmers who did not participate in the promotion. The time window of this promotion was determined by the manufacturer's marketing decision and was unrelated to the soil fertility, climate conditions, and management level of each farmer's melon field, thus satisfying the exogenous property requirement.

[0088] Instrumental variable Z2: In July 2021, a main road in a core production area was closed for bridge repairs, temporarily increasing the transportation distance to the nearest silicon fertilizer supplier from an average of 5 km to 35 km. The closure lasted for 45 days. Affected farmers significantly reduced their silicon fertilizer purchases during the road closure, and some even missed the critical growth period for silicon fertilizer application. The timing and location of the bridge repairs were government engineering decisions, independent of the specific conditions of each farmer's melon field.

[0089] The deep instrumental variable network employs a two-stage architecture. The first stage uses instrumental variables Z (event dummy variables or continuous distance variables) and other exogenous context variables. The input factors are variety, planting density, accumulated temperature during the growing season, and initial soil organic matter content, while the output factor is the amount of silicon fertilizer applied. The conditional probability distribution is used here. A mixture density network is employed to output the mean and variance parameters of the Gaussian mixture model, capturing the differentiated responses of different farmers to the same instrumental variable.

[0090] The second stage is based on the expected conditions of silicon fertilizer application based on the network output from the first stage. As a processing variable for "decontamination", it is compared with the background variable. Together, they are fed into a cascaded network of the quality prediction module and the group preference perception module for longitudinal adaptation, and finally output the predicted group preference score. The key operation here is: in the second stage of input, the actual amount of silicon fertilizer applied is not used. However, using the predictions from instrumental variables .because Independent of confounding factors, The variations in the molecule are solely due to exogenous factors. lead Therefore, the second phase right The regression coefficients are the local average treatment effect of silicon fertilizer on preference, which is an estimate of causal effect without confounding.

[0091] Define instrumental variables as The treatment variable (silicon fertilizer application rate) is: Background covariates are The outcome variable (preference score) is The first-stage mixed density network defines the conditional distribution: ; in normal distribution, The number of Gaussian mixture components. The mixing weights for the k-th component satisfy the following condition: , and Let be the mean and variance of the k-th Gaussian component, respectively, both derived from . and The output of the fully connected network is the input. This yields the decontamination processing variables: ; In the second stage, the cascade function of the quality prediction module and the preference perception module is denoted as... Mapping the treatment variables and covariates to preference scores: ; The causal effect curve is defined as a given hour, right Partial derivatives: ; in The specific values ​​of the background covariates This refers to the specific dosage of silicon fertilizer. The entire network is trained jointly, and the total loss is the weighted sum of the negative log-likelihood in the first stage and the preference prediction error in the second stage. After training, the value is fixed. As a baseline value (e.g., no promotions, no road closures), change The background value is consistent with the target field, and then the input is varied. (by changing) (Under the conditions to achieve this), the preference score for silicon fertilizer application under the target field conditions can be obtained. The causal effect curve shows the magnitude of the increase in the population preference score and its diminishing marginal returns for each additional unit of silicon fertilizer (e.g., 100g of pure silicon per acre). This curve will serve as important prior information for the next stage of Bayesian optimization.

[0092] Example 8: Multi-objective Bayesian optimization scheme for generating silicon fertilizer; Please refer to it again. Figure 4 This embodiment details how the optimization decision module utilizes causal effect priors to perform multi-objective search in the candidate silicon fertilizer scheme space and generate the optimal fertilizer prescription.

[0093] The solution search unit employs a multi-objective Bayesian optimization framework based on Gaussian processes. The objective function for optimization is defined as: F(scheme) = P(scheme is superior to the baseline scheme) – λ Option W Where P (the proposed solution is superior to the baseline solution) is calculated by the preference comparison network: the predicted physicochemical quality vector corresponding to the proposed solution. The predicted physicochemical quality vector corresponding to the baseline scheme As input, output This value ranges from 0 to 1. W(scheme) represents the cost increment of the scheme relative to the baseline, and λ is the cost penalty coefficient, which is adjusted by growers based on market expectations and profit margins. The baseline scheme refers to the conventional planting scheme for the same variety in this production area, obtained by taking the median from farmers' historical data.

[0094] The optimized decision variable space consists of approximately 10 dimensions: root silicon fertilizer type (potassium silicate 0-200ppm, nano silicon 0-50ppm), foliar silicon fertilizer type (potassium silicate 0-0.3%, nano silicon 0-0.1%), root application period (week X after the start of the vine extension period, lasting for Y weeks), foliar application period (day X after the initial fruit expansion stage, day Y after the initial sugar accumulation stage), and number of foliar sprays.

[0095] The optimized constraints include: the total cost of silicon fertilizer should not exceed 20% of the baseline scheme, and the predicted fruit pulp firmness should not be less than 3.0 kg / cm². 2 The predicted TSS is not less than 14%, etc. Specific constraint values ​​can be set by growers as needed.

[0096] The prior mean function for Bayesian optimization is provided by the causal effect curve. The effect curve given by the causal effect estimation unit describes the marginal relationship between silicon fertilizer application rate and preference score improvement. This curve is transformed into the prior mean function m0(x) of a Gaussian process, ensuring that the model's initial guesses are consistent with the causal effect curve at points without observed evaluations. For example, the prior mean is set to a positive slope when the foliar silicon spraying concentration increases from 0 to 0.1%, and the slope tends to zero or even becomes negative after exceeding 0.3%, reflecting the diminishing marginal returns of silicon fertilizer and even the inhibitory effect of excessive application. This prior information provided by causal inference greatly accelerates the convergence speed of the optimization search, reducing the approximately 300 evaluations required for a random "no prior" search to approximately 150.

[0097] Define the decision variable vector as t∈T Rdt, where Let be the dimension of the decision variables. Objective function. Represented as: ; in x(t) = f(t, e) is the physicochemical quality vector output by the quality prediction module. The prior of the Gaussian process surrogate model is: ; Where the prior mean function Derived from the causal effect curve: ; In the formula Let t represent the total silicon fertilizer application rate (cumulative pure silicon) corresponding to scheme t. The integral is performed along the causal effect curve. The covariance function uses the Marton 5 / 2 kernel: ; in , For signal variance, The length scale parameter is a positive hyperparameter. This check has good fitting ability for functions with moderate smoothness and is suitable for agronomic response surfaces with some nonlinearity.

[0098] Let the constraint function be An independent Gaussian process surrogate model is also established for each constraint. The multi-objective acquisition function is a constraint-weighted expected hypervolume improvement: ; Where P is the set of target values ​​on the current Pareto front. The reference point is the minimum value of each target, which is the over-volume index. The probability of constraint satisfaction is given by the posterior distribution of the constraint surrogate model.

[0099] In each iteration, the solution search unit samples in the solution space T. From the candidate points, the next evaluation point is selected by maximizing the acquisition function α(t). This evaluation point is sent to the quality prediction module for simulation evaluation. The returned physicochemical quality vector is converted into a preference score by the preference comparison network and then fed back to the surrogate model to update the posterior.

[0100] The iteration continues until a preset number of evaluations is reached (typically 150-200 times) or the overvolume improvement falls below the convergence threshold (less than 0.1% for 10 consecutive rounds). Finally, the solution search unit outputs a set of non-dominated solutions located at the Pareto front, each representing a trade-off between "preference score improvement" and "cost increment." A typical solution at the Pareto front might be: a 30% increase in preference score with only a 5% increase in cost, achieved through root drip irrigation of 10 ppm nano-silicon during the early fruit expansion stage and foliar spraying of 0.15% potassium silicate during the early sugar accumulation stage; another solution might be: a 45% increase in preference score with an 18% increase in cost, requiring a denser combination of silicon fertilizer varieties and precise staged application. Growers can select the most suitable solution based on their market positioning (mass market or high-end premium) and submit it to the integrated water and fertilizer control system for execution with a single click.

[0101] Example 9: Specific application examples of instrumental variables; This embodiment further details the application methods and data flow of two typical instrumental variables in a real system.

[0102] Scenario 1: Silicon fertilizer promotion event; When analyzing data from approximately 350 melon farmers in a specific production area in a given year, the system automatically identified, by scanning the price and purchase date fields in the agricultural input purchase records, that the price of brand X silicon fertilizer dropped from the usual 25 yuan / kg to 17.5 yuan / kg between June 15th and July 15th, a decrease of 30%, while the prices of other brands of silicon fertilizer remained stable. The system encoded this event as an instrumental variable. For each farmer i, He indicated that he purchased brand X silicon fertilizer during the promotional period. This indicates that they did not purchase or did not buy the brand during the promotional period.

[0103] The system then verified The effectiveness of this study is shown by correlation analysis. The average amount of silicon fertilizer applied by farmers (calculated as pure silicon) is 4.2 kg / mu. The average yield per mu was 3.0 kg / mu, a statistically significant difference. Regarding the exogeneity test, there were no significant differences between the two groups in background variables such as soil organic matter, average yield per mu over the years, and previous nitrogen, phosphorus, and potassium inputs before the promotion event, supporting the exogeneity of the promotion event.

[0104] In deep instrumental variable networks As a binary input in the first stage, it is used together with other background variables to predict silicon fertilizer application rate. Through this design, exogenous variation in silicon fertilizer application caused by promotional events was used to identify the causal effect of silicon fertilizer on population preference. The results showed that, in the "Xizhoumi 25" variety, an increase in silicon fertilizer application from 3.0 kg / mu to 4.2 kg / mu resulted in an average increase in population preference score of approximately 0.21 (normalized units), with a marginal causal effect of approximately 0.175 points / kg pure silicon.

[0105] Formal, for instrumental variables (Promotional event dummy variable), define the conditional expected dosage difference as: ; The marginal causal effect of silicon fertilizer is: ; when This estimate is valid at this time. In this scenario... The preference score difference is 0.21, therefore Pure silicon.

[0106] Scenario 2: Road closure incident; The system identified an event where a main road was closed for 45 days in a certain month due to bridge repairs. By geocoding each farmer's address with the address of their nearest silicon fertilizer supplier, the road network distance between them was calculated. For farmers who needed to purchase silicon fertilizer during the road closure, their road network distance was temporarily increased from 5km to 35km; farmers on the other side of the road were unaffected, and their road network distance remained unchanged at 5km. Instrumental variables Defined as the effective road network distance of farmer i during the closed period (a continuous variable).

[0107] Correlation validation showed that for every 10km increase in road network distance, the average purchase amount of silicon fertilizer decreased by 1.5kg / mu. Regarding exogenous factors, there were no systematic differences between affected and unaffected farmers in terms of soil texture, irrigation conditions, and historical yields. A deep instrumental variable network was used to... As the first-stage input, the causal impact of road closures on exogenous silicon fertilizer reduction on preferences can be estimated.

[0108] For continuous instrumental variables (Road network distance), marginal causal effects are estimated using a continuous form of two-stage least squares: ; in The first phase predicted the amount of silicon fertilizer to be applied. The results showed that for "Xizhoumi 25", the loss of preference score due to the lack of silicon fertilizer during the fruit expansion period was approximately 0.12 units. This quantitative assessment provides data support for setting the weight of silicon fertilizer as indispensable during this growth period in subsequent optimization.

[0109] Example 10: A complete example of closed-loop operation for three consecutive quarters; This embodiment uses a virtual but complete three-season planting scenario to comprehensively illustrate the closed-loop operation process of the system of the present invention, covering the entire chain from base station construction, migration and adaptation, preference field initialization, causal optimization decision-making, consumer feedback feedback to preference field evolution.

[0110] Season 1: System Deployment and Initial Optimization; A 500-mu (approximately 33 hectares) cantaloupe base has deployed the system of this invention. Prior to deployment, the mechanism digital twin submodule had been constructed within the core base station of the artificial climate chamber of a collaborating research institution, covering the mechanism data of the variety throughout its entire growth period.

[0111] On the 10th day after transplanting, the field migration adaptation submodule executed the adaptation process: a drone conducted a hyperspectral flight to acquire canopy images of the entire field; after processing with a silicon state proxy fingerprint model, the first spatial distribution map of leaf silicon state was generated. The image showed that the northwest corner of the field had a lower background silicon value due to stronger soil sandiness, resulting in a weaker intensity of the plant canopy silicon proxy fingerprint. Regarding environmental data, the average available silicon content in the soil was recorded as 85 mg / kg, and the pH value was 7.8. Domain adversarial adaptation was performed between the above data and the source domain data from the core base station, taking approximately 2 hours (on a server equipped with an NVIDIA A100 GPU), completing the localization calibration of the quality prediction module.

[0112] After adaptation, the quality prediction module simulated the default standard treatment (foliar spraying of 0.1% potassium silicate only once during the early stage of fruit expansion), predicting an average TSS of 13.8% and a firmness of 4.2 kg / cm² for mature melons. 2 The preference gradient field of the group preference perception module is initialized based on aggregated data of historical feedback from consumers across the country. Due to insufficient data, the initial gradient field is broadly oriented as "preferring higher sweetness and medium crispness". The gradient direction in the physicochemical space is [+TSS, – hardness (mild), +aroma].

[0113] After accessing historical farmer data from the production area, the causal effect estimation unit of the optimization decision-making module identified a silicon fertilizer promotion event in 2022 as an instrumental variable and estimated the causal effect curve of silicon fertilizer on preference under similar soil conditions. Based on this prior, the scheme search unit performed 200 rounds of Bayesian optimization iterations and output a recommended scheme: based on the conventional scheme, add one application of 12 ppm nano-silicon to the root drip irrigation during the early stage of fruit expansion, increase the foliar silicon concentration from 0.1% to 0.15%, and adjust the spraying time to the early stage of sugar accumulation. Simulations show that this scheme is expected to increase TSS to 15.1% and slightly reduce hardness to 3.9 kg / cm². 2 The concentration of ester aromas increased by 18%. Following evaluation, the plan was adopted and implemented by the growers.

[0114] Harvested in early August, the melons were labeled with traceability codes and sold to six major cities across the country. By the end of September, approximately 8,200 consumer reviews had been collected. The pairwise preference mining unit extracted 1,560 pairwise preference triples from these reviews. The preference comparison network underwent incremental training, and the preference gradient field was updated for the first time.

[0115] Season 2: Preference Drift Perception and Solution Adaptation; Based on the feedback data from the first season, the preference gradient field of the group preference perception module had changed significantly before the second planting season compared to its initial state. The most significant change was a shift in the gradient direction of the hardness dimension. In the first season's feedback, consumers from first-tier cities frequently commented that "a little softer is better" or "too hard to cut." Paired preference data showed that for melons with a TSS greater than 15%, consumers clearly preferred a hardness of 3.5-4.0 kg / cm². 2 Within a certain range, rather than a hardness greater than 4.0 kg / cm². 2 The intensity of the component pointing towards "decreasing stiffness" in the preference gradient field increases.

[0116] The time-varying preference field construction unit, based on data points from the first quarter and historical offline data, started running the Neural ODE model to predict that the preference field for the same period in the following year will further shift towards "medium hardness, high aroma". This prediction was then passed to the optimization decision module.

[0117] Meanwhile, the quality prediction module continued to run in the second quarter, accumulating a full planting season of field measurement comparison data (80 samples sent for testing after harvest), and fine-tuned the domain adaptation network, improving the prediction accuracy by about 5% compared to the first quarter.

[0118] The optimization decision-making module used an updated preference gradient field as the objective when searching for solutions in the second season. Search results showed that, compared to the solutions in the first season, the optimal solution in the second season maintained high TSS (Total Sugar Saturation) while reducing the foliar silica concentration in the later stages of sugar accumulation (from 0.15% to 0.10%) to avoid excessive hardening of the fruit peel. It also advanced the application of silicon fertilizer to the root system at the end of vine elongation to initiate sugar accumulation earlier. After the second harvest, the positive consumer feedback rate (the proportion of paired preferences where this season's melons were superior to the comparison objects) increased by 8 percentage points compared to the first season.

[0119] Third season: Rapid adaptation of new varieties and market leadership; At the end of the second quarter, the company decided to trial-plant a new variety, "Golden Years," in a 200-acre experimental field. This variety is characterized by its naturally soft flesh and rich aroma. Under traditional cultivation methods, this variety is prone to insufficient sugar content (TSS of only around 12%).

[0120] Because the system already has a super-mechanistic twin module, the adaptation process for new varieties is extremely efficient. At the core base station, only a simplified growth cycle cultivation experiment is needed for "Golden Years" (covering only key periods and not repeating full-scale omics) to obtain its variety-specific genotypic parameters and enzyme kinetic constant differences. Using this data as additional input, the quality prediction module can be transferred to the new variety through an extra variety embedding layer in the domain adaptation network. The entire process takes about two weeks (including necessary waiting during the plant growth cycle), and the cost is far lower than rebuilding a complete model.

[0121] When the optimization decision-making module was running for the "Golden Years" variety, the preference gradient field it faced had evolved over two seasons, with the preference for a "soft and glutinous" texture shifting from a trend to a stable state. The Bayesian optimizer explored a novel silicon fertilizer strategy during the search: applying a higher concentration of nano-silicon (15 ppm) to the roots during the early fruit expansion stage to maximize its upregulation effect on SPS enzymes, thus compensating for the sugar content deficiency; simultaneously, completely eliminating later foliar silicon spraying to avoid any operations that might increase fruit peel firmness. This scheme predicted a TSS of 14.8% and a firmness maintained at 3.0 kg / cm² in the simulation. 2 In the soft and glutinous range, the concentration of benzoyl acetate reached a new high.

[0122] Following the third harvest, "Golden Years" melons received an excellent response in the target market. Consumer pairwise preference data showed that this batch of melons had a 67% preference probability advantage compared to its main competitors. The preference gradient field, after incorporating the data from the third season, was further fine-tuned, providing more refined guidance for optimization in the next season.

[0123] After three quarters of closed-loop operation, the system demonstrated its ability to continuously evolve, adapt to shifting market preferences, and rapidly adapt to new varieties. Ultimately, the continuous self-optimization process of this closed loop can be formally summarized as: in each quarter k k At the end, the preference gradient field Updated based on consumer feedback from this season: ; in The consumer feedback dataset collected for season k. Let η be the preference gradient field update function, and η be the learning rate parameter. Simultaneously, the domain adaptation parameters of the quality prediction module are fine-tuned using newly added field measurement data, and the time-varying preference field model corrects its extrapolated predictions through extended time series data. This mathematical structure ensures that the system converges to the optimal silicon fertilizer strategy consistent with real market preferences during continuous operation, demonstrating its value and practicality as a commercial silicon fertilizer quality decision-making tool.

[0124] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality, characterized in that, The system includes: The quality prediction module is used for: Based on the growth environment data of Hami melon and the silicon fertilizer application plan, the physicochemical quality indicators of Hami melon are predicted. The group preference perception module is used for: Obtain consumer feedback data on cantaloupes and construct a preference gradient field that reflects the sensory preference direction of the group of consumers based on the feedback data; The optimization decision-making module is used for: Using the preference gradient field as the optimization objective, the optimal silicon fertilizer application scheme is searched based on the mapping relationship between the physicochemical quality indicators output by the quality prediction module and the candidate silicon fertilizer schemes.

2. The big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality, as described in claim 1, is characterized in that... The group preference perception module includes: The pairwise preference mining unit is used to extract pairwise preference relationships between cantaloupe samples from the feedback data. The pairwise preference relationship indicates which of the two samples consumers prefer more. The preference comparison network training unit is used to train the preference comparison network by taking the physical and chemical quality indicators of the samples as input and the pairwise preference relationship as the supervision signal. The preference comparison network outputs the probability that one sample is better than another sample, so as to learn the direction of the preference gradient in the physical and chemical quality space.

3. The big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality, as described in claim 2, is characterized in that... The feedback data includes consumer natural language evaluations and / or bite voiceprints; The pair preference mining unit automatically constructs pair preference relationships by parsing comparative semantics in natural language and tracing the historical purchase records of the same consumer.

4. The big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality, as described in claim 3, is characterized in that... The group preference perception module also includes: Individual bias decoupling unit is used to assign learnable personal factor embeddings to each consumer when using bite soundprints as feedback data, and to force the physical features output by the acoustic feature encoder to be independent of the consumer identity through adversarial training, so as to separate the pure physical features that reflect the texture of the cantaloupe itself.

5. The big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality, as described in claim 4, is characterized in that... The individual bias decoupling unit also includes an anomaly detection subunit, which is used to determine whether the new input feedback voiceprint exceeds the training distribution range, and to exclude or reduce the weight of the corresponding feedback when it exceeds the range.

6. The big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality, as described in claim 5, is characterized in that... The group preference perception module also includes: The time-varying preference field construction unit is used to treat the pairwise preference relationships at different timestamps as particles in the physicochemical space, and to model the continuous evolution of the preference gradient field over time using neural differential equations in order to capture the trend drift of consumer group preferences.

7. The big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality according to claim 6, characterized in that, The quality prediction module includes: The mechanism digital twin submodule, based on omics experimental data of isotope-labeled silicon fertilizer, constructs a generative neural network constrained by physical information to simulate the internal quality evolution process of cantaloupe under different silicon fertilizer schemes. The field migration adaptation submodule is used to acquire hyperspectral images of the target field and extract silicon state proxy fingerprints. Through domain adversarial transfer learning, the mechanism digital twin submodule is adapted to the variety and microclimate of the target field, and the predicted values ​​of physicochemical quality indicators are output.

8. The big data-based system for analyzing the correlation between silicon fertilizer treatment and cantaloupe quality, as described in claim 7, is characterized in that... The field migration adaptation submodule uses a drone equipped with a hyperspectral sensor to acquire hyperspectral images, and uses a deep learning model to map the narrowband reflectance ratio to the silicon morphology concentration of leaves, which serves as a proxy fingerprint of silicon state.

9. A big data-based correlation analysis system for silicon fertilizer treatment on cantaloupe quality according to claim 8, characterized in that, The optimization decision module includes: The causal effect estimation unit is used to identify instrumental variables from farmers' historical records and to estimate the pure causal effect of silicon fertilizer application on the sensory preferences of group consumers using a deep instrumental variable network, in order to remove confounding factors between environmental and management measures. The scheme search unit is used to run multi-objective Bayesian optimization with causal effects as a prior, and search in the candidate silicon fertilizer scheme space for the silicon fertilizer application scheme that gives the current scheme the greatest preference probability advantage over the benchmark scheme in the preference gradient field.

10. A big data-based correlation analysis system for silicon fertilizer treatment on cantaloupe quality according to claim 9, characterized in that, The instrumental variables include exogenous events that are related to the probability of silicon fertilizer application but have no direct causal relationship with the final quality of Hami melons. Exogenous events include at least one of the following: silicon fertilizer promotional activity timestamp, changes in supplier logistics distance, or occasional stockout records.