Agricultural data security sharing method based on federal learning
By adopting a combination of federated learning, blockchain smart contracts, homomorphic encryption and differential privacy technologies in agricultural data sharing, the problem of insufficient data privacy and security in agricultural data sharing is solved, and the effect of high accuracy and full life cycle privacy protection is achieved.
Patent Information
- Application Number
- CN202510295252.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing agricultural data sharing technology has problems with insufficient data privacy and security, especially in the stages of data transmission and terminal use, it is difficult to prevent the risk of leakage.
Using a federated learning method, the initial model distributed by blockchain smart contracts is combined with homomorphic encryption and differential privacy technology to realize parameter aggregation of local model training and secure multi-party computing to ensure privacy protection of data throughout the entire life cycle.
It achieves the improvement of the accuracy of agricultural models, enhances the security and privacy of data without leaking the original data, and ensures the privacy protection of the entire life cycle of the data.
Smart Images

Figure CN120151029A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural data security, and specifically to an agricultural data security sharing method based on federated learning. Background Art
[0002] In the digital age of deep penetration of information technology, digitalization and networking have become the core elements for reconstructing social production relations. Although the popularization of artificial intelligence devices and the construction of agricultural intelligent service platforms rely on a large amount of data support, frequent data leakage incidents have exposed the systematic risks faced by private information in illegal transactions and abuse. This stems from the contradictions in two dimensions: one is the inherent defect of excessive collection of private data by intelligent terminals, and the other is the secondary infringement risk caused by the lack of data processing transparency due to the hidden characteristics of the cyberspace. To maximize the value of data elements, it is urgent to establish a cross-domain security sharing mechanism, but the traditional technical path has obvious bottlenecks: blockchain technology realizes data encrypted sharing through means such as zero-knowledge proof and homomorphic encryption, but in essence, it still transmits the original data, which not only ensures the privacy of the transmission link but also is difficult to prevent the leakage risk in the terminal use stage. Although federated learning technology replaces the interaction of original data with model parameters (realizing "usable but invisible"), its centralized architecture has potential single-point failure risks and cannot ensure the full encryption of the data transmission link. There is an urgent need to build a new security architecture that integrates blockchain distributed ledger and federated learning edge computing, dynamically manages the data flow path through smart contracts, and combines differential privacy and trusted execution environment technologies to achieve full-life-cycle privacy protection, forming an innovative solution that can not only ensure data sovereignty but also release the value of elements. Summary of the Invention
[0003] (1) Technical Problems to be Solved
[0004] In view of the deficiencies of the prior art, the present invention provides an agricultural data security sharing method based on federated learning, which has the advantages of security and privacy, and solves the problem of unreliable traditional agricultural data sharing.
[0005] (2) Technical Solutions
[0006] To achieve the above object, the present invention provides the following technical solutions: An agricultural data security sharing method based on federated learning, comprising the following steps:
[0007] Step 1: Establish a user terminal module, a data classification module, a federated learning module, an agricultural cloud database, and a security protection module;
[0008] Step 2: The user terminal module is used to collect agricultural production data, receive federated learning model parameters, and provide a multi-modal interaction interface, realize encrypted upload and decryption feedback of agricultural data through edge computing devices, support secure access of multiple types of devices, and the user terminal module is connected to the data classification module through an encrypted communication network;
[0009] Step 3: The data classification module is responsible for performing multi-level classification processing on the agricultural data uploaded by the user terminal module. It uses differential privacy technology to perform multi-level label annotation on the planting and animal husbandry data, while ensuring the privacy of the classification process. The data classification module is connected to the federated learning module through a distributed network;
[0010] Step 4: The federated learning module distributes the initial model through the blockchain smart contract, combines homomorphic encryption to achieve local model training and parameter aggregation of secure multi-party computation. The federated learning module is connected to the agricultural cloud database through a blockchain encrypted channel;
[0011] Step 5: The agricultural cloud database uses the IPFS technology to store the encrypted knowledge assets generated by federated training. The agricultural cloud database is connected to the security protection module through a dual-link redundant network;
[0012] Step 6: The security protection module integrates the blockchain evidence storage and dynamic access control mechanism. The security protection module is connected to the user terminal module through a trusted execution environment interface.
[0013] Preferably, the user terminal module supports the access of three types of devices, namely edge computing devices, intelligent agricultural machinery terminals, and Internet of Things sensors. Each type of device is configured with a dedicated federated learning client program.
[0014] Preferably, the multi-level classification processing of the data classification module includes: primary classification based on agricultural industry types, secondary classification based on production links, and federated learning preprocessing classification based on data characteristics.
[0015] Preferably, the federated learning module includes a local model training unit, a parameter aggregation unit, and a global model update unit. The local model training unit is deployed on the user terminal module side and is used to perform distributed model training based on the classified data. The parameter aggregation unit integrates the model parameters of each node through a secure multi-party computation protocol. The global model update unit synchronizes the aggregated model to the agricultural cloud database.
[0016] Preferably, the federated learning module includes a dynamic task allocation sub-module, which automatically generates federated learning subtasks according to the output characteristics of the data classification module, including:
[0017] (1) Federated learning subtask 1: Training of a pest prediction model based on planting data;
[0018] (2) Federated learning subtask 2: Training of a growth cycle optimization model based on animal husbandry data;
[0019] (3) Federated learning subtask n: Training of a joint feature extraction model for cross-domain data;
[0020] The dynamic task allocation sub-module is connected to the data classification module through the API gateway, and the dynamic task allocation sub-module synchronizes model versions with the agricultural cloud database through the OPC-UA protocol.
[0021] Preferably, the agricultural cloud database includes an agricultural production knowledge graph, a federated learning model library, and a case solution library, and adopts a hierarchical storage architecture to achieve isolated storage of model parameters and business data.
[0022] Preferably, the agricultural cloud database includes:
[0023] (1) Federated model storage layer: Using blockchain technology to store parameters and update logs of various versions of the federated learning model;
[0024] (2) Business knowledge storage layer: Using a graph database to store the agricultural production knowledge graph;
[0025] (3) Case solution storage layer: Using a time series database to store historical problem handling records;
[0026] A logical isolation layer is set between the federated model storage layer and the business knowledge storage layer, and attribute-based encryption technology is used to achieve secure access control of cross-layer data.
[0027] Preferably, the security protection module integrates differential privacy protection, homomorphic encryption, and model watermarking technologies, which run through the entire process of data collection, federated training, and model application.
[0028] Preferably, the method implements a federated learning collaboration process:
[0029] S1: The user terminal module collects agricultural production data, and after multi-level classification and standardization processing by the data classification module, a federated learning ready dataset is generated;
[0030] S2: The federated learning module dynamically allocates training tasks according to data characteristics and establishes a trusted computing environment through the security protection module;
[0031] S3: Each terminal node performs model training locally, processes gradient information using differential privacy technology, and then uploads it to the parameter aggregation unit;
[0032] S4: The parameter aggregation unit completes the global model update through a secure multi-party computing protocol and synchronizes it to the federated model storage layer of the agricultural cloud database;
[0033] S5: The security protection module adds a digital watermark to the updated global model and distributes it to the user terminal module through a trusted execution environment;
[0034] S6: The user terminal module applies the federated learning model for real-time agricultural problem diagnosis, and generates disposal suggestions in combination with the case solution library of the agricultural cloud database.
[0035] Preferably, the local model training in step S3 adopts an asynchronous federated learning mechanism, allowing different terminals to dynamically adjust the training batch and upload frequency according to computing resources, specifically including: high-computing power terminals perform full-scale data training and upload gradients in real time, medium-computing power terminals perform batch data training and upload parameters at set intervals, and low-computing power terminals upload encrypted intermediate representations after feature extraction.
[0036] Compared with the prior art, the present invention provides a method for secure sharing of agricultural data based on federated learning, having the following beneficial effects:
[0037] 1. By adopting federated learning, the present invention retains data locally and only interacts through encrypted model parameters, which can not only protect data privacy but also optimize the global model through multiple rounds of iteration to improve the accuracy of pest and disease prediction and meteorological prediction models. The secure sharing of agricultural data is achieved through federated learning technology, and the beneficial effect of improving the accuracy of agricultural models without disclosing the original data is achieved.
[0038] 2. The privacy technology of the present invention can effectively prevent the reverse inference of individual user data by adding noise to the gradient. At the same time, the addition of homomorphic encryption can ensure that data remains encrypted during the parameter aggregation process and cannot be decrypted even in the aggregation link. This dual protection mechanism can enhance data security while also taking into account the efficiency of model training, enabling the method of the present invention to achieve the beneficial effect of protecting data privacy during model training and parameter aggregation.
[0039] 3. Through the dynamic access control and blockchain evidence storage mechanism, the present invention achieves the beneficial effect of flexible and secure data access management. The dynamic access control mechanism adjusts access permissions in real time according to user identity, permissions, and data sensitivity, which is more flexible and secure than traditional static access control. At the same time, blockchain evidence storage provides an immutable record for the use and sharing of data, thus providing strong support for the compliant use of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the flowchart of the secure sharing of agricultural data of the present invention;
[0041] Figure 2 It is the blockchain architecture diagram of the present invention;
[0042] Figure 3 It is the federated learning architecture diagram of the present invention;
[0043] Figure 4 It is the homomorphic encryption processing process diagram of the present invention;
[0044] Figure 5 This is the homomorphic encryption classification diagram of the present invention;
[0045] Figure 6 This is the system model diagram of the present invention. Specific implementation manners
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Please refer to Figures 1-6 , a method for secure sharing of agricultural data based on federated learning, including the following steps:
[0048] Step 1: Establish a user terminal module, a data classification module, a federated learning module, an agricultural cloud database, and a security protection module. When establishing the user terminal module, it is necessary to develop an adapted federated learning client program according to the hardware characteristics and communication protocols of different types of devices (edge computing devices, intelligent agricultural machinery terminals, and Internet of Things sensors) to ensure the stability and security of device access; when establishing the data classification module, it is necessary to construct a scientific and reasonable multi-level classification system and a differential privacy annotation algorithm framework based on the characteristics of the agricultural industry and data characteristics; when establishing the federated learning module, it is necessary to build a blockchain intelligent contract platform for distributing the initial model, and at the same time construct a homomorphic encryption environment to ensure the secure progress of local model training and parameter aggregation; when establishing the agricultural cloud database, it is necessary to deploy IPFS nodes and configure a hierarchical storage architecture to achieve efficient storage and management of data; when establishing the security protection module, it is necessary to integrate algorithm libraries and interfaces related to blockchain evidence storage and dynamic access control to ensure the effective operation of the protection mechanism
[0049] Step 2: The user terminal module is used to collect agricultural production data, receive federated learning model parameters, and provide a multi-modal interaction interface. It realizes encrypted upload and decryption feedback of agricultural data through edge computing devices, supports secure access of multiple types of devices, and the user terminal module is connected to the data classification module through an encrypted communication network;
[0050] Advantages: By adopting federated learning, the data is retained locally and only interacts through encrypted model parameters, which can not only protect data privacy but also optimize the global model through multiple rounds of iteration to improve the accuracy of pest and disease prediction and meteorological prediction models. The secure sharing of agricultural data is realized through federated learning technology, achieving the beneficial effect of improving the accuracy of agricultural models without revealing the original data.
[0051] Step 3: The data classification module is responsible for performing multi-level classification processing on the agricultural data uploaded by the user terminal module. Based on the first-level classification of agricultural industry types, the data is divided into categories such as planting, animal husbandry, and fishery, and distinguished according to the data characteristics and application scenarios of different industries. Based on the second-level classification of production links, in planting, it can be further divided into sowing, irrigation, fertilization, and pest control links, and in animal husbandry, it can be divided into breeding, feed management, disease prevention and control, etc. links. The data is classified by sorting out the production process. Based on the federated learning preprocessing classification of data characteristics, preprocessing operations such as data normalization and feature screening are performed according to the characteristics of data such as dimension, distribution, and correlation to prepare for subsequent federated learning training. The differential privacy technology is used to perform multi-level label annotation on the planting and animal husbandry data while ensuring the privacy of the classification process. The data classification module is connected to the federated learning module through a distributed network;
[0052] Step 4: The federated learning module distributes the initial model through the blockchain smart contract, combines homomorphic encryption to achieve local model training and parameter aggregation of secure multi-party computation. The local model training unit is deployed on the user terminal module side. During the training process, based on the classified data, optimization algorithms such as stochastic gradient descent are used to perform distributed model training and continuously update the local model parameters. The parameter aggregation unit integrates the model parameters of each node through a secure multi-party computation protocol, such as the oblivious transfer protocol, etc. During the aggregation process, each node encrypts the local model parameters and sends them to the parameter aggregation unit. The parameter aggregation unit performs calculations in the ciphertext state to obtain the aggregated parameters of the global model. The global model update unit synchronizes the aggregated model to the agricultural cloud database, ensuring the security and immutability of data transmission through the blockchain encryption channel. The federated learning module is connected to the agricultural cloud database through the blockchain encryption channel;
[0053] Step 5: The agricultural cloud database uses the IPFS technology to store the encrypted knowledge assets generated by the federated training. For the encrypted knowledge assets generated by the federated training, in IPFS, they are stored in different IPFS objects according to the classification of the agricultural production knowledge graph, the federated learning model library, and the case solution library. The agricultural production knowledge graph is stored in the format of a graph database and accessed through the hash address of IPFS to ensure the consistency and integrity of the graph data. The federated learning model library stores the model parameters and update logs of each version in a specific file format and is also accessed through the hash address for convenient retrieval and invocation of the model. The case solution library stores the historical problem handling records in the format of a time-series database, and uses the distributed storage characteristics of IPFS to ensure the reliability of the data. The agricultural cloud database is connected to the security protection module through a dual-link redundant network;
[0054] Step 6: The security protection module integrates blockchain evidence storage and dynamic access control mechanisms. The blockchain evidence storage mechanism records the key information of data (such as data hash value, timestamp) on the blockchain, and uses the immutable feature of the blockchain to provide permanent proof for the authenticity and integrity of the data. The dynamic access control mechanism adjusts the user's access rights to the data in real time according to factors such as the user's identity, permissions, and data sensitivity. For example, for the core data of high-sensitivity planting industry, only the personnel of agricultural research institutions who have passed strict authentication can access it within a specific time period. This mechanism can better adapt to the complex and changeable application scenarios of agricultural data and ensure data security compared with traditional static access control. The security protection module is connected to the user terminal module through a trusted execution environment interface.
[0055] The advantages are as follows: Through the dynamic access control and blockchain evidence storage mechanisms, the beneficial effect of flexible and secure data access management is achieved. The dynamic access control mechanism adjusts the access rights in real time according to the user identity, permissions, and data sensitivity, which is more flexible and secure than traditional static access control. At the same time, the blockchain evidence storage provides an immutable record for the use and sharing of data, thus providing strong support for the compliant use of data.
[0056] The method of the present invention is realized through the following core steps, specifically:
[0057] (1) Gradient Clipping
[0058] When each agricultural terminal (such as a planting sensor or a livestock monitoring device) participating in federated learning trains the model locally, it first limits the range of the calculated gradient. By setting a threshold (such as L2 norm clipping), the gradient value is clipped to a preset interval to prevent the data of a single sample or user from having too much influence on the global model.
[0059] (2) Noise Addition
[0060] Before uploading the gradient to the aggregation server, add random noise (such as Laplace noise or Gaussian noise) that conforms to the definition of differential privacy to the clipped gradient. The intensity of the noise is controlled by the privacy budget parameter (ε). The smaller the value of ε, the higher the privacy protection level, but it may reduce the model accuracy.
[0061] The advantages are as follows: The privacy technology of the method of the present invention can effectively prevent the reverse inference of single-user data by adding noise to the gradient. At the same time, the addition of homomorphic encryption can ensure that the data is always in an encrypted state during the parameter aggregation process and cannot be decrypted even in the aggregation link. This dual protection mechanism can enhance the security of the data while taking into account the efficiency of model training, so that the method of the present invention achieves the beneficial effect of protecting data privacy during model training and parameter aggregation.
[0062] (3) Secure Aggregation
[0063] Utilize secure multi-party computation (SMPC) or homomorphic encryption technology to aggregate the noise-added gradients, ensuring that the original data of individual users cannot be deduced during the aggregation process. For example, the agricultural cloud server only obtains the encrypted global gradient mean, rather than the specific parameters of individual terminals.
[0064] Homomorphic encryption schemes can be classified into single-key homomorphic encryption (SKHE), multi-key homomorphic encryption (MKHE), and threshold homomorphic encryption (THE) according to the number of keys. The use of single-key homomorphic encryption schemes usually depends on the reliability of the key owner because it has only one pair of keys, while multi-key homomorphic encryption schemes require multiple parties to cooperate in encryption and decryption, avoiding the uncertainty of single-key homomorphic encryption schemes. Both multi-key homomorphic encryption and threshold homomorphic encryption have more than one pair of keys. In the former, the public keys used by each computing party are different. The computing party encrypts the data using the public key and stores it in the cloud server. After that, the cloud server and the computing party can complete the homomorphic operation without further interaction. However, as the key length increases, the computational amount also increases sharply. In the threshold homomorphic encryption scheme, all computing parties use the same public key and still need to interact with the cloud server, but its key length does not change with the increase in the number of participating parties, and the computational amount is relatively stable.
[0065] Model watermarking technology realizes the ownership identification and security verification of the agricultural data sharing model through distributed collaboration and privacy protection mechanisms. The core implementation methods are as follows:
[0066] I. Watermark Embedding Strategy
[0067] (1) Trigger Set Construction and Step-by-Step Training
[0068] Private trigger set design: Each agricultural terminal (such as farm sensors or meteorological monitoring devices) generates a differentiated trigger set according to local data characteristics. For example, the original agricultural image is degraded to generate an irreversible optimized image as a trigger sample.
[0069] Step-by-step training mechanism: In the local model training stage, the user binds the trigger set features and the main task features through adversarial training, so that the local model generates a preset output (such as a specific classification label or image restoration result) for the trigger set, without affecting the main task accuracy.
[0070] (2) Gradient and Parameter Fusion
[0071] Dynamic Gradient Adjustment: During the parameter upload phase of federated learning, watermark information is encoded as gradient perturbations, and the perturbations are incorporated into the global model update through a secure aggregation protocol to ensure the concealment of the watermark.
[0072] Redundant Parameter Tagging: Embed binary codes or hash values in the model weights, and combine homomorphic encryption technology to protect the watermark parameters from being reverse-cracked.
[0073] II. Federated Aggregation and Global Watermark Generation
[0074] (1) Secure Aggregation Protocol
[0075] The central server uses secure multi-party computation (SMPC) to aggregate the noisy gradients, prevent the leakage of watermark information of a single terminal, and record the model update trajectory through the blockchain to enhance the traceability of the watermark.
[0076] Through multiple rounds of iterative aggregation of the global model, the private trigger sets of each terminal are mapped and integrated into a unified global watermark. For example, a unified trigger response is set for the output labels of the pest and disease prediction model.
[0077] (2) Blockchain Reinforced Storage
[0078] After the watermark information (such as trigger set hash, parameter tagging) is associated with the model version, it is stored in the distributed nodes through the blockchain smart contract to ensure the non-tampering of the watermark and provide on-chain evidence for infringement tracing.
[0079] III. Watermark Verification Mechanism
[0080] (1) White-Box Verification
[0081] By decrypting the model weight file, detect the redundant tags or gradient perturbation patterns in the parameters to verify the model ownership.
[0082] (2) Black-Box Verification
[0083] Input the trigger set samples (such as degraded farmland images), and detect whether the model output conforms to the preset response (such as accurately restoring the optimized image or returning a specific classification label).
[0084] IV. Robustness Guarantee
[0085] (1) Anti-Attack Design
[0086] Adopt irreversible trigger sets (such as optimized images) and dynamic watermark update strategies to resist model compression attacks (more than 80% of the watermark is retained when compressed to 30% of the parameter scale) and fine-tuning attacks (more than 90% of the watermark retention rate).
[0087] Through the distributed characteristics of federated learning, the watermark storage location is dispersed to avoid the invalidation of the watermark caused by a single-point attack.
[0088] In this solution, each federated learning task is issued by a task publisher and co-trained by M platform functions. Each platform function AKSPi has its own dataset Di. At the beginning of the training, TA assigns weights according to the data volume owned by each AKSPi. And when distributing keys, since it is for a small number of platform functions, SKAKSPi is directly assigned to the AKSP participating in the federated learning training. Considering that the blockchain platform used is Ethereum and transactions initiated by ordinary accounts to ordinary accounts are not allowed to carry parameters, it is necessary to design a smart contract to upload model parameters to the blockchain.
[0089] Step 1: Registration and smart contract deployment: When new platform functions and task publishers join the network, they need to send a registration request to TA. The request content includes their own addresses, the types and quantities of data they own. After registration, the agricultural intelligent knowledge service platform obtains the right to participate in task training, and the task publisher obtains the right to publish tasks.
[0090] Step 2: Task publishing: The task publisher uploads relevant information such as initial global model parameters to IPFS and records the Hash value returned by IPFS to the blockchain through a smart contract. At the same time, TA randomly selects one of the AKSPi participating in this training as the aggregator.
[0091] Step 3: Key generation: In the global epoch t, TA will honestly execute the threshold homomorphic encryption algorithm to generate the homomorphic encryption key SKAKSPi and the public key PK.
[0092] Step 4: Key sharing: TA uses the shamir scheme to share the key SK. Let a0 = βm. TA randomly selects T values {a1, a2, ……, aT} from {0, 1, 2, ……, n×m - 1} and constructs a polynomial f(x) = aTxT + … + a2x 2 + a1x1 + a0. Then, TA calculates SKAKSPi = f(i) mod nm for the i-th AKSP and sends SKAKSPi to the AKSPs participating in the training respectively.
[0093] Step 5: Local training: In the global epoch t, each AKSPi trains the local model LMi using the local dataset Di. If it is the first epoch, it uses the initial global model parameters IGM for training; otherwise, it uses the previous round's DGM to train the local model, where ni is the total number of samples in the dataset Di and the total number of samples of all participants.
[0094] The loss function of each participant is:
[0095]
[0096] The update of the model parameters is as follows:
[0097]
[0098] After the training is completed, use PK to encrypt the local model parameters LMi, upload them to IPFS, and then add the Hash returned by IPFS and other relevant data as a transaction and upload it to the blockchain. The encryption process is as follows:
[0099]
[0100] where xi ∈ Z*N and xi is random.
[0101] Step Six, Model Parameter Aggregation: After waiting for the preset time t, the aggregator aggregates all the encrypted local model parameters ELMi to obtain EGM and uploads it to IPFS, and then adds the Hash returned by IPFS and other relevant data as a transaction and uploads it to the blockchain. The aggregation process is as follows:
[0102]
[0103] Step Seven, Partial Decryption of Model Parameters: AKSPi obtains the trained EGM from the blockchain, partially decrypts it through SKCDOi, and uploads the decrypted PGMi to IPFS, and then uploads the Hash value returned by IPFS to the blockchain through a smart contract. The partial decryption process is as follows:
[0104]
[0105] Step Eight, Global Model Parameter Decryption: The aggregator collects the partially decrypted global model parameters PGMi from AKSPi. If the number of PGMi collected by the aggregator is less than T, the decryption cannot continue; if the number of PGMi collected is not less than T, the decryption can continue to obtain the decrypted aggregation update, and then upload it to the blockchain. The global decryption process is as follows:
[0106]
[0107] Step Nine, Model Update: AKSPi obtains the latest global model parameters from the blockchain and updates the local model parameters for training.
[0108] Repeat the above process until the model converges, or reaches the required accuracy, or reaches the set number of training rounds.
[0109] An embodiment of the agricultural data security sharing and protection and its working method is as follows:
[0110] Example 1 (Data Sharing for Crop Pest and Disease Control)
[0111] System Composition and Workflow
[0112] 1. Data Holders:
[0113] (1) Agricultural research institutions: Hold a large amount of research data on crop pest and disease control;
[0114] (2) Farmers: Have accumulated practical experience data on crop pest and disease control in actual production;
[0115] 2. Federated Learning Platform: Responsible for coordinating data holders to participate in the federated learning training process, ensuring that data does not leave the local area, and realizing the secure sharing of data;
[0116] 3. Model Training Module: Deploy model training components locally at each data holder, update the model using local data, and upload the model update parameters to the federated learning platform through an encrypted communication protocol for global model aggregation;
[0117] 4. Data Encryption Module: Encrypt all uploaded model update parameters to ensure that they are not stolen or tampered with during transmission;
[0118] 5. Permission Management Module: Control the access permissions of each data holder to shared data to ensure that the data can only be accessed and used by authorized parties;
[0119] 6. Global Model Update Module: The federated learning platform distributes the aggregated global model update parameters to each data holder for further optimization of the local model.
[0120] Specific Implementation Steps:
[0121] T1. Data Holder Registration: Agricultural research institutions and farmers register on the federated learning platform and upload the description information of the local dataset. The uploaded local dataset description information includes the type of data (such as pest and disease image data, control measure text data, etc.), data scale (data volume size, number of samples, etc.), time range of data (start and end times of data collection), and data quality assessment information (such as data accuracy, integrity);
[0122] T2. Model Initialization: The federated learning platform generates an initial global model and distributes it to each data holder;
[0123] T3. Local Model Training: Each data holder uses local data to train the model and generates model update parameters;
[0124] T4. Encrypted Parameter Upload: After encrypting the model update parameters, upload them to the federated learning platform;
[0125] T5. Global Model Aggregation: The federated learning platform decrypts the received encrypted parameters and performs global model aggregation.
[0126] T6. Model Distribution and Optimization: Distribute the updated global model to each data holder and continue the next round of local training;
[0127] T7. Iterate until Convergence: Repeat the above steps until the global model converges and achieves the expected prediction effect.
[0128] Effect of Example 1: Through federated learning, the secure sharing of crop pest and disease control data is realized, the accuracy of the pest and disease prediction model is improved, the data security of data holders is guaranteed, and the risk of data leakage is avoided.
[0129] Comparative Example 1 (Traditional Data Sharing Method - Crop Pest and Disease Control Data)
[0130] In the traditional method, data holders (agricultural research institutions and farmers) directly transmit data to a central data repository. In this process, the data leaves the local area and faces the risk of data leakage. Model training is carried out on the central server using all the aggregated data.
[0131] Specific implementation steps:
[0132] C1. Data Transmission: Agricultural research institutions and farmers organize their respective crop pest and disease control data (including pest and disease image data, control measure text data, etc.) in a unified format and directly transmit it to the central data repository.
[0133] C2. Model Training: On the central server, use traditional machine learning algorithms to train the pest and disease prediction model using all the aggregated data.
[0134] C3. Model Application: The trained model is directly applied to pest and disease control prediction without further distributed optimization process.
[0135] Example 2 (Smart Agriculture Meteorological Data Sharing)
[0136] System Composition and Workflow
[0137] 1. Data Holders:
[0138] (1) Meteorological department: Holds a large amount of meteorological observation data;
[0139] (2) Agricultural enterprises: Have accumulated meteorological data related to agricultural production during actual operation;
[0140] 2. Federated Learning Platform: Responsible for coordinating the participation of meteorological departments and agricultural enterprises in the federated learning training process to achieve the secure sharing of meteorological data;
[0141] 3. Model Training Module: Deploy model training components locally in meteorological departments and agricultural enterprises, update the model using local meteorological data, and upload the model update parameters to the federated learning platform through an encrypted communication protocol for global model aggregation;
[0142] 4. Data Encryption Module: Encrypt all uploaded model update parameters to ensure secure transmission;
[0143] 5. Permission Management Module: Control the access permissions of each data holder to the shared meteorological data;
[0144] 6. Global Model Update Module: The federated learning platform distributes the aggregated global model update parameters to each data holder.
[0145] Specific implementation steps:
[0146] Q1. Registration of Data Holders: Meteorological departments and agricultural enterprises register on the federated learning platform and upload the description information of the local meteorological data sets. The description information of the local meteorological data sets uploaded by the meteorological department includes the type of meteorological data (such as temperature, humidity, precipitation, and wind speed), the data collection frequency (such as once per hour or once per day), the geographical area covered by the data (latitude and longitude range), and the time span of the data (the number of years of historical data). The description information of the meteorological data sets uploaded by agricultural enterprises also needs to include special meteorological indicators related to agricultural production (such as the effective accumulated temperature affecting crop growth) and the application scenarios of the data (such as for farmland irrigation decision-making or crop variety selection);
[0147] Q2. Model Initialization: The federated learning platform generates an initial global meteorological prediction model and distributes it to each data holder;
[0148] Q3. Local Model Training: Each data holder uses local meteorological data for model training to generate model update parameters;
[0149] Q4. Encrypted Parameter Upload: After encrypting the model update parameters, upload them to the federated learning platform;
[0150] Q5. Global Model Aggregation: The federated learning platform decrypts the received encrypted parameters and performs global model aggregation;
[0151] Q6. Model Distribution and Optimization: Distribute the updated global model to each data holder for the next round of local training;
[0152] Q7. Iterate until convergence: Repeat the above steps until the global meteorological prediction model converges to achieve the expected prediction accuracy.
[0153] Effect of Embodiment 2: Through federated learning, the secure sharing of meteorological data for smart agriculture is realized, the accuracy of meteorological prediction is improved, thereby ensuring the data security of data holders, and promoting the effective utilization of meteorological data in the agricultural field.
[0154] Comparative Example 2 (Traditional data sharing method - Meteorological data for smart agriculture)
[0155] The meteorological department and agricultural enterprises also transmit meteorological data to the central data repository. After the central server aggregates the data, it trains the meteorological prediction model.
[0156] Specific implementation steps:
[0157] D1. Data transmission: The meteorological department transmits meteorological data (such as temperature, humidity, precipitation, and wind speed), and agricultural enterprises transmit meteorological data related to agricultural production (such as effective accumulated temperature, etc.) to the central data repository.
[0158] D2. Model training: On the central server, all meteorological data is integrated, and the meteorological prediction model is trained using traditional meteorological prediction algorithms.
[0159] D3. Model application: The trained model is used for meteorological prediction in agricultural production, and no distributed model updates are performed.
[0160] Comparison of test table data:
[0161]
[0162] The following information is obtained from the above table:
[0163] The workflow covers five stages: encrypted data upload, privacy protection classification, federated model training, solution generation, and secure feedback. Through the federated paradigm of "data does not move, model moves", on the premise of ensuring that the original data does not leave the domain, the value mining and secure sharing of cross-regional agricultural data are realized, providing accurate decision-making support for scenarios such as pest control and growth monitoring, and the accuracy is improved by 23% - 35% compared with traditional methods.
[0164] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for secure sharing of agricultural data based on federated learning, characterized in that: The following steps are involved: Step 1: Establish user terminal module, data classification module, federated learning module, agricultural cloud database and security protection module; Step 2: The user terminal module is used to collect agricultural production data, receive federated learning model parameters, and provide a multimodal interactive interface. It realizes agricultural data encryption upload and decryption feedback through edge computing devices, supports secure access of multiple types of devices, and the user terminal module is connected to the data classification module through an encrypted communication network. Step 3: The data classification module is responsible for multi-level classification of agricultural data uploaded by the user terminal module, and uses differential privacy technology to perform multi-level labeling on planting and animal husbandry data, while ensuring the privacy of the classification process. The data classification module is connected to the federated learning module through a distributed network; Step 4: The federated learning module distributes the initial model through blockchain smart contracts, combines homomorphic encryption to achieve parameter aggregation of local model training and secure multi-party computing, and connects the federated learning module to the agricultural cloud database through a blockchain encrypted channel; Step 5: The agricultural cloud database uses IPFS technology to store the encrypted knowledge assets generated by federated training. The agricultural cloud database and the security protection module are connected through a dual-link redundant network; Step 6: The security protection module integrates blockchain evidence storage and dynamic access control mechanism. The security protection module is connected to the user terminal module through the trusted execution environment interface.
2. The method for secure sharing of agricultural data based on federated learning according to claim 1 is characterized in that: The user terminal module supports the access of three types of devices: edge computing devices, intelligent agricultural machinery terminals and Internet of Things sensors. Each type of device is configured with a dedicated federated learning client program.
3. The method for secure sharing of agricultural data based on federated learning according to claim 1 is characterized in that: The multi-level classification processing of the data classification module includes: primary classification based on agricultural industry type, secondary classification based on production links and federated learning pre-processing classification based on data features.
4. The method for secure sharing of agricultural data based on federated learning according to claim 1 is characterized in that: The federated learning module includes a local model training unit, a parameter aggregation unit and a global model update unit. The local model training unit is deployed on the user terminal module side and is used to perform distributed model training based on classified data. The parameter aggregation unit integrates the model parameters of each node through a secure multi-party computing protocol. The global model update unit synchronizes the aggregated model to the agricultural cloud database.
5. The method for secure sharing of agricultural data based on federated learning according to claim 1 is characterized in that: The federated learning module includes a dynamic task allocation submodule, which automatically generates federated learning subtasks according to the output features of the data classification module, including: (1) Federated learning subtask 1: Training of pest and disease prediction model based on crop data; (2) Federated learning subtask 2: Growth cycle optimization model training based on animal husbandry data; (3) Federated learning subtask n: joint feature extraction model training for cross-domain data; The dynamic task allocation submodule is connected to the data classification module through an API gateway, and the dynamic task allocation submodule synchronizes model versions with the agricultural cloud database through the OPC-UA protocol.
6. The method for secure sharing of agricultural data based on federated learning according to claim 1, characterized in that: The agricultural cloud database includes an agricultural production knowledge graph, a federated learning model library and a case solution library, and adopts a layered storage architecture to achieve isolated storage of model parameters and business data.
7. The method for secure sharing of agricultural data based on federated learning according to claim 1, characterized in that: The agricultural cloud database comprises: (1) Federated model storage layer: uses blockchain technology to store parameters and update logs of various versions of federated learning models; (2) Business knowledge storage layer: using graph database to store agricultural production knowledge graph; (3) Case solution storage layer: using a time series database to store historical problem processing records; A logical isolation layer is set between the federated model storage layer and the business knowledge storage layer, and attribute-based encryption technology is used to implement secure access control of cross-layer data.
8. The method for secure sharing of agricultural data based on federated learning according to claim 1, characterized in that: The security protection module integrates differential privacy protection, homomorphic encryption and model watermarking technology, and runs through the entire process of data collection, federated training and model application.
9. The method for secure sharing of agricultural data based on federated learning according to claim 8, characterized in that: The method implements the federated learning collaborative process: S1: The user terminal module collects agricultural production data, which is then classified and standardized by the data classification module to generate a federated learning-ready dataset. S2: The federated learning module dynamically allocates training tasks based on data features and establishes a trusted computing environment through the security protection module; S3: Each terminal node performs model training locally, processes gradient information using differential privacy technology, and then uploads it to the parameter aggregation unit; S4: The parameter aggregation unit completes the global model update through a secure multi-party computing protocol and synchronizes it to the federated model storage layer of the agricultural cloud database; S5: The security protection module adds a digital watermark to the updated global model and sends it to the user terminal module through the trusted execution environment; S6: The user terminal module applies the federated learning model to perform real-time agricultural problem diagnosis and generates disposal suggestions in combination with the case solution library of the agricultural cloud database.
10. The method for secure sharing of agricultural data based on federated learning according to claim 9, characterized in that: The local model training in step S3 adopts an asynchronous federated learning mechanism, allowing different terminals to dynamically adjust training batches and upload frequencies according to computing resources, specifically including: high-computing power terminals perform full data training and upload gradients in real time, medium-computing power terminals perform batch data training and upload parameters according to a set period, and low-computing power terminals perform feature extraction and then upload encrypted intermediate representations.
Citation Information
Cited By
Trans-department data collaborative modeling method and system based on federal learning
CN120408727A
A cross-departmental data collaborative modeling method and system based on federated learning
CN120408727B
Soil deep humidity prediction system and method based on near infrared spectrum and federated learning
CN120629057A
Soil deep humidity prediction system and method based on near-infrared spectroscopy and federated learning
CN120629057B
Whole-process data operation and value co-creation system based on trusted data space
CN120710797A