Agricultural product identification method and system, and storage medium
Through multi-sensor acquisition and deep feature fusion technology, combined with blockchain technology, the problems of low recognition accuracy and insufficient data security in agricultural product recognition are solved, and an agricultural product recognition system with high accuracy, robustness and full traceability are achieved.
Patent Information
- Application Number
- CN202510292330.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems in agricultural product identification with low recognition accuracy, poor noise immunity and insufficient environmental adaptability, especially in multimodal data fusion, deep feature extraction and data secure storage.
Holographic images, multispectral images, weight and environmental data are collected through multiple sensors, and after preprocessing and space-time alignment, features are extracted through the holographic visual Transformer model, spectral convolution network and fully connected neural network respectively. Hypergraph neural network is used to achieve deep feature fusion, digital twin data and digital fingerprints are generated, and the multimodal identification network is used for accurate classification and quality evaluation, while using blockchain technology to ensure the immutability of data.
It significantly improves the accuracy, robustness and full-process traceability of agricultural product identification, solves the shortcomings of traditional methods in environmental adaptability and feature extraction, and ensures the security and credibility of data through blockchain technology.
Smart Images

Figure CN120218950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent production and food safety traceability, and particularly relates to an agricultural product identification method and system, and a storage medium. Background Art
[0002] The identification of agricultural products is of great significance for ensuring food safety, improving product quality, and traceability tracking. Existing technologies mainly use means such as information code scanning, traditional image recognition, and weight detection to identify and classify agricultural products. These technical means usually rely on single-modal data collection and analysis, and there are problems such as low identification accuracy, poor anti-noise performance, and insufficient environmental adaptability. For example, traditional image recognition methods are often limited by environmental factors such as light and occlusion, resulting in incomplete extraction of surface features of agricultural products; relying solely on weight data to judge product quality is difficult to reflect internal defects or changes in maturity; in addition, although information code scanning can provide basic traceability information, the manual records and simple database retrieval it relies on are prone to introducing errors and risks of data tampering. Therefore, there are obvious deficiencies in existing technologies in terms of multi-modal data fusion, deep feature extraction, and data security storage. There is an urgent need for a new identification technology that comprehensively applies multi-sensor data and advanced algorithms to improve the accuracy, robustness, and security of agricultural product identification. Summary of the Invention
[0003] In view of the many problems existing in the above-mentioned prior art, the present invention provides an agricultural product identification method and system, and a storage medium. The present invention collects holographic images, hyperspectral images, weight, and environmental data through multiple sensors, and after preprocessing and spatio-temporal alignment of these data, extracts their respective features through a pre-trained holographic vision Transformer model, a dedicated spectral convolutional network, and a fully connected neural network respectively. Furthermore, a hypergraph neural network is used to achieve deep feature fusion, construct digital twin data, and generate digital fingerprints. Subsequently, a multi-modal recognition network compares the fused data with standard data, and realizes accurate agricultural product classification and quality assessment through Euclidean distance calculation; at the same time, the recognition results are signed and uploaded to the blockchain to ensure the data cannot be tampered with. The system realizes global model adaptive update through federated learning to ensure continuous optimization of the recognition algorithm in different environments. The present invention realizes the comprehensive utilization and secure storage of multi-modal information, and significantly improves the accuracy, robustness, and full-process traceability ability of agricultural product identification.
[0004] An agricultural product identification method includes the following steps:
[0005] When an agricultural product is placed on a smart traceability scale, collect holographic image data, hyperspectral image data, weight data, and environmental data, and attach time stamps to the respective data to form a global original data packet;
[0006] Preprocess the global original data packet within the edge computing unit to obtain an aligned data packet that has undergone basic noise reduction, calibration, and spatio-temporal alignment;
[0007] Based on the aligned data packet, perform feature encoding and fusion on each modality data to construct digital twin data of agricultural products, and use a hypergraph neural network to achieve deep fusion of each modality feature data, generating digital twin data and digital fingerprint data; input the digital twin data into a multi-modal recognition network, compare it with the standard digital twin data, and then output the agricultural product recognition result. After digitally signing the recognition result data, digital fingerprint data, and the aligned data packet summary, upload them to the consortium blockchain to form an immutable traceability block data. At the same time, use federated learning to achieve global model adaptive update.
[0008] Preferably, the feature encoding step of holographic image data includes inputting the preprocessed holographic image data into a pre-trained holographic vision Transformer model. The holographic vision Transformer model uses a 3×3 convolutional layer to extract local features with a stride of 1, and performs deep processing on the extracted local features through eight self-attention encoding layers, thereby outputting holographic image feature data with a fixed dimension of 512.
[0009] Preferably, the feature encoding step of multi-spectral image data includes inputting the preprocessed multi-spectral image data into a dedicated spectral convolutional network. The dedicated spectral convolutional network includes a first convolutional layer with a 3×3 convolutional kernel, a stride of 1, and 64 filters; a second convolutional layer with a 3×3 convolutional kernel, a stride of 1, and 128 filters; a third convolutional layer with a 3×3 convolutional kernel, a stride of 1, and 256 filters; and after global average pooling processing, outputting multi-spectral image feature data with a fixed dimension of 256.
[0010] Preferably, the numerical feature encoding step includes merging the preprocessed weight data and the preprocessed environmental data and inputting them into an encoder composed of a three-layer fully connected neural network. The first fully connected layer maps the merged data to 256 nodes, the second fully connected layer maps the data to 128 nodes, and the third fully connected layer maps the data to 64 nodes, thereby outputting numerical feature data with a fixed dimension of 64.
[0011] Preferably, use the holographic image feature data, the multi-spectral image feature data, and the numerical feature data as nodes respectively to form a hypergraph structure. The hypergraph structure uses two-layer hypergraph convolutional layers for data fusion. Each hypergraph convolutional layer uses a multi-head self-attention mechanism to achieve information interaction between nodes, thereby generating fused digital twin data. The fused digital twin data is processed by L2 normalization and then uses the SHA256 hash algorithm to generate digital fingerprint data.
[0012] Preferably, the multi-modal recognition network includes multiple self-attention modules and consecutive fully-connected layers. The fused digital twin data is processed by at least two layers of self-attention modules. Each layer of self-attention module uses eight attention heads to extract features from the input data, and gradually reduces the dimension through consecutive fully-connected layers and maps it to a feature space consistent with the standard digital twin data. The standard digital twin data is stored in a dedicated storage unit. The comparison uses Euclidean distance calculation, so as to output the recognition result data composed of agricultural product categories, quality evaluation, and status determination.
[0013] Preferably, the recognition result data, digital fingerprint data, and aligned data packet digest are used to generate a digital signature through an asymmetric encryption algorithm. The digital signature uses a preset private key to sign the above data and is transmitted to the consortium blockchain platform through a secure encryption channel, and non-tamperable traceability block data is stored in the blockchain data structure.
[0014] Preferably, the method further includes locally storing the generated recognition result data and abnormal alarm data on the intelligent traceability scale, and using an edge computing unit to perform online training. The online training uses the gradient descent method to calculate the local model update data, and the local model update data is generated based on historical recognition result data and abnormal alarm data;
[0015] The local model update data is periodically uploaded to the central server through a secure encryption protocol. The central server uses the federated averaging algorithm to perform weighted average processing on the local model update data generated by each intelligent traceability scale to generate global update model data, where the federated averaging algorithm includes calculating the average value of the gradients of the local model update data of each device;
[0016] The global update model data is sent to each intelligent traceability scale through a secure network. After each intelligent traceability scale receives the global update model data, it updates the local model parameters, so as to realize the adaptive update and continuous optimization of the global model.
[0017] An agricultural product recognition system for implementing the agricultural product recognition method, the system includes:
[0018] A collection module, configured to collect holographic image data, multi-spectral image data, weight data, and environmental data when the agricultural product is placed on the intelligent traceability scale, and attach a high-precision time stamp to the above data to form a global original data packet;
[0019] An edge processing unit, configured to preprocess the global original data packet, including performing noise reduction, color correction, and spatio-temporal alignment processing on the holographic image data and multi-spectral image data, so as to generate an aligned data packet;
[0020] Feature Encoding and Fusion Unit, including a holographic image feature encoding module, a hyperspectral image feature encoding module, and a numerical feature encoding module. Each module extracts features from the preprocessed holographic image data, hyperspectral image data, and weight and environmental data respectively. The extracted holographic image feature data, hyperspectral image feature data, and numerical feature data form a hypergraph structure as nodes, and are deeply fused through a hypergraph neural network to generate fused digital twin data and digital fingerprint data;
[0021] Multimodal Recognition Unit, which is used to input the fused digital twin data into a multimodal recognition network, compare it with the standard digital twin data, and output the agricultural product recognition result;
[0022] Secure Upload Module, which is used to upload the recognition result data, the digital fingerprint data, and the aligned data packet digest to the consortium blockchain platform after digital signature to form tamper-proof traceability block data;
[0023] Federated Learning Unit, which is used to collect local model update data on each intelligent traceability scale and transmit it to the central server periodically through a secure encryption channel. The central server uses the federated averaging algorithm to integrate the update data of each device to generate global update model data, and sends the global update model data to each intelligent traceability scale to achieve adaptive update of the global model.
[0024] A storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the agricultural product recognition method.
[0025] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:
[0026] By adopting the hypergraph neural network technology means, the present invention realizes the deep fusion and information interaction of multimodal data (holographic image, hyperspectral image, and numerical data), thereby generating highly discriminative fused digital twin data, and generating unique digital fingerprint data using the SHA256 hash algorithm;
[0027] By adopting the multi-head self-attention module and the continuous fully connected layer, the present invention realizes the accurate comparison of the fused digital twin data and the standard digital twin data in the Euclidean space, and outputs the category, quality evaluation, and status determination data of agricultural products, thereby solving the deficiencies of traditional methods in environmental adaptability and feature extraction;
[0028] The present invention uses asymmetric encryption technology to digitally sign the recognition results, digital fingerprints, and data digests, and combines the consortium blockchain technology to achieve the immutability and traceability of the recognition data; through the means of federated learning technology, each terminal device can adaptively update the global model without sharing the original data, ensuring that the system continuously improves the recognition accuracy and robustness in a dynamic environment;
[0029] In addition, the present invention comprehensively utilizes technical means such as hardware acceleration, real-time data synchronization, and secure transmission, significantly improving the overall performance and application value of the agricultural product recognition system, and meeting the strict requirements of modern agricultural production for food safety, quality control, and traceability tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a schematic flowchart of the method of the present invention;
[0031] Figure 2 is a flowchart of federated learning and global model update in the present invention;
[0032] Figure 3 is a block diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0034] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0035] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0036] As Figure 1 shown, an agricultural product recognition method includes the following steps:
[0037] When agricultural products are placed on a smart traceability scale, holographic image data, multispectral image data, weight data, and environmental data are collected, and a timestamp is attached to each piece of data to form a global raw data packet;
[0038] When agricultural products are placed on a smart traceability scale, it is first necessary to synchronously collect multi-source information for the agricultural products. In order to achieve a comprehensive analysis of the appearance, internal characteristics, and external environmental conditions of agricultural products in subsequent steps, it is necessary to record holographic image data, multispectral image data, weight data, and environmental data at the same time point, and attach timestamps to these data and encapsulate them into a global raw data packet.
[0039] The collection of holographic image data relies on a combination of high-speed cameras from multiple perspectives. Triggered by precise hardware synchronization signals, the appearance images of agricultural products at different angles can be completely captured at the same moment, and each perspective image contains relatively rich lighting and geometric information. Multispectral image data is obtained using a camera module with specific band response capabilities to acquire the image information of agricultural products under different spectra. This part of the information is often related to the internal moisture, sugar content, or lesion areas of agricultural products, thus providing a key basis for subsequent quality assessment and anomaly detection. The collection of weight data uses a high-precision digital load sensor to monitor the actual load value exerted by the agricultural products on the scale surface, and accurately records this load value as numerical data. The collection of environmental data mainly includes temperature, humidity, and illuminance information, which is obtained by arranging multiple sensors on the scale body or around it to obtain the real-time environmental state. In order to ensure the corresponding relationship of these data in the time dimension, it is necessary to synchronize all collection devices with the same high-precision clock at the hardware level, and attach a millisecond-level or higher-precision timestamp to each collection result at the software level, so as to ensure that the appearance, multispectral characteristics, weight, and environmental conditions of agricultural products at any moment can correspond to the same real-world state.
[0040] For example, if an instantaneous fluctuation in the recorded weight value is detected at a certain moment, and the multispectral image shows an abnormal reflection intensity in the near-infrared channel, then these two anomalies can be correlated and analyzed based on the same timestamp. Through this multi-sensor multi-angle synchronous collection and timestamp integration, the final global raw data packet contains both the appearance characteristics of agricultural products at the visual and spectral levels, as well as supplementary information such as weight and environmental status, laying a data foundation for subsequent multi-modal fusion and deep recognition. At the operational level, automatic exposure adjustment can be performed on the holographic image data to ensure that the image brightness is appropriate, while the timing of band filter switching of the multispectral camera module is set to ensure that images in different bands can be obtained within the same sampling period. The weight sensor part realizes zero-point automatic calibration in the hardware to reduce the no-load error.
[0041] In addition, the deployment locations of environmental sensors can also be optimized. For example, temperature and humidity sensors can be installed at the edge of the weighing surface to be close to the actual surrounding environment of agricultural products, while light sensors can be arranged above the weighing body to ensure better sensitivity to surrounding light changes. After uniformly labeling these data, they are packaged and stored, facilitating the alignment and correlation of multi-modal features based on timestamps during subsequent reading and analysis.
[0042] Preprocess the global raw data packet within the edge computing unit to obtain an aligned data packet that has undergone basic noise reduction, calibration, and spatio-temporal alignment;
[0043] In the edge computing unit, it is necessary to perform basic noise reduction, calibration, and spatio-temporal alignment on the above-mentioned global raw data packet to obtain a higher-quality aligned data packet that meets the requirements of multi-modal fusion. Holographic image data and multi-spectral image data may be affected by factors such as uneven illumination, image sensor noise, and lens distortion, and need to be processed through filtering and calibration means.
[0044] Common noise reduction methods include mean filtering and median filtering. However, to better retain the detailed information on the surface of agricultural products, bilateral filtering can be selected. This filter retains edges while smoothing the image. When operating specifically, filtering parameters need to be set separately for different bands or for the visible light and infrared channels. For calibration, a geometric distortion correction method based on a calibration board can be used to compensate for the geometric distortion in multi-view images and perform coordinate alignment at the pixel level for multi-spectral images to prevent spatial offsets in images of different bands.
[0045] If color correction or spectral reflectance calibration is required, a standard color card or spectral reflectance plate with known reflectance can be used. By collecting images of this standard sample and comparing the difference between the theoretical reflectance value and the actual collected value, a calibration coefficient matrix is calculated, thereby performing unified color or spectral correction on agricultural product images. To ensure temporal synchronization consistency, interpolation or truncation processing needs to be performed on each frame of image and its corresponding weight value and environmental data. Assuming that the image acquisition frame rate and the weight sensor update frequency are inconsistent, the weight data can be mapped to the time points corresponding to the image frames through interpolation, so that the data of all modalities at the same timestamp can be correctly corresponded.
[0046] If the linear interpolation method is used, the weight value at a certain time point can be calculated by weighting the weight values at two adjacent sampling times according to the time difference, so as to achieve multi-modal synchronization without affecting the overall accuracy. After processing, all the data that have completed noise reduction, calibration, and spatio-temporal alignment will be encapsulated into a new data structure, that is, the alignment data packet. This data packet can be recorded in the form of hierarchical indexing. For example, the first-level index indicates the timestamp, and the second-level index indicates the storage location of each modal data at this time point. Through such a design, subsequent algorithms can quickly retrieve the holographic image, hyperspectral image, weight, and environmental data at a specific moment, so as to perform multi-modal fusion and in-depth analysis. This alignment data packet can also record meta-information such as filtering parameters and calibration matrices used during processing, which is convenient for traceability and reuse. In this way, the whole process from agricultural products contacting the intelligent traceability scale to forming the alignment data packet not only ensures the integrity and consistency of multi-source data, but also lays a solid data foundation for subsequent digital twin construction, deep learning recognition, and blockchain traceability records.
[0047] As Figure 2 shown, based on the alignment data packet, feature encoding and fusion are performed on each modal data to construct the digital twin data of agricultural products, and a hypergraph neural network is used to achieve the deep fusion of each modal feature data, generating digital twin data and digital fingerprint data; the digital twin data is input into a multi-modal recognition network, and after comparison with the standard digital twin data, the agricultural product recognition result is output, and the recognition result data, digital fingerprint data, and alignment data packet summary are uploaded to the consortium blockchain after digital signature to form an immutable traceability block data, and at the same time, federated learning is used to achieve global model adaptive update.
[0048] After the basic processing of the alignment data packet is completed, it is necessary to perform feature encoding and fusion on each modal data contained therein to obtain digital twin data that can comprehensively reflect the state of agricultural products. The construction of digital twin data depends on the deep extraction and correlation analysis of different modal features, including holographic image data, hyperspectral image data, and numerical information (such as weight data and environmental data).
[0049] In order to achieve a high degree of coupling of each modal feature in the same representation space, feature encoding can be performed on each modal separately first. For example, key visual features in the holographic image are extracted through a vision Transformer model, the difference features of the surface and internal structure of agricultural products under the hyperspectral image are obtained through a spectral convolutional network, and a fully connected network is used to perform normalization mapping on the weight and environmental values. After unifying these feature vectors of different dimensions, a hypergraph neural network needs to be used to construct the interaction relationship between multi-modal data.
[0050] The hypergraph neural network will achieve cross-modal aggregation of features on the abstract structures of nodes and hyperedges. Each node represents a feature vector of a certain modality, while the hyperedges characterize the connections between different modalities. To achieve deep fusion, a self-attention mechanism can be introduced into the hypergraph neural network, enabling the network to focus on those feature dimensions that contribute more to agricultural product recognition when updating node representations. A unique index needs to be assigned to each modality node, and its feature vector should be placed in the node table of the network during initialization. Then, based on the complementary relationship between modalities, the hyperedge structure is defined, making the image feature nodes and multi-spectral feature nodes of the same agricultural product closely connected. The numerical feature nodes can also interact with other nodes through weight adjustment for information exchange.
[0051] After the network completes multi-layer convolutional updates, a fused high-dimensional vector can be output. This vector can comprehensively reflect the characteristics of agricultural products in terms of appearance, internal properties, and environmental conditions. This fused vector is the digital twin data. To ensure the efficiency of subsequent retrieval or comparison, the fused vector can be normalized first, and then digital fingerprint data can be generated through a hashing algorithm (such as SHA256). The digital fingerprint data can be understood as a hash value of a fixed length, used to identify the unique digital twin of the agricultural product. Through this process, the system can not only extract key information from multi-source data but also represent it uniformly in the same representation space, thus significantly improving the accuracy and stability of subsequent recognition.
[0052] After the digital twin data is constructed, it needs to be input into the multi-modal recognition network and compared with the standard digital twin data, and finally the agricultural product recognition result is output. The multi-modal recognition network can be a deep network containing self-attention and hypergraph learning structures, used to further explore the similarity or difference between the fused vector and known standard samples. The standard digital twin data can be stored in a dedicated database, and each standard entry corresponds to the feature representation of a specific agricultural product variety or quality grade. When the network compares, it calculates the similarity between the input vector and the standard vector. If the similarity is higher than the preset threshold, it is determined to be of the same variety or quality grade; if the deviation is large, corresponding abnormal prompts are output or more refined category divisions are made. After the recognition is completed, the recognition result data, digital fingerprint data, and aligned data packet summary will be digitally signed and uploaded to the consortium blockchain. Digital signatures usually use asymmetric encryption algorithms to ensure the uniqueness of the signatory and the integrity of the data, and the consortium blockchain guarantees that these records cannot be tampered with through a multi-node consensus mechanism.
[0053] At the same time, the federated learning mechanism can also be used to achieve global model adaptive updates. Each intelligent traceability scale securely transmits the model update data obtained from local training to the central node. The central node performs federated averaging on the updates of each device and distributes the global updated model, thereby continuously improving the overall recognition performance in different environments and different agricultural product variety scenarios.
[0054] For example, when a new special variety of agricultural product is added in a certain area, the intelligent traceability scales in this area accumulate a large amount of new data locally after multiple identifications. They can report the updated model parameters through federated learning. The central node integrates them and then broadcasts them to other devices, enabling the entire system to have the ability to identify this new agricultural product. During this process, the existence of digital twin data enables each identification to be judged based on the high-dimensional representation after multi-modal fusion, while the immutable characteristic of the blockchain provides a traceable security guarantee for the identification process and results.
[0055] Preferably, the feature encoding step of the holographic image data includes inputting the preprocessed holographic image data into a pre-trained holographic vision Transformer model. The holographic vision Transformer model uses a 3×3 convolutional layer to extract local features with a stride set to 1, and performs in-depth processing on the extracted local features through eight self-attention encoding layers, thereby outputting holographic image feature data with a fixed dimension of 512.
[0056] The feature encoding step of the holographic image data includes inputting the preprocessed and normalized holographic image data into a pre-trained holographic vision Transformer model. Before receiving the input data, the model first normalizes the input image to ensure the consistency of the images under different shooting conditions in terms of brightness, color, and contrast, thereby reducing the influence of environmental light and device noise on subsequent feature extraction. Subsequently, the model uses a 3×3 convolutional layer (with a convolutional kernel size of 3×3 and a stride set to 1) to extract local features from the image. At this time, the convolutional operation can capture the subtle texture, edges, and local geometric structure information in the image, ensuring that the subtle changes on the surface of the agricultural product in the image are effectively retained. The extracted local features are then fed into eight self-attention encoding layers, and each layer uses a multi-head self-attention mechanism to achieve global feature interaction. The formula is as follows:
[0057]
[0058] Among them, Q (query vector), K (key vector), and V (value vector) are respectively generated by linear mapping, d is the dimension of the key vector, and the softmax function ensures that each attention score can play a key role in the weighted sum after normalization. These eight self-attention modules enable the relationships between different local regions to be fully modeled. Even if the same agricultural product shows large differences at different angles, its internal correlation can still be learned through the attention weights. After self-attention encoding, the model finally outputs holographic image feature data with a fixed dimension of 512. To further ensure the numerical stability of the feature vector, the output holographic image feature data will also undergo L2 normalization processing, and its standard form is:
[0059]
[0060] Among them, x is the output vector, and ‖x‖2 represents its Euclidean norm. This normalization step not only helps to reduce the inconsistency of the data distribution but also provides a stable input for subsequent hashing operations. Through the above operations, the pre-trained holographic vision Transformer model can fully extract and fuse local and global information in the agricultural product images, generating holographic image feature data with high discrimination ability and robustness, providing an accurate visual basis for the entire agricultural product recognition system. In the embodiment, in actual application, the convolutional layer parameters and the number of self-attention layers can be fine-tuned according to the characteristics of specific agricultural products to achieve the optimal recognition effect.
[0061] Preferably, the feature encoding step of the multispectral image data includes inputting the preprocessed multispectral image data into a dedicated spectral convolutional network, and the dedicated spectral convolutional network includes a first convolutional layer, using a 3×3 convolutional kernel, a stride of 1, and 64 filters; a second convolutional layer, using a 3×3 convolutional kernel, a stride of 1, and 128 filters; a third convolutional layer, using a 3×3 convolutional kernel, a stride of 1, and 256 filters; and performing global average pooling processing, thereby outputting multispectral image feature data with a fixed dimension of 256.
[0062] The feature encoding step of the multispectral image data includes inputting the preprocessed multispectral image data into a dedicated spectral convolutional network, which extracts features for different band information contained in the multispectral data. First, the preprocessed multispectral image data is processed by a first convolutional layer, which uses a 3×3 convolutional kernel, a stride of 1, and sets the number of filters to 64 to initially extract low-level spectral features such as edges, textures, and local spectral reflectance information; subsequently, the data is passed to a second convolutional layer, which also uses a 3×3 convolutional kernel, a stride of 1, but the number of filters is increased to 128 to further capture more detailed spectral changes and local structural differences in the image; then, the data enters a third convolutional layer, using the same size convolutional kernel and stride settings, and at the same time the number of filters is increased to 256, aiming to fuse and abstract the features extracted by the previous two layers at a deeper level.
[0063] After passing through the convolutional layer, the data is processed by a global average pooling layer. This step integrates the spatial dimension information of each channel into a feature vector with a fixed dimension, and outputs multispectral image feature data with a fixed dimension of 256. Global average pooling not only helps to reduce the dimension of the feature vector but also can smooth the noise, thereby improving the stability of the model.
[0064] In this way, the reflectance, absorptance, and other band feature information contained in the multi-spectral image data can be comprehensively captured and integrated into an efficient feature representation. In the embodiments, the number of filters and the convolution kernel size of each convolutional layer can be determined through experiments to adapt to the specific spectral characteristics of different agricultural products. At the same time, the parameters of global average pooling are adjusted by using the cross-validation method between different spectral channels, so as to further improve the expression ability and robustness of the multi-spectral feature data and provide key support for subsequent multi-modal fusion.
[0065] Preferably, the numerical feature encoding step includes inputting the preprocessed weight data and the preprocessed environmental data into an encoder composed of a three-layer fully connected neural network after merging them. The first fully connected layer maps the merged data to 256 nodes, the second fully connected layer maps the data to 128 nodes, and the third fully connected layer maps the data to 64 nodes, so as to output numerical feature data with a fixed dimension of 64.
[0066] The numerical feature encoding step includes numerically merging the preprocessed weight data and the preprocessed environmental data and inputting them into an encoder composed of a three-layer fully connected neural network. First, the merged data is used as the input and sent to the first fully connected layer, which maps the input data to 256 nodes to capture the preliminary interaction information between the weight and environmental data. Subsequently, the output passes through the second fully connected layer, which maps the data to 128 nodes, and further extracts the deep features in the data through a non-linear activation function (such as ReLU). Then, the data enters the third fully connected layer, which maps the data to 64 nodes and outputs numerical feature data with a fixed dimension of 64. This fully connected neural network structure can effectively reduce the dimension of the original numerical data while retaining the key information, facilitating the consistency of data dimensions when fusing with visual and spectral features subsequently.
[0067] To ensure network stability, the Batch Normalization technique can be used to normalize the output of each layer during the training process, reducing the training fluctuations caused by differences in data distributions. In the embodiments, for different agricultural products, specific normalization strategies can be adjusted during the preprocessing of weight and environmental data. For example, maximum-minimum normalization or Z-score standardization can be used for temperature, humidity, and illuminance data respectively to ensure that their numerical ranges are consistent. The finally output 64-dimensional numerical feature data provides a concise and effective numerical representation for subsequent multi-modal fusion. This representation can not only reflect the physical quality information of agricultural products but also enhance the robustness and accuracy of overall recognition as an auxiliary feature. Through the layer-by-layer mapping and dimension reduction of numerical data by the fully connected network, the present invention effectively reduces noise interference while improving the compactness and discriminability of data representation, providing indispensable numerical support for the overall agricultural product recognition system.
[0068] Preferably, the holographic image feature data, the hyperspectral image feature data, and the numerical feature data are respectively used as nodes to form a hypergraph structure. The hypergraph structure uses two layers of hypergraph convolutional layers for data fusion. Each hypergraph convolutional layer uses a multi-head self-attention mechanism to achieve information interaction between nodes, thereby generating fused digital twin data. The fused digital twin data is subjected to L2 normalization processing and then uses the SHA256 hash algorithm to generate digital fingerprint data.
[0069] In the present invention, the holographic image data is input into a pre-trained holographic vision Transformer model after normalization processing. The model first uses a three-by-three convolutional layer (convolution kernel size is 3×3, stride is set to 1) to extract local features of the input image, capturing the texture and edge information on the surface of agricultural products in the image; subsequently, the model uses eight layers of self-attention encoding layers to perform global feature interaction on the local features output by the convolutional layer, thereby obtaining holographic image feature data with a fixed dimension of 512. The hyperspectral image data is processed by a dedicated spectral convolutional network, which consists of three convolutional layers. The first convolutional layer uses a three-by-three convolutional kernel, a stride of 1, and 64 filters. The second convolutional layer uses a three-by-three convolutional kernel, a stride of 1, and 128 filters. The third convolutional layer uses a three-by-three convolutional kernel, a stride of 1, and 256 filters. Immediately afterwards, the global average pooling operation integrates the spatial information output by each layer into a feature vector with a fixed dimension of 256. The numerical feature data is obtained by merging the preprocessed weight data and environmental data (temperature, humidity, light, etc.) and input into a three-layer fully connected neural network. The first layer maps the merged data to 256 nodes, the second layer maps to 128 nodes, and the third layer maps to 64 nodes, finally outputting numerical feature data with a fixed dimension of 64.
[0070] After initializing the above three types of feature data as nodes respectively, according to the complementary relationship between different modality data of agricultural products, hyperedges are constructed in the hypergraph structure to connect each node. The hyperedge weights can be set according to the feature similarity between nodes to ensure that different modality information can complement each other. Subsequently, the present invention uses two layers of hypergraph convolutional layers for data fusion. Each layer uses a multi-head self-attention mechanism to achieve information interaction between nodes, and its core calculation formula is
[0071]
[0072] Among them, Q (query vector), K (key vector), and V (value vector) are respectively generated by linear mappings. d is the dimension of the key vector. In actual operations, the features of each node are mapped to the query, key, and value spaces through linear transformations, and the attention scores are calculated using the above formula. Then, the information from neighboring nodes is fused in a weighted summation manner. After the first layer of hypergraph convolution, the features of each node are preliminarily updated, and the second layer of hypergraph convolution further refines the fusion effect. The finally output fusion vector is the fused digital twin data. To ensure the consistency of the output data in terms of scale and numerical stability, the present invention performs L2 normalization on the fused digital twin data, and its normalization formula is:
[0073]
[0074] where x represents the fused digital twin data vector, and ‖x‖2 represents its Euclidean norm, and the calculation method is The normalized vector is processed by the SHA256 hashing algorithm to generate digital fingerprint data of a fixed length, and this digital fingerprint data is used as a unique identifier in subsequent identification and traceability processes. Through the above steps, the present invention uses a hypergraph neural network to achieve the deep fusion of visual, spectral, and numerical data, forming a stable and highly discriminative digital representation, providing strong data support for the accurate identification of agricultural products, and at the same time ensuring the consistency and robustness of the data in subsequent processing.
[0075] Preferably, the multi-modal recognition network includes multiple self-attention modules and consecutive fully connected layers. The fused digital twin data is processed by at least two layers of self-attention modules. Each layer of self-attention module uses eight attention heads to extract features from the input data, and gradually reduces the dimension through consecutive fully connected layers and maps it to a feature space consistent with the standard digital twin data. The standard digital twin data is stored in a dedicated storage unit, and the comparison uses the Euclidean distance calculation, so as to output the recognition result data composed of the agricultural product category, quality assessment, and status determination.
[0076] The multi-modal recognition network is used to compare the above-mentioned fused digital twin data with the pre-constructed standard digital twin data, so as to output the agricultural product recognition result data. This recognition network is mainly composed of multiple self-attention modules and consecutive fully connected layers. Its design concept is to use the self-attention mechanism to dynamically weight key features in the high-dimensional feature space to achieve fine extraction of information between different modalities. Specifically, the fused digital twin data first enters at least two layers of self-attention modules. Each layer of the module uses eight attention heads to extract the correlation features of each dimension in the input data in parallel. The attention calculation follows the formula In this formula, Q, K, and V are respectively the query, key, and value vectors obtained through linear transformation, and d is the dimension of the key vector.
[0077] The output of the self-attention module is gradually reduced in dimension through successive fully connected layers, which adopt a combination of linear mapping and ReLU activation functions to gradually map the high-dimensional feature vectors to the feature space consistent with the standard digital twin data. The standard digital twin data is obtained by training a large number of agricultural product samples and stored in a dedicated storage unit. In the comparison stage, the similarity is evaluated by calculating the Euclidean distance between the fused digital twin data and the standard digital twin data. The Euclidean distance calculation formula is:
[0078]
[0079] where x represents the input fused digital twin data vector, y represents the standard digital twin data vector, and (x i -y i ) 2 represents the square of the difference between the corresponding components of the two vectors. When the calculated distance is lower than the preset threshold, the system determines that the agricultural product matches the corresponding standard category and outputs the recognition result data including the agricultural product category, quality assessment, and status determination. Through the design of this multi-modal recognition network, the present invention can adapt to different types and quality differences of agricultural products while ensuring high accuracy, and further improve the recognition robustness by dynamically adjusting network parameters, forming a set of efficient and widely applicable agricultural product recognition solutions.
[0080] Preferably, the recognition result data, digital fingerprint data, and aligned data packet digest are used to generate a digital signature through an asymmetric encryption algorithm. The digital signature signs the above data using a preset private key and is transmitted to the consortium blockchain platform through a secure encrypted channel, and non-tamperable traceability block data is stored in the blockchain data structure.
[0081] To ensure the security and non-tamperability of the agricultural product recognition process and result data during transmission and storage, the present invention uses an asymmetric encryption algorithm to digitally sign the recognition result data, digital fingerprint data, and aligned data packet digest, and transmits them to the consortium blockchain platform for storage through a secure encrypted channel.
[0082] In the specific implementation process, the recognition result data output by the multi-modal recognition network contains key information such as agricultural product category, quality assessment, and status determination; the digital fingerprint data is generated by performing L2 normalization and SHA256 hashing algorithm on the fused digital twin data; the aligned data packet digest is the extraction of the key verification information of all preprocessed multi-modal data. When using an asymmetric encryption algorithm, the present invention can select the RSA algorithm, and its signature process can be expressed as:
[0083] S = Sign(private key, R‖F‖D)
[0084] Where R represents the recognition result data, F represents the digital fingerprint data, D represents the alignment data packet digest, "∥" represents the data concatenation operation, and the private key is only held by the data generation device. After the digital signature S is generated, the system transmits R, F, D, and S to the consortium blockchain platform through security protocols such as TLS. Each node in the platform verifies the legality of the signature according to the consensus mechanism and records the verified data in the blockchain data structure to form an immutable traceability block data. For example, if the recognition of a batch of agricultural products is completed, the recognition result data, digital fingerprint data, and alignment data packet digest are all signed and uploaded to the chain. Any third party can use the preset public key to verify the signature to ensure the authenticity and reliability of the data source. Through this digital signature and blockchain evidence storage mechanism, the present invention not only improves the security of data transmission, but also provides a credible basis for the entire agricultural product traceability system, ensuring the integrity and anti-tampering ability of data during subsequent traceability or supervision processes.
[0085] Preferably, the method further includes locally storing the generated recognition result data and abnormal alarm data on the intelligent traceability scale, and using the edge computing unit to perform online training. The online training uses the gradient descent method to calculate the local model update data, and the local model update data is generated based on historical recognition result data and abnormal alarm data;
[0086] The local model update data is periodically uploaded to the central server through a secure encryption protocol. The central server uses the federated averaging algorithm to perform weighted average processing on the local model update data generated by each intelligent traceability scale to generate global update model data, where the federated averaging algorithm includes calculating the average value of the gradients of the local model update data of each device;
[0087] The global update model data is sent to each intelligent traceability scale through a secure network. After each intelligent traceability scale receives the global update model data, it updates the local model parameters, thereby realizing the adaptive update and continuous optimization of the global model.
[0088] In the intelligent traceability scale of the present invention, the recognition result data generated after the agricultural product recognition is locally stored together with the abnormal alarm data, and the edge computing unit is used to perform online training to generate local model update data. To achieve this process, first, each intelligent traceability scale is equipped with a high-performance edge computing device, which is responsible for preliminarily processing the real-time recognition results collected by the sensor, and storing the result data and abnormal alarm data in the local storage medium in chronological order. The recognition result data includes the category of agricultural products, quality evaluation indicators, and status determination information, while the abnormal alarm data records the abnormal situations found during the recognition process, such as a large deviation between the input data and the standard data, inconsistent image or spectral features, etc. The stored data is not only marked with timestamps but also undergoes preliminary formatting to ensure that the data format and content meet the requirements of subsequent online training. After the data storage is completed, the edge computing unit starts the online training module, which updates the local model using the gradient descent method. Specifically, the online training module will first load the currently stored historical recognition result data and abnormal alarm data and organize them into batches of training samples. The training samples usually consist of input data and target labels, where the input data includes the currently collected fused digital twin data and other auxiliary features, and the target labels correspond to the standard agricultural product categories and quality evaluation results. To ensure model convergence and improve generalization ability during training, batch normalization and learning rate decay techniques are usually adopted during edge training to balance the pace of model parameter updates. During training, the local model update data is calculated using the gradient descent method, that is, the gradient of the loss function with respect to the model parameters is calculated through the backpropagation algorithm, and then the model parameters are updated at a fixed or adaptive learning rate. The core calculation expression is:
[0089]
[0090] where θ (t) represents the current model parameters, η represents the learning rate, represents the loss function with respect to the gradient of the model parameters θ. After being updated by the gradient descent method, local model update data is generated. Here, the loss function can choose cross-entropy loss or mean squared error loss according to the specific objectives of agricultural product recognition, and the learning rate η is determined by experimental tuning to ensure a stable and efficient update process.
[0091] In actual operation, the edge computing unit of each intelligent traceability scale periodically recalculates the model gradient according to the latest collected recognition result data and abnormal alarm data and generates corresponding local model update data. To prevent a single device from causing model instability due to data fluctuations, a regularization term, such as the L2 regularization term, is also introduced during the online training process, and its form is:
[0092]
[0093] Where λ is the regularization coefficient, and this term helps to constrain excessive changes in model parameters, thereby improving the robustness of the model. During the local training process, the edge computing unit will also monitor the data of each training batch, record key metrics such as training loss and accuracy, and adjust the training strategy according to real-time feedback to ensure that the local model update data reflects the optimal parameter update information in the current acquisition environment. In this way, the local model update data not only contains the model parameter gradients, but also can record the abnormal data situations that occur during the training process, providing data support and reference basis for the subsequent global model integration. Finally, these local model update data will be stored locally and used as an important input in the federated learning stage to ensure that each intelligent traceability scale can generate accurate and timely model update information according to the latest local training situation, thereby laying a foundation for the adaptive update of the global model.
[0094] This local online training and model update process is highly operable in practice. For example, in actual deployment, a fixed training period (such as every 30 minutes or every hour) can be set, and an appropriate batch size and learning rate can be selected according to the device computing power and network conditions to ensure that the update data is both accurate and efficient. Through the above measures, the present invention effectively realizes the automatic storage and online training of local recognition results and abnormal data, generates local model update data, provides a reliable data source and computing basis for the global model update of the overall system, and at the same time ensures the real-time adaptive ability and efficient operation performance of the system in the face of dynamic environmental changes.
[0095] The local model update data is periodically uploaded to the central server through a secure encryption protocol. The central server uses the federated averaging algorithm to perform weighted averaging on the local model update data generated by each intelligent traceability scale, thereby generating global update model data. In the present invention, in order to realize the secure aggregation of multi-terminal model update data, first, after each intelligent traceability scale completes local online training, it will encrypt the local model update data through a dedicated encryption module, use encryption algorithms such as AES or RSA to encrypt and protect the data, and periodically upload the encrypted local model update data to the central server through a secure encryption protocol (such as TLS or IPSec). After receiving the data uploaded by each device, the central server first decrypts and verifies the integrity of the data to ensure that the uploaded data has not been tampered with or damaged.
[0096] Subsequently, the central server uses the federated averaging algorithm to integrate the local model update data of all devices. The core idea of this algorithm is to perform weighted averaging on the model gradients or parameter updates calculated by each device to obtain a global model update. The specific calculation formula is as follows:
[0097]
[0098] where θ i represents the local model update data uploaded by the i-th intelligent traceability scale, and N represents the total number of devices participating in the update. In the implementation of this formula, each θ i contains the gradient information obtained by the device through online training. The central server calculates the global updated model parameter θ global by weighted averaging the gradient data of each device. To ensure that the contribution of each device to the global model can reflect the real data distribution, the number of device samples is usually used as the weighting factor in the federated averaging algorithm, that is, the improved formula is:
[0099]
[0100] where n i represents the number of training samples collected by the i-th device. This weighting method can effectively avoid the overinfluence of devices with a small amount of data on the global model. In actual operation, the central server will also record the timestamps and error metrics of the data uploaded by each device, dynamically monitor the data, and adjust the weighting strategy according to the detection results to cope with situations such as network latency or data anomalies.
[0101] After the global updated model data is generated, the central server saves it in a dedicated database and distributes it to each intelligent traceability scale through a secure network to ensure that all terminal devices obtain the latest model parameters. To verify the effectiveness of the upload and integration process, the central server usually conducts a model evaluation test at the end of each update cycle, calculates the loss function value and accuracy of the global model on the validation dataset, and compares the historical data trends to ensure that the performance of the global model is continuously improving.
[0102] Through this federated learning mechanism, the present invention can not only make full use of the local update data of each terminal device, but also achieve the adaptive update of the global model without directly sharing the original data, which not only protects the data privacy of each device, but also improves the generalization ability and robustness of the overall recognition system in a dynamic environment. In the embodiment, the update cycle can be set to a fixed time interval (for example, once per hour), and the upload frequency and batch size can be adjusted according to the stability of the data uploaded by the device and the network bandwidth conditions, so as to achieve the optimal global integration effect. Through the above measures, the global updated model data generated by the central server provides the latest and optimal model parameters for the entire agricultural product recognition system, ensuring that each intelligent traceability scale can utilize the continuously optimized model in the subsequent recognition process, improving the recognition accuracy and response speed, and realizing the continuous optimization and dynamic adaptation of the system.
[0103] The global updated model data is sent to each intelligent traceability scale through a secure network. After receiving the global updated model data, each intelligent traceability scale replaces the local model parameters according to the updated data, so as to realize the adaptive update and continuous optimization of the global model. In the present invention, each intelligent traceability scale is pre-installed with a local model, which is responsible for real-time identification of the collected agricultural product data and outputting the identification result. The global model data updated through the federated learning mechanism is embodied as a set of new model parameter sets, and these parameters reflect the more comprehensive data characteristics of each terminal device after being weighted and integrated by the central server.
[0104] After receiving the global updated model data, the terminal device will transmit the data through a secure network (such as using the TLS protocol) and replace and update the original model parameters in the local storage medium. To ensure the smooth progress of the update process, each device usually verifies and tests the local model before and after the update. For example, the recognition accuracy and loss function value are calculated on a set of standard data sets, and the test results are compared with the evaluation results of the global model on the central server to confirm that the updated model can indeed improve the recognition performance. The updated local model parameters will directly participate in the subsequent agricultural product identification tasks, ensuring that the system always uses the latest and optimal model for processing, thereby improving the recognition accuracy and robustness.
[0105] To ensure the integrity of data transmission and parameter replacement during the model update process, each intelligent traceability scale is built-in with a security module (such as a TPM chip) for digital signature and verification, so as to prevent malicious tampering. In the embodiment, the model update frequency can be set to be automatically triggered after each training ends or batch updated according to a preset time interval; at the same time, during the update process, each device can record the model performance indicators before and after the update and report the update effect to the central server through a feedback mechanism to further optimize the global model integration strategy. Through this global update mechanism, the present invention realizes the real-time adaptive update of the model parameters, enabling the system to flexibly cope with the diversity of agricultural product types, environmental conditions, and data collection changes, so as to maintain a high level of recognition accuracy and response speed during long-term operation, while ensuring the continuous and stable optimization of the system.
[0106] As Figure 3 shown, an agricultural product identification system for implementing the agricultural product identification method, the system includes:
[0107] The acquisition module is used to collect holographic image data, multispectral image data, weight data, and environmental data when agricultural products are placed on the intelligent traceability scale, and attach a high-precision timestamp to the above data to form a global raw data packet. The hardware implementation of the acquisition module is designed for multi-modal data acquisition when agricultural products are placed on the intelligent traceability scale. This module consists of multiple groups of high-performance sensors, mainly including a holographic light field camera system, a multispectral camera module, a digital load sensor, and an environmental monitoring sensor. The holographic light field camera system uses multiple high-resolution (e.g., 12 megapixels) high-speed cameras, which are installed at different positions on the scale body according to a predetermined geometric arrangement to achieve synchronous shooting of agricultural products from multiple perspectives.
[0108] Each camera supports a collection rate of at least 60 frames per second and achieves precise synchronization through an internal hardware trigger signal to ensure that the image data collected at the same time point can reflect the details of agricultural products from different angles. The multispectral camera module integrates near-infrared and short-wave infrared sensors, and its filter and optical components are strictly calibrated to ensure accurate reflectance data in specific bands, thereby capturing the spectral characteristics of the surface and internal structure of agricultural products. The digital load sensor uses a high-precision measuring device with a measurement accuracy of up to 0.1 grams and has an automatic zero calibration function to eliminate long-term drift and environmental interference. In addition, the environmental monitoring part includes temperature, humidity, and light intensity sensors, and all sensors convert analog signals into digital signals through independent A / D converters.
[0109] To ensure the consistency of all sensor data in the time dimension, the system is built-in with a high-precision real-time clock module, and all acquisition devices are synchronized with this clock. A millisecond-level timestamp is attached during data acquisition through embedded firmware. The signals of each sensor in the acquisition module are transmitted to the main control unit through a high-speed bus (e.g., SPI or LVDS), and preliminary data encapsulation is performed by an embedded processor, finally forming a global raw data packet containing holographic image data, multispectral image data, weight data, and environmental data. In the actual hardware design, to ensure data stability and anti-interference ability, the system shields sensor signals and uses differential transmission technology on the data bus. In addition, each sensor obtains a stable power supply through an independent power supply module to avoid affecting data acquisition accuracy due to power fluctuations. The entire acquisition module fully considers the diversity of agricultural products and the complexity of the acquisition environment in its structural design, and realizes efficient data acquisition and preprocessing through modular design, providing a solid, accurate, and synchronous raw data foundation for subsequent multi-modal data fusion.
[0110] Edge processing unit, which is used to preprocess the global raw data packets, including denoising, color correction, and spatio-temporal alignment processing on holographic image data and multispectral image data, so as to generate aligned data packets; the hardware implementation of the edge processing unit aims to efficiently preprocess the global raw data packets and generate aligned data packets that meet the requirements of subsequent feature encoding. This unit is mainly composed of a high-speed embedded processor (such as based on the ARM Cortex-A series or NVIDIA Jetson AGX Xavier platform), and integrates a dedicated digital signal processor (DSP) and FPGA module to accelerate image processing and data correction tasks. During implementation, the holographic image data and multispectral image data first undergo denoising processing through a digital filter. To reduce the influence of uneven illumination and random noise, the system adopts a bilateral filtering algorithm. Next, the processing unit performs color correction and geometric correction on the image data, generating a correction matrix (for example, using a perspective transformation model) with the pre-acquired calibration plate image to correct the geometric distortion caused by lens distortion; color correction calculates the correction parameter matrix through the color deviation between the standard color card image and the actually acquired image, mapping the color space of the acquired image to the standard color space. At the same time, the edge processing unit performs Kalman filtering smoothing on the weight data to eliminate sensor jitter. All processed image, weight, and environmental data are spatio-temporally aligned using an interpolation algorithm based on their respective timestamps to ensure that each data sample corresponds to an exact match of multimodal data at the same time point. The aligned data packets are stored in a hierarchical data structure, with the first layer being the timestamp index and the second layer being the storage unit for each modality data. The edge processing unit manages various processing tasks through a real-time operating system (RTOS) to ensure the minimization of data processing latency, and transfers the processing results to the subsequent feature encoding module through a high-speed data bus. This processing unit adopts a heat dissipation design and redundant power supply to ensure the stable operation of the system, and can still ensure the efficient and accurate output of data under high-load data processing conditions, laying a solid foundation for subsequent multimodal feature encoding and fusion.
[0111] Feature Encoding and Fusion Unit, including a holographic image feature encoding module, a multispectral image feature encoding module, and a numerical feature encoding module. Each module respectively extracts features from the preprocessed holographic image data, multispectral image data, as well as weight data and environmental data. The extracted holographic image feature data, multispectral image feature data, and numerical feature data are used as nodes to form a hypergraph structure, and deep fusion is achieved through a hypergraph neural network to generate fused digital twin data and digital fingerprint data. In hardware implementation, the Feature Encoding and Fusion Unit integrates the holographic image feature encoding module, the multispectral image feature encoding module, and the numerical feature encoding module. Each module operates on a dedicated accelerator platform to achieve real-time feature extraction and efficient data fusion. The holographic image feature encoding module uses a pre-trained holographic vision Transformer model for processing. This module is accelerated by GPU in hardware. By loading the pre-trained model parameters, the normalized holographic image data is sent into the convolutional layer. Local features are extracted using a 3x3 convolutional kernel (stride of 1), and global feature interaction is achieved through eight self-attention encoding layers, finally outputting holographic image feature data with a fixed dimension of 512. The multispectral image feature encoding module is processed through a dedicated spectral convolutional network, which is implemented on FPGA. Three consecutive convolutional operations are adopted, with the number of filters being 64, 128, and 256 in sequence, and finally outputting multispectral image feature data with a fixed dimension of 256 through a global average pooling layer. The numerical feature encoding module merges the preprocessed weight data and environmental data and inputs them into an encoder composed of a three-layer fully connected neural network. Matrix operations are accelerated by ASIC in hardware. The first layer maps the input data to 256 nodes, the second layer maps to 128 nodes, and the third layer maps to 64 nodes, outputting numerical feature data with a fixed dimension of 64. After each feature data is extracted, the Feature Encoding and Fusion Unit uses the above-mentioned holographic image feature data, multispectral image feature data, and numerical feature data as nodes to form a hypergraph structure. The hardware platform manages the hypergraph data structure through an application-specific integrated circuit (ASIC), and connects each node according to its modal relevance based on the predefined hyperedge connection rules. Next, a two-layer hypergraph convolutional network is used for deep fusion, and each layer integrates a multi-head self-attention mechanism. In hardware, the multi-head self-attention module performs parallel computing on the GPU, and eight attention heads simultaneously process the input features. After passing through the two-layer hypergraph convolutional network, fused digital twin data is output. To ensure the numerical stability of the feature fusion result, the fused digital twin data undergoes L2 normalization processing. The normalized data is hashed through an embedded SHA256 hash engine to generate digital fingerprint data with a fixed length, which is used for subsequent data comparison and traceability records.The entire feature encoding and fusion unit adopts high-bandwidth memory and high-speed data buses in its hardware design to ensure the real-time performance and accuracy of large-scale matrix operations and data transmission, fully meeting the high requirements of the agricultural product recognition system for data processing speed and precision.
[0112] The multi-modal recognition unit is used to input the fused digital twin data into the multi-modal recognition network, compare it with the standard digital twin data, and output the agricultural product recognition result. In its hardware implementation, the multi-modal recognition unit is composed of a high-performance GPU and a dedicated AI accelerator. Its main task is to input the fused digital twin data into the multi-modal recognition network, compare it with the pre-constructed standard digital twin data, and thus output the agricultural product recognition result. The multi-modal recognition network is composed of multiple self-attention modules and consecutive fully-connected layers. Its design is based on the dynamic weighting of high-dimensional features in the self-attention modules and the gradual dimensionality reduction in the fully-connected layers. First, the fused digital twin data is transmitted by the hardware to the recognition unit and processed by at least two layers of self-attention modules, with each self-attention module using eight attention heads. In this process, the node features are linearly mapped to generate query, key, and value vectors, and the attention weights are obtained through parallel computing, thereby realizing the adaptive reconstruction of features. Next, the output data after passing through the self-attention modules enters the consecutive fully-connected layers, where linear transformations and non-linear activation functions (such as ReLU) are executed layer by layer to gradually reduce the high-dimensional fused vector to a low-dimensional feature space that matches the standard digital twin data. The standard digital twin data is obtained by training with a large number of agricultural product samples and stored in a dedicated storage unit. During the recognition process, the hardware system calculates the Euclidean distance between the fused digital twin data and the standard digital twin data, and determines the agricultural product category, quality assessment, and status determination based on the distance threshold. In the hardware implementation, the multi-modal recognition unit utilizes the parallel computing ability and efficient memory caching technology to achieve large-scale matrix operations and distance calculations, ensuring real-time recognition and high accuracy. After the comparison, the recognition result data output by the system is processed by the AI accelerator and then transmitted to the next link. At the same time, the key intermediate data during the recognition process is used for subsequent traceability records to ensure the traceability and data integrity of the entire recognition process.
[0113] The secure blockchain module is used to upload the recognition result data, the digital fingerprint data, and the aligned data packet digest to the consortium blockchain platform after digital signature to form tamper-proof traceability block data. In hardware implementation, the secure blockchain module consists of an embedded security chip and a dedicated encryption processor. Its main function is to digitally sign the recognition result data, digital fingerprint data, and aligned data packet digest output by the multi-modal recognition network, and upload them to the consortium blockchain platform through a secure encryption channel to form tamper-proof traceability block data. The secure blockchain module first digitally signs the recognition result data locally using an asymmetric encryption algorithm (such as RSA or ECC). In the hardware platform, the secure encryption processor is responsible for real-time encryption and data integrity verification to ensure that the data is not tampered with during the transmission process. Nodes on the blockchain platform use a preset public key to verify the digital signature and write the verified data into the blockchain through a multi-node consensus mechanism. In this way, whether during the data transmission process or after storage, the authenticity and non-forgery of agricultural product recognition information can be guaranteed, providing a strong security guarantee for supervision and subsequent traceability. In the embodiment, the blockchain operation can be automatically triggered after each recognition to ensure that the data is securely stored with the shortest delay.
[0114] The federated learning unit is used to collect local model update data on each intelligent traceability scale and periodically transmit it to the central server through a secure encrypted channel. The central server uses the federated averaging algorithm to integrate the update data of each device to generate global update model data, and then sends the global update model data to each intelligent traceability scale to achieve the adaptive update of the global model. In hardware implementation, the federated learning unit is composed of the edge computing units of each intelligent traceability scale and the central server. Its core lies in achieving the adaptive update of global model parameters while protecting the data privacy of each terminal. After each intelligent traceability scale completes the identification of agricultural products, it stores the local identification result data and abnormal alarm data locally, and uses the edge computing unit to perform online training through the gradient descent method to calculate the local model update data. The local model update data is periodically uploaded to the central server through a dedicated encryption module (using AES or RSA algorithm) via a secure encryption protocol. After receiving the local model update data from each intelligent traceability scale, the central server uses the federated averaging algorithm to perform weighted integration on the update data of each device. After generating the global update model data, the central server sends it to each intelligent traceability scale through a secure network (such as TLS). After each terminal device receives the update data, it replaces the original model parameters and conducts verification tests to ensure that the new model has higher accuracy and robustness in actual identification tasks. In hardware implementation, the federated learning unit needs to fully consider the communication delay between devices, the computational load of data encryption and decryption, and the storage and processing capabilities of the central server. Therefore, each intelligent traceability scale is usually equipped with a dedicated accelerator to handle local training tasks, while the central server uses a high-performance distributed computing platform for data integration and model update. Through this federated learning mechanism, the present invention realizes the adaptive optimization of the global model in different acquisition environments while maintaining data privacy, and finally ensures that the entire agricultural product identification system has the ability to continuously improve and efficient processing performance during long-term operation.
[0115] A storage medium, on which a computer program is stored, and when the program is executed by a processor, it realizes the steps of the agricultural product identification method.
[0116] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0117] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for identifying agricultural products, characterized in that: The following steps are involved: When agricultural products are placed on the smart traceability scale, holographic image data, multispectral image data, weight data and environmental data are collected, and timestamps are added to the data to form a global original data packet; Preprocessing the global original data packet in the edge computing unit to obtain an aligned data packet that has undergone basic noise reduction, correction, and spatiotemporal alignment; Based on the aligned data packet, feature encoding and fusion are performed on each modal data to construct the digital twin data of the agricultural product, and a hypergraph neural network is used to realize deep fusion of feature data of each modality to generate digital twin data and digital fingerprint data; the digital twin data is input into a multimodal recognition network, and the agricultural product identification result is output after comparison with the standard digital twin data, and the identification result data, digital fingerprint data and alignment data packet summary are uploaded to the alliance blockchain after digital signature to form tamper-proof traceability block data, and federated learning is used to realize global model adaptive update.
2. The method according to claim 1, characterized in that: The feature encoding step of the holographic image data includes inputting the preprocessed holographic image data into a pre-trained holographic vision Transformer model, wherein the holographic vision Transformer model uses a three-by-three convolutional layer to extract local features, with a stride of one, and performs deep processing on the extracted local features through eight self-attention encoding layers, thereby outputting holographic image feature data with a fixed dimension of 512.
3. The method according to claim 2, characterized in that: The feature encoding step of the multispectral image data includes inputting the preprocessed multispectral image data into a dedicated spectral convolutional network, wherein the dedicated spectral convolutional network includes a first convolutional layer, using a three-by-three convolution kernel, a stride of one, and a filter number of 64; a second convolutional layer, using a three-by-three convolution kernel, a stride of one, and a filter number of 128; a third convolutional layer, using a three-by-three convolution kernel, a stride of one, and a filter number of 256; and undergoing global average pooling processing, thereby outputting multispectral image feature data with a fixed dimension of 256.
4. The method according to claim 3, characterized in that: The numerical feature encoding step includes merging the preprocessed weight data with the preprocessed environmental data and inputting the data into an encoder composed of a three-layer fully connected neural network, wherein the first fully connected layer maps the merged data to 256 nodes, the second fully connected layer maps the data to 128 nodes, and the third fully connected layer maps the data to 64 nodes, thereby outputting numerical feature data with a fixed dimension of 64.
5. The method according to claim 4, characterized in that: The holographic image feature data, the multispectral image feature data and the numerical feature data are respectively used as nodes to form a hypergraph structure. The hypergraph structure adopts two layers of hypergraph convolution layers for data fusion. Each hypergraph convolution layer adopts a multi-head self-attention mechanism to realize information interaction between nodes, thereby generating fused digital twin data. The fused digital twin data is L2 normalized and then the SHA256 hash algorithm is used to generate digital fingerprint data.
6. The method according to claim 5, characterized in that: The multimodal recognition network includes multiple self-attention modules and continuous fully connected layers. The fused digital twin data is processed by at least two layers of self-attention modules. Each layer of self-attention modules uses eight attention heads to extract features from the input data, and gradually reduces the dimension through continuous fully connected layers and maps it to a feature space consistent with the standard digital twin data. The standard digital twin data is stored in a dedicated storage unit, and the comparison is calculated using Euclidean distance to output recognition result data consisting of agricultural product category, quality assessment and status determination.
7. The method according to claim 1, characterized in that: The identification result data, digital fingerprint data and aligned data packet summary are encrypted using an asymmetric algorithm to generate a digital signature. The digital signature uses a preset private key to sign the above data and is transmitted to the alliance blockchain platform through a secure encrypted channel. The generated tamper-proof traceability block data is stored in the blockchain data structure.
8. The method according to claim 1, characterized in that: The method further includes locally storing the generated recognition result data and abnormal alarm data on the smart traceability scale, and performing online training using an edge computing unit, wherein the online training uses a gradient descent method to calculate local model update data, and the local model update data is generated based on the historical recognition result data and abnormal alarm data; The local model update data is periodically uploaded to the central server through a secure encryption protocol. The central server uses a federated average algorithm to perform weighted average processing on the local model update data generated by each smart traceability scale to generate global update model data, wherein the federated average algorithm includes calculating the average value of the local model update data gradient of each device; The global update model data is sent to each smart traceability scale through a secure network. After receiving the global update model data, each smart traceability scale updates the local model parameters, thereby achieving adaptive update and continuous optimization of the global model.
9. An agricultural product identification system, used to implement the agricultural product identification method according to any one of claims 1 to 8, characterized in that: The system includes: The acquisition module is used to collect holographic image data, multispectral image data, weight data and environmental data when agricultural products are placed on the smart traceability scale, and to attach high-precision timestamps to the above data to form a global original data packet; An edge processing unit, used for preprocessing the global raw data packet, including performing noise reduction, color correction and spatiotemporal alignment processing on the holographic image data and the multispectral image data, so as to generate an aligned data packet; A feature coding and fusion unit, including a holographic image feature coding module, a multispectral image feature coding module and a numerical feature coding module, wherein each module extracts features from the preprocessed holographic image data, multispectral image data, weight data and environmental data, respectively, and the extracted holographic image feature data, multispectral image feature data and numerical feature data are used as nodes to form a hypergraph structure, and deep fusion is achieved through a hypergraph neural network to generate fused digital twin data and digital fingerprint data; A multimodal recognition unit, used to input the fused digital twin data into a multimodal recognition network, compare it with the standard digital twin data, and output the agricultural product recognition result; A secure on-chain module, used to upload the identification result data, the digital fingerprint data and the summary of the aligned data packet to the alliance blockchain platform after digital signature, to form tamper-proof traceability block data; The federated learning unit is used to collect local model update data on each smart traceability scale and periodically transmit it to the central server through a secure encrypted channel. The central server uses a federated averaging algorithm to integrate the update data of each device to generate global update model data, and then sends the global update model data to each smart traceability scale to achieve adaptive update of the global model.
10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the agricultural product identification method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Jujube variety discrimination method and system based on artificial intelligence
CN120995315A
Red, green and blue (RGB) image and weight information fusion-based test method and system
CN121877663A