Intelligent data acquisition and processing method for whole industry chain of shredded rice

By building a terminal-edge node-cloud collaborative architecture in the Siliao rice industry chain, data is collected and processed in real time. Machine learning and AI algorithms are used to extract feature information, which solves the problem of unreasonable data collection and processing. This enables efficient and accurate data transmission and analysis across the entire industry chain, supporting quality control and safety traceability.

CN121901697APending Publication Date: 2026-04-21GUANGDONG MODERN AGRI EQUIP RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG MODERN AGRI EQUIP RES INST
Filing Date
2025-12-05
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In the Silky Rice industry chain, data collection is scattered, formats are incompatible, data types are complex, transmission delays are not suitable, data feature extraction is incomplete, and processing is unreasonable, which makes data linkage analysis difficult and affects data value mining and industry chain efficiency.

Method used

A collaborative architecture of terminal-edge node-cloud is built to collect and preprocess data in real time. Latency parameters and security factors are determined through machine learning models. Feature information is extracted by combining clustering and AI algorithms, and weights are assigned for data processing and compression.

Benefits of technology

It enables timely, secure transmission and accurate processing of data across the entire industry chain, eliminates redundant data, improves data processing efficiency and accuracy, provides reliable support for quality control and safety traceability, and promotes the intelligent upgrading of the industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901697A_ABST
    Figure CN121901697A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent data acquisition and processing method for a whole industry chain of shredded rice, and belongs to the technical field of data acquisition and processing, and the method comprises the steps: building a terminal-edge node-cloud collaborative architecture, and collecting historical industry chain data to form a sample library; collecting data in real time, preprocessing the data, analyzing the change to obtain sensitive indexes and characteristics, and learning a fixed delay parameter by using a machine in combination with an architecture parameter; calculating a time factor and a safety factor, clustering a sample library and matching similar sample cluster preliminary screening data; and extracting features by using CNN-Transform to calculate a difference degree, distributing weights according to link importance and the difference degree, generating a space correlation sequence in combination with smoothing parameters, and uploading compressed data. The data processing efficiency and accuracy are improved, and the intelligence of the Yumiao rice industry is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data acquisition and processing technology, and in particular to an intelligent data acquisition and processing method for the entire industrial chain of high-quality rice. Background Technology

[0002] Silky Rice, a high-quality indica rice variety in my country, encompasses multiple stages of its industrial chain, including planting, processing, storage, transportation, and sales. Accurate data collection and efficient processing at each stage are crucial for ensuring the quality of Silky Rice, improving production efficiency, and achieving safety traceability. However, current data management within the Silky Rice industrial chain suffers from several pain points: First, data collection is fragmented, with each stage employing independent collection equipment and methods, lacking a unified collaborative architecture. This leads to incompatible data formats, incomplete collection scope, and difficulty in achieving interconnected analysis of data across the entire industrial chain. Second, data types are complex, with invalid and redundant data consuming significant storage and transmission resources. Third, data transmission latency lacks dynamic adaptability; different stages have varying demands for timely transmission, and fixed latency settings can lead to delays in critical data transmission or waste of resources in non-critical data transmission. Fourth, data feature extraction is insufficient; traditional methods struggle to simultaneously capture both local details and global correlations, resulting in inadequate data value mining. Fifth, the data processing process is poorly allocated, failing to fully consider the differences in importance and data characteristics across stages, impacting the accuracy of comprehensive analysis of data across the entire industrial chain.

[0003] To address the aforementioned technical problems, this invention provides an intelligent data acquisition and processing method for the entire Siliao rice industry chain. Summary of the Invention

[0004] To address at least one of the aforementioned technical problems, this invention provides an intelligent data acquisition and processing method for the entire silk rice industry chain.

[0005] In a first aspect, the present invention provides an intelligent data collection and processing method for the entire industrial chain of high-quality rice, comprising: Step 1: Deploy sensing terminals in all links of the entire Siliao rice industry chain, build a collaborative architecture of terminal-edge node-cloud platform, and simultaneously collect multiple sets of historical sample data of the industry chain links and corresponding optimal data type records to form a historical sample library; Step 2: Collect multiple types of data from each link of the current industrial chain in real time and preprocess them to form a data sequence. Perform change analysis on the data sequence to obtain the sensitivity index and change characteristics. Input the relevant parameters based on the collaborative architecture and the change characteristics based on each data sequence into the machine learning model to determine the data transmission delay parameters. Step 3: Based on the delay parameters, the preprocessing time, and the sensitivity index, determine the timeliness factor of data interaction. At the same time, determine the security factor of data interaction based on the data sequence collection period of the deployment node of the corresponding link and the frequency of security events triggered during the collection period. Step 4: Based on the clustering algorithm, the historical sample library is divided into multiple sample clusters, and the optimal data type corresponding to each sample cluster is recorded. Based on the characteristics of each link in the entire industrial chain of the rice, and combined with the timeliness factor and safety factor, at least one similar sample cluster is matched from the multiple sample clusters to the corresponding link. The multi-type data collected by the current terminal is initially screened with reference to the optimal data type. Step 5: Use an AI algorithm model combining convolutional neural networks and Transformers to extract feature information of multiple types of data after initial screening, calculate the mean difference of data features at different stages, and divide the mean difference by the product of the standard deviation of the feature value and the preset adjustment coefficient to obtain the feature difference degree between the data at each stage. Step 6: Assign weights to each link according to their relative importance and characteristic differences in the whole industry chain analysis. Combine the smoothing parameter to process the weights of each link to obtain the spatial correlation feature sequence of the whole industry chain data. Compress the multi-type data after the initial screening and upload it to the cloud platform based on the corresponding edge nodes.

[0006] Preferably, the data sequence is subjected to variation analysis to obtain a sensitivity index and variation characteristics, including: Extract the neighboring data set and trend change characteristics of each data point in each data sequence, calculate the dynamic fluctuation index of each data point, and divide the corresponding data sequence into normal sequence and mutation sequence to obtain the sensitivity index of each data sequence. Simultaneously, adjacency analysis is performed on all trend change characteristics to obtain the change features.

[0007] Preferably, the relevant parameters based on the collaborative architecture include: the interaction data between the terminal and the edge node, and between the edge node and the cloud platform, transmission adaptation parameters, and transmission component parameters.

[0008] Preferably, the terminal includes at least one of a soil moisture sensor, a temperature and humidity sensor, a high-definition image acquisition camera, a processing parameter sensor, a storage environment sensor, and a transportation positioning sensor.

[0009] Preferably, matching at least one similar sample cluster from multiple sample clusters to the corresponding stage includes: The link features, timeliness factors, and safety factors are each assigned a weight that matches the importance of the link, and the overall matching degree between each link and each sample cluster is calculated. Based on the comprehensive matching degree, at least one similar sample cluster is matched from multiple sample clusters to the corresponding link of the current entire industry chain.

[0010] Preferably, an AI algorithm model combining convolutional neural networks and Transformers is used to extract feature information from multiple types of data after initial screening, including: A convolutional neural network was used to identify plant height, tiller texture, and particle uniformity features of finished product appearance in the growth images of glutinous rice, and output local visual feature vectors of image data. The Transformer model is used to construct sequences of data from each stage according to the time dimension and the industrial chain stage dimension. Based on the attention mechanism of the weight of the industrial chain stage of the fragrant rice, the temporal correlation and cross-stage dependency between data from different stages are learned, and a global correlation feature vector is output. The feature fusion layer concatenates the local visual feature vectors with the global associated feature vectors to obtain the feature information corresponding to the multiple types of data after initial screening.

[0011] Preferably, the spatial correlation feature sequence of the entire industry chain data is obtained by processing the weights of each link with a smoothing parameter, including: Based on the value contribution of each link to the quality control, production efficiency, and safety traceability of the entire Silky Rice industry chain, the basic importance is calculated using the Analytic Hierarchy Process (AHP). At the same time, the basic importance is dynamically adjusted by combining the reliability, completeness, and real-time nature of the data in each link, so as to obtain the relative importance of each link. By combining the variance of the characteristic differences between the corresponding link and the other links, and taking into account the relative importance, weights are assigned to the corresponding links. The fluctuation characteristics of data in each link of the entire Siliao rice industry chain are determined. The corresponding smoothing parameters are retrieved from the characteristic-parameter comparison table. The smoothing parameters of each link are integrated with the weights of the corresponding links. According to the temporal and spatial correlation logic of the links in the entire Siliao rice industry chain, a spatial correlation feature sequence is generated.

[0012] Preferably, the data of various types after initial screening is compressed, including: Based on the correlation strength and feature contribution of multiple types of data in the spatial correlation feature sequence of the whole industrial chain of high-quality rice after initial screening, the importance level of each data is determined, and the data is divided according to the importance level and the corresponding compression operation is performed.

[0013] In a second aspect, the present invention also provides an electronic device, comprising: a processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein when the processor executes the computer instructions, the electronic device performs a method as described in the first aspect above and any possible implementation thereof.

[0014] Thirdly, the present invention also provides a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor of an electronic device, cause the processor to perform a method as described in the first aspect above and any possible implementation thereof.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: By building a terminal-edge node-cloud collaborative architecture, comprehensive data collection and hierarchical processing of the entire industry chain are achieved. Combined with historical sample databases and AI algorithms, data is accurately screened, features are extracted, and parameters are optimized. This ensures the timeliness and security of data transmission, improves data processing efficiency and accuracy, effectively eliminates redundant data, and reduces storage and transmission costs. It provides reliable data support for quality control, production optimization, and safety traceability of the entire Silky Rice industry chain, and promotes the intelligent upgrading of the Silky Rice industry.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the background art, the accompanying drawings used in the embodiments of the present invention or the background art will be described below.

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.

[0019] Figure 1 This is a flowchart illustrating an intelligent data acquisition and processing method for the entire industrial chain of high-quality rice, provided in a certain embodiment of the present invention. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0022] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0023] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0024] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art will understand that the present invention can be practiced without certain specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art have not been described in detail in order to highlight the spirit of the invention.

[0025] This invention provides an intelligent data collection and processing method for the entire industrial chain of high-quality rice, such as... Figure 1 As shown, it includes: Step 1: Deploy sensing terminals in all links of the entire Siliao rice industry chain, build a collaborative architecture of terminal-edge node-cloud platform, and simultaneously collect multiple sets of historical sample data of the industry chain links and corresponding optimal data type records to form a historical sample library; In this embodiment, each link in the entire industrial chain of Silky Rice refers to the complete process of Silky Rice from the start of production to the end sales, covering core links such as planting (seedling cultivation, field growth), processing (hulling, milling, sorting), warehousing (warehousing, storage, and delivery), transportation (trunk transportation, short-distance delivery), and sales (terminal display, traceability query).

[0026] In this embodiment, the sensing terminal refers to hardware devices deployed in all links of the entire industry chain to collect various types of data such as environmental parameters, physical parameters, and image data. For example, in the planting stage, soil moisture sensors, temperature and humidity sensors, and high-definition image acquisition cameras are deployed; in the processing stage, processing parameter sensors such as mill speed sensors and sorting accuracy sensors are deployed; in the warehousing stage, warehousing environment sensors such as temperature and humidity sensors and gas concentration sensors are deployed; and in the transportation stage, transportation positioning sensors are deployed.

[0027] In this embodiment, a collaborative architecture of terminal-edge node-cloud platform is implemented. The edge nodes are deployed in core areas of each link (such as field operation areas, processing workshops, and storage centers). Communication with the terminal adopts the LoRa / WiFi protocol. The LoRa protocol is used in the field to ensure long-distance transmission, while the WiFi protocol is used in the workshop to improve the transmission rate. The edge nodes communicate with the cloud platform through 5G / fiber optic networks to ensure the stability and efficiency of data transmission.

[0028] In this embodiment, the historical sample industry chain link data refers to the actual data generated at each stage of the Silky Rice industry chain that has completed the entire process operation, including environmental data, process parameter data, product testing data, etc. For example, for a Silky Rice industry chain in a planting season of 2024, the soil moisture data in the planting stage is 25% moisture content at 10:00 am and 22% moisture content at 2:00 pm; the temperature and humidity data is 28°C and 75% daily average temperature; the milling speed data in the processing stage is 3000 r / min and the sorting accuracy data is 98%; the storage temperature data in the storage stage is 15°C and humidity data. The final product quality grade is Grade 1 rice. All relevant data are used as historical sample data.

[0029] In this embodiment, the optimal data type refers to the data type in the historical sample industry chain that plays a key supporting role in improving operational effectiveness (such as product quality and production efficiency).

[0030] In this embodiment, the historical sample library refers to a database that stores multiple sets of historical sample industry chain link data and corresponding optimal data type records. For example, it can collect 100 sets of historical data of the Siliao rice industry chain with different planting areas and different operating models. Each set of data contains multiple types of data of each link and corresponding optimal data type.

[0031] Step 2: Collect multiple types of data from each link of the current industrial chain in real time and preprocess them to form a data sequence. Perform change analysis on the data sequence to obtain the sensitivity index and change characteristics. Input the relevant parameters based on the collaborative architecture and the change characteristics based on each data sequence into the machine learning model to determine the data transmission delay parameters. In this embodiment, preprocessing refers to cleaning, format conversion, and noise reduction of the raw data collected in real time. For example, noise reduction is performed on the collected soil moisture data (using the moving average method to remove instantaneous fluctuation outliers); the plant image data is format converted (JPG format is converted to PNG format) and the size is normalized (uniformly adjusted to 800×600 pixels); and the processing parameter data is standardized in units (the rotation speed unit is unified to r / min).

[0032] In this embodiment, the data sequence refers to a continuous set of data of a certain type or a certain stage arranged in chronological order after preprocessing. For example, the soil moisture data sequence from time Q to time P in the planting stage: [22%, 23%, 25%, 24%, 23%, 22%, 21%].

[0033] In this embodiment, trend analysis (determining whether it is an upward, downward, or stable trend), fluctuation analysis (calculating the standard deviation to reflect the fluctuation amplitude), and mutation analysis (detecting whether there are any abnormal points of sudden increase or decrease) are performed on the soil moisture data sequence.

[0034] In this embodiment, the sensitivity index ranges from 0 to 1. It is calculated by analyzing the impact of mutations in the data sequence on the operational performance indicators. If a mutation in a data sequence causes the operational performance indicator to change by ≥10%, the sensitivity index is ≥0.7; if the change rate is 5%-10%, the sensitivity index is 0.4-0.7; and if the change rate is <5%, the sensitivity index is <0.4.

[0035] In this embodiment, adjacency analysis (calculating the difference and rate of change between adjacent data points) is performed on the trend change characteristics of the data sequence to extract parameters such as trend type, fluctuation frequency, number of mutation points, and mutation amplitude to form variation characteristics. For example, the variation characteristics of the soil moisture data sequence are: stable trend + low-frequency fluctuation + no mutation; the variation characteristics of the milling speed data sequence in the processing stage are: slight upward trend + medium-frequency fluctuation + 1 mutation (9:30 mutation amplitude 20r / min).

[0036] In this embodiment, a random forest regression model is constructed using the Scikit-learn library in Python. The training dataset consists of 100,000 sets of historical data, including the co-architecture parameters, data sequence variation characteristics, and corresponding actual transmission delay data over the past three years. After the model is trained, it is optimized using 5-fold cross-validation.

[0037] In this embodiment, the delay parameter ranges from 1 to 60 seconds and is dynamically adjusted according to the sensitivity of the data and the transmission requirements. The delay parameter refers to the maximum allowable transmission delay from data collection at the terminal to processing at the edge node and uploading to the cloud platform.

[0038] Step 3: Based on the delay parameters, the preprocessing time, and the sensitivity index, determine the timeliness factor of data interaction. At the same time, determine the security factor of data interaction based on the data sequence collection period of the deployment node of the corresponding link and the frequency of security events triggered during the collection period. In this embodiment, the timeliness factor = sensitivity index / ((delay parameter + preprocessing time) × 10).

[0039] In this embodiment, the data collection period is determined based on the data type and application requirements. For example, the data collection period for the warehouse environment sensor is 24 hours a day.

[0040] In this embodiment, a security event refers to the number of events that affect data security that occur during the data collection period, including events such as data leakage, data tampering, and abnormal device offline. For example, if the sensing terminal of a certain planting field does not experience any security events during the data collection period, the security event frequency is 0.

[0041] In this embodiment, the security factor = 1 - (frequency of security events / number of hours in the data collection period) × 0.5.

[0042] Step 4: Based on the clustering algorithm, the historical sample library is divided into multiple sample clusters, and the optimal data type corresponding to each sample cluster is recorded. Based on the characteristics of each link in the entire industrial chain of the rice, and combined with the timeliness factor and safety factor, at least one similar sample cluster is matched from the multiple sample clusters to the corresponding link. The multi-type data collected by the current terminal is initially screened with reference to the optimal data type. In this embodiment, the K-means clustering algorithm is used to divide 100 sets of historical sample data into 10 sample clusters, each containing 10 sets of historical samples with high similarity. For example, sample cluster 1 contains 10 sets of historical samples of first-grade rice quality from large-scale planting in rainy areas of southern China, with the optimal data types being soil moisture data, plant tillering stage image data, and milling speed data; sample cluster 2 contains 10 sets of historical samples of second-grade rice quality from family farms in arid areas of northern China, with the optimal data types being irrigation amount data, plant grain-filling stage image data, and storage humidity data.

[0043] In this embodiment, the link characteristics refer to the inherent attributes and operational characteristics of each link in the entire industrial chain of high-quality rice, including link type, production scale, environmental conditions, and technological level. For example, the link characteristics of the planting link are: rainy areas in the south, large-scale planting (1,000 mu), and conventional planting technology; the link characteristics of the processing link are: automated production line, daily processing capacity of 50 tons, and milling precision level 2.

[0044] In this embodiment, the number of matching similar sample clusters is 1-3, and the top 3 sample clusters with the highest overall matching degree and a matching degree ≥ 0.7 are selected as similar sample clusters.

[0045] In this embodiment, the initial screening operation is completed at the edge node. By comparing the current data type with the optimal data type set of similar sample clusters, the intersection data type is retained.

[0046] Step 5: Use an AI algorithm model combining convolutional neural networks and Transformers to extract feature information of multiple types of data after initial screening, calculate the mean difference of data features at different stages, and divide the mean difference by the product of the standard deviation of the feature value and the preset adjustment coefficient to obtain the feature difference degree between the data at each stage. In this embodiment, the AI ​​algorithm model combining convolutional neural networks and Transformers is built on the PyTorch framework. The CNN part uses a pre-trained ResNet50 model (pre-trained on the ImageNet dataset), and the Transformer model has 6 encoder layers and 8 attention heads. The model training efficiency and feature extraction accuracy are improved through transfer learning and fine-tuning.

[0047] In this embodiment, feature information refers to the quantized vector extracted from the data that reflects the essential attributes and relationships of the data, including local feature vectors and global feature vectors. For example, the local visual feature vector of plant image data in the planting stage is [0.23, 0.45, 0.18, ..., 0.31] (256 dimensions); the global correlation feature vector of data from each stage of the entire industry chain is [0.52, 0.37, 0.61, ..., 0.48] (512 dimensions); the fused feature information is a 256-dimensional + 512-dimensional = 768-dimensional vector.

[0048] In this embodiment, the mean difference refers to the absolute difference between the means of the feature vectors of data from different stages.

[0049] In this embodiment, the eigenvalue standard deviation refers to the standard deviation of the eigenvector elements of all stages of data.

[0050] In this embodiment, the preset adjustment coefficient is determined through experimental verification. Ten typical data samples are selected to test the matching degree between the feature difference degree and the actual data difference under different adjustment coefficients. The coefficient with the highest matching degree is selected as the preset adjustment coefficient. For example, combined with the characteristics of the Siliao rice industry chain data, the preset adjustment coefficient is set to 1.5. If the feature difference degree calculation result is found to be large in subsequent practical applications, it can be adjusted to a value between 1.2 and 1.8.

[0051] Step 6: Assign weights to each link according to their relative importance and characteristic differences in the whole industry chain analysis. Combine the smoothing parameter to process the weights of each link to obtain the spatial correlation feature sequence of the whole industry chain data. Compress the multi-type data after the initial screening and upload it to the cloud platform based on the corresponding edge nodes.

[0052] In this embodiment, relative importance refers to the degree of importance of each link in the whole industry chain analysis.

[0053] In this embodiment, a fluctuation characteristic-smoothing parameter comparison table is pre-constructed, such as stable = 0.1 to 0.2, low-frequency fluctuation = 0.3 to 0.5, and mid-to-high frequency fluctuation = 0.6 to 0.9. The fluctuation characteristics are determined by analyzing the standard deviation of the data in each stage, and then the corresponding smoothing parameters are retrieved from the comparison table.

[0054] In this embodiment, for example, according to the time sequence of planting-processing-warehousing-transportation-sales, after integrating weights and smoothing parameters, a spatial correlation feature sequence is formed: [planting stage (weight 0.32, smoothing parameter 0.8), processing stage (weight 0.23, smoothing parameter 0.3), warehousing stage (weight 0.18, smoothing parameter 0.1), transportation stage (weight 0.13, smoothing parameter 0.4), sales stage (weight 0.04, smoothing parameter 0.2)].

[0055] In this embodiment, compression processing refers to the importance level of the data, and different compression algorithms are used to compress the data after the initial screening.

[0056] Preferably, the relevant parameters based on the collaborative architecture include: the interaction data between the terminal and the edge node, and between the edge node and the cloud platform, transmission adaptation parameters, and transmission component parameters.

[0057] Preferably, the terminal includes at least one of a soil moisture sensor, a temperature and humidity sensor, a high-definition image acquisition camera, a processing parameter sensor, a storage environment sensor, and a transportation positioning sensor.

[0058] The beneficial effects of the above technical solution are as follows: by building a terminal-edge node-cloud collaborative architecture, comprehensive data collection and hierarchical processing of the entire industry chain can be achieved. By combining historical sample databases and AI algorithms, accurate data screening, feature extraction and parameter optimization can be completed, ensuring the timeliness and security of data transmission, improving data processing efficiency and accuracy, effectively eliminating redundant data, reducing storage and transmission costs, providing reliable data support for quality control, production optimization and safety traceability of the entire Siliao rice industry chain, and promoting the intelligent upgrading of the Siliao rice industry.

[0059] This invention provides an intelligent data acquisition and processing method for the entire jasmine rice industry chain, which performs variation analysis on the data sequence to obtain sensitivity indices and variation characteristics, including: Extract the neighboring data set and trend change characteristics of each data point in each data sequence, calculate the dynamic fluctuation index of each data point, and divide the corresponding data sequence into normal sequence and mutation sequence to obtain the sensitivity index of each data sequence. Simultaneously, adjacency analysis is performed on all trend change characteristics to obtain the change features.

[0060] In this embodiment, the size of the neighboring data set is determined according to the length of the data sequence (2 data points before and after for length ≤ 10, 5 data points before and after for length 11-50, and 10 data points before and after for length > 50), and the neighboring data set of each data point is extracted; the data sequence is divided into multiple segments by 5 consecutive data points, and a linear regression algorithm (Python SciPy library) is used to fit each segment to obtain the trend direction (rising / falling / stable), trend slope and duration of each segment, and the trend information of all segments is integrated to obtain the trend change characteristics of the data sequence.

[0061] In this embodiment, for each data point, the mean and standard deviation of its neighboring dataset are calculated using the formula: Dynamic fluctuation index = (target data point - mean of neighboring dataset) / standard deviation of neighboring dataset. A normal fluctuation range is set at ±3. Data points with a dynamic fluctuation index within ±3 are classified as normal sequences, while those outside this range are classified as mutation sequences. The proportion of mutation sequences to the total number of data points is calculated; this proportion is the sensitivity index. For example, with 100 total data points and 25 mutation sequences, the sensitivity index is 0.25; with 40 mutation sequences, the sensitivity index is 0.4.

[0062] In this embodiment, all trend change features are arranged in chronological order, and the absolute value of the slope difference between two adjacent trend features (similarity) is calculated. If the similarity is ≥0.8, it is determined to be a combination of related features. All parameters such as trend direction, trend slope, fluctuation frequency (number of trend changes per unit time), number of mutation sequences, mutation amplitude (maximum difference between mutation data point and neighboring mean), and combination of related features are extracted and formed into a change feature vector in the format of [trend direction, trend slope, fluctuation frequency, number of mutations, mutation amplitude, combination of related features].

[0063] The beneficial effects of the above technical solution are: through the implementation of refined change analysis, the sensitivity index and change characteristics of the data sequence are accurately extracted, providing an accurate basis for the determination of subsequent delay parameters and timely factors, effectively distinguishing between normal fluctuations and abnormal mutations in the data, clarifying the degree of impact of the data on the operational effectiveness of the industrial chain, improving the pertinence and accuracy of data processing, and laying the foundation for the optimized processing of data across the entire industrial chain.

[0064] This invention provides an intelligent data collection and processing method for the entire hyacinth rice industry chain, which matches at least one similar sample cluster from multiple sample clusters to the corresponding link, including: The link features, timeliness factors, and safety factors are each assigned a weight that matches the importance of the link, and the overall matching degree between each link and each sample cluster is calculated. Based on the comprehensive matching degree, at least one similar sample cluster is matched from multiple sample clusters to the corresponding link of the current entire industry chain.

[0065] In this embodiment, 8 experts (5 industry technology experts and 3 data processing experts) were invited to score the importance of process characteristics, timeliness factors, and safety factors (1-10 points). It is assumed that the average score of process characteristics is 9 points, timeliness factors are 7 points, and safety factors are 6 points, with a total average score of 22 points. The normalized weights are: process characteristics 9 / 22≈0.41, timeliness factors 7 / 22≈0.32, and safety factors 6 / 22≈0.27, with a total weight of 1.0 (which can be fine-tuned to 0.5, 0.3, or 0.2 according to the actual scenario).

[0066] In this embodiment, the process feature vector, timeliness factor, and safety factor of the current process are processed by Min-Max normalization (the values ​​are mapped to 0-1); the Euclidean distance between the current parameter and the central parameter of each sample cluster is calculated, and the similarity of each parameter is calculated according to the formula "similarity = 1 / (1 + Euclidean distance)" (e.g., process feature similarity 0.9, timeliness factor similarity 0.8, safety factor similarity 0.95).

[0067] In this embodiment, the overall matching degree is calculated according to the formula: Overall matching degree = Similarity of process features × Weight of process features + Similarity of timeliness factor × Weight of timeliness factor + Similarity of safety factor × Weight of safety factor.

[0068] In this embodiment, the overall matching degree of all sample clusters is sorted, and the top 3 sample clusters with an overall matching degree ≥ 0.7 are selected as the similar sample clusters in the current stage; if there are fewer than 3 sample clusters with an overall matching degree ≥ 0.7, then all sample clusters with an overall matching degree ≥ 0.6 (at least 1) are selected to ensure that there are enough reference samples.

[0069] The beneficial effects of the above technical solution are: through scientific weight allocation and comprehensive matching degree calculation, accurate matching of similar sample clusters can be achieved, ensuring that the matched sample clusters are highly relevant to the current industrial chain links, and the corresponding optimal data type can provide a reliable reference for the current data screening, improve the targeting and effectiveness of data screening, and reduce invalid and redundant data.

[0070] This invention provides an intelligent data acquisition and processing method for the entire jasmine rice industry chain. It employs an AI algorithm model combining convolutional neural networks and Transformer to extract feature information from multiple types of data after initial screening, including: A convolutional neural network was used to identify plant height, tiller texture, and particle uniformity features of finished product appearance in the growth images of glutinous rice, and output local visual feature vectors of image data. The Transformer model is used to construct sequences of data from each stage according to the time dimension and the industrial chain stage dimension. Based on the attention mechanism of the weight of the industrial chain stage of the fragrant rice, the temporal correlation and cross-stage dependency between data from different stages are learned, and a global correlation feature vector is output. The feature fusion layer concatenates the local visual feature vectors with the global associated feature vectors to obtain the feature information corresponding to the multiple types of data after initial screening.

[0071] In this embodiment, the images of the rice after initial screening (such as images of plants during the tillering stage in the planting process and images of the finished rice appearance in the processing stage) are standardized and uniformly adjusted to 224×224 pixels (to meet the input requirements of ResNet50). The pixel values ​​are also normalized (converting RGB values ​​of 0-255 into floating-point numbers of 0-1) to eliminate the influence of lighting and size differences on feature extraction.

[0072] A pre-trained ResNet50 model based on the PyTorch framework (pre-trained on the ImageNet dataset) was used. The last fully connected layer of the original model was removed, and a new fully connected layer with an output dimension of 256 (using ReLU activation function) was added to output local visual feature vectors.

[0073] The preprocessed image is input into the ResNet50 model. Through the model's 50 convolutional layers (such as Conv1, Conv2_x to Conv5_x) and pooling layers, local detail features of the image are captured step by step (such as the plant height can be extracted by comparing the plant pixel height with the background in the image, the tiller texture can be captured by the edge detection convolution kernel, and the uniformity of the finished rice grains can be extracted by the standard deviation feature of the grain contour). Finally, a 256-dimensional local visual feature vector is output through the newly added fully connected layer. Each element in the vector corresponds to the quantized value of a local detail feature (e.g., element 0.45 corresponds to the density feature of tiller texture, and element 0.23 corresponds to the particle uniformity feature).

[0074] Transformer model extracts global correlation features: Multi-dimensional sequence construction: Collect non-image data from each link of the entire industrial chain after initial screening (such as soil moisture in the planting link, milling speed in the processing link, temperature and humidity in the storage link, and location and temperature in the transportation link), and construct a two-dimensional data sequence according to the time dimension and the industrial chain link dimension.

[0075] Time dimension: Data from each stage is arranged in ascending order by collection timestamp, with a uniform time granularity of 1 hour to ensure timeline alignment.

[0076] Industry chain segment dimension: Data for each segment is encoded in a fixed order of planting-processing-warehousing-transportation-sales (planting=1, processing=2, warehousing=3, transportation=4, sales=5) to form a segment identifier vector, which is combined with time dimension data to form a triple sequence of (time step, segment identifier, data value), such as (8:00, 1, 25%), (9:00, 2, 2800r / min).

[0077] Transformer model configuration: Construct a 6-layer Transformer model (based on the nn.TransformerEncoder module of PyTorch), set the number of attention heads to 8, the hidden layer dimension to 512, the FeedForward network dimension to 2048, and the activation function to GELU; use the relative importance of each step calculated in claim 7 (e.g., planting 0.35, processing 0.25) as the initial attention weights, and input them into the attention layer of the model so that the model prioritizes the data of important steps.

[0078] Global association learning: The constructed two-dimensional data sequence is input into the Transformer encoder, and the association weights between data at different time steps and different stages are calculated through a self-attention mechanism. Temporal correlation learning: Calculate the attention weight of adjacent time steps within the same stage (e.g., the weight of soil moisture at 8:00 and 9:00 in the planting stage). The higher the weight, the stronger the temporal correlation (e.g., during a continuous drought period, the weight of soil moisture at adjacent time steps is ≥0.7, reflecting strong temporal dependence).

[0079] Cross-stage dependency learning: Calculate the attention weights of data at the same or similar time steps between different stages (such as the weight of soil moisture at 8:00 in the planting stage and the weight of milling speed at 9:00 in the processing stage). When the weight is ≥0.6, it is judged as strong cross-stage dependency (such as when the soil moisture is low, the weight increases, reflecting the impact of planting on processing).

[0080] Global feature output: Through layer-by-layer feature aggregation by a 6-layer encoder, the final output is a global association feature vector with a dimension of 512. Each element in the vector corresponds to the quantized value of a global association feature (e.g., element 0.52 corresponds to the cross-stage dependency feature of planting-processing, and element 0.37 corresponds to the temporal association feature of warehousing-transportation).

[0081] The feature fusion layer generates comprehensive feature information: Feature vector alignment: Dimension verification is performed on the local visual feature vector (256-dimensional) and the global associated feature vector (512-dimensional) to ensure that both are floating-point vectors and have no missing values ​​(if there are missing values, they are filled with the mean of the feature dimension; for example, if a certain element of the local feature vector is missing, it is filled with the mean of all elements of the vector, which is 0.32).

[0082] Vector concatenation and fusion: A direct concatenation method is adopted, using local visual feature vectors as the first 256 dimensions and global correlation feature vectors as the last 512 dimensions, connecting the first and last to form a 768-dimensional comprehensive feature vector, which is the feature information corresponding to the multiple types of data after initial screening; for example, the element 0.45 (tillering texture) in the first 256 dimensions and the element 0.52 (planting-processing dependence) in the last 512 dimensions together constitute a comprehensive feature, which contains both local image details and global correlation information.

[0083] Feature normalization: The fused 768-dimensional feature vector is subjected to Min-Max normalization to map all element values ​​to the 0-1 range (formula: normalized value = (original value - minimum value) / (maximum value - minimum value)). This avoids the impact of the difference in the dimensions of features of different dimensions on subsequent calculations (e.g., the maximum value of local feature vector elements is 0.9 and the minimum value is 0.1, while the maximum value of global feature vector elements is 0.8 and the minimum value is 0.2, and the range is unified after normalization).

[0084] The beneficial effects of the above technical solution are as follows: By using a hybrid AI algorithm model of convolutional neural network + Transformer, comprehensive extraction of features of multiple types of data after initial screening is achieved. It not only uses convolutional neural network to accurately capture local visual details of image data, but also uses Transformer model to effectively learn the temporal correlation and cross-link dependency of data in the entire industry chain. Then, it integrates the features through feature fusion layer to form comprehensive feature information, avoiding the limitation of a single model that can only extract local or global features. It significantly improves the completeness and accuracy of feature information, provides high-quality basic data for the calculation of the difference in data features in each subsequent link, and thus ensures the scientificity and reliability of data processing in the entire industry chain.

[0085] This invention provides an intelligent data acquisition and processing method for the entire jasmine rice industry chain. It combines smoothing parameters to process the weights of each stage to obtain a spatial correlation feature sequence of the entire industry chain data, including: Based on the value contribution of each link to the quality control, production efficiency, and safety traceability of the entire Silky Rice industry chain, the basic importance is calculated using the Analytic Hierarchy Process (AHP). At the same time, the basic importance is dynamically adjusted by combining the reliability, completeness, and real-time nature of the data in each link, so as to obtain the relative importance of each link. By combining the variance of the characteristic differences between the corresponding link and the other links, and taking into account the relative importance, weights are assigned to the corresponding links. The fluctuation characteristics of data in each link of the entire Siliao rice industry chain are determined. The corresponding smoothing parameters are retrieved from the characteristic-parameter comparison table. The smoothing parameters of each link are integrated with the weights of the corresponding links. According to the temporal and spatial correlation logic of the links in the entire Siliao rice industry chain, a spatial correlation feature sequence is generated.

[0086] In this embodiment, the relative importance of each step is calculated: Basic Importance Calculation (Based on Analytic Hierarchy Process): The hierarchical structure is as follows: the target layer is the importance assessment of the entire industrial chain of high-quality rice; the criteria layer is quality control (C1), production efficiency (C2), and safety traceability (C3); and the solution layer is planting (P1), processing (P2), warehousing (P3), transportation (P4), and sales (P5).

[0087] Construct a judgment matrix and verify consistency: Judgment matrix of the criterion layer to the target layer: According to expert scores, the importance ratio of C1 (quality control) to C2 (production efficiency) is 3:1, the ratio of C1 to C3 (safety traceability) is 5:1, and the ratio of C2 to C3 is 3:1. The judgment matrix is ​​shown in Table 1.

[0088] The weights of the criterion layer were calculated using the eigenvalue method: C1 weight was 0.637, C2 weight was 0.258, and C3 weight was 0.105. The consistency test CR=0.003<0.1, indicating that the matrix is ​​reasonable.

[0089] The judgment matrix of the scheme layer to the criterion layer: Taking C1 (quality control) as an example, the ratio of P1 (planting) to P2 (processing) is 3:1, the ratio of P1 to P3 is 5:1, the ratio of P1 to P4 is 7:1, the ratio of P1 to P5 is 9:1, the ratio of P2 to P3 is 3:1, and so on, to construct the matrix and calculate the weight of the scheme layer under each criterion (e.g., the weight of P1 under C1 is 0.539, under C2 is 0.157, and under C3 is 0.089).

[0090] Calculate basic importance: Basic importance = Σ (criteria layer weight × scheme layer weight under that criterion). Dynamic adjustment (combined with data quality indicators): Data quality indicators are obtained by monitoring the edge nodes and obtaining the average data reliability (P1=0.9, P2=0.85, P3=0.8, P4=0.75, P5=0.7), average data integrity (P1=0.92, P2=0.88, P3=0.85, P4=0.8, P5=0.78), and average data real-time performance (P1=0.8, P2=0.75, P3=0.7, P4=0.65, P5=0.6) for each stage over the past month.

[0091] Revised formula: Relative importance = Basic importance × (Reliability + Integrity + Real-time performance) / 3, where all three indicators have been normalized to 0-1.

[0092] Allocate weights for each stage: Calculate the variance of the feature differences; The weight of a certain link is calculated as (relative importance of the link × (1 + variance)) / Σ (relative importance of all links × (1 + variance)), where 1 + variance is used to adjust the weights of links with uneven differences.

[0093] In this embodiment, a spatial correlation feature sequence of the entire industry chain data is generated: Determine the data fluctuation characteristics and smoothing parameters: Fluctuation characteristics assessment: The coefficient of variation of data for each stage over the past month was statistically analyzed. The coefficient of variation for soil moisture in P1 (planting) was 0.13 (medium to high frequency fluctuations), the coefficient of variation for milling speed in P2 (processing) was 0.08 (low frequency fluctuations), the coefficient of variation for temperature and humidity in P3 (storage) was 0.03 (stable), the coefficient of variation for temperature in P4 (transportation) was 0.09 (low frequency fluctuations), and the coefficient of variation for traceability query volume in P5 (sales) was 0.06 (low frequency fluctuations).

[0094] Retrieve smoothing parameters: Retrieve the corresponding parameters from the characteristic-parameter lookup table. P1=0.8 (medium-high frequency fluctuations), P2=0.4 (low frequency fluctuations), P3=0.15 (stable), P4=0.45 (low frequency fluctuations), P5=0.35 (low frequency fluctuations). Integrate weights and smoothing parameters: For each stage, combine the weights and smoothing parameters into a triplet of (stage identifier, weight, smoothing parameter). For example, P1=(planting, 0.35, 0.8), P2=(processing, 0.25, 0.4), P3=(warehousing, 0.18, 0.15), P4=(transportation, 0.13, 0.45), P5=(sales, 0.09, 0.35). The data is ordered according to the temporal and spatial relationships of each stage: Triples are arranged according to the temporal sequence of planting → processing → warehousing → transportation → sales, while ensuring that directly related stages are adjacent (e.g., planting and processing are adjacent, processing and warehousing are adjacent, warehousing and transportation are adjacent, transportation and sales are adjacent), forming a spatial relationship feature sequence: [(planting, 0.35, 0.8), (processing, 0.25, 0.4), (warehousing, 0.18, 0.15), (transportation, 0.13, 0.45), (sales, 0.09, 0.35)]. This sequence reflects the chronological order of the industry process and reflects the importance and data stability of each stage through weights and smoothing parameters, providing a basis for subsequent data compression.

[0095] The beneficial effects of the above technical solution are as follows: By combining the analytic hierarchy process (AHP) with data quality indicators, the relative importance of each link can be accurately calculated, avoiding the one-sidedness of single-factor evaluation. Furthermore, by combining the variance of feature difference degree to allocate weights, it is ensured that the weights not only reflect the importance of the links but also adapt to the balance of feature differences between each link and other links. Finally, by retrieving smoothing parameters through data fluctuation characteristics, a spatial correlation feature sequence that conforms to the temporal and spatial correlation logic of the links is generated. This provides a structured and accurate basis for data compression processing and comprehensive analysis of the entire industry chain data, significantly improving the logic and reliability of data processing and ensuring the accuracy of subsequent in-depth analysis on the cloud platform.

[0096] This invention provides an intelligent data collection and processing method for the entire jasmine rice industry chain, which compresses various types of data after initial screening, including: Based on the correlation strength and feature contribution of multiple types of data in the spatial correlation feature sequence of the whole industrial chain of high-quality rice after initial screening, the importance level of each data is determined, and the data is divided according to the importance level and the corresponding compression operation is performed.

[0097] In this embodiment, the correlation strength = link weight × contribution ratio; In this embodiment, the feature contribution is calculated as (original norm - new norm) / original norm.

[0098] For example, by substituting the correlation strength and feature contribution into the hierarchical criteria, soil moisture change data (correlation strength 0.85, feature contribution 0.8) meets the criteria of correlation strength ≥ 0.7 and feature contribution ≥ 0.7, and is classified as the core layer; conventional mill rotation speed data (correlation strength 0.6, feature contribution 0.6) meets the criteria of 0.5 ≤ correlation strength < 0.7 and 0.5 ≤ feature contribution < 0.7, and is classified as the key layer; and display image data (correlation strength 0.3, feature contribution 0.2) meets the criteria of correlation strength < 0.5, and is classified as the ordinary layer.

[0099] In this embodiment, differentiated compression operations are performed: Core layer data compression: Algorithm selection: The LZ77 lossless compression algorithm (implemented based on the zlib library) is adopted. This algorithm finds repeated data sequences by sliding window and replaces the repeated sequences with offset + length, without information loss, and is suitable for core layer critical data.

[0100] Compression parameter settings: Sliding window size is set to 32KB, matching length threshold is set to 3 bytes (replace when the length of the repeating sequence is ≥3 bytes), and compression ratio is controlled between 1.5:1 and 2:1.

[0101] For example, soil moisture change data during the planting process (original data is a 10MB numerical sequence file), after being compressed by LZ77, the file size becomes 6.7MB, with a compression ratio of 1.5:1. After decompression, the data is completely consistent with the original data, with no information loss.

[0102] Key layer data compression: Numerical data (such as mill speed, temperature and humidity): Huffman coding lossy compression is adopted (implemented based on Python's heapq library). By statistically analyzing the frequency of occurrence of data values, short codes are assigned to high-frequency values ​​and long codes are assigned to low-frequency values. A small amount of coding deviation of low-frequency data is allowed (error rate ≤2%). The compression ratio is set to 5:1-8:1.

[0103] For example, the milling speed data for 1 hour in the processing stage (original 5MB, containing 3600 data points) becomes 1MB after Huffman coding compression, with a compression ratio of 5:1. The error rate between the decompressed data and the original data is 1.5%, which is within the allowable range (≤2%).

[0104] Image-type data (such as images of regular plants): JPEG lossy compression algorithm (implemented based on OpenCV library) is used, with quality factor set to 70 (0-100, 70 is medium quality with less distortion), and compression ratio set to 8:1-10:1.

[0105] For example, the standard appearance image of finished rice during the processing stage (original 20MB, 2592×1944 pixels) becomes 2.2MB after JPEG compression (quality factor 70), with a compression ratio of 9:1. The decompressed image can still clearly identify the particle outline, meeting the needs of subsequent analysis.

[0106] Normal layer data compression: Numerical data (such as routine transaction data in the sales process): arithmetic encoding lossy compression (implemented based on the PyArmor library) is used, the compression precision is set to 0.01 (data is retained to two decimal places), the allowable error rate is ≤5%, and the compression ratio is set to 15:1-20:1.

[0107] For example, a day's routine transaction data in the sales process (originally 8MB, containing 10,000 records) becomes 0.5MB after arithmetic encoding compression, with a compression ratio of 16:1. The data error rate after decompression is 3%, which meets the needs of statistical analysis (high precision is not required).

[0108] Image-based data (such as images displayed in the sales process): The JPEG2000 lossy compression algorithm (based on the OpenJPEG library) is used, with a compression ratio of 20:1-30:1. A certain degree of detail loss (such as blurred background texture) is allowed, but it does not affect the recognition of the subject.

[0109] For example, the supermarket display image (original 30MB, 3840×2160 pixels) is compressed to 1.2MB using JPEG2000 (compression ratio 25:1). After decompression, the location and quantity of the rice packaging can still be identified, meeting basic traceability requirements.

[0110] Compressed Data Verification and Storage: Compression Effect Verification: For the compressed data, calculate the compression ratio (original size / compressed size) and error rate (lossless compression only) to ensure that the compression ratio meets the requirements of each level, and the error rate is ≤ the corresponding level threshold (core layer 0%, key layer ≤ 2%, ordinary layer ≤ 5%). If not, readjust the compression parameters (e.g., increase the JPEG quality factor, decrease the arithmetic coding precision) and recompress. Data Identification and Storage: Add level identifiers to the compressed data (e.g., core layer - planting - soil moisture - 202405010800), and store it in the temporary storage area of ​​the edge node (e.g., / edge_storage / core layer / planting / 20240501 / ) according to the directory structure of level-stage-time, waiting to be uploaded to the cloud platform.

[0111] The beneficial effects of the above technical solution are as follows: by classifying data importance levels based on correlation strength and feature contribution, differentiated compression operations are achieved: lossless compression is used in the core layer to ensure no loss of key data, while lossy compression of different degrees is used in the key and ordinary layers to balance data quality and storage / transmission costs, effectively reducing the storage footprint and transmission bandwidth requirements of the entire industry chain; at the same time, compression effect verification ensures that data quality meets the needs of subsequent analysis, avoids information failure caused by over-compression, provides key support for edge nodes to efficiently upload data to the cloud platform, and improves the economy and efficiency of data processing across the entire industry chain.

[0112] The present invention also provides an electronic device, comprising: a processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein when the processor executes the computer instructions, the electronic device performs a method as described in any of the above possible implementations.

[0113] The present invention also provides a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor of an electronic device, cause the processor to perform a method as described in any of the above possible implementations.

[0114] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0115] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will also readily understand that the various embodiments of the present invention have different focuses, and for the sake of convenience and brevity, the same or similar parts may not be repeated in different embodiments. Therefore, parts not described or not described in detail in one embodiment can be referred to in other embodiments.

[0116] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0117] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0118] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0119] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0120] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for intelligent data collection and processing across the entire silk rice industry chain, characterized in that, include: Step 1: Deploy sensing terminals in all links of the entire Siliao rice industry chain, build a collaborative architecture of terminal-edge node-cloud platform, and simultaneously collect multiple sets of historical sample data of the industry chain links and corresponding optimal data type records to form a historical sample library; Step 2: Collect multiple types of data from each link of the current industrial chain in real time and preprocess them to form a data sequence. Perform change analysis on the data sequence to obtain the sensitivity index and change characteristics. Input the relevant parameters based on the collaborative architecture and the change characteristics based on each data sequence into the machine learning model to determine the data transmission delay parameters. Step 3: Based on the delay parameters, the preprocessing time, and the sensitivity index, determine the timeliness factor of data interaction. At the same time, determine the security factor of data interaction based on the data sequence collection period of the deployment node of the corresponding link and the frequency of security events triggered during the collection period. Step 4: Based on the clustering algorithm, the historical sample library is divided into multiple sample clusters, and the optimal data type corresponding to each sample cluster is recorded. Based on the characteristics of each link in the entire industrial chain of the rice, and combined with the timeliness factor and safety factor, at least one similar sample cluster is matched from the multiple sample clusters to the corresponding link. The multi-type data collected by the current terminal is initially screened with reference to the optimal data type. Step 5: Use an AI algorithm model combining convolutional neural networks and Transformers to extract feature information of multiple types of data after initial screening, calculate the mean difference of data features at different stages, and divide the mean difference by the product of the standard deviation of the feature value and the preset adjustment coefficient to obtain the feature difference degree between the data at each stage. Step 6: Assign weights to each link according to their relative importance and characteristic differences in the whole industry chain analysis. Combine the smoothing parameter to process the weights of each link to obtain the spatial correlation feature sequence of the whole industry chain data. Compress the multi-type data after the initial screening and upload it to the cloud platform based on the corresponding edge nodes.

2. The intelligent data acquisition and processing method for the entire industrial chain of high-quality rice as described in claim 1, characterized in that, The data sequence is subjected to variation analysis to obtain sensitivity indices and variation characteristics, including: Extract the neighboring data set and trend change characteristics of each data point in each data sequence, calculate the dynamic fluctuation index of each data point, and divide the corresponding data sequence into normal sequence and mutation sequence to obtain the sensitivity index of each data sequence. Simultaneously, adjacency analysis is performed on all trend change characteristics to obtain the change features.

3. The intelligent data acquisition and processing method for the entire industrial chain of high-quality rice as described in claim 1, characterized in that, The relevant parameters based on the collaborative architecture include: the interaction data between the terminal and the edge node, as well as between the edge node and the cloud platform, transmission adaptation parameters, and transmission component parameters.

4. The intelligent data acquisition and processing method for the entire silk rice industry chain according to claim 1, characterized in that, The terminal includes at least one of the following: soil moisture sensor, temperature and humidity sensor, high-definition image acquisition camera, processing parameter sensor, storage environment sensor, and transportation positioning sensor.

5. The intelligent data acquisition and processing method for the entire industrial chain of high-quality rice as described in claim 1, characterized in that, Matching at least one similar sample cluster from multiple sample clusters to the corresponding stage includes: The link features, timeliness factors, and safety factors are each assigned a weight that matches the importance of the link, and the overall matching degree between each link and each sample cluster is calculated. Based on the comprehensive matching degree, at least one similar sample cluster is matched from multiple sample clusters to the corresponding link of the current entire industry chain.

6. The intelligent data acquisition and processing method for the entire industrial chain of high-quality rice as described in claim 1, characterized in that, An AI algorithm model combining convolutional neural networks and Transformers was used to extract feature information from multiple types of data after initial screening, including: A convolutional neural network was used to identify plant height, tiller texture, and particle uniformity features of finished product appearance in the growth images of glutinous rice, and output local visual feature vectors of image data. The Transformer model is used to construct sequences of data from each stage according to the time dimension and the industrial chain stage dimension. Based on the attention mechanism of the weight of the industrial chain stage of the Siliao rice, the temporal correlation and cross-stage dependency between data from different stages are learned, and a global correlation feature vector is output. The feature fusion layer concatenates the local visual feature vectors with the global associated feature vectors to obtain the feature information corresponding to the multiple types of data after initial screening.

7. The intelligent data acquisition and processing method for the entire industrial chain of high-quality rice as described in claim 1, characterized in that, By combining smoothing parameters to process the weights of each stage, a spatial correlation feature sequence of the entire industry chain data is obtained, including: Based on the value contribution of each link to the quality control, production efficiency, and safety traceability of the entire Silky Rice industry chain, the basic importance is calculated using the Analytic Hierarchy Process (AHP). At the same time, the basic importance is dynamically adjusted by combining the reliability, completeness, and real-time nature of the data in each link, so as to obtain the relative importance of each link. By combining the variance of the characteristic differences between the corresponding link and the other links, and taking into account the relative importance, weights are assigned to the corresponding links. The fluctuation characteristics of data in each link of the entire Siliao rice industry chain are determined. The corresponding smoothing parameters are retrieved from the characteristic-parameter comparison table. The smoothing parameters of each link are integrated with the weights of the corresponding links. According to the temporal and spatial correlation logic of the links in the entire Siliao rice industry chain, a spatial correlation feature sequence is generated.

8. The intelligent data acquisition and processing method for the entire industrial chain of high-quality rice as described in claim 7, characterized in that, Compression processing is performed on the various types of data after the initial screening, including: Based on the correlation strength and feature contribution of multiple types of data in the spatial correlation feature sequence of the whole industrial chain of high-quality rice after initial screening, the importance level of each data is determined, and the data is divided according to the importance level and the corresponding compression operation is performed.

9. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein when the processor executes the computer instructions, the electronic device performs the intelligent data acquisition and processing method for the entire industrial chain of high-quality rice as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor of an electronic device, the program instructions cause the processor to perform the intelligent data acquisition and processing method for the entire industrial chain of jasmine rice as described in any one of claims 1 to 8.