Agricultural products big data service system based on trusted data space

Through the agricultural product big data service system based on trusted data space, the integration and analysis of multi-dimensional data in the entire chain of agricultural products is realized, accurate decision-making characteristics and risk warning indicators are generated, data dispersion and quality problems are solved, users are met, and the digital development of the industry is promoted.

CN120338854BActive Publication Date: 2025-08-22CHENGDU YUANBEN INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510800880.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-08-22
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The agricultural product industry has scattered data and inconsistent formats, lack of effective integration mechanisms, uneven data quality, traditional analysis methods are difficult to integrate multi-dimensional data, and service interfaces lack flexibility, which cannot meet the personalized needs of different user groups.

Method used

The agricultural product big data service system based on trusted data space, through multi-source data integration, data cleaning and orchestration, data association map construction, multi-dimensional analysis and configurable service interface, the integration, cleaning and analysis of multi-dimensional data in the entire chain can be realized, accurate decision-making characteristics and risk warning indicators are generated, and flexible and configurable data services are provided.

Benefits of technology

Break the data silos, improve data availability and accuracy, provide accurate market analysis and risk warning, meet the diversified needs of different user groups, and promote the digital development of the agricultural product industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338854B_ABST
    Figure CN120338854B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of agricultural product big data processing and application technology, and discloses an agricultural product big data service system based on a trusted data space. The system includes a multi-source data integration module, which collects multi-dimensional data of the entire chain of agricultural products; a data cleaning and arrangement module, which cleans the data and dynamically arranges it; a data association map construction module, which generates a cross-domain data association map; a multi-dimensional analysis engine module, which extracts data evolution patterns, generates classification decision features and risk warning indicators; and a service interface generation module, which constructs a configurable service interface set. Multi-dimensional data covers production bases, processes, product attributes, market conditions and cold chain logistics data. The system performs normalization, smoothing correction and other processing on the data, and realizes functions such as data association analysis and feature fusion through a variety of algorithms, providing adaptive data service solutions for different scenarios, helping the digital development of the agricultural product industry, and achieving accurate decision-making and risk warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of agricultural product big data processing and application technology, and specifically to an agricultural product big data service system based on a trusted data space. Background Art

[0002] In the agricultural products industry, data volumes have exploded with the development of information technology. However, this data presents numerous challenges, severely hindering its digital development and refined management. Data sources are diverse but fragmented. The agricultural products industry chain encompasses multiple links, including production, processing, distribution, and sales, each involving distinct data types. In the production process, production base data includes base certification information, scale, and output indicators, while production process data includes environmental monitoring parameters and operational records. In the distribution process, cold chain logistics data records information such as temperature and humidity during transportation. At the sales end, market data reflects price fluctuations, supply and demand, and other factors. These data are stored separately in different systems, with inconsistent formats and a lack of effective integration mechanisms. This makes data sharing and integrated utilization difficult, creating numerous "data silos." For example, agricultural product producers struggle to obtain market data for accurate production planning, while logistics companies struggle to allocate transportation resources based on actual production base output.

[0003] Data quality varies widely. Due to differences in data collection equipment and human operations, the data contains a significant amount of noise, errors, and missing values. For example, in environmental monitoring data from production processes, sensor failures can lead to outliers, and manual recording of operational records can result in clerical errors. If these erroneous data are not addressed, they can seriously affect the accuracy of subsequent data analysis. Adjusting agricultural production measures based on erroneous environmental monitoring data can affect crop growth and cause economic losses.

[0004] Traditional data analysis methods struggle to meet these demands. The agricultural products industry faces complex and volatile market environments and production conditions, requiring in-depth analysis of the patterns underlying data for accurate market forecasting, risk warnings, and decision support. However, existing analysis methods, mostly based on single data types or simple statistical models, fail to fully integrate the characteristics of multidimensional data and struggle to uncover deep temporal and spatial connections and semantic dependencies between data. When analyzing market data, these methods only consider price fluctuations while ignoring correlations with production processes, product attributes, and logistics factors, making it impossible to accurately predict price trends and market supply and demand fluctuations.

[0005] Service interfaces lack flexibility and configurability. Different user groups, such as agricultural product producers, distributors, and consumers, have varying needs for data services. However, current agricultural product data service system interfaces are often fixed and cannot be customized to meet specific user needs. This results in limited targeted data services and an inability to fully leverage the data's value. Agricultural product producers require risk warning information and recommendations for optimizing production decisions during the production process, but existing service interfaces may not be able to provide these personalized data services.

[0006] In order to solve the above problems, the present invention proposes an agricultural product big data service system based on a trusted data space, which aims to realize the effective integration, cleaning and analysis of multi-dimensional data of the entire agricultural product chain, construct a data association map, generate accurate decision-making features and early warning indicators, and provide a flexible and configurable data service interface to promote the digital transformation and sustainable development of the agricultural product industry. Summary of the Invention

[0007] The purpose of the present invention is to provide an agricultural product big data service system based on a trusted data space to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: an agricultural product big data service system based on a trusted data space, the system comprising:

[0009] Multi-source data integration module: used to obtain multi-dimensional data sets of the entire agricultural product chain;

[0010] Data cleaning and arrangement module: cleans the multi-dimensional data set, generates standardized data sequences, and dynamically arranges the sequences through data association rules;

[0011] Data association graph construction module: Based on the cleaned and standardized data sequence, a cross-domain data association graph is generated through a distributed graph computing algorithm. The graph contains the spatiotemporal association weights and semantic dependencies between entity nodes.

[0012] Multidimensional analysis engine module: inputs the data association map into the hybrid analysis model, extracts the data evolution pattern through multimodal feature fusion, and generates classification decision features and risk warning indicators;

[0013] Service interface generation module: Based on the classification decision features and risk warning indicators, a configurable service interface set is constructed through a strategy mapping algorithm to output data service solutions adapted to different scenarios.

[0014] Preferably, the multi-dimensional data includes production base data, production process data, product attribute data, market data and cold chain logistics data; wherein, the production base data includes base certification information, scale and output indicators, the production process data includes environmental monitoring parameters and operation records, and the product attribute data includes test results and nutritional ingredients.

[0015] Preferably, the data cleaning and arrangement module includes:

[0016] Perform entity normalization on the text description fields in the production base data to eliminate naming ambiguity;

[0017] A sliding window mechanism is used to smooth and correct time series anomalies in production process data;

[0018] Construct quality confidence labels based on the test results of product attribute data and associate them with standardized data sequences;

[0019] The rule engine dynamically adjusts data association weights to generate a hierarchical orchestration structure.

[0020] Preferably, the implementation of the distributed graph computing algorithm includes:

[0021] Spatial grid coding of the geographic coordinates of production base data and cold chain logistics data;

[0022] Extract device identifiers and operation timestamps from production process data to build event chains;

[0023] Calculate the dynamic association strength between entity nodes through incremental graph embedding algorithm;

[0024] The spatiotemporal coding, event chain and dynamic correlation strength are integrated to generate a multidimensional data correlation map.

[0025] Preferably, the hybrid analysis model includes:

[0026] Perform periodic decomposition on the price fluctuation series in market data to extract trend characteristics and residual components;

[0027] Matrix-join the product testing data and nutritional data to generate a quality feature vector;

[0028] A multi-dimensional decision space is constructed by fusing trend features, residual components and quality feature vectors through a hierarchical attention mechanism.

[0029] Divide risk level thresholds in the decision space to generate classification decision features and warning trigger conditions.

[0030] Preferably, the implementation of the hierarchical attention mechanism includes: aligning the trend features in the time dimension to generate a first attention weight matrix; extracting high-frequency mutation features in the residual component to construct a second attention weight matrix; mapping the quality feature vector into a spatial distribution heat map to generate a third attention weight matrix; and outputting multimodal decision features by weighted superposition and fusing the three types of weight matrices.

[0031] Preferably, the method for constructing the strategy mapping algorithm includes:

[0032] Define interface function templates and parameter constraints based on data service scenarios;

[0033] Semantically match the classification decision features with the function template to generate the initial interface configuration plan;

[0034] The compatibility of the cost function evaluation scheme with the risk warning indicators is evaluated, and the parameter combination is iteratively optimized;

[0035] Generates a set of service interfaces that support dynamic expansion and binds data access permission policies.

[0036] Preferably, the specific method of entity normalization processing includes: constructing a standard vocabulary and alias mapping table for agricultural product base names, using a fuzzy matching algorithm to identify synonymous entities in text descriptions, filtering low-probability matching results through a confidence threshold, merging duplicate entity records, and adding unique identifiers and data traceability labels to the normalized entities.

[0037] Preferably, the optimization method of the incremental graph embedding algorithm includes:

[0038] Divide the graph structure into static and dynamic partitions based on data update frequency;

[0039] A batch embedding learning algorithm is used to generate reference vectors in static partitions;

[0040] Perform local gradient propagation optimization on newly added nodes in the dynamic partition and update the association strength;

[0041] Static and dynamic partition features are fused through residual connections.

[0042] Preferably, the method for constructing the cost function includes:

[0043] Define the balance coefficient between interface response delay and data calculation complexity, and use the compatibility score of the initialization parameter combination as the benchmark value;

[0044] The back propagation algorithm is used to calculate the gradient of the impact of parameter adjustment on the early warning indicators, and the optimal parameter configuration sequence is generated according to the gradient direction and balance coefficient.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] At the data processing level, the system uses a multi-source data integration module to obtain a multi-dimensional data set covering the entire agricultural product chain, breaking down data silos. Production base data, production process data, product attribute data, market data, and cold chain logistics data are aggregated, providing a rich data foundation for comprehensive analysis of the agricultural product industry. The data cleaning and orchestration module refines this data, normalizing the text description fields in the production base data to eliminate naming ambiguities, making the data presentation more standardized and uniform, and improving data usability. A sliding window mechanism is used to smooth out time series outliers in production process data, ensuring data accuracy and providing a reliable basis for subsequent analysis. Quality confidence labels are constructed based on the test results of product attribute data, associating standardized data sequences to strengthen the logical relationships between data and lay the foundation for in-depth mining of data value.

[0047] In terms of data association analysis, the data association graph construction module generates a cross-domain data association graph based on cleaned and standardized data sequences using a distributed graph computing algorithm. The geographic coordinates of production base data and cold chain logistics data are spatially gridded and encoded. Equipment identifiers and operation timestamps are extracted from production process data to construct an event chain. An incremental graph embedding algorithm is used to calculate the dynamic association strength between entity nodes, which are then integrated to generate a multidimensional data association graph. This graph, which includes spatiotemporal association weights and semantic dependencies between entity nodes, clearly presents the complex connections between data from all links in the entire agricultural product chain, providing a new perspective for a deeper understanding of the production, circulation, and sales processes of agricultural products, and assisting enterprises in making more scientific decisions.

[0048] From the perspective of data analysis and early warning, the multidimensional analysis engine module feeds the data correlation graph into a hybrid analysis model. It extracts data evolution patterns through multimodal feature fusion, generating classification decision features and risk warning indicators. It performs periodic decomposition on price fluctuation sequences in market data, extracts trend features and residual components, and combines product testing data with nutritional data to generate quality feature vectors. It then fuses these features through a hierarchical attention mechanism to construct a multidimensional decision space and assign risk level thresholds. This enables the system to accurately analyze market dynamics and promptly identify potential risks, such as significant price fluctuations in agricultural products and product quality issues, helping companies take proactive measures to mitigate operational risks.

[0049] In terms of data service provision, the service interface generation module builds a configurable service interface set based on classification decision features and risk warning indicators through a strategy mapping algorithm, and outputs data service solutions that are suitable for different scenarios. According to the data service scenario, the interface function template and parameter constraints are defined, the classification decision features are semantically matched with the function template, and an initial interface configuration plan is generated. The cost function is used to evaluate the compatibility of the solution with the risk warning indicators, and the parameter combination is iteratively optimized to generate a service interface set that supports dynamic expansion, while binding data access permission policies. This meets the diverse needs of different user groups. Agricultural product producers can obtain production optimization suggestions and risk warnings, distributors can obtain market trend analysis and inventory management suggestions, and consumers can also obtain product quality traceability information. This comprehensively improves the quality and value of data services and promotes the digital development and management level of the agricultural product industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a working principle diagram of the agricultural product big data service system based on the trusted data space according to the present invention;

[0051] Figure 2 Flowchart for distributed graph computing algorithm implementation;

[0052] Figure 3 It is the workflow diagram of the hybrid analysis model;

[0053] Figure 4 Flowchart implemented for the layered attention mechanism. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0055] See also Figures 1-4 The present invention provides an agricultural product big data service system based on a trusted data space. This system aims to utilize advanced data processing technologies and algorithms to integrate multi-dimensional data across the entire agricultural product supply chain, providing accurate and efficient data services for related fields. The following are specific implementations of the present invention.

[0056] The multi-source data integration module is responsible for acquiring multi-dimensional data sets across the entire agricultural product supply chain. This data covers every aspect of agricultural production, processing, circulation, and sales, and forms the foundation for the system's subsequent analysis and services. By establishing connections with various data sources, such as sensors at production bases, market trading platform databases, and logistics enterprise information systems, it enables the automated collection and aggregation of multi-dimensional data.

[0057] The data cleaning and orchestration module cleans the acquired multi-dimensional data sets to generate standardized data sequences. Because raw data may contain noise, erroneous values, and duplicate data, it requires a series of data cleaning techniques. Furthermore, this module dynamically orchestrates standardized data sequences based on data association rules, enabling more efficient subsequent data analysis and utilization.

[0058] The Data Association Graph Construction module uses a distributed graph computing algorithm to generate a cross-domain data association graph based on cleaned and standardized data sequences. This graph not only contains the spatiotemporal association weights between entity nodes but also reflects semantic dependencies, providing a powerful tool for in-depth exploration of potential connections between data.

[0059] The multidimensional analysis engine module feeds the data association graph into a hybrid analysis model and extracts data evolution patterns using multimodal feature fusion technology. Based on these patterns, it generates classification decision features and risk warning indicators, providing users with valuable decision support information.

[0060] The service interface generation module uses a strategy mapping algorithm to construct a set of configurable service interfaces based on classification decision features and risk warning indicators. These interfaces can output data service solutions adapted to different scenarios to meet the diverse needs of users.

[0061] Next, the specific implementation details of the present invention are described in detail through six embodiments.

[0062] Example 1:

[0063] During the operation of the agricultural product big data service system, the specific workflow of the multi-source data integration module is as follows. For production base data, the system connects with the information management systems of each production base to obtain data such as base certification information, scale, and yield indicators. For example, by establishing a data transmission channel with the internal system of a large vegetable production base, real-time data is obtained on the base's organic certification, planting area, and annual yield of different vegetable varieties. For production process data, various sensors installed at the production base, such as temperature and humidity sensors and soil pH sensors, collect environmental monitoring parameters. Furthermore, production management software records operator operations, including sowing time, fertilizer type and amount, and irrigation schedule. For product attribute data, the system collaborates with professional agricultural product testing agencies to obtain product test results, such as pesticide residue and heavy metal content data. Nutritional data for the agricultural products is also obtained from the Agricultural Product Nutrition Research Database. For market data, web crawlers are used to collect agricultural product price information and transaction volume data from major agricultural product trading platforms and industry information websites. For cold chain logistics data, it is connected with the logistics management system of the logistics company to obtain temperature, humidity data of agricultural products during transportation and storage, as well as transportation trajectory, delivery time and other information.

[0064] The data cleaning and orchestration module cleans and orchestrates the acquired multi-dimensional data. Taking production base data as an example, for text description fields, a standard vocabulary and alias mapping table for agricultural product base names are constructed, and a fuzzy matching algorithm is used to identify synonymous entities within text descriptions. For example, if two names, "XX Organic Farm" and "XX Ecological Agriculture Base," exist, the standard vocabulary, alias mapping table, and fuzzy matching algorithm can be used to identify them as likely referring to the same entity. Low-probability matches are filtered out by setting a confidence threshold, duplicate entity records are merged, and unique identifiers and data traceability tags are attached to the normalized entities to facilitate subsequent data query and management. For production process data, a sliding window mechanism is used to smooth out time series outliers. For example, if temperature monitoring data during the production process exhibits abnormal fluctuations, a sliding window mechanism is used to select data within a certain time range and calculate the mean. The calculated mean replaces the outliers, making the data more stable and reliable. Quality confidence labels are constructed based on the test results of product attribute data and associated with the standardized data series. If the pesticide residue test result for an agricultural product falls below the national standard, a higher quality confidence label can be assigned; otherwise, a lower quality confidence label can be assigned and associated with other relevant data. The rule engine dynamically adjusts data association weights to generate a hierarchical orchestration structure. Based on the closeness of associations between different data and business needs, corresponding rules are set and data association weights are adjusted to give important data associations a more prominent position in the hierarchical structure.

[0065] Example 2:

[0066] When generating a cross-domain data association map, the data association graph construction module performs spatial grid encoding on the geographic coordinates of production base data and cold chain logistics data. For example, the latitude and longitude coordinates of a production base are divided into grids with a certain accuracy, and each grid is assigned a unique code. Assuming a grid unit of 100 meters x 100 meters, and the coordinates of a production base are located within a specific grid, the grid code is "010101", then the geographic coordinates of the production base are encoded as this grid number. In this way, geographic coordinates are converted into a digital code that is easy to process, facilitating subsequent analysis.

[0067] Device identifiers and operation timestamps are extracted from production process data to construct an event chain. When an irrigation device starts operating, its device identifier and operation start timestamp are recorded; when the device stops operating, its stop timestamp is recorded. These device identifiers and timestamps are linked sequentially according to the operation sequence to form an event chain. For example, if device A starts operating at time t1 and stops operating at time t2, and device B starts operating at time t3, etc., an event chain is constructed that reflects the operation sequence and temporal relationship of the devices in the production process.

[0068] An incremental graph embedding algorithm is used to calculate the dynamic association strength between entity nodes. The graph structure is divided into static and dynamic partitions based on data update frequency. Data with low update frequency, such as basic production site information (land area, geographic location, etc.), is classified into static partitions; data with high update frequency, such as real-time monitoring data from the production process, is classified into dynamic partitions. A batch embedding learning algorithm is used to generate reference vectors in the static partitions. By learning and analyzing a large amount of data in the static partitions, a reference vector representing the characteristics of each entity node is generated. Local gradient propagation optimization is performed on newly added nodes in the dynamic partitions to update the association strength. When new production process data or logistics data is added, the local gradient propagation algorithm is used to optimize the association strength of these newly added nodes with other nodes. Residual connections are used to fuse the features of the static and dynamic partitions, combining the reference vectors of the static partitions with the optimized association strengths of the dynamic partitions to obtain more accurate dynamic association strengths between entity nodes.

[0069] The multidimensional data association map is generated by integrating spatiotemporal coding, event chains, and dynamic association strength. The spatially gridded geographic coordinate information, the constructed event chains, and the calculated dynamic association strength are integrated to construct a multidimensional data association map that includes the spatiotemporal association weights and semantic dependencies between entity nodes. For example, a node in the map represents a production base. The spatiotemporal coding can reflect its geographic location, the event chains reflect the production operations of the base, and the dynamic association strength shows the closeness of the connection between the base and other entities (such as logistics nodes and market nodes).

[0070] Example 3:

[0071] The hybrid analysis model of the multi-dimensional analysis engine module performs periodic decomposition on the price fluctuation sequence in the market data. Assume that the price fluctuation sequence in the market data is , using the cycle decomposition algorithm to decompose it into trend characteristics and the residual component ,Right now . Where t represents time, represents the price of agricultural products at time t, The trend part of the price reflects the long-term trend of price changes. The residual component represents the remaining volatility after removing the trend, including information such as short-term random fluctuations and outliers. By analyzing historical price data and using appropriate cycle decomposition algorithms, such as seasonality-to-trend (STL) decomposition, trend characteristics and residual components can be extracted.

[0072] Matrix the product test data and nutritional data to generate a quality feature vector. Arrange the product test data (such as pesticide residues, heavy metal content, etc.) and nutritional data (such as protein content, vitamin content, etc.) in a certain order into a matrix form. Assuming that the test data has m indicators and the nutritional data has n indicators, then construct a This matrix is ​​the quality feature vector. For example, for a certain fruit, there are 3 items of pesticide residue detection data and 5 items of nutritional component data. These 8 items of data are arranged in sequence into a vector , as the quality feature vector of the fruit.

[0073] The trend features, residual components and quality feature vectors are integrated through the hierarchical attention mechanism to construct a multi-dimensional decision space. The trend features are aligned in the time dimension to generate the first attention weight matrix. Since the importance of trend features may be different at different time points, by aligning the time dimension, the weight corresponding to each time point is calculated according to the characteristics of the data at each time point to form the first attention weight matrix . Extract high-frequency mutation features from the residual component and construct the second attention weight matrix The high-frequency mutation features in the residual component often reflect abnormal market fluctuations or sudden changes in product quality. These features are extracted through a specific algorithm, and the second attention weight matrix is ​​constructed according to their importance. Map the quality feature vector to a spatial distribution heat map to generate the third attention weight matrix According to the numerical value of each element in the quality feature vector, a visual mapping is performed in space to form a heat map. Then, according to the heat distribution of different areas in the heat map, the third attention weight matrix is ​​generated. By weighted superposition and fusion of three types of weight matrices, multimodal decision features are output, namely , where M represents the multimodal decision feature and Q represents the quality feature vector.

[0074] Risk level thresholds are assigned within the decision space to generate classification decision features and early warning trigger conditions. Different risk level thresholds are set based on historical agricultural product market data and industry experience. For example, price volatility risk can be categorized as low, medium, and high. When price fluctuations exceed a certain threshold, a corresponding early warning is triggered. Furthermore, by combining quality feature vectors with other relevant data, classification decision features are generated to determine, for example, whether a batch of agricultural products is suitable for entering a particular market or whether a special marketing strategy is necessary.

[0075] Example 4:

[0076] When constructing the strategy mapping algorithm, the service interface generation module defines interface function templates and parameter constraints based on the data service scenario. For example, for a decision support scenario for agricultural product production enterprises, an interface function template is defined that provides functions such as agricultural product yield forecasting and market demand analysis. Parameter constraints are also set, such as the time range parameter for yield forecasting (forecasting yield for the next one, three, or six months) and the regional parameters for market demand analysis (specifying a specific region or nationwide).

[0077] Semantically match the classification decision features with the function template to generate an initial interface configuration. When the system generates a classification decision feature, such as predicting a significant increase in market demand for a certain agricultural product in a certain region over the next three months, this decision feature is semantically matched with the previously defined interface function template. Based on the matching results, the appropriate function module is selected and the corresponding parameters are set to generate the initial interface configuration. Assuming the interface function template includes a "Market Demand Forecast" function module, the forecast time and region parameters are set to values ​​that match the decision feature to obtain the initial interface configuration.

[0078] The compatibility between the cost function evaluation scheme and the risk warning indicator is evaluated and the parameter combination is iteratively optimized. The balance coefficient between the interface response delay and the data calculation complexity is defined and the compatibility score of the initial parameter combination is the benchmark value. Assume that the balance coefficient is , the interface response delay is D, the data calculation complexity is C, and the initial compatibility score is The back propagation algorithm is used to calculate the gradient of the impact of parameter adjustment on the early warning indicators, and the optimal parameter configuration sequence is generated according to the gradient direction and the balance coefficient. When adjusting a certain parameter (such as the prediction time range parameter), the back propagation algorithm is used to calculate the degree of impact of the parameter adjustment on the risk early warning indicators (such as the accuracy and timeliness of the prediction) to obtain the impact gradient. According to this gradient direction, combined with the balance coefficient , adjust the parameters to continuously optimize the compatibility score and finally obtain the optimal parameter configuration sequence.

[0079] Generate a set of dynamically scalable service interfaces and bind them to data access permission policies. Based on the optimized parameter configuration sequence, specific service interfaces are generated. These interfaces can be dynamically expanded based on user needs. For example, when users request new functionality, new functional modules can be easily added to the interfaces. Furthermore, to ensure data security and privacy, data access permission policies are bound. Different data access permissions are set for different user roles (such as manufacturers, distributors, and regulators). Only users with the appropriate permissions can access specific data and use the corresponding interface functions.

[0080] Example 5:

[0081] In the data cleaning and arrangement module, the specific process of entity normalization is further explained. Constructing a standard vocabulary and alias mapping table for agricultural product base names is the basis for entity normalization. By collecting a large number of agricultural product base names, a database containing standard names and common aliases is established. For example, "XX Green Vegetable Base" is a standard name, and "XX Green Farm" and "XX Vegetable Plantation" are its aliases. When actually processing the text description field, a fuzzy matching algorithm is used to identify synonymous entities in the text description. Assuming that the input text contains "XX Ecological Farm", the fuzzy matching algorithm will search for similar names in the standard vocabulary and alias mapping table. By calculating the text similarity (such as using the edit distance algorithm), when the similarity exceeds the set threshold, it is considered that a synonymous entity has been found.

[0082] Filter low-probability matching results through the confidence threshold. Since fuzzy matching may result in some less accurate matching results, set a confidence threshold. For example, when the confidence corresponding to the matching similarity calculation result is lower than 0.8, it is considered a low-probability matching result, and it is filtered out without subsequent processing. Merge duplicate entity records. When multiple text descriptions are identified pointing to the same entity, merge these records into one record. For example, it is found that "XX Organic Farm 1" and "XX Organic Farm 2" are actually different records of the same farm. Merge them into one record and update the relevant data information. Attach a unique identifier and data traceability label to the normalized entity. Give each normalized entity a globally unique identifier to facilitate accurate identification and management in the system. At the same time, add a data traceability label to record the source of the entity data, such as the time of data collection, the system of collection, and other information, so that the accuracy and reliability of the data can be traced later.

[0083] Example 6:

[0084] In the Data Association Graph Construction module, the optimization process of the incremental graph embedding algorithm is discussed in depth. Dividing the graph structure into static and dynamic partitions based on data update frequency is a key optimization step. The partitioning criteria are determined by statistically analyzing the update frequency of various types of data in the system. For example, basic information about a production site (such as its geographic location and land area) typically remains unchanged for a long time and is classified as a static partition. Real-time monitoring data from the production process (such as temperature and humidity data, equipment operating status data, etc.) and real-time market price data, which are updated more frequently, are classified as dynamic partitions.

[0085] In static partitions, a batch embedding learning algorithm is used to generate benchmark vectors. This algorithm centrally learns and analyzes the large amount of data in the static partitions, extracting data features and generating benchmark vectors that represent the characteristics of each entity node. For example, for a production base node, a benchmark vector representing the production base's characteristics is generated by learning relevant data such as land area, crop varieties, and historical production. This benchmark vector reflects the production base's basic attributes and characteristics, providing a foundation for subsequent calculations of dynamic association strength.

[0086] Local gradient propagation optimization is performed on newly added nodes in the dynamic partition to update the association strength. When new data is added to the dynamic partition, such as production process monitoring data from a certain point in time, these newly added nodes are optimized using the local gradient propagation algorithm. The local gradient propagation algorithm calculates the gradient of the impact on the association strength based on the relationship between the newly added node and the surrounding nodes, and adjusts the association strength based on this gradient. For example, when new temperature and humidity monitoring data is added, the local gradient propagation algorithm calculates the impact of this data on the association strength of surrounding production base nodes, equipment nodes, etc., and then updates the values ​​of these association strengths.

[0087] Static and dynamic partition features are fused through residual connections. Residual connections are an effective fusion method that combines the baseline vectors generated by static partitions with the optimized association strengths of dynamic partitions. Residual connections preserve the fundamental features of static partitions while also integrating the latest changes in dynamic partitions. For example, residual connections are performed on the baseline vectors of production base nodes in the static partition and the association strengths in the dynamic partition, updated based on the latest monitoring data. This yields a more accurate node feature representation that better reflects actual conditions, thereby improving the accuracy and real-time performance of the data association graph.

[0088] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0089] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. An agricultural product big data service system based on a trusted data space, characterized by: include: Multi-source data integration module: used to obtain multi-dimensional data sets of the entire agricultural product chain; Data cleaning and arrangement module: cleans the multi-dimensional data set, generates standardized data sequences, and dynamically arranges the sequences through data association rules; Data association graph construction module: Based on the cleaned and standardized data sequence, a cross-domain data association graph is generated through a distributed graph computing algorithm. The graph contains the spatiotemporal association weights and semantic dependencies between entity nodes. Multidimensional analysis engine module: inputs the data association map into the hybrid analysis model, extracts the data evolution pattern through multimodal feature fusion, and generates classification decision features and risk warning indicators; Service interface generation module: Based on the classification decision features and risk warning indicators, a configurable service interface set is constructed through a strategy mapping algorithm to output data service solutions adapted to different scenarios; The data cleaning and arrangement module includes: Perform entity normalization on the text description fields in the production base data to eliminate naming ambiguity; A sliding window mechanism is used to smooth and correct time series anomalies in production process data; Construct quality confidence labels based on the test results of product attribute data and associate them with standardized data sequences; Dynamically adjust data association weights through the rule engine to generate a hierarchical orchestration structure; The implementation of the distributed graph computing algorithm includes: Spatial grid coding of the geographic coordinates of production base data and cold chain logistics data; Extract device identifiers and operation timestamps from production process data to build event chains; Calculate the dynamic association strength between entity nodes through incremental graph embedding algorithm; Fusion of spatiotemporal coding, event chains, and dynamic correlation strength to generate multidimensional data correlation maps; The hybrid analysis model includes: Perform periodic decomposition on the price fluctuation series in market data to extract trend characteristics and residual components; Matrix-join the product testing data and nutritional data to generate a quality feature vector; A multi-dimensional decision space is constructed by fusing trend features, residual components and quality feature vectors through a hierarchical attention mechanism. Divide risk level thresholds in the decision space to generate classification decision features and warning trigger conditions; The construction method of the strategy mapping algorithm includes: Define interface function templates and parameter constraints based on data service scenarios; Semantically match the classification decision features with the function template to generate the initial interface configuration plan; The compatibility of the cost function evaluation scheme with the risk warning indicators is evaluated, and the parameter combination is iteratively optimized; Generates a set of service interfaces that support dynamic expansion and binds data access permission policies.

2. The agricultural product big data service system based on a trusted data space according to claim 1, characterized in that: The multi-dimensional data includes production base data, production process data, product attribute data, market data and cold chain logistics data; among them, the production base data includes base certification information, scale and output indicators, the production process data includes environmental monitoring parameters and operation records, and the product attribute data includes test results and nutritional ingredients.

3. The agricultural product big data service system based on a trusted data space according to claim 1 is characterized in that: The implementation of the hierarchical attention mechanism includes: aligning the trend features in the time dimension to generate a first attention weight matrix; extracting high-frequency mutation features in the residual component to construct a second attention weight matrix; mapping the quality feature vector into a spatial distribution heat map to generate a third attention weight matrix; and outputting multimodal decision features by weighted superposition and fusion of the three types of weight matrices.

4. The agricultural product big data service system based on a trusted data space according to claim 1, characterized in that: The specific method of entity normalization processing includes: constructing a standard vocabulary and alias mapping table for agricultural product base names, using a fuzzy matching algorithm to identify synonymous entities in text descriptions, filtering low-probability matching results through a confidence threshold, merging duplicate entity records, and adding unique identifiers and data traceability labels to the normalized entities.

5. The agricultural product big data service system based on a trusted data space according to claim 1 is characterized in that: The optimization method of the incremental graph embedding algorithm includes: Divide the graph structure into static and dynamic partitions based on data update frequency; A batch embedding learning algorithm is used to generate reference vectors in static partitions; Perform local gradient propagation optimization on newly added nodes in the dynamic partition and update the association strength; Static and dynamic partition features are fused through residual connections.

6. The agricultural product big data service system based on a trusted data space according to claim 1, characterized in that: The method for constructing the cost function includes: Define the balance coefficient between interface response delay and data calculation complexity, and use the compatibility score of the initialization parameter combination as the benchmark value; The back propagation algorithm is used to calculate the gradient of the impact of parameter adjustment on the early warning indicators, and the optimal parameter configuration sequence is generated according to the gradient direction and balance coefficient.

Citation Information

Patent Citations

  • Platform for promoting intelligent development of industrial internet of things system

    CN114424167A

  • Mirror image type industry linkage engine system and method based on knowledge graph

    CN119537864A