Agricultural product data real-time analysis method and system based on big data
The method optimizes agricultural product supply chain management by employing big data analysis techniques to handle heterogeneous data efficiently, improving scheduling precision and reducing communication overhead, thus enhancing the accuracy of transportation and sales strategies.
Patent Information
- Application Number
- CN202510529747.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The lack of refined management capabilities in agricultural product supply chain management, making it difficult to quickly respond to the superposition of multiple dynamic factors, limited computing resources lead to time-consuming and possible leakage of big data analysis, traditional scheduling methods are inefficient, making it difficult to achieve global optimization.
By obtaining multi-source heterogeneous data, the task DAG is constructed, the nodes on the critical path are clustered using linear clustering method, nodes with high communication overhead are merged, the DAG granularity is calculated and the remaining nodes are clustered, the feature weights are updated using the prediction accuracy of the feature sequence to generate more accurate prediction results to adjust transportation and sales strategies.
It improves the efficiency and accuracy of big data analysis, reduces communication overhead, realizes more efficient resource allocation and optimized transportation route planning, and improves the scheduling and management capabilities of the agricultural product supply chain.
Smart Images

Figure CN120317622A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data, and specifically to a real-time analysis method and system for agricultural product data based on big data. Background Art
[0002] In the management of the agricultural product supply chain, especially in terms of transportation and sales scheduling, most rely on empirical decisions, static plans or single data sources. For example, only based on historical sales, there is generally a lack of refined management capabilities. It is difficult to quickly evaluate the superposition effects of multiple dynamic factors globally and make an optimal systematic response. Moreover, agricultural products have obvious seasonal and perishable characteristics, and have extremely high requirements for transportation timeliness. Once there is a delay in the transportation link, it will directly affect the product quality and market value. Traditional empirical scheduling is difficult to cope with this complex and changeable market environment, easily causing overstock or short supply. Moreover, when dealing with large-scale, multi-variety, and multi-mode transportation demands, it is often inefficient and difficult to achieve global optimization. Especially when facing massive and heterogeneous agricultural product-related data, it is impossible to dig out deep-level rules and trends from it, thus making it difficult to achieve precise scheduling and management. Big data technology can better achieve precise scheduling management of agricultural product transportation and sales through the collection, storage, processing, and analysis of massive data. However, when the data volume is too large, processing these data is time-consuming. General agricultural enterprises have limited computing resources. If computing is performed on the cloud platform, data leakage may occur. Therefore, it is very important to be able to complete big data analysis quickly and well, and make predictions based on the big data analysis results and then adjust transportation and sales strategies. Summary of the Invention
[0003] In view of the problem that it is difficult for enterprises with limited computing resources to quickly complete big data analysis, in the first aspect of the present invention, a real-time analysis method for agricultural product data based on big data is provided, including the following steps: Obtain multi-source heterogeneous data related to agricultural products and perform preprocessing to obtain a task DAG according to the analysis task; Obtain all paths from the entrance to the exit in the DAG, take the sum of the communication times of all nodes on each path as the length of the path, take the path with the maximum length as the critical path, cluster the nodes on the critical path by a linear clustering method, calculate the DAG granularity based on the linear clustering result and cluster the remaining nodes, and schedule the nodes in the DAG to execution nodes according to the clustering result to obtain the analysis task result.
[0004] Preferably, the clustering of the nodes on the critical path by a linear clustering method is specifically: Calculate the ratio of the communication overhead from the previous node to the current node to the computing overhead of the current node to obtain the local granularity of the current node. Both the current node and the previous node are nodes on the critical path, and the communication overhead of the first node on the critical path is 0; For each node on the critical path with a local granularity less than 1, if the local granularity of the next node is not less than 1, add the next node to the set of nodes with a local granularity less than 1, and repeat continuously; Take the set of each node with a local granularity less than 1 as a cluster for linear clustering.
[0005] Preferably, calculating the DAG granularity based on the linear clustering result and clustering the remaining nodes specifically includes: Update the DAG by taking each cluster in the linear clustering as a node in the DAG, and calculate the granularity of the DAG; If the granularity of the DAG is less than 1, merge the two nodes with the largest communication overhead, recalculate the granularity of the DAG until the granularity of the DAG is greater than or equal to 1; when the granularity of the DAG is greater than or equal to 1, for each path, cluster each path in the way of linear clustering.
[0006] Preferably, the method further includes: Calculate the features of the analysis task result, and then obtain a feature sequence. Determine the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence, update the features according to the weights, input the updated features into the prediction model to obtain a prediction result, and use the prediction result to adjust the transportation and sales scheduling strategy.
[0007] Preferably, determining the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence specifically includes: Set the weight of the first feature in the feature sequence to 1; For the i-th element in the feature sequence, take the prediction accuracy of the i-th element or the average value of the prediction accuracies from the second element to the i-th element as the weight of the i-th element; where i is a positive integer and i≥2.
[0008] Preferably, updating the features according to the weights specifically includes: For the features in the feature sequence, if the weight is not greater than the average value of the feature weights in the feature sequence, delete the feature from the feature sequence; Obtain the weighted average feature by weighting the features according to the weights in the feature sequence, and take the weighted average feature as the updated feature.
[0009] In addition, the present invention also provides a real-time analysis system for agricultural product data based on big data, including the following modules: A task acquisition module, which is used to acquire multi-source heterogeneous data related to agricultural products, perform preprocessing, and obtain a task DAG according to the analysis task; A data analysis module, which is used to obtain all paths from the entrance to the exit in the DAG, take the sum of the communication times of all nodes on each path as the length of the path, take the path with the maximum length as the critical path, cluster the nodes on the critical path by linear clustering, calculate the DAG granularity based on the linear clustering result, cluster the remaining nodes, and schedule the nodes in the DAG to the execution nodes according to the clustering result to obtain the analysis task result.
[0010] Preferably, the clustering of the nodes on the critical path by linear clustering is specifically as follows: Calculate the ratio of the communication overhead from the previous node to the current node to the computing overhead of the current node to obtain the local granularity of the current node. Both the current node and the previous node are nodes on the critical path, and the communication overhead of the first node on the critical path is 0; For each node on the critical path with a local granularity less than 1, if the local granularity of the next node is not less than 1, add the next node to the set of nodes with a local granularity less than 1, and repeat continuously; Take the set of each node with a local granularity less than 1 as a cluster of linear clustering.
[0011] Preferably, the calculation of the DAG granularity based on the linear clustering result and the clustering of the remaining nodes are specifically as follows: Take each cluster in the linear clustering as a node in the DAG to update the DAG, and calculate the granularity of the DAG; If the granularity of the DAG is less than 1, merge the two nodes with the largest communication overhead, recalculate the granularity of the DAG until the granularity of the DAG is greater than or equal to 1; when the granularity of the DAG is greater than or equal to 1, for each path, cluster each path by linear clustering.
[0012] Preferably, the system further includes: A prediction module, which is used to calculate the features of the analysis task result, thereby obtaining a feature sequence, determine the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence, update the features according to the weights, input the updated features into a prediction model to obtain a prediction result, and use the prediction result to adjust the transportation and sales scheduling strategies.
[0013] Preferably, the determination of the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence is specifically as follows: Set the weight of the first feature in the feature sequence to 1; For the i-th element in the feature sequence, the prediction accuracy of the i-th element or the average value of the prediction accuracies from the 2nd element to the i-th element is used as the weight of the i-th element; where i is a positive integer and i ≥ 2.
[0014] Preferably, the updating of the features according to the weights is specifically as follows: For the features in the feature sequence, if the weight is not greater than the average value of the feature weights in the feature sequence, the feature is deleted from the feature sequence; The weighted average feature is obtained by weighting the features according to the weights in the feature sequence, and the weighted average feature is used as the updated feature.
[0015] Finally, the present invention provides a real-time analysis program for agricultural product data based on big data, including the following modules: A task acquisition module, configured to acquire multi-source heterogeneous data related to agricultural products and perform preprocessing, and obtain a task DAG according to the analysis task; A data analysis module, configured to acquire all paths from the entrance to the exit in the DAG, use the sum of the communication times of all nodes on each path as the length of the path, use the path with the maximum length as the critical path, cluster the nodes on the critical path by linear clustering, calculate the DAG granularity based on the linear clustering result and cluster the remaining nodes, and schedule the nodes in the DAG to execution nodes according to the clustering result to obtain the analysis task result.
[0016] Aiming at the problems of long calculation time and inaccurate prediction results when analyzing agricultural product data due to the large amount of original agricultural product data used, the present invention reduces the communication overhead by allocating related tasks to the same execution node, improves the parallelism and overall throughput; in addition, an adaptive feature update mechanism is adopted to improve the accuracy of the prediction model, and these more accurate prediction results are used to guide the generation of transportation and sales scheduling strategies, realizing more efficient resource allocation, more optimized transportation route planning, etc. Description of the Drawings
[0017] Figure 1 Is the flowchart of Embodiment 1; Figure 2 Is a typical DAG graph; Figure 3 Is the schematic diagram of the critical path in the typical DAG graph; Figure 4 Is the updated DAG graph; Figure 5 Is the schematic diagram of the feature sequence. Detailed Embodiments
[0018] In the embodiments of the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner for ease of understanding.
[0019] It can be understood that the "embodiments" mentioned throughout the specification mean that specific features, structures, or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the various embodiments mentioned throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It can be understood that in the various embodiments of the present application, the magnitude of the serial numbers of the various processes does not mean the order of execution, and the order of execution of the various processes should be determined by their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0020] In the present invention, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of the present invention, and in each implementation manner / implementation method / realization method in each embodiment, if there is no special specification and logical conflict, the terms and / or descriptions between different embodiments, and between each implementation manner / implementation method / realization method in each embodiment are consistent and can be mutually referred to. The technical features in different embodiments, and in each implementation manner / implementation method / realization method in each embodiment, can be combined to form new embodiments, implementation manners, implementation methods, or realization methods according to their internal logical relationships. The implementation manners of the present application described below do not constitute a limitation to the protection scope of the present application.
[0021] Figure 1 The first embodiment of the present invention is shown, as Figure 1 shown, a real-time analysis method for agricultural product data based on big data, including the following steps: S1, obtain multi-source heterogeneous data related to agricultural products and perform preprocessing, and obtain a task DAG according to the analysis task; The transportation and sales scheduling of agricultural products involves multiple influencing factors, such as weather, supply and demand relationship, vehicle transportation, sales channels, etc. Before performing big data analysis, data is obtained. The data includes, but is not limited to, GPS positioning data of agricultural product transportation vehicles, wholesale prices, retail prices, supply and demand information, consumer preference data, sales volume, sales amount, return rate, customer feedback, order information, demand quantity, delivery time, location, etc. of channels. Since these data come from different sources, such as vehicle information from transportation companies, order information from in-house systems, and market conditions from third parties, etc., and the data formats are also different, the original data is cleaned to remove incorrect, incomplete, duplicate or inconsistent data. The data cleaning includes, but is not limited to, missing value processing, outlier processing, format conversion, etc.
[0022] When the data volume is large, the resources of a single machine system are limited. To improve the processing speed, the overall analysis task is decomposed into smaller subtasks that can be executed independently or in parallel, and the execution order and dependency relationship between these subtasks are analyzed. If the start of subtask B needs to wait for the completion of subtask A, there is a directed edge from subtask A to subtask B. For example, the data aggregation task usually depends on the completion of data extraction and cleaning tasks for each channel, and the report generation task usually depends on the completion of the data aggregation task. According to the dependencies or execution order of the subtasks, a DAG (Directed Acyclic Graph) of the analysis task is obtained. In one embodiment, the DAG is represented as G=(V, E), where V is the nodes in the DAG, which are subtasks in the present invention, and E is the edges in the DAG. The size of the edge is the communication overhead between two subtasks. A typical DAG is as Figure 2 shown, n1 - n8 are nodes, and the values on the edges are communication overheads. Obtain all paths from the entrance to the exit in the DAG, and take the sum of the communication times of all nodes on each path as the length of the path. Figure 2 In it, the length of the path of n1, n2, n6, n8 is 10. Figure 2 In it, the critical path is n1, n2, n5, n8, as Figure 3 shown.
[0023] S2. Obtain all paths from the entrance to the exit in the DAG, take the sum of the communication times of all nodes on each path as the length of the path, take the path with the maximum length as the critical path, cluster the nodes on the critical path in a linear clustering manner, calculate the DAG granularity based on the linear clustering result and cluster the remaining nodes, and schedule the nodes in the DAG to execution nodes according to the clustering result to obtain the analysis task result.
[0024] Since the communication overhead of nodes on the critical path is the largest, in order to reduce the impact of communication overhead, linear clustering is used for the nodes on the critical path. In the linear clustering, the nodes or tasks in each cluster are directly dependent on each other in sequence, and there are no indirectly dependent or independent tasks. For example, for the critical path n1, n2, n5, n8, a result of linear clustering is cluster 1 (n1, n2, n5) and cluster 2 (n8). If the clustering result is cluster 1 (n1, n2, n8) and cluster 2 (n5), it is not linear clustering because n2 and n8 are not dependent on each other in sequence.
[0025] In one embodiment, the linear clustering of the nodes on the critical path is specifically as follows: Calculate the ratio of the communication overhead from the previous node to the current node to the computational overhead of the current node to obtain the local granularity of the current node. Both the current node and the previous node are nodes on the critical path, and the communication overhead of the first node on the critical path is 0. For each node on the critical path with a local granularity less than 1, if the local granularity of the next node is not less than 1, add the next node to the set of nodes with a local granularity less than 1, and repeat continuously. Take the set of each node with a local granularity less than 1 as a cluster in the linear clustering.
[0026] The current node on the critical path has a local granularity which is the ratio of the communication overhead from the previous node to the current node to the computational overhead of the current node. For the first node on the critical path , since it has no previous node on the critical path, its communication overhead is 0.
[0027] For each node on the critical path with a local granularity less than 1 , construct a set. At the initial moment , the set only includes this one element. Check the next node on the critical path of . If has a local granularity not less than 1, add to the set of . Then, check the next node on the critical path of . If has a local granularity not less than 1, add to the set of . Stop checking until a node with a local granularity less than 1 is encountered, and the node with a local granularity less than 1 will not be added to the set of . Each node with a local granularity less than 1 The set forms a cluster of linear clustering. Figure 4 It shows that n1 and n2 on the critical path are merged into one cluster to obtain a new node n1', and n5 and n8 are merged into one cluster to obtain a new node n8'.
[0028] After linearly clustering the nodes on the critical path, further cluster the remaining nodes of the DAG. In one embodiment, calculating the DAG granularity based on the linear clustering result and clustering the remaining nodes is specifically as follows: Update the DAG by taking each cluster in the linear clustering as a node in the DAG, and calculate the granularity of the DAG; If the granularity of the DAG is less than 1, merge the two nodes with the largest communication overhead, recalculate the granularity of the DAG until the granularity of the DAG is greater than or equal to 1; when the granularity of the DAG is greater than or equal to 1, for each path, cluster each path in the way of linear clustering.
[0029] For the updated DAG, calculate the fork granularity and join granularity of each node. Take the minimum value of the calculation overhead and communication overhead of all nodes in the fork task set as the fork granularity, and take the minimum value of the calculation overhead and communication overhead of all nodes in the join task set as the join granularity.
[0030] Then take the minimum value of the fork granularity and join granularity as the node granularity, and take the minimum value of the node granularity among the remaining nodes as the DAG granularity. If the DAG granularity is less than 1 and the communication overhead is greater than the calculation overhead, at this time, merge the two nodes with the largest communication overhead in the updated DAG into one cluster or node, recalculate the DAG granularity, and repeat continuously until the DAG granularity is greater than or equal to 1. Finally, for each path in the updated DAG, continue to use linear clustering, that is, merge the dependent nodes on the non-critical path into clusters. Preferably, for each path in the updated DAG, still use the same method as the critical path to cluster the nodes on the same path.
[0031] After clustering the critical path and the remaining nodes, multiple clusters will be obtained. Each cluster is used as a scheduling whole and scheduled to the same execution node, and the execution node is different servers or virtual machines or containers, etc.
[0032] In one embodiment, if after any DAG node merge, it does not meet the definition of DAG, then exit. For the remaining DAG nodes, schedule the merged clusters and remaining tasks in the DAG according to the affinity between tasks and execution nodes in the DAG.
[0033] In yet another embodiment, the method further includes calculating the features of the analysis task result to obtain a feature sequence, determining the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence, updating the features according to the weights, inputting the updated features into a prediction model to obtain a prediction result, and using the prediction result to generate a transportation and sales scheduling strategy.
[0034] After the analysis task is completed, multiple analysis results will be obtained, such as overall sales volume, recent production volume, average price, etc. These analysis results constitute the features of the analysis task result. For example, a feature of the analysis task result is [10, 5, 2.1], where 10 is the overall sales volume, 5 is the recent production volume, and 2.1 is the market average price. The current analysis task corresponds to one feature. Similarly, the previous analysis task will also obtain one feature. The features of the current analysis task and the previous n analysis tasks constitute a feature sequence, as Figure 5 shown. For the previous analysis tasks of the current analysis task, a predicted value will be obtained after being input into the prediction model. For example, the predicted sales volume of each channel or store. Comparing the predicted sales volume with the actual sales volume to obtain the prediction accuracy of the analysis task. The higher the prediction accuracy, the more important the feature of the analysis task is. Taking the prediction accuracy as the weight of the analysis task feature. In one embodiment, the determining the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence is specifically: Set the weight of the first feature in the feature sequence to 1; For the i-th element in the feature sequence, take the prediction accuracy of the i-th element or the average value of the prediction accuracies from the 2nd element to the i-th element as the weight of the i-th element; where i is a positive integer and i≥2.
[0035] There is no corresponding actual data for the first element in the feature sequence. For example, the sales volume of the next day. Set the weight of the first element in the feature sequence to 1; for other elements in the feature sequence, take the prediction accuracy as the weight of the i-th element, or take the average value of the prediction accuracies from the 2nd element to the i-th element as the weight of the i-th element. For example, for the 3rd element in the feature sequence, its weight is equal to the average value of the prediction accuracy of the 2nd element plus the prediction accuracy of the 3rd element. Where the element in the feature sequence is the analysis result corresponding to an analysis task.
[0036] After obtaining the feature sequence, update the features using the weights. In one embodiment, multiply the elements in the feature sequence by their corresponding weights, and then add the multiplication results to obtain the updated feature sequence. In yet another embodiment, in order to reduce the influence of elements with inaccurate predictions, the updating the features according to the weights is specifically: For a feature in the feature sequence, if the weight is not greater than the average value of the feature weights in the feature sequence, the feature is deleted from the feature sequence; The features are weighted according to the weights in the feature sequence to obtain a weighted average feature, and the weighted average feature is used as the updated feature.
[0037] After obtaining the prediction results, such as the sales prediction value, production prediction value, price prediction value, etc. of a certain sales point, adjust the transportation and sales strategies according to the inventory and transportation costs of each sales point. For example, determine the allocation quantity for each city or sales point according to the predicted inventory backlog, allocated weight, and predicted price. Obtain the transportation time, cost, and loss rate of multiple transportation methods, construct an objective function, and add constraints, such as the maximum transportation volume per vehicle, the fastest transportation time, etc., and solve the objective function to obtain a transportation plan. Similarly, determine the sales price according to the allocated weight and predicted price.
[0038] In the second embodiment of the present invention, a real-time analysis system for agricultural product data based on big data is also provided, including the following modules: A task acquisition module for acquiring multi-source heterogeneous data related to agricultural products and performing preprocessing, and obtaining a task DAG according to the analysis task; A data analysis module for obtaining all paths from the entrance to the exit in the DAG, taking the sum of the communication times of all nodes on each path as the length of the path, taking the path with the maximum length as the critical path, clustering the nodes on the critical path by a linear clustering method, calculating the DAG granularity based on the linear clustering result and clustering the remaining nodes, and scheduling the nodes in the DAG to execution nodes according to the clustering result to obtain the analysis task result.
[0039] Preferably, the clustering of the nodes on the critical path by a linear clustering method is specifically: Calculate the ratio of the communication overhead from the previous node to the current node to the computational overhead of the current node to obtain the local granularity of the current node. Both the current node and the previous node are nodes on the critical path, and the communication overhead of the first node on the critical path is 0; For each node on the critical path with a local granularity less than 1, if the local granularity of the next node is not less than 1, add the next node to the set of nodes with a local granularity less than 1, and repeat continuously; Take the set of each node with a local granularity less than 1 as a cluster of the linear clustering.
[0040] Preferably, the calculation of the DAG granularity based on the linear clustering result and the clustering of the remaining nodes is specifically: Take each cluster in the linear clustering as a node in the DAG to update the DAG, and calculate the granularity of the DAG; If the granularity of the DAG is less than 1, merge the two nodes with the largest communication overhead, recalculate the granularity of the DAG until the granularity of the DAG is greater than or equal to 1; when the granularity of the DAG is greater than or equal to 1, for each path, cluster each path in a linear clustering manner.
[0041] Preferably, the system further includes: A prediction module, configured to calculate the features of the analysis task result, thereby obtaining a feature sequence, determine the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence, update the features according to the weights, input the updated features into a prediction model to obtain a prediction result, and use the prediction result to adjust the transportation and sales scheduling strategy.
[0042] Preferably, the determining the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence is specifically: Set the weight of the first feature in the feature sequence to 1; For the i-th element in the feature sequence, use the prediction accuracy of the i-th element or the average value of the prediction accuracies from the second element to the i-th element as the weight of the i-th element; where i is a positive integer and i≥2.
[0043] Preferably, the updating the features according to the weights is specifically: For the features in the feature sequence, if the weight is not greater than the average value of the feature weights in the feature sequence, delete the feature from the feature sequence; Obtain a weighted average feature by weighting the features according to the weights in the feature sequence, and use the weighted average feature as the updated feature.
[0044] The above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0045] The steps of the methods or algorithms described in the embodiments of the present application can be directly embedded in hardware, a software unit executed by a processor, or a combination of the two. The software unit can be stored in a RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, register, hard disk, removable disk, CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be provided in an ASIC.
[0046] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 the steps of the functions specified in one block or multiple blocks.
[0047] Although the present application has been described in connection with specific features and their embodiments, it will be apparent that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are merely exemplary illustrations of the present application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to cover these changes and modifications.
Claims
1. A real-time analysis method for agricultural product data based on big data, characterized in that, Including the following steps: Obtain multi-source heterogeneous data related to agricultural products and perform preprocessing, and obtain a task DAG according to the analysis task; Obtain all paths from the entrance to the exit in the DAG, take the sum of the communication times of all nodes on each path as the length of the path, take the path with the maximum length as the critical path, cluster the nodes on the critical path in a linear clustering manner, calculate the DAG granularity based on the linear clustering result and cluster the remaining nodes, and schedule the nodes in the DAG to the execution nodes according to the clustering result to obtain the analysis task result.
2. The method according to claim 1, characterized in that, The clustering of the nodes on the critical path in a linear clustering manner is specifically as follows: Calculate the ratio of the communication overhead from the previous node to the current node to the computing overhead of the current node to obtain the local granularity of the current node. Both the current node and the previous node are nodes on the critical path, and the communication overhead of the first node on the critical path is 0; For each node on the critical path with a local granularity less than 1, if the local granularity of the next node is not less than 1, then add the next node to the set of nodes with a local granularity less than 1, and repeat continuously; Take the set of each node with a local granularity less than 1 as a cluster of linear clustering.
3. The method according to claim 1, characterized in that The calculating the DAG granularity based on the linear clustering result and clustering the remaining nodes is specifically as follows: Take each cluster in the linear clustering as a node in the DAG to update the DAG, and calculate the granularity of the DAG; If the granularity of the DAG is less than 1, then merge the two nodes with the largest communication overhead, recalculate the granularity of the DAG until the granularity of the DAG is greater than or equal to 1; When the granularity of the DAG is greater than or equal to 1, for each path, cluster each path in a linear clustering manner.
4. The method according to claim 1, wherein The method further includes: Calculate the features of the analysis task result, and then obtain a feature sequence. Determine the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence, update the features according to the weights, input the updated features into the prediction model to obtain a prediction result, and use the prediction result to adjust the transportation and sales scheduling strategies.
5. The method according to claim 4, wherein The determining the weights of the features in the feature sequence based on the prediction accuracy of the features in the feature sequence is specifically as follows: Set the weight of the first feature in the feature sequence to 1; For the i-th element in the feature sequence, take the prediction accuracy of the i-th element or the average value of the prediction accuracies from the 2nd element to the i-th element as the weight of the i-th element; where i is a positive integer and i≥2.
6. The method according to claim 4, characterized in that, The updating the features according to the weights is specifically as follows: For the features in the feature sequence, if the weight is not greater than the average value of the feature weights in the feature sequence, then delete the feature from the feature sequence; Obtain a weighted average feature by weighting the features according to the weights in the feature sequence, and take the weighted average feature as the updated feature.
7. An agricultural product data real-time analysis system based on big data, characterized in that Including the following modules: A task acquisition module, configured to obtain multi-source heterogeneous data related to agricultural products and perform preprocessing, and obtain a task DAG according to the analysis task; The data analysis module is used to obtain all paths from the entrance to the exit in the DAG, take the sum of the communication times of all nodes on each path as the length of the path, take the path with the maximum length as the critical path, cluster the nodes on the critical path in a linear clustering manner, calculate the DAG granularity based on the linear clustering result and cluster the remaining nodes, and schedule the nodes in the DAG to the execution nodes according to the clustering result to obtain the analysis task result.
8. The system according to claim 7, wherein The clustering of the nodes on the critical path in a linear clustering manner is specifically as follows: Calculate the ratio of the communication overhead from the previous node to the current node to the computing overhead of the current node to obtain the local granularity of the current node. Both the current node and the previous node are nodes on the critical path, and the communication overhead of the first node on the critical path is 0. For each node on the critical path with a local granularity less than 1, if the local granularity of the next node is not less than 1, then add the next node to the set of nodes with a local granularity less than 1, and repeat continuously. Take the set of each node with a local granularity less than 1 as a cluster of linear clustering.
9. The system according to claim 7, wherein The calculation of the DAG granularity based on the linear clustering result and the clustering of the remaining nodes is specifically as follows: Take each cluster in the linear clustering as a node in the DAG to update the DAG, and calculate the granularity of the DAG. If the granularity of the DAG is less than 1, then merge the two nodes with the largest communication overhead, recalculate the granularity of the DAG until the granularity of the DAG is greater than or equal to 1. When the granularity of the DAG is greater than or equal to 1, for each path, cluster each path in a linear clustering manner.
10. A real-time analysis program for agricultural product data based on big data, characterized in that, It includes the following modules: The task acquisition module is used to acquire multi-source heterogeneous data related to agricultural products and perform preprocessing, and obtain the task DAG according to the analysis task. The data analysis module is used to obtain all paths from the entrance to the exit in the DAG, take the sum of the communication times of all nodes on each path as the length of the path, take the path with the maximum length as the critical path, cluster the nodes on the critical path in a linear clustering manner, calculate the DAG granularity based on the linear clustering result and cluster the remaining nodes, and schedule the nodes in the DAG to the execution nodes according to the clustering result to obtain the analysis task result.
Citation Information
Cited By
Customer information acquisition method and system based on multi-source data fusion
CN120910038A
Customer information collection method and system based on multi-source data fusion
CN120910038B