Facility agriculture metadata model construction method based on big data analysis
By building an open environmental control data model architecture and using big data analysis technology, the problem of inconsistent data structures between environmental control systems in the facility agriculture is solved, data interoperability and system interoperability are achieved, data security is ensured, and the production efficiency and resource utilization of facility agriculture are improved.
Patent Information
- Application Number
- CN202411776863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-09
AI Technical Summary
The lack of data structure standards among the environmental control systems of facilities agriculture has led to difficulties in data interoperability, resulting in data silos in the construction of informationization of facilities agriculture, reducing the complexity of data integration in agricultural production and the interoperability of agricultural environmental control systems.
The construction method of the facility agricultural metadata model based on big data analysis is adopted, including building an open environmental control data model architecture, using data encryption technology and secure transmission protocol to achieve the interconnection of different environmental control systems, and monitoring and automatically adjusting environmental parameters in real time through intelligent sensors and machine learning algorithms.
The data structure standardization between different environmental control systems has been realized, data interoperability and system interoperability have been improved, data security has been ensured, manual intervention has been reduced, and the production efficiency and resource utilization of facility agriculture have been improved.
Smart Images

Figure CN119963114A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital agriculture, and in particular to a method for constructing a metadata model of facility agriculture based on big data analysis. Background Art
[0002] With the continuous development of digital technology, facility agriculture has made remarkable progress on the road of standardization and precision. As one of the basic modules of digital management of facility agriculture, the environmental control system plays a key role, especially in temperature, However, the widespread application of the current environmental control system still faces some problems and challenges to be solved, mainly in terms of data standardization and interoperability; At present, there is a lack of data structure standards between different environmental control systems. With the acceleration of agricultural localization, more and more environmental control systems based on Internet of Things technology are promised. However, these systems lack unified standards in data structure. By defining data structures by themselves, data intercommunication between different systems becomes difficult. Standardized data structures not only reduce the complexity of data integration in the agricultural production process, but also improve the interoperability of agricultural environmental control systems, avoiding the problem of inefficiency of data sharing and collaborative work due to inconsistent data structures. The digital management of facility agriculture mainly includes multiple modules such as environmental control system, crop growth monitoring system, irrigation control system, etc. However, the definition and classification of basic data of these systems often lead to data island phenomenon in the information construction of facility agriculture due to the lack of unified data metamodel, data exchange and interoperability differences between systems. To achieve efficient management of facility agriculture, it is necessary to first unify and standardize the metamodel of these data, so as to reduce the complexity of data and improve the compatibility and compatibility between systems. In view of this, we propose a method for constructing facility agriculture metadata model based on big data analysis. Summary of the invention
[0003] The purpose of the present invention is to provide a method for constructing a facility agriculture metadata model based on big data analysis to solve the problems raised in the above-mentioned background technology.
[0004] To achieve the above object, the present invention provides a method for constructing a facility agriculture metadata model based on big data analysis, comprising the following steps: S1. Build an open environmental control data model architecture to define the basic data structure and data fields of various agricultural environmental control equipment, and describe the relationship and dependency between agricultural environmental control equipment, crop growth and environmental parameters; S2. Use data encryption technology to encrypt data, and after identity authentication and permission control authorization, use a secure transmission protocol to interconnect the environmental control systems formed by various agricultural environmental control equipment on the same platform; S3. Deploy smart sensors to collect environmental data in real time, and use time series analysis and machine learning algorithms to predict future trends in environmental parameters, and automatically adjust environmental parameters based on the prediction results.
[0005] As a further improvement of the present technical solution, the basic data structure and data fields are used to determine key environmental parameters commonly found in facility agriculture, including temperature, humidity, light intensity, water pump status, irrigation system status, equipment failure status and data timestamp.
[0006] As a further improvement of the present technical solution, the description of the relationship and dependence between agricultural environmental control equipment, crop growth and environmental parameters in S1 includes the following steps: Based on different agricultural environmental control equipment, crop growth data under different key environmental parameters are collected, and the data are cleaned and standardized; Divide the crop growth data into a training set and a test set. The training set is used to train the model, and the test set is used to evaluate the performance of the model. A multivariate linear regression model was used to train the effects of multiple key environmental parameters on crop growth, and the linear relationship between key environmental parameters and crop growth was verified using a test set.
[0007] As a further improvement of the technical solution, the multivariate linear regression model is used to predict the relationship between crop growth and multiple environmental parameters, which is expressed as: in, is the response variable of crop growth, is a key environmental parameter, is the regression coefficient, is the error term.
[0008] As a further improvement of the technical solution, the data encryption technology in S2 includes the following steps: From the ciphertext encrypted by MD5, two parts are intercepted according to the predefined number of bits, including intercepting a string of specified bits from the beginning of the ciphertext, which is defined as password A, and intercepting the remaining string of specified bits from the position after password A, which is defined as password B; Use a random function to generate a fixed-length key to fill in the password B part; The password A and the generated key are combined to form a final transformed password value, and the final transformed password value is used to compress the data packet, wherein the data packet is used to store the multivariate linear regression model data corresponding to different agricultural environmental control equipment to form multiple confidential data packets.
[0009] As a further improvement of the technical solution, the secure transmission protocol is adopted in S2 to interconnect the environmental control systems formed by various agricultural environmental control equipment on the same platform, including the following steps: After identity authentication and permission control authorization, the application layer of the environmental control data model architecture perceives the data packets that need to be transmitted; At the transport layer, packets are encapsulated into transport layer data units and transmitted down the protocol stack, with each layer adding a header; At the network layer, packets are encapsulated into data units at the network layer. At the data link layer, packets are encapsulated into data units at the data link layer. The encapsulated content is passed as data to the next layer until it reaches the physical layer, where the data is converted into a bit stream and transmitted through the medium.
[0010] As a further improvement of the present technical solution, the intelligent sensors in S3 include but are not limited to temperature sensors, humidity sensors, light sensors, soil moisture sensors, carbon dioxide sensors and wind speed sensors; Use the data logger as the central node of the data acquisition system.
[0011] As a further improvement of the technical solution, the S3 uses time series analysis and machine learning algorithms to predict the changing trend of future environmental parameters, including the following steps: Collect historical environmental data, extract time-related features, and introduce lag features, that is, data from previous time points, to capture the autocorrelation of time series. Use sliding window technology to convert time series data into a supervised learning problem. Divide the dataset into a training set and a test set, using a time order to ensure that the training set comes first and the test set comes last. Use the training set to train the machine learning model, adjust the hyperparameters to optimize the model performance, and use cross-validation techniques based on the test set to evaluate the generalization ability of the model. Input time points to the machine learning model to predict environmental parameters.
[0012] As a further improvement of the technical solution, the automatic adjustment of environmental parameters based on the prediction results includes the following steps: The adjustment time point is preset, and the environmental parameters at this time point are received and input into the multivariate linear regression model, crop growth is output, and the crop growth status is judged. If the crop growth is normal, the agricultural environmental control equipment will not be triggered. If the crop growth is abnormal, the PID controller will be used to trigger the agricultural environmental control equipment to adjust the environmental parameters. The environmental control equipment is adjusted according to the predicted results until the crop grows normally. Through the feedback mechanism, when the environmental parameters are adjusted to a stage value, the crop growth status is re-predicted to ensure that the final crop growth condition is normal.
[0013] Compared with the prior art, the present invention has the following beneficial effects: In the facility agriculture metadata model construction method based on big data analysis, an open environmental control data model architecture is constructed, which is conducive to data structure standards between different environmental control systems, realizing unified standards in data structure, avoiding the impact of data intercommunication between different systems caused by manufacturers defining their own data structures, and the standardized data structure not only reduces the complexity of data integration in the agricultural production process, but also improves the interoperability of agricultural environmental control systems, avoiding the problem of inefficiency of data sharing and collaborative work due to inconsistent data structures; And the use of secure transmission protocols enables the environmental control systems formed by various agricultural environmental control equipment to be interconnected on the same platform, ensuring the security of data during interconnection. Whether it is data storage, exchange or sharing, it is carried out under a strict security mechanism, effectively preventing data leakage and tampering, and providing reliable data protection for agricultural production management; At the same time, it is conducive to real-time monitoring of agricultural environmental parameters, automatic adjustment of environmental control equipment, optimization of production management, effective improvement of production efficiency and resource utilization of facility agriculture, reduction of human intervention, and improvement of automation and precision of agricultural production. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is the overall flow chart of the present invention. DETAILED DESCRIPTION
[0015] The following will be combined with the accompanying drawings in the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0016] See also Figure 1 As shown, this embodiment provides a method for constructing a facility agriculture metadata model based on big data analysis, comprising the following steps: S1. Build an open environmental control data model architecture to define the basic data structure and data fields of various agricultural environmental control equipment, and describe the relationship and dependence between agricultural environmental control equipment, crop growth and environmental parameters. This is conducive to the establishment of data structure standards between different environmental control systems, and to achieve unified standards in data structure, so as to avoid manufacturers defining their own data structures and affecting data interoperability between different systems. Standardized data structures not only reduce the complexity of data integration in the agricultural production process, but also improve the interoperability of agricultural environmental control systems, and avoid the problem of inefficiency of data sharing and collaborative work due to inconsistent data structures. Among them, the basic data structure and data fields are used to determine the key environmental parameters commonly used in facility agriculture, including temperature, humidity, light intensity, water pump status, irrigation system status, equipment fault status and data timestamp. The following is an example of this data structure and data field table, showing the common data types and structures of the environmental control system:
[0017] Next, S1 describes the interrelationships and dependencies among agricultural environmental control equipment, crop growth, and environmental parameters, including the following steps: Based on different agricultural environmental control equipment, crop growth data under different key environmental parameters are collected, and the data are cleaned and standardized; Divide the crop growth data into a training set and a test set. The training set is used to train the model, and the test set is used to evaluate the performance of the model. A multivariate linear regression model was used to train the effects of multiple key environmental parameters on crop growth, and the linear relationship between key environmental parameters and crop growth was verified using a test set.
[0018] Specifically, the multivariate linear regression model is used to predict the relationship between crop growth and multiple environmental parameters, expressed as: in, is the response variable of crop growth (such as crop height, weight, etc.), are key environmental parameters (such as temperature, humidity, light intensity, etc.), is the regression coefficient, It is an error term. When using the test set to evaluate the performance of the model, commonly used evaluation indicators include mean square error (MSE), determination coefficient (R²), etc., which can determine the accuracy of the relationship between the predicted crop growth and multiple environmental parameters. When the environmental parameters are input into the multiple linear regression model, the output result is compared with the crop growth conditions recorded in the training set. If the deviation is lower than the error threshold, an incorrect signal is output. When the environmental parameters in multiple training sets are input, if the proportion of incorrect signals is higher than the proportion threshold, it means that the previous multiple linear regression model is inaccurate. More crop growth data under different key environmental parameters can be collected and the multiple linear regression model can be retrained until the proportion of incorrect signals is lower than the proportion threshold, thereby improving the accuracy of predicting crop growth.
[0019] S2. Use data encryption technology to encrypt data, and after identity authentication and permission control authorization, use a secure transmission protocol to interconnect the environmental control system formed by various agricultural environmental control equipment on the same platform to ensure data security during interconnection. Whether it is data storage, exchange or sharing, it is carried out under a strict security mechanism, effectively preventing data leakage and tampering, and providing reliable data protection for agricultural production management; In addition, the data encryption technology in S2, MD5 is a widely used hash function that can convert input (plaintext) of any length into output (ciphertext) of fixed length (128 bits / 16 bytes). The output generated by the MD5 algorithm is usually expressed in 32-bit hexadecimal numbers, including the following steps: From the ciphertext encrypted by MD5, two parts are intercepted according to the predefined number of bits, including intercepting a string of specified bits from the beginning of the ciphertext, which is defined as password A, and intercepting the remaining string of specified bits from the position after password A, which is defined as password B; Use a random function to generate a fixed-length key to fill in the password B part; Combining the password A and the generated key to form a final transformed password value, and using the final transformed password value to compress a data packet, wherein the data packet is used to store the multivariate linear regression model data corresponding to different agricultural environmental control equipment to form multiple confidential data packets; The specific principle is: for example, the encrypted bit value is six digits, and the user uses the last six digits of the ID number 165826 as the plain text password, then the plain text password is encrypted with md5 to obtain the ciphertext md5 password, and the bit value is intercepted from the starting position of the ciphertext. The bit value can be four digits, then the intercepted password A is 1658, and then the value of the remaining two digits after intercepting the four-digit bit value is defined as B, and a random function is used to generate a key to fill B. If the key generated by the random function is 42, passwords A and B are combined to obtain the transformed password value 165842, so that the password can be obtained. Through interception and combination operations, the complexity of the password is increased, making the password more difficult to guess, ensuring the randomness and unpredictability of the random number generator to improve security.
[0020] Furthermore, S2 uses a secure transmission protocol to interconnect the environmental control systems formed by various agricultural environmental control equipment on the same platform, including the following steps: After identity authentication and permission control authorization, the application layer of the environmental control data model architecture perceives the data packets that need to be transmitted; At the transport layer, packets are encapsulated into transport layer data units and transmitted down the protocol stack, with each layer adding a header (the header contains information such as source port, destination port, sequence number, acknowledgment number, etc.); At the network layer, the data packet is encapsulated into a data unit of the network layer (such as an IP datagram). The network layer adds a header `H_N`, which contains information such as the source IP address, destination IP address, TTL (time to live), and protocol type. At the data link layer, the data packet is encapsulated into a data unit of the data link layer (such as an Ethernet frame). The data link layer adds a header `H_L` and a trailer `F_L`. The header contains information such as the source MAC address, destination MAC address, and type, and the trailer contains a frame check sequence (FCS); The encapsulated content is passed as data to the next layer until it reaches the physical layer. At the physical layer, the data is converted into a bit stream and transmitted through the medium. The physical layer does not add headers or trailers, but only converts the data unit of the data link layer into a bit stream. The data packet passes through the application layer, transport layer, network layer, data link layer, and finally reaches the physical layer, where it is converted into a bit stream for transmission. The header and trailer of each layer contain the necessary control information to ensure the correctness and security of the data during transmission, which is conducive to the interconnection and interoperability of data packets on the same platform.
[0021] It is worth noting that the smart sensors in S3 include but are not limited to temperature sensors, humidity sensors, light sensors, soil moisture sensors, carbon dioxide sensors and wind speed sensors. Choose a suitable installation location to ensure that the sensors can accurately collect data. For example, the temperature sensor should avoid direct exposure to sunlight, and the humidity sensor should be installed close to the crops. According to the sensor type and environmental conditions, choose a suitable installation method. For example, the soil moisture sensor can be buried in the soil, and the light sensor can be installed above the crops. At the same time, use wired or wireless methods to connect the sensor to the data acquisition system. Common wireless communication protocols include Wi-Fi, LoRa, Zigbee, etc., and ensure that the sensor has a stable power supply. It can be powered by batteries or solar energy. Use a data logger (such as Raspberry Pi, Arduino, etc.) as the central node of the data acquisition system.
[0022] S3. Deploy smart sensors to collect environmental data in real time, and use time series analysis and machine learning algorithms to predict future trends in environmental parameters. Automatically adjust environmental parameters based on the prediction results, which is conducive to real-time monitoring of agricultural environmental parameters, automatic adjustment of environmental control equipment, optimization of production management, and effective improvement of facility agriculture production efficiency and resource utilization, reducing human intervention, and improving the automation and precision of agricultural production.
[0023] Specifically, S3 uses time series analysis and machine learning algorithms to predict future trends in environmental parameters, including the following steps: Collect historical environmental data, including but not limited to temperature, humidity, wind speed, precipitation, etc. The data can be obtained from sources such as weather stations, satellite data, sensors, etc., extract time-related features such as hours, days, weeks, months, quarters, etc., and introduce lag features, that is, data from the previous few time points, to capture the autocorrelation of the time series. Use sliding window technology to convert time series data into a supervised learning problem; Divide the dataset into a training set and a test set, using a time order to ensure that the training set comes first and the test set comes last. Use the training set to train the machine learning model, adjust the hyperparameters to optimize the model performance, and use cross-validation techniques based on the test set to evaluate the generalization ability of the model. By inputting time points into the machine learning model to predict environmental parameters, time series analysis and machine learning algorithms can be used effectively to predict future trends in environmental parameters. Time series graphs of predicted and actual values can also be plotted to intuitively demonstrate the model's prediction effect.
[0024] In addition, the environmental parameters are automatically adjusted based on the prediction results, including the following steps: The adjustment time point is preset, and the environmental parameters at this time point are input into the multivariate linear regression model, the crop growth is output, and the crop growth status is judged. If the crop growth is normal, the agricultural environmental control equipment is not triggered. If the crop growth is abnormal, the PID controller is used to trigger the agricultural environmental control equipment to adjust the environmental parameters. The environmental control equipment is adjusted according to the prediction results until the crop growth is normal. Through the feedback mechanism, when the environmental parameters are adjusted to a stage value, the crop growth status is re-predicted to ensure that the final crop growth status is normal, which is conducive to the realization of automatic environmental adjustment based on the data model; When judging the growth status of crops, a threshold is set. If the crop growth is ≥ the threshold, the output is normal growth. If the crop growth is < the threshold, the output is abnormal growth. This enables more accurate adaptation to crop growth when adjusting environmental control equipment and provides precise environmental adjustment suggestions.
[0025] Building an open environmental control data model architecture is conducive to having data structure standards between different environmental control systems, achieving unified standards in data structure, and using secure transmission protocols to enable environmental control systems formed by various agricultural environmental control equipment to be interconnected on the same platform, ensuring data security during interconnection. Whether it is data storage, exchange or sharing, it is carried out under a strict security mechanism, effectively preventing data leakage and tampering, and providing reliable data protection for agricultural production management; at the same time, it is conducive to real-time monitoring of agricultural environmental parameters, automatic adjustment of environmental control equipment, optimization of production management, effectively improving the production efficiency and resource utilization of facility agriculture, reducing human intervention, and improving the automation and precision level of agricultural production.
[0026] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A method for constructing a facility agriculture metadata model based on big data analysis, characterized in that: The following steps are involved: S1. Build an open environmental control data model architecture to define the basic data structure and data fields of various agricultural environmental control equipment, and describe the relationship and dependency between agricultural environmental control equipment, crop growth and environmental parameters; S2. Use data encryption technology to encrypt data, and after identity authentication and permission control authorization, use a secure transmission protocol to interconnect the environmental control systems formed by various agricultural environmental control equipment on the same platform; S3. Deploy smart sensors to collect environmental data in real time, and use time series analysis and machine learning algorithms to predict future trends in environmental parameters, and automatically adjust environmental parameters based on the prediction results.
2. The method for constructing a facility agriculture metadata model based on big data analysis according to claim 1, characterized in that: The basic data structure and data fields are used to determine key environmental parameters commonly found in facility agriculture, including temperature, humidity, light intensity, water pump status, irrigation system status, equipment fault status, and data timestamp.
3. The method for constructing a facility agriculture metadata model based on big data analysis according to claim 2, characterized in that: The S1 describes the relationship and dependence between agricultural environmental control equipment, crop growth and environmental parameters, including the following steps: Based on different agricultural environmental control equipment, crop growth data under different key environmental parameters are collected, and the data are cleaned and standardized; Divide the crop growth data into a training set and a test set. The training set is used to train the model, and the test set is used to evaluate the performance of the model. A multivariate linear regression model was used to train the effects of multiple key environmental parameters on crop growth, and the linear relationship between key environmental parameters and crop growth was verified using a test set.
4. The method for constructing a facility agriculture metadata model based on big data analysis according to claim 3 is characterized in that: The multivariate linear regression model is used to predict the relationship between crop growth and multiple environmental parameters, which is expressed as: in, is the response variable of crop growth, is a key environmental parameter, is the regression coefficient, is the error term.
5. The method for constructing a facility agriculture metadata model based on big data analysis according to claim 4 is characterized in that: The data encryption technology in S2 includes the following steps: From the ciphertext encrypted by MD5, two parts are intercepted according to the predefined number of bits, including intercepting a string of specified bits from the beginning of the ciphertext, which is defined as password A, and intercepting the remaining string of specified bits from the position after password A, which is defined as password B; Use a random function to generate a fixed-length key to fill in the password B part; The password A and the generated key are combined to form a final transformed password value, and the final transformed password value is used to compress the data packet, wherein the data packet is used to store the multivariate linear regression model data corresponding to different agricultural environmental control equipment to form multiple confidential data packets.
6. The method for constructing a facility agriculture metadata model based on big data analysis according to claim 5 is characterized in that: The S2 adopts a secure transmission protocol to interconnect the environmental control systems formed by various agricultural environmental control equipment on the same platform, including the following steps: After identity authentication and permission control authorization, the application layer of the environmental control data model architecture perceives the data packets that need to be transmitted; At the transport layer, packets are encapsulated into transport layer data units and transmitted down the protocol stack, with each layer adding a header; At the network layer, packets are encapsulated into data units at the network layer. At the data link layer, packets are encapsulated into data units at the data link layer. The encapsulated content is passed as data to the next layer until it reaches the physical layer, where the data is converted into a bit stream and transmitted through the medium.
7. The method for constructing a facility agriculture metadata model based on big data analysis according to claim 6 is characterized in that: The smart sensors in S3 include but are not limited to temperature sensors, humidity sensors, light sensors, soil moisture sensors, carbon dioxide sensors and wind speed sensors; Use the data logger as the central node of the data acquisition system.
8. The method for constructing a facility agriculture metadata model based on big data analysis according to claim 7 is characterized in that: The S3 uses time series analysis and machine learning algorithms to predict the changing trends of future environmental parameters, including the following steps: Collect historical environmental data, extract time-related features, and introduce lag features, i.e., data from previous time points, to capture the autocorrelation of the time series. Use sliding window technology to convert time series data into a supervised learning problem. Divide the dataset into a training set and a test set, using a time order to ensure that the training set comes first and the test set comes last. Use the training set to train the machine learning model, adjust the hyperparameters to optimize the model performance, and use cross-validation techniques based on the test set to evaluate the generalization ability of the model. Input time points to the machine learning model to predict environmental parameters.
9. The method for constructing a facility agriculture metadata model based on big data analysis according to claim 8, characterized in that: The method of automatically adjusting the environmental parameters based on the prediction results comprises the following steps: The adjustment time point is preset, and the environmental parameters at this time point are received and input into the multivariate linear regression model, crop growth is output, and the crop growth status is judged. If the crop growth is normal, the agricultural environmental control equipment will not be triggered. If the crop growth is abnormal, the PID controller will be used to trigger the agricultural environmental control equipment to adjust the environmental parameters. The environmental control equipment is adjusted according to the predicted results until the crop growth is normal. Through the feedback mechanism, when the environmental parameters are adjusted to a stage value, the crop growth status is re-predicted to ensure that the final crop growth condition is normal.