Aggregation access method, system and device for irrigation multi-source heterogeneous data and medium
By classifying and preprocessing the irrigated multi-source heterogeneous data, converged multi-source heterogeneous data and storing them into different types of databases, the problems of difficulty and low efficiency of data fusion in traditional methods are solved, efficient data storage and reading are achieved, and the operation efficiency of the irrigation system is improved.
Patent Information
- Application Number
- CN202510176333.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional methods cannot effectively fuse irrigation multi-source heterogeneous data, resulting in low fusion efficiency and the inability to directly find matching characteristics between data, which can easily lead to difficulty in data fusion, and hard fusion will lose or destroy part of the feature information of the data, affecting the operation of the irrigation system.
By classifying and preprocessing the irrigated multi-source heterogeneous data, the data is deduplicated, integrated and formatted according to the data type and correlation, forming converged multi-source heterogeneous data, and depositing them into the relational database MySQL and the non-relational database MongoDB respectively. Screen and extract relevant data according to business needs and separate data.
It realizes efficient storage and reading of irrigated multi-source heterogeneous data, solves the problems of difficulty and low efficiency in traditional methods, retains the characteristic information of the data, and improves the operation efficiency of the irrigation system.
Smart Images

Figure CN120030078A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing technology, and in particular to a method, system, device and medium for converging and accessing multi-source heterogeneous irrigation data. Background Art
[0002] With the development of information technology and the upsurge of digital transformation in the industry, the business data required by irrigation districts is increasing. Since the irrigation data required in irrigation districts come from a wide range of sources and has a variety of data types, the storage and access problems of multi-source heterogeneous irrigation data are gradually emerging.
[0003] Traditional data access methods mainly use machine learning or deep learning algorithms to fuse data, optimize the fused data and store it in the database. However, irrigation area data involves more multi-source heterogeneous data, such as meteorological data, sensor data, model prediction data, geographic information data, image data, etc. The formats and sizes of these data vary greatly. The traditional method of fusion is inefficient and cannot directly find matching features between data, which easily leads to data fusion difficulties. Hard fusion will also lose or destroy some feature information of the data, ultimately affecting the operation of the irrigation system. Summary of the invention
[0004] The purpose of the present invention is to provide a method, system, device and medium for the aggregation and access of multi-source heterogeneous irrigation data, which can solve the problem that traditional methods cannot fuse heterogeneous data and have low fusion efficiency.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A method for converging and accessing multi-source heterogeneous irrigation data, comprising:
[0007] Acquire multi-source heterogeneous irrigation data;
[0008] The irrigation multi-source heterogeneous data is stored, and the specific process is as follows:
[0009] Classifying the irrigation multi-source heterogeneous data according to data types and relevance to obtain classified data;
[0010] Preprocessing the classified data, deduplicating and integrating the same type of data with set correlation strength, and converting and integrating the different types of data with the same correlation strength to obtain aggregated multi-source heterogeneous data;
[0011] According to data types and relevance, the aggregated multi-source heterogeneous data are stored in the relational database MySQL and the non-relational database MongoDB respectively;
[0012] The irrigation multi-source heterogeneous data is extracted, and the specific process is as follows:
[0013] Determine the corresponding data screening conditions and data sources according to current business needs;
[0014] Query the data records or indexes in the data source, extract relevant data according to the data screening conditions, and separate the relevant data according to the business requirements.
[0015] Optionally, the classified data includes data of the same type with strong correlation, data of different types with strong correlation, data of the same type with weak correlation, and data of different types with weak correlation.
[0016] Optionally, deduplication and integration of data of the same type with set correlation strength specifically includes:
[0017] For the highly correlated data of the same type, the objects described by the data are extracted, and the different data are mapped to the various attributes of the object. During the mapping process, data deduplication and integration are performed to aggregate the data of the same type into a whole.
[0018] Optionally, the format conversion and data integration of different types of data with the same correlation strength specifically includes:
[0019] For the same type of data with strong correlation, for geojson format data and text data, the text data is converted into json format and inserted into the properties attribute in the geojson format data; for shapefile format data, it is converted into geojson format data; for text data and png image data, the two are combined and converted into tif format image files, the text data is written into the file header, and the png image data is written into the file body.
[0020] Optionally, determining corresponding data screening conditions and data sources according to current business needs specifically includes:
[0021] According to current business needs, determine whether the data source is MySQL or MongoDB, and set corresponding data filtering conditions according to the corresponding data source;
[0022] If the data source is MySQL, determine the object to which the data belongs based on the corresponding data table in MySQL; if the data source is MongoDB, determine the collection in which the data is stored based on the data format.
[0023] Optionally, querying the data records or indexes in the data source, extracting relevant data according to the data screening condition, and separating the relevant data according to the business requirements specifically includes:
[0024] When data needs to be obtained, the server uses conditional query SQL statements to query the specified data source and returns the queried data field results. Since the data has been aggregated before storage, if the business needs require the separate use of data, the data is separated or the fields are copied according to the reverse process of data integration to meet the requirements.
[0025] The present invention also provides a converged access system for irrigation multi-source heterogeneous data, comprising:
[0026] Data acquisition unit, used to obtain irrigation multi-source heterogeneous data;
[0027] The storage unit is used to store the irrigation multi-source heterogeneous data. The specific process is as follows:
[0028] Classifying the irrigation multi-source heterogeneous data according to data types and relevance to obtain classified data;
[0029] Preprocessing the classified data, deduplicating and integrating the same type of data with set correlation strength, and converting and integrating the different types of data with the same correlation strength to obtain aggregated multi-source heterogeneous data;
[0030] According to data types and relevance, the aggregated multi-source heterogeneous data are stored in the relational database MySQL and the non-relational database MongoDB respectively;
[0031] The extraction unit is used to extract the irrigation multi-source heterogeneous data. The specific process is as follows:
[0032] Determine the corresponding data screening conditions and data sources according to current business needs;
[0033] Query the data records or indexes in the data source, extract relevant data according to the data screening conditions, and separate the relevant data according to the business requirements.
[0034] The present invention also provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the above-mentioned aggregation access method for irrigation multi-source heterogeneous data.
[0035] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the above-mentioned method for convergent access of multi-source heterogeneous irrigation data.
[0036] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0037] The present invention discloses a method, system, device and medium for the aggregation access of multi-source heterogeneous irrigation data. The method includes obtaining multi-source heterogeneous irrigation data; further storing the data in the following steps: classifying the multi-source heterogeneous irrigation data according to data type and relevance, and preprocessing the data; performing data deduplication and integration on the same type of data with set relevance strength, and format conversion and data integration on different types of data with the same relevance strength, so as to obtain the aggregated multi-source heterogeneous data; storing the aggregated multi-source heterogeneous data in relational database MySQL and non-relational database MongoDB respectively according to data type and relevance; further extracting the data in the following steps: determining the corresponding data screening conditions and the data source to which it belongs according to the current business needs; querying the data records or indexes in the data source to which it belongs, extracting the relevant data according to the data screening conditions, and separating the relevant data according to the business needs. The present invention classifies the multi-source heterogeneous data according to factors such as data type and the corresponding relationship between the data, thereby filtering out the data that can be integrated for data aggregation, and finally accessing the irrigation data separately according to different data characteristics, so as to realize the storage and reading of the multi-source heterogeneous irrigation data. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0039] Figure 1 It is a flow chart of the method for converging and accessing multi-source heterogeneous irrigation data according to the present invention;
[0040] Figure 2 It is a specific application block diagram in this embodiment. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] The purpose of the present invention is to provide a method, system, device and medium for the aggregation and access of multi-source heterogeneous irrigation data, which can solve the problem that traditional methods cannot fuse heterogeneous data and have low fusion efficiency.
[0043] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] like Figure 1-Figure 2 As shown, the present invention provides a method for converging and accessing multi-source heterogeneous irrigation data, comprising:
[0045] Acquire multi-source heterogeneous irrigation data;
[0046] The irrigation multi-source heterogeneous data is stored, and the specific process is as follows:
[0047] Classifying the irrigation multi-source heterogeneous data according to data types and relevance to obtain classified data;
[0048] Preprocessing the classified data, deduplicating and integrating the same type of data with set correlation strength, and converting and integrating the different types of data with the same correlation strength to obtain aggregated multi-source heterogeneous data;
[0049] According to data types and relevance, the aggregated multi-source heterogeneous data are stored in the relational database MySQL and the non-relational database MongoDB respectively;
[0050] The irrigation multi-source heterogeneous data is extracted, and the specific process is as follows:
[0051] Determine the corresponding data screening conditions and data sources according to current business needs;
[0052] Query the data records or indexes in the data source, extract relevant data according to the data screening conditions, and separate the relevant data according to the business needs
[0053] As a specific implementation method, the specific processing process of the above steps is as follows:
[0054] Step 1: Obtain the collected irrigation multi-source heterogeneous data. Irrigation multi-source heterogeneous data include irrigation area meteorological and soil moisture data, sensor collection data, equipment status data, model prediction result data, geographic information data, drone aerial satellite image data, etc. Since the sources, types and structures of these data are different, they are called multi-source heterogeneous data.
[0055] Multi-source heterogeneous data is mainly collected through various terminal devices such as sensors, satellites, and drones. Since the communication transmission protocols supported by each terminal device are different, a message queue is used for asynchronous data transmission from the terminal to the server. Common message queues include RocketMQ, Kafka, RabbitMQ, etc. This method intends to use RocketMQ as the message queue. Each terminal device sends the collected data to the message queue for transfer through the communication protocol it supports (such as HTTP, MQTT, AMQP, etc.). The server reads the collected data through the message queue to complete the acquisition of multi-source heterogeneous data.
[0056] Step 2: Classify multi-source heterogeneous data by data type and correlation strength. Multi-source heterogeneous data usually contains various data types, such as text data, json data, shapefile geographic information data, geojson geographic information data, png images, tif images, etc. The above different types of data are likely to be descriptions of different characteristics and different levels of the same thing, that is, there is correlation between the data. For example, for the experimental fields in the irrigation area, there are text data and json data used to describe its basic information, geographic information data describing its geographical location, geojson data describing the specific plot outlines, data describing the crops planted in the experimental fields, data reflecting the meteorological changes in the experimental fields, and real satellite tif format images of the experimental fields. The above data are strongly correlated. The irrigation multi-source heterogeneous data can be classified according to data type and correlation strength, and can be initially divided into four categories: data of the same type with strong correlation, data of different types with strong correlation, data of the same type with weak correlation, and data of different types with weak correlation.
[0057] Step 3: Preprocess the classified data to obtain aggregated multi-source heterogeneous data. In this method, "preprocessing" mainly refers to data format conversion, data integration and other operations on the classified data. Specifically, for data of the same type with strong correlation, the object described by the data can be extracted, and different data can be mapped to the various attributes of the object. Data deduplication is performed during the mapping process, and finally the data of the same type are aggregated into a whole data through data integration. For example, for the object of the experimental field in the irrigation area, there are two json data describing the basic information of the experimental field and the planting conditions of the experimental field crops. Both data include the name and number of the experimental field. Then, only one copy of the same part of the two data can be retained, and the different parts can be aggregated and integrated together to form a more detailed data for easy storage; for multiple picture files with the same format, the features of these pictures are extracted and compared through the deep neural network structure, and the duplicate features are removed and data fusion is performed. For different types of data with strong correlation, in addition to the necessary data deduplication and data integration, data format conversion should also be performed. For example, for geojson format data and text data, the text data can be converted to json format and inserted into the properties attribute of geojson format data; for shapefile format data, it can be converted to geojson format data for easy storage and call by the web server; for text data and png image data, the two can be combined and converted into tif format image files, the file header is written with text data, and the file body is written with png image data, so as to achieve the purpose of data preprocessing. For data with weak correlation, it can be regarded as integrated and aggregated data, so it will not be processed in this step.
[0058] Step 4: Store the preprocessed data in the relational database MySQL and the non-relational database MongoDB according to different data types and correlation strengths. The preprocessed multi-source heterogeneous irrigation data is basically unified aggregate data. According to different data formats, it can be divided into relational data (such as digital and text description information of an object) and non-relational data (such as pictures, audio, video, geographic information data and other large files in specific formats). Objects composed of relational data can be converted into tables containing different attribute columns in a relational database. The corresponding relationship between objects is the foreign key or index in the relational database; due to its special data type and format, non-relational data is more suitable for storage using data models such as key-value pairs, documents, and column families in NoSQL databases.
[0059] For relational data, MySQL database is selected for storage. MySQL is a popular relational database management system (RDBMS) that manages databases based on SQL (Structured Query Language). MySQL supports ACID (atomicity, consistency, isolation, and durability) transactions to ensure the reliability of transaction processing; it supports multiple index types, such as B-tree, Hash, R-tree, etc., to improve query efficiency. MySQL is one of the most popular database systems in the world.
[0060] For non-relational data, MongoDB database is selected for storage. MongoDB is a document-based NoSQL database, which is famous for its high performance, high availability and easy scalability. MongoDB converts data into BSON (binary JSON) documents for storage, which can contain a variety of data types, such as strings, numbers, arrays, objects, etc. Documents in MongoDB are organized in collections, similar to tables in relational databases, but they do not need to have a fixed schema, which means that documents in a collection can have different structures. Due to its above characteristics, MongoDB is suitable for storing semi-structured or nested data. In addition, MongoDB also has a GridFS module for storing and retrieving large files (such as pictures) in blocks.
[0061] Step 5: According to the current business needs, determine the data screening conditions and data sources that meet the requirements. Business mainly refers to the specific business areas of the irrigation district in reality, and the data requirements required by each business are different. For a specific business, determine the specific data range required by the business, determine the data source according to the data type, and set the data screening conditions. Specifically, first determine whether the data is stored in MySQL or MongoDB. If it is placed in MySQL, further determine the object to which the data belongs and find the corresponding data table in MySQL; if it is placed in MongoDB, determine the collection (Collections) where the data is stored according to the data format. Then determine the screening conditions based on the specific business needs, so as to selectively filter out the specified data in the data source, such as which part of the records or the values of which attribute columns in a certain MySQL data table; which specific one or several BSON files are stored in the MongoDB collection, etc.
[0062] Step 6: Query the data records or indexes in the data source, obtain relevant data based on the filtering conditions of the required data, and further separate the data according to business needs. Establish a corresponding relationship between the data filtering conditions and data sources determined by specific business needs. When data needs to be obtained, the server uses conditional query SQL statements to query the specified data source and returns the query data field results. Since the data has been aggregated before storage, if the business needs to use the data separately, data separation or field replication is performed according to the reverse process of data integration to meet the requirements.
[0063] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0064] This article uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only used to help understand the core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A method for accessing multi-source heterogeneous irrigation data, characterized in that: include: Acquire multi-source heterogeneous irrigation data; The irrigation multi-source heterogeneous data is stored, and the specific process is as follows: Classifying the irrigation multi-source heterogeneous data according to data types and relevance to obtain classified data; Preprocessing the classified data, deduplicating and integrating the same type of data with set correlation strength, and converting and integrating the different types of data with the same correlation strength to obtain aggregated multi-source heterogeneous data; According to data types and relevance, the aggregated multi-source heterogeneous data are stored in the relational database MySQL and the non-relational database MongoDB respectively; The irrigation multi-source heterogeneous data is extracted, and the specific process is as follows: Determine the corresponding data screening conditions and data sources according to current business needs; Query the data records or indexes in the data source, extract relevant data according to the data screening conditions, and separate the relevant data according to the business requirements.
2. The method for converging and accessing multi-source heterogeneous irrigation data according to claim 1, characterized in that: The classified data includes data of the same type with strong correlation, data of different types with strong correlation, data of the same type with weak correlation, and data of different types with weak correlation.
3. The method for converging and accessing multi-source heterogeneous irrigation data according to claim 2, characterized in that: The deduplication and integration of the same type of data with set correlation strength specifically includes: For the highly correlated data of the same type, the objects described by the data are extracted, and the different data are mapped to the various attributes of the object. During the mapping process, data deduplication and integration are performed to aggregate the data of the same type into a whole.
4. The method for converging and accessing multi-source heterogeneous irrigation data according to claim 2, characterized in that: The format conversion and data integration of different types of data with the same correlation strength specifically includes: For the same type of data with strong correlation, for geojson format data and text data, the text data is converted into json format and inserted into the properties attribute in the geojson format data; for shapefile format data, it is converted into geojson format data; for text data and png image data, the two are combined and converted into tif format image files, the text data is written into the file header, and the png image data is written into the file body.
5. The method for converging and accessing multi-source heterogeneous irrigation data according to claim 1, characterized in that: Determining the corresponding data screening conditions and data sources according to the current business needs specifically includes: According to current business needs, determine whether the data source is MySQL or MongoDB, and set corresponding data filtering conditions according to the corresponding data source; If the data source is MySQL, determine the object to which the data belongs based on the corresponding data table in MySQL; if the data source is MongoDB, determine the collection in which the data is stored based on the data format.
6. The method for converging and accessing multi-source heterogeneous irrigation data according to claim 1, characterized in that: The querying of the data records or indexes in the data source, extracting relevant data according to the data screening conditions, and separating the relevant data according to the business requirements specifically includes: When data needs to be obtained, the server uses conditional query SQL statements to query the specified data source and returns the queried data field results. Since the data has been aggregated before storage, if the business needs require the separate use of data, the data is separated or the fields are copied according to the reverse process of data integration to meet the requirements.
7. A converged access system for irrigation multi-source heterogeneous data, characterized in that: include: Data acquisition unit, used to obtain irrigation multi-source heterogeneous data; The storage unit is used to store the irrigation multi-source heterogeneous data. The specific process is as follows: Classifying the irrigation multi-source heterogeneous data according to data types and relevance to obtain classified data; Preprocessing the classified data, deduplicating and integrating the same type of data with set correlation strength, and converting and integrating the different types of data with the same correlation strength to obtain aggregated multi-source heterogeneous data; According to data types and relevance, the aggregated multi-source heterogeneous data are stored in the relational database MySQL and the non-relational database MongoDB respectively; The extraction unit is used to extract the irrigation multi-source heterogeneous data. The specific process is as follows: Determine the corresponding data screening conditions and data sources according to current business needs; Query the data records or indexes in the data source, extract relevant data according to the data screening conditions, and separate the relevant data according to the business requirements.
8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the method for converging and accessing irrigation multi-source heterogeneous data according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The device stores a computer program, which, when executed by a processor, implements the method for converging and accessing irrigation multi-source heterogeneous data as described in any one of claims 1 to 6.