Crop phenotype data storage and retrieval method, device, equipment and storage medium
By classifying and storing crop phenotypic data through a hybrid architecture storage and retrieval system, the problems of low storage efficiency and slow retrieval speed in existing technologies are solved, achieving efficient data storage and retrieval.
Patent Information
- Application Number
- CN202511212009.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing centralized storage and retrieval methods suffer from low storage efficiency and slow retrieval speed when processing large-scale and multi-type crop phenotypic data, and data consistency is difficult to guarantee.
A hybrid architecture storage and retrieval system is adopted, including a distributed database, a distributed search and analysis engine, and an in-memory database. Multimodal crop phenotypic data is classified and stored, with structured and unstructured data stored in different blocks. When a user searches, the search results are retrieved from different blocks.
It improves the storage efficiency and retrieval speed of multimodal crop phenotypic data, and ensures data consistency and flexibility.
Smart Images

Figure CN120723776B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data storage technology, and in particular to a method, apparatus, device, and storage medium for storing and retrieving crop phenotypic data. Background Technology
[0002] Crop phenotypic data is characterized by its large volume, diverse types, varying storage requirements, and different access characteristics. Existing centralized storage and retrieval methods suffer from low storage efficiency, slow retrieval speed, and difficulty in ensuring data consistency when processing large-scale and multi-type crop phenotypic data. Therefore, achieving efficient storage and rapid retrieval of crop phenotypic data has become an urgent technical problem to be solved. Summary of the Invention
[0003] This invention provides a method, apparatus, device, and storage medium for storing and retrieving crop phenotypic data, in order to solve the problems of low storage efficiency and slow retrieval speed in existing data storage and retrieval methods.
[0004] This invention provides a method for storing and retrieving crop phenotypic data, applied to a hybrid architecture storage and retrieval system; the hybrid architecture storage and retrieval system includes a distributed database, a distributed search and analysis engine, a distributed file system, and a memory database; the method includes:
[0005] Acquire multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; the crop phenotypic data includes structured data and unstructured data; the unstructured data includes small file data and large file data;
[0006] The structured data and metadata are stored in the distributed database, and the unstructured data is stored in the distributed file system.
[0007] The retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data are synchronously stored in the distributed search and analysis engine;
[0008] If no data matching the user's search criteria is found in the in-memory database, the search results are determined based on the user's search criteria and the distributed search analysis engine; the in-memory database is used to cache the structured data and the metadata information of the unstructured data.
[0009] According to a crop phenotypic data storage and retrieval method provided by the present invention, the metadata information of the structured data includes crop identifiers and other crop information; storing the structured data and the metadata information in the distributed database includes:
[0010] The first row key and the first column family are filled into the first data table of the distributed database; the first row key is determined based on the crop identifier; the first column family is determined based on the structured data and the other crop information.
[0011] According to a crop phenotypic data storage and retrieval method provided by the present invention, the step of storing the unstructured data in the distributed file system includes:
[0012] Iterate through the unstructured data to be stored to obtain the file size of the currently uploaded file;
[0013] If the file size is greater than a first threshold, the currently uploaded file is compressed and stored in the distributed file system; if the file size is less than the first threshold but greater than a second threshold, the currently uploaded file is placed in an empty merge queue; if the file size is less than or equal to the second threshold, the currently uploaded file is placed in a temporary queue.
[0014] Select the largest file from the temporary queue and determine the maximum remaining space in each of the merge queues;
[0015] If the largest file is greater than the maximum remaining space, the largest file is placed in the merge queue obtained by converting the spare queue; if the largest file is less than or equal to the maximum remaining space, the largest file is placed in the merge queue corresponding to the maximum remaining space.
[0016] If the files in the merge queue meet the preset conditions, the files in the merge queue are merged and stored in the distributed file system.
[0017] According to a crop phenotypic data storage and retrieval method provided by the present invention, the crop phenotypic data storage and retrieval method further includes:
[0018] Establish a mapping relationship between the files before merging and the files after merging;
[0019] The second row key and the second column family are filled into the second data table of the distributed database; the second row key is determined based on the file before merging; the second column family is determined based on the metadata information of the unstructured data and the mapping relationship.
[0020] According to a crop phenotypic data storage and retrieval method provided by the present invention, the step of synchronously storing the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data into the distributed search analysis engine includes:
[0021] When the structured data is retrieved based on user search criteria and the search results exist in the in-memory database, the search results are loaded into the local machine.
[0022] When the structured data is retrieved based on user search criteria and no search results are found in the in-memory database, the target row key corresponding to the user search criteria is determined by the distributed search analysis engine, the search results are obtained from the distributed database based on the target row key, the search results are loaded into the local database, and the in-memory database is updated based on the loaded results.
[0023] In the case of retrieving the unstructured data based on user search criteria, metadata results are retrieved from the in-memory database, and search data is obtained from the distributed file system based on the metadata results, and the search data is loaded locally.
[0024] According to a crop phenotypic data storage and retrieval method provided by the present invention, the crop phenotypic data storage and retrieval method further includes:
[0025] Determine the access information for each cache key in the memory database; the access information includes access frequency, last most recent access time, and current access time;
[0026] Based on the access information of each cache key, adjust the dynamic weight value of each cache key;
[0027] The cached data in the memory database is dynamically adjusted based on the adjusted weight values.
[0028] The present invention also provides a crop phenotypic data storage and retrieval device, comprising the following modules:
[0029] A multimodal data acquisition module is used to acquire multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; the crop phenotypic data includes structured data and unstructured data; the unstructured data includes small file data and large file data;
[0030] The first storage module is used to store the structured data and the metadata information in a distributed database, and to store the unstructured data in a distributed file system;
[0031] The second storage module is used to synchronously store the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data to the distributed search and analysis engine.
[0032] The retrieval module is used to determine retrieval results based on the user's retrieval conditions and the distributed search analysis engine when no data matching the user's retrieval conditions exists in the in-memory database; the in-memory database is used to cache the metadata information of the structured data and the unstructured data.
[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the crop phenotypic data storage and retrieval method as described above.
[0034] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the crop phenotypic data storage and retrieval method as described above.
[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the crop phenotypic data storage and retrieval method as described above.
[0036] The crop phenotypic data storage and retrieval method, apparatus, device, and storage medium provided by this invention classify and store multimodal crop phenotypic data by constructing a hybrid architecture storage and retrieval system. Structured and unstructured crop phenotypic data are stored in different blocks, and when a user retrieves data, the system queries results from different blocks based on the user's search criteria. By classifying and storing multimodal crop phenotypic data according to different data types, the storage efficiency and retrieval speed of multimodal crop phenotypic data are improved. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0038] Figure 1 This is one of the flowcharts illustrating the crop phenotypic data storage and retrieval method provided by the present invention.
[0039] Figure 2 This is the second flowchart of the crop phenotypic data storage and retrieval method provided by the present invention.
[0040] Figure 3 This is a schematic diagram of the structure of the crop phenotypic data storage and retrieval device provided by the present invention.
[0041] Figure 4This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0043] The following is combined with Figures 1-4 This invention describes a method, apparatus, device, and storage medium for storing and retrieving crop phenotypic data.
[0044] Figure 1 This is one of the flowcharts illustrating the crop phenotypic data storage and retrieval method provided by the present invention, such as... Figure 1 As shown, this method is applied to a hybrid architecture storage and retrieval system; the hybrid architecture storage and retrieval system includes a distributed database, a distributed search and analysis engine, a distributed file system, and a memory database; the crop phenotypic data storage and retrieval method provided by this invention includes the following:
[0045] Step 100: Obtain the multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; the crop phenotypic data includes structured data and unstructured data; the unstructured data includes small file data and large file data;
[0046] Specifically, before storing the multimodal crop phenotypic data, it is preprocessed. The preprocessing process includes cleaning, classifying, denoising, and format standardization. Multimodal crop phenotypic data includes structured crop phenotypic data (e.g., sowing date, maturity date, plant height, yield, etc.) and unstructured crop phenotypic data (e.g., images, point clouds, spectra, and thermal imaging, etc.). The multimodal data to be stored includes the aforementioned multimodal crop phenotypic data and its corresponding metadata (e.g., crop name, plant ID, and variety, etc.). The multimodal data to be stored also includes environmental data (e.g., maximum temperature, minimum temperature, sunshine duration, and precipitation, etc.). Unstructured crop phenotypic data can be further divided into small file data and large file data.
[0047] Step 200: Store the structured data and the metadata information in the distributed database, and store the unstructured data in the distributed file system;
[0048] Specifically, the crop phenotypic data storage and retrieval method provided by this invention is applied to a hybrid architecture storage and retrieval system including distributed databases, distributed search and analysis engines, distributed file systems, and in-memory databases.
[0049] The structured crop phenotypic data and metadata are stored in a distributed database (such as HBase, a distributed, column-oriented open-source database); unstructured crop phenotypic data is stored differently based on file size. Files of unstructured crop phenotypic data larger than a certain threshold are compressed and directly uploaded to a distributed file system, such as the core distributed file system of the Apache Hadoop ecosystem (Hadoop Distributed File System, HDFS); files of unstructured crop phenotypic data smaller than a certain threshold are merged into smaller files before being uploaded to the distributed file system.
[0050] Step 300: Synchronously store the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data into the distributed search analysis engine;
[0051] Specifically, the distributed search and analysis engine (e.g., Elasticsearch, a distributed search and analysis engine that supports real-time data processing, full-text search, and multidimensional aggregation analysis) in the hybrid architecture storage and retrieval system is used to store the retrieval fields (and their corresponding row keys) of structured crop phenotypic data, the retrieval fields (and their corresponding row keys) of small file data, and the metadata information of large file data, so as to realize diversified retrieval of multimodal crop phenotypic data.
[0052] By using a real-time synchronization tool between the distributed database and the distributed search and analysis engine, retrieval fields of structured crop phenotypic data (such as crop name, variety, sowing date, and plant ID) can be synchronized from the distributed database to the distributed search and analysis engine. The row key of the distributed database table serves as the unique identifier for the document in the distributed search and analysis engine (i.e., synchronizing the row key of the HBase table to Elasticsearch). This not only ensures data consistency and uniqueness but also enables multi-condition combined queries based on crop phenotypic data, thereby improving the efficiency and flexibility of data retrieval and analysis.
[0053] Step 400: If no data matching the user's search criteria exists in the in-memory database, the search results are determined based on the user's search criteria and the distributed search analysis engine; the in-memory database is used to cache the metadata information of the structured data and the unstructured data.
[0054] Specifically, an in-memory database (e.g., Redis) can store structured crop phenotypic data itself, as well as metadata information for unstructured data. Based on the user's search criteria, if the in-memory database cache is hit, the data is returned directly; otherwise, a diversified search is performed.
[0055] The distributed search and analysis engine offers diverse retrieval capabilities: it returns a list of row keys that match the user's diverse search criteria; and it performs batch retrieval from a distributed database: it loads and returns complete data from a table (crop_structuring) in the distributed database based on the row key. Finally, the loaded data is stored in an in-memory database. Other retrieval methods are described in detail below.
[0056] This embodiment constructs a hybrid architecture storage and retrieval system to classify and store multimodal crop phenotypic data. Structured and unstructured crop phenotypic data are stored in different blocks. When a user searches, the system retrieves results from different blocks based on the user's search criteria. By classifying and storing multimodal crop phenotypic data according to different data types, the storage efficiency and retrieval speed of multimodal crop phenotypic data are improved.
[0057] In one embodiment, the metadata information of the structured data includes crop identifiers and other crop information. The crop phenotypic data storage and retrieval method provided in this embodiment of the invention may further include:
[0058] Step 210: Fill the first row key and the first column family into the first data table of the distributed database; the first row key is determined based on the crop identifier; the first column family is determined based on the structured data and the other crop information.
[0059] Specifically, the distributed database includes structured data tables (crop_structuring) and unstructured data tables (crop_unstructuring). The structured data table (i.e., the first data table in this embodiment): the first row key is the hashed crop ID (i.e., the crop identifier in this embodiment), and the first column family stores structured crop phenotypic data (e.g., sowing date, maturity date, plant height, and yield). The unstructured data table is described in detail below.
[0060] This embodiment uses structured data tables to classify and store structured crop phenotypic data.
[0061] In one embodiment, the crop phenotypic data storage and retrieval method provided by this invention may further include:
[0062] Step 220: Traverse the unstructured data to be stored to obtain the file size of the currently uploaded file;
[0063] Step 230: If the file size is greater than the first threshold, compress and store the currently uploaded file in the distributed file system; if the file size is less than the first threshold but greater than the second threshold, place the currently uploaded file in an empty merge queue; if the file size is less than or equal to the second threshold, place the currently uploaded file in a temporary queue.
[0064] Step 240: Select the largest file from the temporary queue and determine the maximum remaining space in each of the merge queues;
[0065] Step 250: If the largest file is greater than the maximum remaining space, place the largest file into the merge queue obtained by converting the spare queue; if the largest file is less than or equal to the maximum remaining space, place the largest file into the merge queue corresponding to the maximum remaining space.
[0066] Step 260: If the files in the merge queue meet the preset conditions, merge the files in the merge queue and store them in the distributed file system.
[0067] Specifically, the storage process for small file data in unstructured crop phenotypic data is as follows.
[0068] Iterate through the set of unstructured crop phenotypic data files to be uploaded to the distributed file system. If the size of the currently uploaded file in the set is greater than the data block threshold (e.g., 128MB), proceed accordingly. ( If the file size of the currently uploaded file exceeds the first threshold by a factor of 1, then the currently uploaded file is directly stored in the distributed file system, and the remaining files in the file set are traversed.
[0069] Based on the total file size of the file set to be uploaded and the file merging threshold, determine the number of merging queues A, the number of standby queues B, and the number of temporary queues C. B and C are much smaller than A. Then, initialize the corresponding merging queues, standby queues, and temporary queues as needed.
[0070] If the size of the currently uploaded file is greater than the minimum file size threshold (e.g., 1MB) (i.e., the second threshold) but less than the first threshold, then the currently uploaded file is added to an empty merge queue, and the remaining files in the file set are traversed. If the size of the currently uploaded file is not greater than the second threshold, then the currently uploaded file is placed in a temporary queue.
[0071] The largest file in the temporary queue (the largest file in this embodiment) is compared with the maximum remaining space in the current merge queue (the maximum remaining space in this embodiment). If the largest file is larger than the maximum remaining space, it means that the remaining space in all current merge queues cannot accommodate the file. The largest file is then added to an empty standby queue, which is converted into a new merge queue, and the new merge operation continues. If the largest file is less than or equal to the maximum remaining space, the largest file is added to the merge queue corresponding to that maximum remaining space.
[0072] Calculate the total file size occupied by the current merge queue. If the total size reaches the data block threshold... ( When the number of files in the merge queue reaches 10 times the preset condition (i.e., the files in the merge queue meet the preset condition), all files in the merge queue are packaged and stored in the data block of the distributed file system cluster node, and then the merge queue is destroyed; otherwise, a new file merge operation is performed.
[0073] This embodiment uses a large file storage area and a small file merging area in a distributed file system to classify and store unstructured crop phenotypic data, thereby improving the flexibility of crop phenotypic data storage.
[0074] In one embodiment, the crop phenotypic data storage and retrieval method provided by this invention may further include:
[0075] Step 270: Establish the mapping relationship between the files before merging and the files after merging;
[0076] Step 280: Fill the second row key and the second column family into the second data table of the distributed database; the second row key is determined based on the file before merging; the second column family is determined based on the metadata information of the unstructured data and the mapping relationship.
[0077] Specifically, a mapping relationship is established between small files (i.e., files before merging) and merged files (i.e., files after merging). This mapping relationship includes filename, file position, offset, and file size. An Elastic Hash Table (EHT) is used to construct n buckets, each storing the index information of k small files. Simultaneously, the file index information in each bucket is sorted lexicographically according to the hash values of the filenames. Then, a Minimal Perfect Hashing with Chaining (MWHC) hash function is used to store the position information of each small file's index information. The MWHC hash function is written at the beginning of each index file.
[0078] An unstructured data table (an HBase table named crop_unstructuring, i.e., the second data table in this embodiment) stores small file index information and metadata information. The index data after merging small files (e.g., filename hash, file position, file size, and offset) is used to locate small file data; metadata information (e.g., crop name, plant ID, variety, and modality identifier) is used to retrieve small file data. Modality identifiers are mainly used to efficiently obtain small file data of the same modality type.
[0079] The second row key of the second data table can be a small filename; the second column family of the unstructured data table includes Metadata_info and Small_info. The Metadata_info column family contains crop name, plant ID, variety, and modality identifier, etc., and is used to assist in multi-condition combined retrieval of small file data; the Small_info column family contains filename (hash), file location (HDFS path), offset, and file size, and is mainly used to determine the specific location of small files in the merged file.
[0080] This embodiment uses unstructured data tables to classify and store unstructured crop phenotypic data and its metadata.
[0081] In one embodiment, the crop phenotypic data storage and retrieval method provided by this invention may further include:
[0082] Step 500: If the structured data is retrieved based on the user's search criteria and the in-memory database contains search results, the search results are loaded into the local machine.
[0083] Step 600: When the structured data is retrieved based on the user's search criteria and no search results are found in the in-memory database, the target row key corresponding to the user's search criteria is determined by the distributed search analysis engine, the search results are obtained from the distributed database based on the target row key, the search results are loaded into the local database, and the in-memory database is updated based on the loaded results.
[0084] Step 700: When retrieving the unstructured data based on user search criteria, retrieve metadata results from the in-memory database, obtain search data from the distributed file system based on the metadata results, and load the search data into the local machine.
[0085] Specifically, the user's search criteria are analyzed. When searching structured data based on the analysis results, the data is first retrieved from the in-memory database. If a match is found, the data is directly loaded into the local database. If no match is found, the row key corresponding to the specific search criteria (the target row key in this embodiment) is obtained through the distributed search analysis engine. Then, the data is retrieved from the distributed database using the target row key and loaded into the local database. At the same time, the in-memory database is updated based on the retrieved data.
[0086] When retrieving large file data (e.g., point cloud data and spectral data), the process first retrieves metadata information from the in-memory database. If a match is found, the system directly interacts with the DataNode (a node in the distributed file system) based on the metadata information to obtain the data and loads it locally. If a match is not found, the system retrieves metadata information from the distributed search and analysis engine, interacts with the DataNode again to obtain the data, loads it locally, and then loads the metadata information of that data into the in-memory database.
[0087] When retrieving small file data (e.g., image data and thermal imaging data), the corresponding metadata information is retrieved from the in-memory database based on specific search criteria. If a match is found, the data is directly loaded locally via interaction with the DataNode. If no match is found, the distributed search and analysis engine uses a diversified search to return the corresponding row key. Metadata information is then loaded from the unstructured data table (crop_unstructuring) using the row key, and the location information of the merged file blocks is obtained through interaction with the NameNode (a node in the distributed file system). Next, phenotypic data is retrieved through interaction with the DataNode and loaded locally. Finally, the corresponding merged file block information and small file metadata information are loaded into the in-memory database. The retrieved multimodal crop phenotypic data is then integrated and returned to the user. The metadata information of small files is mainly used to locate which merged file the small file is in and its position within the merged file. Interaction with the NameNode is primarily to obtain the location of the merged file (including how many blocks the merged file is divided into and which DataNode each block resides on).
[0088] This embodiment improves the retrieval efficiency of multimodal crop phenotypic data by using a hybrid architecture storage and retrieval system with categorized storage to perform diversified retrieval.
[0089] Figure 2 This is the second flowchart illustrating the crop phenotypic data storage and retrieval method provided by the present invention, as shown below. Figure 2 As shown, the method may further include:
[0090] Step 10: Determine the access information for each cache key in the memory database; the access information includes access frequency, last most recent access time, and current access time;
[0091] Step 20: Adjust the dynamic weight value of each cache key based on the access information of each cache key;
[0092] Step 30: Dynamically adjust the cached data in the memory database based on the adjusted weight values.
[0093] Specifically, the present invention also provides an adaptive weighting mechanism for a Least Frequently Used (LFU) cached data eviction strategy, which is applied to the in-memory database of the present invention.
[0094] Redis stores both structured data and metadata about unstructured data. Each stored data exists as a key-value pair. The cache key is a unique index for the stored data in Redis, allowing direct location of the stored data through its key-value pair. For example, a cache key named point_cloud_2024 could have an HDFS path of / large_files / point2024.pcd.
[0095] For each cache key in the in-memory database, an attribute set is established. This attribute set includes a dynamic weight value and access information (including access frequency, last access time, and current access time). After each retrieval, the dynamic weight value (which is continuously updated based on the attributes to assess the popularity of each cache key, i.e., the popularity of the table data corresponding to the cache key) and access information are continuously updated. The initial value of the access frequency can be set to 1, incremented by 1 after each retrieval, and the dynamic weight value of the cache key is dynamically calculated according to the following weight formula 1. Cache keys with a small difference between the current access time and the last access time and a high number of accesses are considered high-frequency retrieval data, and their weight value increments are large; cache keys with a large difference between the current access time and the last access time and a low number of accesses are considered low-frequency retrieval data, and their weight value increments are small.
[0096] For sporadic, high-frequency, discontinuous access data, the access characteristics are characterized by a high frequency of searches in the early stages and a significant decrease in access frequency in the later stages, resulting in a significant gap between the current access time and the time interval of the most recent historical access. By introducing access frequency as a weight calculation factor, this type of sporadic, high-frequency, discontinuous access data has accumulated a large number of searches in the early high-frequency access stage, preventing incremental deviations in its weight value.
[0097] (1)
[0098] in, These are dynamic weight values; For access frequency; This is the current access time; The last time accessed; It is a compound assignment operator that represents the addition and assignment of values.
[0099] The process of evicting cached data is as follows:
[0100] During the weighted sorting phase, when the memory of the in-memory database reaches a certain threshold, the eviction process is triggered. All cached keys are sorted in descending order by their weight values, generating an ordered list of keys. The sorting is based on the weight attribute maintained by each cached key; the higher the weight value, the earlier it appears in the list.
[0101] High-weight enhancement: Select the top-ranked cached data keys (i.e. hot data) as enhancement objects based on their weight values, trigger the access counter update of the underlying storage system's memory database by forcing data reload operations (i.e. rewriting the complete data), evict cached data according to the LFU algorithm, and finally reset the weight value of the cache key.
[0102] This embodiment uses dynamic weight value adjustment to form a thermal data maintenance scheme with adaptive characteristics.
[0103] The crop phenotypic data storage and retrieval device provided by the present invention is described below. The crop phenotypic data storage and retrieval device described below can be referred to in correspondence with the crop phenotypic data storage and retrieval method described above.
[0104] Please refer to Figure 3 The present invention also provides a crop phenotypic data storage and retrieval device, comprising:
[0105] The multimodal data acquisition module 301 is used to acquire multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; the crop phenotypic data includes structured data and unstructured data; the unstructured data includes small file data and large file data;
[0106] The first storage module 302 is used to store the structured data and the metadata information in a distributed database, and to store the unstructured data in a distributed file system;
[0107] The second storage module 303 is used to synchronously store the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data into the distributed search and analysis engine.
[0108] The retrieval module 304 is used to determine the retrieval result based on the user's retrieval conditions and the distributed search analysis engine when no data matching the user's retrieval conditions exists in the in-memory database; the in-memory database is used to cache the metadata information of the structured data and the unstructured data.
[0109] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communications bus 440. The processor 410 can invoke logical instructions in the memory 430 to execute a crop phenotypic data storage and retrieval method. This method includes: acquiring multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; the crop phenotypic data includes structured data and unstructured data; the unstructured data includes small file data and large file data; storing the structured data and the metadata information in a distributed database, and storing the unstructured data in a distributed file system; synchronously storing the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data in a distributed search analysis engine; determining retrieval results based on the user's retrieval conditions and the distributed search analysis engine when no data matching the user's retrieval conditions exists in the memory database; the memory database is used to cache the structured data and the metadata information of the unstructured data.
[0110] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0111] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the crop phenotypic data storage and retrieval method provided by the above methods. The method includes: acquiring multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; the crop phenotypic data includes structured data and unstructured data; the unstructured data includes small file data and large file data; storing the structured data and the metadata information in a distributed database, and storing the unstructured data in a distributed file system; synchronously storing the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data in a distributed search analysis engine; determining the retrieval result based on the user retrieval conditions and the distributed search analysis engine when no data matching the user retrieval conditions exists in the memory database; the memory database is used to cache the structured data and the metadata information of the unstructured data.
[0112] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the crop phenotypic data storage and retrieval method provided by the above methods. The method includes: acquiring multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; the crop phenotypic data includes structured data and unstructured data; the unstructured data includes small file data and large file data; storing the structured data and the metadata information in a distributed database, and storing the unstructured data in a distributed file system; synchronously storing the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data in a distributed search analysis engine; determining retrieval results based on the user retrieval conditions and the distributed search analysis engine when no data matching the user retrieval conditions exists in the in-memory database; the in-memory database is used to cache the structured data and the metadata information of the unstructured data.
[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for storing and retrieving crop phenotypic data, characterized in that, The method is applied to a hybrid architecture storage and retrieval system, which includes a distributed database, a distributed search and analysis engine, a distributed file system, and an in-memory database. Acquire multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; the crop phenotypic data includes structured data and unstructured data; the unstructured data includes small file data and large file data; The structured data and metadata are stored in the distributed database, and the unstructured data is stored in the distributed file system. The retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data are synchronously stored in the distributed search and analysis engine; The search results are determined based on the user's search criteria and the distributed search analysis engine; the in-memory database is used to cache the structured data and the metadata information of the unstructured data; The process of synchronously storing the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data into the distributed search analysis engine includes: When the structured data is retrieved based on user search criteria and the search results exist in the in-memory database, the search results are loaded into the local machine. When the structured data is retrieved based on user search criteria and no search results are found in the in-memory database, the target row key corresponding to the user search criteria is determined by the distributed search analysis engine, the search results are obtained from the distributed database based on the target row key, the search results are loaded into the local database, and the in-memory database is updated based on the loaded results. In the case of retrieving the unstructured data based on user search criteria, metadata results are retrieved from the in-memory database, and search data is obtained from the distributed file system based on the metadata results, and the search data is loaded locally.
2. The crop phenotypic data storage and retrieval method according to claim 1, characterized in that, The metadata information of the structured data includes crop identifiers and other crop information; The step of storing the structured data and the metadata information in the distributed database includes: The first row key and the first column family are filled into the first data table of the distributed database; the first row key is determined based on the crop identifier; the first column family is determined based on the structured data and the other crop information.
3. The crop phenotypic data storage and retrieval method according to claim 1, characterized in that, The step of storing the unstructured data in the distributed file system includes: Iterate through the unstructured data to be stored to obtain the file size of the currently uploaded file; If the file size is greater than a first threshold, the currently uploaded file is compressed and stored in the distributed file system; if the file size is less than the first threshold but greater than a second threshold, the currently uploaded file is placed in an empty merge queue; if the file size is less than or equal to the second threshold, the currently uploaded file is placed in a temporary queue. Select the largest file from the temporary queue and determine the maximum remaining space in each of the merge queues; If the largest file is greater than the maximum remaining space, the largest file is placed in the merge queue obtained by converting the spare queue; if the largest file is less than or equal to the maximum remaining space, the largest file is placed in the merge queue corresponding to the maximum remaining space. If the files in the merge queue meet the preset conditions, the files in the merge queue are merged and stored in the distributed file system.
4. The crop phenotypic data storage and retrieval method according to claim 3, characterized in that, The crop phenotypic data storage and retrieval method further includes: Establish a mapping relationship between the files before merging and the files after merging; The second row key and the second column family are filled into the second data table of the distributed database; the second row key is determined based on the file before merging; the second column family is determined based on the metadata information of the unstructured data and the mapping relationship.
5. The crop phenotypic data storage and retrieval method according to claim 1, characterized in that, The crop phenotypic data storage and retrieval method further includes: Determine the access information for each cache key in the memory database; the access information includes access frequency, last most recent access time, and current access time; Based on the access information of each cache key, adjust the dynamic weight value of each cache key; The cached data in the memory database is dynamically adjusted based on the adjusted weight values.
6. A crop phenotypic data storage and retrieval device, characterized in that, include: A multimodal data acquisition module is used to acquire multimodal data to be stored; the multimodal data includes crop phenotypic data and metadata information of the crop phenotypic data; The crop phenotypic data includes structured and unstructured data; the unstructured data includes small file data and large file data. The first storage module is used to store the structured data and the metadata information in a distributed database, and to store the unstructured data in a distributed file system; The second storage module is used to synchronously store the retrieval fields of the structured data, the retrieval fields of the small file data, and the metadata information of the large file data to the distributed search and analysis engine. The retrieval module is used to determine retrieval results based on user search criteria and the distributed search analysis engine; the in-memory database is used to cache the metadata information of the structured data and the unstructured data. The device is also used for: When the structured data is retrieved based on user search criteria and the search results exist in the in-memory database, the search results are loaded into the local machine. When the structured data is retrieved based on user search criteria and no search results are found in the in-memory database, the target row key corresponding to the user search criteria is determined by the distributed search analysis engine, the search results are obtained from the distributed database based on the target row key, the search results are loaded into the local database, and the in-memory database is updated based on the loaded results. In the case of retrieving the unstructured data based on user search criteria, metadata results are retrieved from the in-memory database, and search data is obtained from the distributed file system based on the metadata results, and the search data is loaded locally.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the crop phenotypic data storage and retrieval method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the crop phenotypic data storage and retrieval method as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the crop phenotypic data storage and retrieval method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Distributed cloud native storage oriented small file merging optimization method
CN120162006A
File-based data management method and system
CN120216605A