A distributed storage and management system for field specimen data based on a grid architecture
By adopting a distributed storage and management system based on a grid-based architecture in the field specimen data storage and management system, the traditional centralized system has solved the problems of limited capacity, low processing efficiency and single point failure when facing massive data, and achieved efficient, scalable and reliable data management.
Patent Information
- Application Number
- CN202510429159.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-08
AI Technical Summary
When facing massive field specimen data, traditional centralized data storage and management systems have problems such as limited capacity, low processing efficiency and single point failure, which is difficult to meet the needs of scientific research.
A distributed storage and management system based on a grid architecture is adopted, and the data collection module, species identification module, association analysis module, resource allocation module, data cache module and data management module are connected through cloud communication to realize distributed storage and intelligent management of data.
It improves the data processing speed and system scalability and fault tolerance, realizes intelligent allocation of resources and efficient storage and management of data, and improves the efficiency of scientific research work and data reliability.
Smart Images

Figure CN119938717B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage and management, and specifically to a distributed storage and management system for field specimen data based on a grid architecture. Background Art
[0002] CN117539841B "A distributed file system metadata management system and its operation method" includes a key-value storage module, a consistency and concurrency control module, a distributed architecture module, and a metadata operation module. The key-value storage module implements a key-value storage mechanism for storing and accessing key-value pairs of metadata. The consistency and concurrency control module implements a concurrency control mechanism. The distributed architecture module is used to design and implement a distributed architecture to achieve the scalability and fault tolerance of the system. The metadata operation module is used to support various metadata operations. The key-value storage module and the consistency and concurrency control module are interrelated.
[0003] CN115203177A "A distributed data storage system and storage method" includes storage nodes internally provided with a processor and a memory. The storage nodes are interconnected through a network; a monitoring module that monitors and records the capacity occupancy rate and resource utilization rate of each storage node; a calculation module that calculates the migration time period of each storage node based on the historical resource utilization rate for each storage node; an evaluation module that screens the storage nodes according to the data provided by the calculation module and the capacity occupancy rate to determine the storage nodes that need to be migrated in or out; and a migration module for migrating stored data.
[0004] With the continuous deepening of scientific research, the scale of field specimen data has grown explosively. Traditional centralized data storage and management systems have gradually exposed many defects when faced with massive field specimen data. For example, the capacity of centralized storage is limited and difficult to meet the growing data storage requirements; data processing is concentrated on a single server, resulting in low processing efficiency, especially in large-scale data query and analysis, the response time is too long. In addition, the centralized system has a single point of failure problem. Once the server fails, it may cause the entire data storage and management system to collapse, resulting in data loss or inaccessibility, seriously affecting the normal progress of scientific research work. Therefore, there is an urgent need for a new technical solution to solve the problem of field specimen data storage and management. Summary of the Invention
[0005] In order to solve the above technical problems, the purpose of the present invention is to provide a distributed storage and management system for field specimen data based on a grid architecture, including a cloud, which is communicatively connected to a data acquisition module, a species identification module, an association analysis module, a resource allocation module, a data cache module, and a data management module;
[0006] The data acquisition module is used to obtain the specimen data collected by several grid nodes, mark the collection time, and set the collection period;
[0007] The species identification module is used to identify the specimen data collected by the grid nodes and obtain the hierarchical classification relationship between the species corresponding to the specimen data and other species;
[0008] The correlation analysis module is used to obtain the correlation coefficient of each grid node and perform extended storage of specimen data according to the correlation coefficient;
[0009] The resource allocation module is used to construct a data flow prediction model and set the preset cloud computing resources of each grid node in each time period;
[0010] The data caching module obtains the popularity coefficient of each species of each grid node and performs specimen data caching operations according to the popularity coefficient of each species;
[0011] The data management module matches keywords and attributes for the query statement input by the user, and filters out the best grid node for specimen data reading operations.
[0012] Furthermore, the process by which the species identification module identifies the specimen data collected by the grid nodes and obtains the hierarchical classification relationship between the species corresponding to the specimen data and other species includes:
[0013] Extract the morphological characteristics of the specimen data collected by the grid nodes to obtain the morphological characteristics corresponding to the specimen data. Preset a species database in the cloud. The species database includes the morphological characteristics and classification knowledge graphs of several types of species. The classification knowledge graph includes species classification system information (such as detailed records of the positions of species in classification levels such as kingdom, phylum, class, order, family, genus, and species, and clarifies the upper and lower relationships between species) and species distribution information. Associate the species classification system information and species distribution information in the classification knowledge graph with the morphological characteristics of several types of species in the species database;
[0014] Obtain the geographical location characteristics of the grid nodes. Input the morphological characteristics and geographical location characteristics corresponding to the specimen data into the species database for feature similarity matching, obtain the feature similarities corresponding to different types of species, select the species with the highest feature similarity as the species corresponding to the specimen data. At the same time, when performing similarity feature matching, set the recognition weights of the morphological characteristics of each type of species in the species database according to the classification knowledge graph. For example, if the knowledge graph indicates that a certain morphological characteristic is of great significance in species classification, assign a higher weight to this morphological characteristic in the similarity calculation. At the same time, obtain the hierarchical classification relationship between the species corresponding to the specimen data and other species, and store the species corresponding to the specimen data associated with the specimen data in the grid node.
[0015] Furthermore, the process of the association analysis module obtaining the association coefficients of each grid node and performing extended storage of specimen data based on the association coefficients includes:
[0016] Each grid node in the target area is communicatively connected to the cloud in a distributed manner, and each grid node is communicatively connected to each other. When a grid node stores specimen data and the species corresponding to the specimen data, the usage records of other grid nodes except the grid node are obtained;
[0017] Statistically analyze the usage records of other grid nodes to obtain the usage frequencies of each species of other grid nodes. At the same time, according to the hierarchical classification relationship between the species corresponding to the specimen data and other species, set the weight coefficients of other species. According to the usage frequencies of each species of other grid nodes and the weight coefficients of other species, obtain the association coefficients of other grid nodes;
[0018] Among them, the calculation formula for obtaining the association coefficients of other grid nodes according to the usage frequencies of each species of other grid nodes and the weight coefficients of other species is:
[0019] ;
[0020] Among them, represents the association coefficient of other grid nodes, represents the usage frequency of the species corresponding to the specimen data, traverse all other species, represents other species of the usage frequency, represents other species of the weight coefficient; The above formulas are all calculated by removing the dimension and taking their numerical values. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and preset thresholds in the formulas are set by those skilled in the art according to the actual situation or obtained by simulating a large amount of data;
[0021] Compare the association coefficients of other grid nodes with the preset association coefficient threshold. Mark the other grid nodes with association coefficients greater than the association coefficient threshold as associated grid nodes, and send the specimen data of the grid node and the species corresponding to the specimen data to the associated grid nodes for storage at the same time.
[0022] Furthermore, the process of the resource allocation module constructing a data flow prediction model includes:
[0023] Build a data flow prediction model based on deep learning, obtain the specimen data of each grid node in several historical collection cycles for traffic analysis, obtain the data traffic sequences of each grid node in several historical collection cycles, use the data traffic sequences of each grid node in several historical collection cycles as the training set and the test set, input the training set into the data flow prediction model for training until the loss function is trained stably, save the model parameters, test the data flow prediction model with the test set until it meets the preset requirements, and output the data flow prediction model.
[0024] Building a data flow prediction model based on deep learning is a complex process that involves multiple steps such as model selection, training, validation, and testing. The following is a detailed supplementary description of this process:
[0025] Select a convolutional neural network (CNN) suitable for time series analysis as the deep learning architecture, choose binary cross-entropy loss as the optimization objective, and then input the prepared training set into the selected deep learning model to start training. During the training process, the weights are continuously updated through the backpropagation algorithm to gradually reduce the loss function until it reaches a stable state. During this period, techniques such as early stopping are used to avoid overfitting. In addition to the basic training process, various parameters of the model are tuned through grid search, and the parameters include learning rate, batch size, regularization coefficient, etc.
[0026] When the model training is completed and the parameters are adjusted, the final evaluation is carried out through the test set to obtain the evaluation results of the model. The evaluation results include classification metrics such as accuracy, recall rate, and F1 score. According to the evaluation results on the test set, it is judged whether the model meets the expected standards. If the requirements are met, the model parameters are saved and deployment is prepared; if not, it is necessary to return to a previous stage to re-examine issues such as data quality, model structure, or training strategy.
[0027] Furthermore, the process of the resource allocation module setting the preset cloud computing resources of each grid node in each time period includes:
[0028] Obtain the predicted data traffic sequences of each grid node in the current collection cycle according to the data flow prediction model, extract the features of the predicted data traffic sequences of each grid node in the current collection cycle, obtain the predicted average traffic and predicted traffic fluctuation amplitude of each grid node in each time period, and set the preset cloud computing resources of each grid node in each time period according to the predicted average traffic and predicted traffic fluctuation amplitude of each grid node.
[0029] Further, the process of the data cache module obtaining the popularity coefficients of each species of each grid node and performing specimen data caching operations according to the popularity coefficients of each species includes:
[0030] Preset a cache area in each grid node, obtain the usage frequencies of each species in each grid node in a number of historical collection cycles, and at the same time obtain the interval between each historical collection cycle and the current collection cycle, and set the historical forgetting coefficient of each historical collection cycle according to the interval;
[0031] Obtain the popularity coefficient of each species according to the usage frequency of each species in each historical collection cycle and the historical forgetting coefficient of each historical collection cycle , , where represents the usage frequency of the nth historical collection cycle, represents the historical forgetting coefficient of the nth historical collection cycle, n traverses each historical collection cycle, compare the popularity coefficient of each species with a preset popularity coefficient threshold, and transfer the specimen data corresponding to the species with a popularity coefficient greater than the popularity coefficient threshold in each grid node to the cache area.
[0032] Further, the process of the data management module performing keyword and attribute matching on the query statement input by the user and screening out the best grid node for specimen data reading operations includes:
[0033] When the user logs in to the cloud and enters a query statement, perform keyword and attribute matching on the query statement input by the user. First, parse the query statement, use the natural language processing tool NLTK to split the query statement input by the user into individual words. For example, for the query statement "Find flower specimens collected in 2020 with red petals", it will be split into words such as "Find", "2020", "collected", "with", "red", "petals", "flower", "specimens", etc. Perform part-of-speech tagging on the word segmentation result to determine the part of speech of each word, such as noun, verb, adjective, adverb, etc. For example, "2020" is a time noun, "red" is an adjective, "flower" and "specimen" are nouns. According to the part of speech and semantics, extract the keywords in the query statement, such as "2020", "red petals", "flower specimens" as keywords;
[0034] Subsequently, attribute recognition and extraction are carried out (including time attributes, morphological attributes, and category attributes) to obtain species description keywords. For example, information related to time is extracted from the query statement to determine the time range. "Collected in 2020" clarifies that the collection time of the specimen is 2020. Keywords describing the morphological characteristics of the specimen are identified, such as color, shape, size, etc. In the "flower specimen with red petals", "red petals" is a description of the morphological attribute. The category information to which the specimen belongs is determined. For example, "flower specimen" indicates that the flower specimen of the plant category is to be queried. The system will filter out the specimens that meet the category requirements according to the classification field in the specimen data. Keyword retrieval is performed on each grid node according to the species description keywords to obtain the specimen data that meets the keyword retrieval conditions. The grid nodes storing the specimen data that meets the keyword retrieval conditions are obtained, and the grid nodes are marked as key grid nodes. Dynamic load analysis is carried out on each key grid node, and the best grid node is selected for the specimen data reading operation.
[0035] Further, the process of performing dynamic load analysis on each key grid node includes:
[0036] An identification data packet is allocated for transmission between each key grid node and the cloud. After the transmission of the identification data packet is completed, the data transmission delay of the identification data packet of each key grid node is obtained. The key grid nodes with a data transmission delay less than the preset delay threshold are filtered out, and the key grid nodes are marked as the first grid nodes. The data transmission delay of the current identification data packet of the first grid nodes is subjected to amplitude variation analysis to obtain the delay amplitude variation of the first grid nodes , , denotes the data transmission delay of the previous current identification data packet. The first grid nodes with a delay amplitude variation less than the preset delay amplitude variation threshold are filtered out, and the first grid nodes are marked as the best grid nodes.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] 1. Distributed architecture improves performance: Through the grid architecture and distributed storage, data is dispersed and stored in multiple grid nodes, changing the drawbacks of traditional centralized storage. When performing data storage and reading, each node can process in parallel, greatly improving the data processing speed. For example, when entering a large amount of specimen data, multiple nodes can receive data simultaneously, avoiding the storage bottleneck that may occur in a centralized system and improving the overall data processing efficiency.
[0039] 2. Intelligent resource allocation to ensure efficiency: The resource allocation module can accurately predict the data traffic of each grid node at different time periods by building a data flow prediction model. According to the predicted average traffic and traffic fluctuation range, it reasonably allocates cloud computing resources for each node. During peak traffic periods, more resources are allocated to high-traffic nodes to ensure the efficiency of data processing; during low-traffic periods, resource allocation is reduced to avoid resource waste, realizing the dynamic and optimal utilization of resources and ensuring that the system always maintains an efficient operating state.
[0040] 3. High precision in multi-dimensional feature recognition: The species recognition module uses dual information of morphological features and geographical location features to match in a species database containing rich morphological features and classification knowledge graphs. Compared with single-feature recognition, this method greatly improves the accuracy of species recognition. For example, for species with similar morphologies but different distribution regions, combining geographical location features can effectively distinguish them, providing a reliable species data basis for scientific research.
[0041] 4. Association analysis to achieve data expansion: The association analysis module can discover potential data connections between nodes by analyzing the correlation coefficients of each grid node. According to the weight coefficients determined by the usage records of other grid nodes, the species usage frequency, and the hierarchical classification relationship, it identifies the associated grid nodes. It stores important specimen data and their corresponding species in the associated grid nodes to achieve extended storage of data. This not only improves the availability of data but also enables the associated data to be quickly obtained from other nodes when the data of a certain node is frequently accessed, reducing data transmission latency and improving data access efficiency.
[0042] 5. Enhancing data integrity and collaboration: This extended storage mechanism helps to ensure data integrity, with data between different nodes complementing each other to form a more comprehensive dataset. Moreover, the data association between nodes enhances the collaboration of the system, facilitating cross-regional and cross-node comprehensive research and promoting data sharing and cooperation between different research teams or institutions.
[0043] 6. Heat-driven caching strategy: The data caching module performs specimen data caching operations according to the heat coefficients of different species in each grid node. By analyzing the usage frequency of species during historical collection periods and setting historical forgetting coefficients, it accurately evaluates the heat of species. It caches the data of high-heat species in a preset cache area, enabling users to quickly obtain the data directly from the cache when they request it again, significantly shortening the data reading time and improving the system response speed.
[0044] 7. Precise Keyword and Attribute Matching: The data management module performs precise keyword and attribute matching on the query statements input by users. By extracting the keyword descriptions of species, targeted retrieval is carried out in each grid node to quickly locate the specimen data and storage nodes that meet the conditions, namely the key grid nodes. This precise matching avoids blind retrieval and greatly improves the query efficiency.
[0045] 8. Dynamic Load Analysis for Optimizing Node Selection: Conduct dynamic load analysis on the key grid nodes, considering the data transmission delay and its variation range, and filter out the best grid nodes for specimen data reading operations. Ensure that data is obtained from the node with the fastest response speed and the highest stability, reduce the delay and fluctuation during data transmission, and ensure that users can obtain the required specimen data efficiently and stably, improving the quality and efficiency of the entire query process. Brief Description of the Drawings
[0046] Figure 1 It is a schematic diagram of a distributed storage and management system for field specimen data based on a grid architecture according to an embodiment of the present application. Detailed Embodiments
[0047] Next, in combination with the drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0048] As Figure 1 shown, a distributed storage and management system for field specimen data based on a grid architecture includes a cloud, and the cloud is communicatively connected to a data acquisition module, a species identification module, an association analysis module, a resource allocation module, a data cache module, and a data management module;
[0049] The data acquisition module is used to obtain the specimen data collected by a number of grid nodes, mark the collection time, and set the collection period;
[0050] The species identification module is used to identify the specimen data collected by the grid nodes and obtain the hierarchical classification relationship between the species corresponding to the specimen data and other species;
[0051] The association analysis module is used to obtain the correlation coefficients of each grid node and perform extended storage of specimen data according to the correlation coefficients;
[0052] The resource allocation module is used to construct a data flow prediction model and set the preset cloud computing resources of each grid node in each time period;
[0053] The data cache module obtains the popularity coefficients of each species of each grid node and performs specimen data caching operations according to the popularity coefficients of each species;
[0054] The data management module performs keyword and attribute matching on the query statement input by the user, and filters out the best grid node for specimen data reading operations.
[0055] It should be further noted that in the specific implementation process, the process by which the species identification module identifies the specimen data collected by the grid node and obtains the hierarchical classification relationship between the species corresponding to the specimen data and other species includes:
[0056] Extract the morphological characteristics of the specimen data collected by the grid node to obtain the morphological characteristics corresponding to the specimen data. Preset a species database in the cloud. The species database includes the morphological characteristics and classification knowledge graphs of several types of species. The classification knowledge graph includes species classification system information (such as detailed records of the positions of species in classification levels such as kingdom, phylum, class, order, family, genus, and species, and clarifies the upper and lower relationships between species) and species distribution information. Associate the species classification system information and species distribution information in the classification knowledge graph with the morphological characteristics of several types of species in the species database;
[0057] Obtain the geographical location characteristics of the grid node, input the morphological characteristics and geographical location characteristics corresponding to the specimen data into the species database for feature similarity matching, obtain the feature similarities corresponding to different types of species, select the species with the highest feature similarity as the species corresponding to the specimen data. At the same time, when performing similarity feature matching, set the recognition weights of the morphological characteristics of each type of species in the species database according to the classification knowledge graph. For example, if the knowledge graph indicates that a certain morphological characteristic is of great significance in species classification, assign a higher weight to this morphological characteristic in the similarity calculation. At the same time, obtain the hierarchical classification relationship between the species corresponding to the specimen data and other species, and store the species corresponding to the specimen data in association with the specimen data in the grid node.
[0058] It should be further noted that in the specific implementation process, the process by which the association analysis module obtains the correlation coefficients of each grid node and performs extended storage of specimen data according to the correlation coefficients includes:
[0059] Each grid node in the target area is communicatively connected to the cloud in a distributed manner, and each grid node is communicatively connected to each other. When a grid node stores specimen data and the species corresponding to the specimen data, obtain the usage records of other grid nodes except the grid node;
[0060] Statistically analyze the usage records of other grid nodes to obtain the usage frequencies of various species of other grid nodes. At the same time, according to the hierarchical classification relationship between the species corresponding to the specimen data and other species, set the weight coefficients of other species. According to the usage frequencies of various species of other grid nodes and the weight coefficients of other species, obtain the correlation coefficients of other grid nodes;
[0061] Among them, the calculation formula for obtaining the correlation coefficient of other grid nodes according to the usage frequencies of various species of other grid nodes and the weight coefficients of other species is:
[0062] ;
[0063] Among them, represents the correlation coefficient of other grid nodes, represents the usage frequency of the species corresponding to the specimen data, traverse all other species, represents other species 's usage frequency, represents other species 's weight coefficient; the above formulas are all calculated by removing the dimension and taking their numerical values. The formulas are obtained by collecting a large amount of data and performing software simulation to get a formula closest to the real situation. The preset parameters and preset thresholds in the formulas are set by those skilled in the art according to the actual situation or obtained by a large amount of data simulation;
[0064] Compare the correlation coefficient of other grid nodes with the preset correlation coefficient threshold. Mark the other grid nodes with a correlation coefficient greater than the correlation coefficient threshold as associated grid nodes, and send the specimen data of the grid nodes and the species corresponding to the specimen data to the associated grid nodes for storage at the same time.
[0065] It should be further noted that in the specific implementation process, the process of the resource allocation module constructing the data flow prediction model includes:
[0066] Construct a data flow prediction model based on deep learning. Obtain the specimen data of each grid node in several historical collection cycles for traffic analysis, obtain the data flow sequences of each grid node in several historical collection cycles, and use the data flow sequences of each grid node in several historical collection cycles as the training set and the test set. Input the training set into the data flow prediction model for training until the loss function is trained stably, and save the model parameters. Test the data flow prediction model through the test set until it meets the preset requirements, and output the data flow prediction model.
[0067] Building a data flow prediction model based on deep learning is a complex process that involves multiple steps such as model selection, training, validation, and testing. The following is a detailed supplementary description of this process:
[0068] Select a convolutional neural network (CNN) suitable for time series analysis as the deep learning architecture, and choose binary cross-entropy loss as the optimization objective. Then, input the prepared training set into the selected deep learning model to start training. During the training process, the weights are continuously updated through the backpropagation algorithm, making the loss function gradually decrease until it reaches a stable state. During this period, techniques such as early stopping are used to avoid overfitting. In addition to the basic training process, various parameters of the model are tuned through grid search, including the learning rate, batch size, regularization coefficient, etc.
[0069] When the model training is completed and the parameters are adjusted, the final evaluation is carried out through the test set to obtain the evaluation results of the model. The evaluation results include classification metrics such as accuracy, recall rate, and F1 score. According to the evaluation results on the test set, it is judged whether the model meets the expected standards. If the requirements are met, the model parameters are saved and prepared for deployment; if not, it is necessary to return to a previous stage to re-examine issues such as data quality, model structure, or training strategy.
[0070] It should be further noted that in the specific implementation process, the process of the resource allocation module setting the preset cloud computing resources of each grid node in each time period includes:
[0071] Obtain the predicted data traffic sequences of each grid node in the current collection period according to the data flow prediction model, extract the features of the predicted data traffic sequences of each grid node in the current collection period, obtain the predicted average traffic and predicted traffic fluctuation amplitude of each grid node in each time period, and set the preset cloud computing resources of each grid node in each time period according to the predicted average traffic and predicted traffic fluctuation amplitude of each grid node.
[0072] It should be further noted that in the specific implementation process, the process of the data caching module obtaining the popularity coefficients of each species of each grid node and performing specimen data caching operations according to the popularity coefficients of each species includes:
[0073] Preset a cache area in each grid node, obtain the usage frequencies of each species of each grid node in several historical collection periods, and at the same time obtain the time interval between each historical collection period and the current collection period, and set the historical forgetting coefficient of each historical collection period according to the time interval.
[0074] Obtain the popularity coefficient of each species according to the usage frequency of each species in each historical collection period and the historical forgetting coefficient of each historical collection period. , , where represents the usage frequency in the nth historical collection period, represents the historical forgetting coefficient in the nth historical collection period, n traverses each historical collection period, compare the popularity coefficient of each species with a preset popularity coefficient threshold, and transfer the specimen data corresponding to the species with a popularity coefficient greater than the popularity coefficient threshold in each grid node to the cache area.
[0075] It should be further noted that in the specific implementation process, the process of the data management module performing keyword and attribute matching on the query statement input by the user and screening out the best grid node for specimen data reading operations includes:
[0076] When the user logs in to the cloud and inputs a query statement, perform keyword and attribute matching on the query statement input by the user. First, parse the query statement. Use the natural language processing tool NLTK to split the query statement input by the user into individual words. For example, for the query statement "Find the flower specimens collected in 2020 with red petals", it will be split into words such as "Find", "2020", "collected", "with", "red", "petals", "flower", "specimens", etc. Perform part-of-speech tagging on the tokenization result to determine the part of speech of each word, such as noun, verb, adjective, adverb, etc. For example, "2020" is a time noun, "red" is an adjective, "flower" and "specimen" are nouns. According to the part of speech and semantics, extract the keywords in the query statement, such as "2020", "red petals", and "flower specimens" as keywords;
[0077] Subsequently, attribute recognition and extraction are carried out (including time attributes, morphological attributes, and category attributes) to obtain species description keywords. For example, information related to time is extracted from the query statement to determine the time range. "Collected in 2020" clarifies that the collection time of the specimen is 2020. Keywords describing the morphological characteristics of the specimen are identified, such as color, shape, size, etc. In the "flower specimen with red petals", "red petals" is a description of the morphological attribute. The category information to which the specimen belongs is determined. For example, "flower specimen" indicates that the flower specimen of the plant category is to be queried. The system will screen out the specimens that meet the requirements of this category according to the classification field in the specimen data. Keyword retrieval is performed on each grid node according to the species description keywords to obtain the specimen data that meets the keyword retrieval conditions. The grid node storing the specimen data that meets the keyword retrieval conditions is obtained, and the grid node is marked as a key grid node. Dynamic load analysis is carried out on each key grid node, and the best grid node is selected for the specimen data reading operation.
[0078] It should be further noted that in the specific implementation process, the process of performing dynamic load analysis on each key grid node includes:
[0079] An identification data packet is allocated for transmission between each key grid node and the cloud. After the transmission of the identification data packet is completed, the data transmission delay of the identification data packet of each key grid node is obtained. The key grid nodes with a data transmission delay less than the preset delay threshold are screened out, and the key grid nodes are marked as the first grid nodes. The data transmission delay of the current identification data packet of the first grid nodes is subjected to amplitude variation analysis to obtain the delay amplitude variation of the first grid nodes , , represents the data transmission delay of the previous current identification data packet. The first grid nodes with a delay amplitude variation less than the preset delay amplitude variation threshold are screened out, and the first grid nodes are marked as the best grid nodes.
[0080] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A distributed storage and management system for field specimen data based on a grid architecture, characterized in that: It includes a cloud, wherein the cloud is communicatively connected with a data acquisition module, a species identification module, an association analysis module, a resource allocation module, a data cache module and a data management module; The data acquisition module is used to obtain the sample data collected by several grid nodes and mark the collection time and set the collection cycle; The species identification module is used to identify the specimen data collected by the grid nodes and obtain the hierarchical classification relationship between the species corresponding to the specimen data and other species; The correlation analysis module is used to obtain the correlation coefficient of each grid node and perform sample data expansion storage according to the correlation coefficient; The resource allocation module is used to build a data flow prediction model and set the preset cloud computing resources for each grid node in each time period; The data cache module obtains the heat coefficient of each species of each grid node and performs specimen data cache operation according to the heat coefficient of each species; The data management module matches keywords and attributes of the query statements input by the user, and selects the best grid nodes for specimen data reading operations; The correlation analysis module obtains the correlation coefficient of each grid node, and the process of expanding and storing the specimen data according to the correlation coefficient includes: Each grid node in the target area is connected to the cloud through distributed communication, and each grid node is connected to each other through communication. When a grid node stores specimen data and the species corresponding to the specimen data, the use records of other grid nodes except the grid node are obtained; Perform statistical analysis on the usage records of other grid nodes to obtain the usage frequency of each species of other grid nodes. At the same time, according to the hierarchical classification relationship between the species corresponding to the specimen data and other species, set the weight coefficient of other species. According to the usage frequency of each species of other grid nodes and the weight coefficient of other species, obtain the correlation coefficient of other grid nodes. Among them, according to the usage frequency of each species of other grid nodes and the weight coefficient of other species, the calculation formula for obtaining the correlation coefficient of other grid nodes is: ; in, represents the correlation coefficient of other grid nodes, Indicates the usage frequency of the species corresponding to the specimen data, Go through all other species, represents the usage frequency of other species k, Represents the weight coefficient of other species k; the above formulas are all dimensionless and numerical calculations, and the formula is a formula closest to the actual situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and preset thresholds in the formula are set by technicians in this field according to actual conditions or obtained by simulating a large amount of data; The correlation coefficients of other grid nodes are compared with a preset correlation coefficient threshold, other grid nodes whose correlation coefficients are greater than the correlation coefficient threshold are marked as associated grid nodes, and the specimen data of the grid nodes and the species corresponding to the specimen data are simultaneously sent to the associated grid nodes for storage.
2. According to claim 1, a distributed storage and management system for field specimen data based on a grid architecture is characterized in that: The species identification module identifies the specimen data collected by the grid node, and the process of obtaining the hierarchical classification relationship between the species corresponding to the specimen data and other species includes: Extract morphological features from the specimen data collected by the grid nodes to obtain morphological features corresponding to the specimen data, preset a species database, the species database includes several morphological features of several types of species and a classification knowledge graph, the classification knowledge graph includes species classification system information and species distribution information, and associate the species classification system information and species distribution information in the classification knowledge graph with several morphological features of several types of species in the species database; The geographic location characteristics of the grid node are obtained, and the morphological characteristics and geographic location characteristics corresponding to the specimen data are input into the species database for feature similarity matching. The feature similarities corresponding to different types of species are obtained, and the species with the highest feature similarity is selected as the species corresponding to the specimen data. At the same time, the hierarchical classification relationship between the species corresponding to the specimen data and other species is obtained, and the species corresponding to the specimen data is associated with the specimen data and stored in the grid node.
3. According to claim 2, a distributed storage and management system for field specimen data based on a grid architecture is characterized in that: The process of building a data flow prediction model in the resource allocation module includes: A data flow prediction model is constructed based on deep learning, and sample data of each grid node in several historical collection cycles is obtained for flow analysis. The data flow sequence of each grid node in several historical collection cycles is obtained, and the data flow sequence of each grid node in several historical collection cycles is used as a training set and a test set. The training set is input into the data flow prediction model for training until the loss function training is stable, and the model parameters are saved. The data flow prediction model is tested by the test set until it meets the preset requirements, and the data flow prediction model is output.
4. The distributed storage and management system for field specimen data based on a grid architecture according to claim 3 is characterized in that: The process of setting the preset cloud computing resources of each grid node in each time period by the resource allocation module includes: According to the data flow prediction model, the predicted data flow sequence of each grid node in the current collection period is obtained, and the predicted data flow sequence of each grid node in the current collection period is subjected to feature extraction to obtain the predicted average flow and predicted flow fluctuation range of each grid node in each time period. According to the predicted average flow and predicted flow fluctuation range of each grid node, the preset cloud computing resources of each grid node in each time period are set.
5. The distributed storage and management system for field specimen data based on a grid architecture according to claim 4 is characterized in that: The data cache module obtains the heat coefficient of each species of each grid node, and the process of performing specimen data cache operation according to the heat coefficient of each species includes: Preset a cache area in each grid node, obtain the usage frequency of each species in each grid node in several historical collection cycles, and simultaneously obtain the period interval between each historical collection cycle and the current collection cycle, and set the historical forgetting coefficient of each historical collection cycle according to the period interval; According to the usage frequency of each species in each historical collection cycle and the historical forgetting coefficient of each historical collection cycle, the heat coefficient of each species is obtained, the heat coefficient of each species is compared with the preset heat coefficient threshold, and the specimen data corresponding to the species with a heat coefficient greater than the heat coefficient threshold in each grid node is transmitted to the cache area.
6. The distributed storage and management system for field specimen data based on a grid architecture according to claim 5 is characterized in that: The data management module matches keywords and attributes of the query statement entered by the user and selects the best grid node for specimen data reading operation. The process includes: When a user enters a query statement, the query statement entered by the user is matched with keywords and attributes; Obtain species description keywords, perform keyword search on each grid node according to the species description keywords, obtain specimen data that meets the keyword search conditions, obtain grid nodes that store the specimen data that meets the keyword search conditions, mark the grid nodes as key grid nodes, perform dynamic load analysis on each key grid node, and screen out the best grid node for specimen data reading operations.
7. The distributed storage and management system for field specimen data based on a grid architecture according to claim 6 is characterized in that: The process of dynamic load analysis for each key grid node includes: Obtain the data transmission delay of each key grid node, filter out the key grid node whose data transmission delay is less than the preset delay threshold, mark the key grid node as the first grid node, and A variation analysis is performed to obtain the delay variation of the first grid node, a first grid node whose delay variation is less than a preset delay variation threshold is screened out, and the first grid node is marked as the best grid node.
Citation Information
Patent Citations
Distributed data storage system and storage method
CN115203177A
A distributed file system metadata management system and operation method thereof
CN117539841B
Image recognition method, device and system
CN111931835A
Data storage method and device, data query method and device, equipment and storage medium
CN115878513A