Distributed storage method and system for emergency command and rescue data based on artificial intelligence
By performing feature identification processing and distributed storage strategy optimization on emergency command and rescue data, a storage solution matching the rescue scenario is generated, which solves the problem of data and node mismatch in existing technologies and improves the adaptability and responsiveness of emergency command and rescue data storage.
Patent Information
- Application Number
- CN202510983592.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing emergency command and rescue data storage methods fail to effectively consider the dynamic relationship between data and rescue scenarios, resulting in a mismatch between storage nodes and data requirements, increased data access latency and resource waste, and are unable to meet the dynamic response capability requirements of emergency command and rescue.
By acquiring rescue scene perception data collected in real time by multiple types of emergency terminal devices, feature identification processing is performed to generate scene context features and storage demand features, and the distributed storage strategy model is called to perform node adaptation analysis, generate a distributed storage mapping solution, and optimize and adjust it in real time to form a multi-node collaborative storage cluster.
The data storage solution is matched with the actual needs of the rescue process, which improves the adaptability, reliability and dynamic response capability of the storage system, ensuring that the data storage system continues to operate efficiently during the rescue process.
Smart Images

Figure CN120492546B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage, and in particular to an artificial intelligence-based distributed storage method and system for emergency command and rescue data. Background Art
[0002] With the advancement of information technology in emergency command and rescue, the storage and management of rescue scene data has become a critical component in ensuring efficient command and decision-making. Emergency command and rescue data storage ensures efficient and reliable storage of environmental monitoring, equipment status, and personnel location data collected in real time by various emergency terminal devices by rationally allocating storage resources. Currently, most rescue data storage methods simply allocate storage nodes based on the data's basic physical properties (e.g., file size and update frequency). For example, large files are stored on high-capacity nodes, and frequently updated files are stored on high-computing nodes. However, such methods focus solely on the static properties of the data and fail to consider the dynamic relationship between the data and the rescue scenario (e.g., requiring low-latency access to fire source location data at a certain stage, or prioritizing storage of medical equipment status data during patient transport). This leads to a mismatch between storage nodes and actual data needs. When the rescue process changes (e.g., from on-site reconnaissance to material deployment) or when node status suddenly changes (e.g., a surge in storage node load), this can lead to increased data access latency and wasted storage resources. These methods fail to meet the requirements of emergency command and rescue for data storage adaptability and dynamic responsiveness. Summary of the Invention
[0003] The present invention provides an artificial intelligence-based distributed storage method and system for emergency command and rescue data.
[0004] In a first aspect, an embodiment of the present invention provides a method for distributed storage of emergency command and rescue data based on artificial intelligence, the method comprising:
[0005] Acquire a rescue scene perception data set collected in real time by multiple types of emergency terminal devices, wherein the rescue scene perception data set includes environmental monitoring data, equipment status data, and personnel positioning data at different rescue stages;
[0006] Performing feature identification processing on the rescue scene perception data set to generate an emergency data feature set including scene context features and storage requirement features, wherein the scene context features are used to characterize the association attributes between the data and the rescue scene, and the storage requirement features are used to characterize the storage priority and access frequency attributes of the data;
[0007] Calling a pre-built distributed storage strategy model to perform storage node adaptation analysis on the emergency data feature set, generating a distributed storage mapping scheme that matches the storage node computing power, storage space capacity, and access latency characteristics, wherein the distributed storage mapping scheme includes information on the correspondence between data and target storage nodes;
[0008] Distributing each data in the rescue scene perception data set to a corresponding target storage node according to the distributed storage mapping scheme to obtain a distributed data storage cluster with a multi-node collaborative storage structure;
[0009] Real-time operating status information of the distributed data storage cluster is collected, and the distributed storage mapping scheme is dynamically optimized and adjusted based on the real-time operating status information and the emergency data feature set to generate an updated distributed storage mapping scheme.
[0010] In a second aspect, an embodiment of the present invention provides a computer system, including:
[0011] a memory storing a computer program;
[0012] A processor is used to load the computer program to implement the above-mentioned artificial intelligence-based distributed storage method for emergency command and rescue data.
[0013] The artificial intelligence-based distributed storage method for emergency command and rescue data provided by the present invention can comprehensively cover the environment, equipment and personnel-related data in different rescue stages by acquiring the rescue scene perception data set collected in real time by multiple types of emergency terminal devices, and provide a multi-source data basis for subsequent storage decisions; the data set is processed by feature identification to generate an emergency data feature set containing scene context features and storage demand features, so that the storage decision not only focuses on the storage attributes of the data itself, but also combines the dynamic correlation attributes of the data and the rescue scene to ensure that the storage plan matches the actual needs of the rescue process; the pre-built distributed storage strategy model is called to perform storage segmentation on the feature set Point adaptation analysis and generation of mapping schemes can accurately match data with node capabilities based on multi-dimensional characteristics of storage nodes such as computing power, storage space, and access delay, avoiding the problem of low storage efficiency caused by considering only a single dimension in general storage; distributing data according to the mapping scheme to form a multi-node collaborative storage cluster, realizing a distributed storage layout of data, and improving the reliability and scalability of the storage system; collecting the real-time operating status of the cluster and dynamically optimizing and adjusting the mapping scheme in combination with the characteristics of emergency data, which can promptly respond to sudden changes in data demand or node status in the rescue scene (such as a sudden increase in node load or a change in data access frequency), and ensure that the storage system continues to be in an efficient operating state. Through the synergistic effect of the above steps, the present invention can effectively improve the adaptability, reliability and dynamic response capabilities of emergency command and rescue data storage, so that the data storage scheme not only meets the actual needs of the rescue scene, but also can flexibly respond to the uncertainty in the rescue process, and provide stable data storage support for the efficient implementation of emergency command and rescue. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flowchart of an artificial intelligence-based distributed storage method for emergency command and rescue data provided by an embodiment of the present invention.
[0015] Figure 2 It is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0017] See also Figure 1 , Figure 1A flowchart of a method for distributed storage of emergency command and rescue data based on artificial intelligence provided by an embodiment of the present invention is provided. The method can be executed by a computer system and includes the following steps:
[0018] Step S100: Acquire a rescue scene perception data set collected in real time by multiple types of emergency terminal devices. The rescue scene perception data set includes environmental monitoring data, equipment status data, and personnel positioning data at different rescue stages.
[0019] Rescue scene perception data consists of various types of data collected from the rescue scene during emergency rescue scenarios. Environmental monitoring data is the monitoring results of indicators related to the natural and physical environment of the rescue site, such as temperature, humidity, air quality, and wind speed. This data can reflect the environmental conditions at the rescue site. For example, high temperatures may indicate a fire hazard, while high humidity may affect the normal operation of rescue equipment. Equipment status data contains operating parameters and status information about the equipment used during the rescue process, such as the engine speed and water pump pressure of fire trucks and the power, speed, and direction of rescue robots. This equipment status data allows real-time monitoring of equipment operating conditions, determining whether the equipment is operating properly and whether maintenance or replacement is required. Personnel location data is the specific location of rescue personnel and trapped people at the rescue site, obtained using positioning technology.
[0020] Various types of emergency terminal devices are used to collect sensory data at the rescue scene. Environmental sensors are specifically designed to collect environmental monitoring data. For example, temperature sensors can measure the temperature at the rescue site in real time, and air quality sensors can detect the concentration of harmful gases in the air. Equipment monitoring devices are connected to rescue equipment and collect operating status data through various sensors and monitoring modules. For example, sensors installed on fire truck engines can monitor engine speed and temperature in real time. Personnel positioning tags can be worn by rescuers and trapped personnel, using satellite positioning, Bluetooth positioning, WiFi positioning, and other technologies to track personnel's location in real time.
[0021] Step S200: Perform feature identification processing on the rescue scene perception data set to generate an emergency data feature set including scene context features and storage requirement features. The scene context features are used to characterize the association attributes between the data and the rescue scene, and the storage requirement features are used to characterize the storage priority and access frequency attributes of the data.
[0022] Feature identification processing involves in-depth analysis and processing of rescue scene perception data sets, assigning specific feature identifiers to the data for better understanding and management. Scene context features are a set of characteristic information that reflects the relationship between the data and the rescue scenario. Storage requirement features are determined based on the data's storage requirements and usage frequency. Data size, update frequency, and access priority are key components of storage requirement features. Data size determines the storage space required, update frequency reflects the speed of data changes, and access priority reflects the data's importance within the emergency command system.
[0023] As an implementation method, step S200 performs feature identification processing on the rescue scene perception data set to generate an emergency data feature set including scene context features and storage requirement features. Specifically, the following steps S210 to S250 may be implemented:
[0024] Step S210: Standardize the formats of the environmental monitoring data, equipment status data, and personnel positioning data in the rescue site perception data set, eliminate the differences in data types and encoding rules among the data collected by different emergency terminal devices, and obtain a standardized data set.
[0025] Format standardization addresses the differences in data types and encoding rules among data collected by different emergency terminal devices. Different emergency terminal devices may employ different data collection methods and encoding standards due to their design objectives, manufacturers, and technical standards. For example, some environmental sensors may collect temperature data as decimal integers ranging from 0 to 100, while others may use binary floating-point numbers for higher precision. These data format differences create significant challenges in subsequent data processing and analysis, as data in different formats cannot be directly compared or operated on. Format standardization specifically involves data type conversion and encoding rule unification. Data type conversion involves converting different types of data to a unified data type, such as converting all temperature data to decimal floating-point representation, to ensure data consistency and comparability. Encoding rule unification involves converting data encoded in different formats to a common encoding, such as converting all data to UTF-8, to avoid data parsing errors caused by encoding differences. Format standardization ensures that the data in the resulting standardized data set has consistent data types and encoding rules, facilitating subsequent feature extraction and analysis.
[0026] Step S220: Perform scene semantic analysis on each data in the standardized data set, and extract descriptive information related to the rescue scene type, rescue stage and key scene elements in the data as scene context features.
[0027] Scene semantic parsing and processing involves in-depth mining and analysis of data within standardized datasets to understand the semantic information underlying the rescue scenario. Rescue scenario types refer to the specific scenarios targeted by rescue activities, such as fire rescue, earthquake rescue, flood rescue, and chemical spill rescue. Different rescue scenario types have different characteristics and requirements. For example, fire rescue focuses on the location of the fire source, fire growth, and smoke spread, while earthquake rescue focuses on building collapse, the location of trapped personnel, and injuries. Rescue phases refer to the chronological stages of the rescue process, generally including the early-stage reconnaissance phase, the mid-stage implementation phase, and the late-stage conclusion phase. Data requirements and processing methods vary across different rescue phases. For example, the reconnaissance phase focuses on collecting basic information and environmental data about the rescue site to formulate a rescue plan. The implementation phase requires real-time monitoring of the operating status of rescue equipment and the location of personnel to ensure the smooth progress of the rescue operation. Key elements of a scene refer to factors in a rescue scene that have an important impact on rescue activities. For example, in a fire rescue scene, the location of the fire source, the distribution of flammable materials, and the smooth flow of fire passages are all key elements of the scene.
[0028] Through scene semantic parsing, descriptive information related to the rescue scene type, rescue stage and key elements of the scene is extracted from the standardized data, and this information is used as scene context features.
[0029] As an implementation manner, step S220 may be specifically implemented as the following steps S221 to S226:
[0030] Step S221: Perform semantic keyword extraction processing on the environmental monitoring data in the standardized data set to extract environmental feature description words related to the rescue scene.
[0031] Semantic keyword extraction involves analyzing the text of environmental monitoring data within standardized datasets to identify representative terms related to rescue scenarios. Environmental descriptors are words that describe the characteristics of the rescue scene. These words reflect the state and changes of the environment and are crucial for decision-making.
[0032] In a fire rescue scenario, environmental feature descriptors may include "high temperature", "thick smoke", "lack of oxygen", "fire spread", etc. In an earthquake rescue scenario, environmental feature descriptors may include "building collapse", "ground cracks", "aftershocks", etc. To extract these environmental feature descriptors, keyword extraction algorithms in natural language processing can be used, such as the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm. First, perform text preprocessing on the environmental monitoring data, removing stop words (such as "of", "is", "in", etc.), punctuation marks, and meaningless characters, and converting the text to lowercase to improve the accuracy of keyword extraction. Then, calculate the TF-IDF values of each word in the environmental monitoring data. The TF value represents the frequency of the word in the environmental monitoring data. The higher the frequency, the more important the word. The IDF value represents the rarity of the word in the entire dataset. The higher the rarity, the more representative the word. According to the magnitude of the TF-IDF values, select words with higher values as environmental feature descriptors related to the rescue scenario. By extracting environmental feature descriptors, the key information in the environmental monitoring data can be refined, providing an important basis for subsequent scene type classification and rescue phase identification.
[0033] Step S222: Perform semantic keyword extraction processing on the equipment status data in the standardized data set to extract equipment feature descriptors related to the rescue scenario.
[0034] For the equipment status data in the standardized data set, semantic keyword extraction processing is also required to find the equipment feature descriptors related to the rescue scenario. Equipment feature descriptors are words that can describe the status and performance of rescue equipment. These words reflect the working conditions and usability of the equipment and have an important impact on the implementation of rescue operations.
[0035] In fire rescue, equipment feature descriptors may include "pump normal", "fire truck malfunction", "rescue robot low battery", "fire hose insufficient pressure", etc. In earthquake rescue, equipment feature descriptors may include "crane lifting capacity up to standard", "life detector working properly", "demolition tool damaged", etc. Use a semantic keyword extraction method similar to that for environmental monitoring data, such as the TF-IDF algorithm. First, preprocess the equipment status data, remove noise information, and convert the data to text form. Then calculate the TF-IDF values of each word and select words with higher representativeness according to the magnitude of the values as equipment feature descriptors. By extracting equipment feature descriptors, the status of rescue equipment can be understood in a timely manner, and it can be judged whether the equipment can work normally and whether maintenance or replacement is required, thus ensuring the smooth progress of rescue operations.
[0036] Step S223: performing semantic keyword extraction processing on the personnel positioning data in the standardized data set to extract personnel feature description words related to the rescue scene.
[0037] Personnel descriptors are terms that describe the location and status of rescuers and trapped individuals in a rescue scenario. At large rescue sites, accurate personnel descriptors are crucial for properly deploying rescue forces and ensuring personnel safety.
[0038] At an earthquake rescue site, character descriptions might include phrases like "rescuers near the rubble," "trapped people in the basement," and "injured people at a temporary medical center." At a fire rescue site, character descriptions might include phrases like "firefighters extinguishing the fire at the source" and "trapped people calling for help near the window." Personnel location data is processed using semantic keyword extraction algorithms, such as the TF-IDF algorithm. First, the personnel location data is organized and cleaned, converting it into text. For example, a person's latitude and longitude coordinates are converted into a specific location description. The TF-IDF values of the terms are then calculated, and those with higher values are selected as character descriptions.
[0039] Step S224: Input the environmental feature descriptors, equipment feature descriptors and personnel feature descriptors into the scene type classification model for classification processing to generate a scene type identifier corresponding to the rescue scene type.
[0040] A scene type classification model is a classification model based on machine learning or deep learning. It determines the type of rescue scene based on input environmental, equipment, and personnel descriptors, and generates a corresponding scene type identifier. For example, a scene type classification model can be a support vector machine (SVM), a decision tree, or a convolutional neural network (CNN). Taking the SVM as an example, during the training phase, a large amount of historical rescue data is used as training samples. The environmental, equipment, and personnel descriptors are vectorized and converted into vector representations. For example, a bag-of-words model or word embedding techniques can be used to convert text into vectors. These vectors are used as input features, and the rescue scene type is used as a label to train the SVM model. The SVM model distinguishes different types of rescue scene data by finding the optimal classification hyperplane. After training is complete, the current environmental, equipment, and personnel descriptors are input into the trained SVM model. The model calculates the distance between the input feature vector representations and the classification hyperplanes for different scene types, thereby determining the most likely rescue scene type and generating a corresponding scene type identifier. For example, for a fire rescue scenario, the scenario type identifier might be "01," and for an earthquake rescue scenario, the scenario type identifier might be "02." The scenario type identifier is a concise representation that facilitates subsequent data processing and management. For example, the scenario type identifier can be used to quickly filter out data related to a specific rescue scenario type.
[0041] Step S225: Input the environmental feature descriptors, equipment feature descriptors and personnel feature descriptors into the rescue stage identification model for time series analysis and processing to generate a rescue stage identifier corresponding to the data collection time point.
[0042] The rescue phase recognition model is used to identify the different phases of the rescue process. Because the rescue process is a dynamic process that changes over time, the rescue phase recognition model needs to consider the temporal information of the data. This model can use temporal analysis models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), and gated recurrent units (GRUs).
[0043] Taking the LSTM model as an example, during the training phase, historical rescue data is used to construct training samples. Descriptors of environmental, equipment, and personnel characteristics at different time points are arranged in chronological order to form sequence data. This sequence data is then vectorized and converted into a vector sequence. The LSTM model is trained using the vector sequence as input and the corresponding rescue stage as a label. The LSTM model, with its unique memory unit structure, effectively handles long-term dependencies in sequence data and learns the data characteristics and changing patterns at different stages of the rescue process. In practical applications, the environmental, equipment, and personnel characteristics descriptors at the current data collection time point are input into the trained LSTM model. The model performs time series analysis based on the characteristics of the input sequence and previous state information, determines the current rescue stage, and generates a corresponding rescue stage identifier. For example, the identifier for the early rescue stage of investigation might be "01," the identifier for the mid-rescue stage of implementation might be "02," and the identifier for the late rescue stage of closing might be "03." By generating rescue stage identifiers, we can better understand the rescue process, rationally arrange resources and make decisions. For example, during the reconnaissance stage, we can focus on collecting data to provide a basis for subsequent rescue operations; during the implementation stage, we can monitor the status of rescue equipment and personnel in real time to ensure the smooth progress of the rescue operation.
[0044] Step S226: Merge and integrate the scene type identifier, rescue stage identifier, and scene key element description words to obtain scene context features.
[0045] Scenario key element descriptors provide detailed descriptions of key factors in a rescue scenario. These descriptors further refine and supplement the information contained in the scenario type identifier and rescue phase identifier. For example, in a fire rescue scenario, scenario key element descriptors might include "the fire source is located in the northwest corner of the first floor of the shopping mall," "flammable materials are primarily textiles and plastics," and "the fire escape is blocked by debris."
[0046] In order to merge and integrate the scene type identifier, rescue phase identifier, and scene key element descriptors, this information must first be vectorized. For the scene type identifier and rescue phase identifier, one-hot encoding can be used to convert them into vector representations. One-hot encoding is a method for converting discrete variables into vectors. Each identifier corresponds to a unique vector in which only one element is 1 and the rest are 0. For example, the scene type identifier "01" (representing a fire rescue scene) can be converted into the vector [1,0,0,…,0], where the length of the vector is equal to the total number of scene types. For the scene key element descriptors, word embedding techniques such as Word2Vec, GloVe, or FastText are used to convert them into vector representations. Word embedding techniques can map words in a text into a low-dimensional vector space, so that semantically similar words are closer in the vector space.
[0047] After converting the scene type identifier, rescue phase identifier, and scene key element descriptors into vectors, these vectors are concatenated. Concatenation involves sequentially connecting multiple vectors to form a longer vector. For example, the one-hot encoded vector of the scene type identifier, the one-hot encoded vector of the rescue phase identifier, and the word embedding vector of the scene key element descriptors are concatenated in sequence to produce a comprehensive vector representation. This comprehensive vector is the scene context feature.
[0048] Step S230: Perform storage attribute analysis on each data in the standardized data set, and extract attribute information related to data size, update frequency, and access priority from the data as storage demand characteristics.
[0049] Storage attribute analysis and processing involves detailed analysis and evaluation of data within standardized datasets to determine storage requirements. Data size refers to the amount of storage space occupied by the data. Different data sizes have different storage device requirements. For example, large data volumes may require high-capacity storage devices, such as disk arrays or cloud storage services; whereas smaller data volumes can be stored on smaller but faster-access devices, such as solid-state drives (SSDs). Update frequency refers to the number of times data is updated within a given period. Data with a high update frequency requires more frequent storage and processing to ensure data timeliness. For example, real-time environmental monitoring data may be updated every second, requiring high-speed storage devices and efficient data processing algorithms to ensure real-time data. Access priority refers to the urgency and importance of data in the emergency command system. Data with high access priority requires faster access speeds and higher storage reliability. For example, in a fire rescue scenario, data on the location and spread of a fire are high-priority data.
[0050] Through storage attribute analysis and processing, attribute information related to data size, update frequency, and access priority is extracted from standardized data and used as storage demand characteristics. For example, real-time environmental monitoring data may have a high update frequency and a high access priority. On the other hand, historical rescue data may have a low update frequency and a relatively low access priority.
[0051] As an implementation manner, step S230 may be specifically implemented as the following steps S231 to S234:
[0052] Step S231: Perform data volume statistics processing on each data in the standardized data set, and generate a corresponding data volume size classification identifier based on the comparison result of the number of bytes of the data with the preset data volume threshold.
[0053] Data volume statistics accurately count the bytes of each data item in a standardized data set. Preset data volume thresholds are pre-defined standards for classifying data volume into different levels. These thresholds can be adjusted based on actual storage requirements and storage device characteristics. For example, the small data volume threshold can be set at 1KB, the medium data volume threshold at 10MB, and the large data volume threshold at 100MB.
[0054] Based on the comparison of the data's byte count with a preset data volume threshold, the data is classified into different data volume levels and corresponding data volume classification identifiers are generated. If the data's byte count is less than the small data volume threshold, it is classified as small data volume, and the data volume classification identifier can be "01." If the data's byte count is between the small and medium data volume thresholds, it is classified as medium data volume, and the data volume classification identifier can be "02." If the data's byte count is greater than the medium data volume threshold but less than the large data volume threshold, it is classified as large data volume, and the data volume classification identifier can be "03." If the data's byte count is greater than the large data volume threshold, it is classified as very large data volume, and the data volume classification identifier can be "04." By generating data volume classification identifiers, data storage requirements can be quickly determined and storage resources can be allocated appropriately. For example, small data volumes can be stored in smaller but faster-accessible storage devices, such as memory or SSDs; whereas large data volumes require larger storage devices, such as disk arrays or cloud storage services.
[0055] Step S232: Perform update frequency statistics on each data in the standardized data set, calculate the number of times the data is updated within a preset time window, and generate a corresponding data update frequency classification identifier based on the comparison result of the update number and the preset frequency threshold.
[0056] Update frequency statistics are calculated by counting the number of updates to each data point in the standardized data set within a preset time window. The preset time window is a period of time set based on the characteristics of the data and actual needs, for example, it can be set to 1 minute, 1 hour, or 1 day. The preset frequency thresholds are pre-set standards for classifying data update frequency levels. These thresholds can be adjusted appropriately based on the importance and real-time requirements of the data. For example, the low-frequency update threshold can be set to 1 update per day, the medium-frequency update threshold to 1 update per hour, and the high-frequency update threshold to 1 update per minute.
[0057] Calculate the number of data updates within a preset time window and compare the update count to a preset frequency threshold. If the update count is less than the low-frequency update threshold, the data is classified as low-frequency update data, and the data update frequency classification identifier can be "01." If the update count is between the low-frequency update threshold and the medium-frequency update threshold, the data is classified as medium-frequency update data, and the data update frequency classification identifier can be "02." If the update count is between the medium-frequency update threshold and the high-frequency update threshold, the data is classified as high-frequency update data, and the data update frequency classification identifier can be "03." If the update count is greater than the high-frequency update threshold, the data is classified as ultra-high-frequency update data, and the data update frequency classification identifier can be "04."
[0058] Step S233: Analyze and process the access records of each data in the standardized data set, and collect the number of times the data is called by the emergency command system and the urgency of the call within the historical time period. Generate the corresponding data access priority classification identifier based on the comparison result between the statistical results and the preset priority threshold.
[0059] The access record analysis and processing is to conduct a detailed analysis of the access records of each data in the standardized data set within the historical time period. The data is counted for the number of times it is called by the emergency command system, while taking into account the urgency of the call. The urgency of the call can be evaluated based on the scenario and needs at the time of the call. For example, the data called in emergency rescue situations has a higher urgency; while the data called in daily data analysis and decision support situations has a relatively lower urgency. The preset priority thresholds are pre-set standards for dividing data access priority levels. These thresholds can be adjusted according to the importance of the data and the actual needs of the emergency command system. For example, the low access priority threshold can be set as less than 10 calls per month with a low urgency, the medium access priority threshold can be set as between 10-100 calls per month or a general urgency, and the high access priority threshold can be set as more than 100 calls per month with a high urgency.
[0060] Based on the comparison between the statistical results and the preset priority threshold, the data is classified into different access priority levels and a corresponding data access priority classification identifier is generated. If the number of data calls and the urgency level are lower than the low access priority threshold, it is classified as low access priority data, and the data access priority classification identifier can be "01"; if the number of data calls and the urgency level are between the low access priority threshold and the medium access priority threshold, it is classified as medium access priority data, and the data access priority classification identifier can be "02"; if the number of data calls and the urgency level are between the medium access priority threshold and the high access priority threshold, it is classified as high access priority data, and the data access priority classification identifier can be "03"; if the number of data calls and the urgency level are higher than the high access priority threshold, it is classified as ultra-high access priority data, and the data access priority classification identifier can be "04".
[0061] Step S234: Merge and integrate the data size classification identifier, the data update frequency classification identifier, and the data access priority classification identifier to obtain storage demand characteristics.
[0062] To merge and integrate the data size classification identifiers, data update frequency classification identifiers, and data access priority classification identifiers, a vectorization approach is also used. Each classification identifier is converted into a vector representation using one-hot encoding. For example, the data size classification identifier "01" (indicating a small amount of data) can be converted into the vector [1, 0, 0, ..., 0], where the length of the vector equals the total number of data size classifications. Data update frequency classification identifiers and data access priority classification identifiers are converted using a similar approach.
[0063] After converting all three classification identifiers into vectors, these vectors are concatenated in sequence to form a comprehensive vector representation, and the storage demand characteristics are obtained, which comprehensively reflects the storage requirements of the data, including information such as data size, update frequency, and access priority.
[0064] Step S240: Inputting the scene context features and the storage requirement features into the feature association model for collaborative analysis and processing to generate a feature association descriptor that reflects the intrinsic association relationship between the scene context features and the storage requirement features.
[0065] The feature association model is a model used to analyze the intrinsic relationship between scene context features and storage requirement features. Scene context features reflect the association attributes between data and rescue scenarios, while storage requirement features reflect the storage requirements of data. There is a certain intrinsic connection between the two. For example, the storage requirements for environmental monitoring data may vary at different stages of fire rescue. In the early stages of rescue, the investigation phase may focus more on real-time environmental data, and the update frequency and access priority of its storage requirements are higher. In the later stages of rescue, the storage and analysis of historical environmental data may be more important. The data volume and access frequency of its storage requirements may change, and more emphasis is placed on long-term data preservation and data analysis.
[0066] By inputting the scene context features and storage demand features into the feature association model for collaborative analysis and processing, the intrinsic correlation between the two can be mined to generate a feature association descriptor. The feature association descriptor is a comprehensive information representation that can reflect the degree of correlation and the direction of influence between the scene context features and the storage demand features. For example, the feature association descriptor may include an association strength value and an association direction identifier. The association strength value indicates the degree of correlation between the scene context features and the storage demand features, and the value range may be 0-1. The closer the value is to 1, the closer the correlation is. The association direction identifier indicates the positive or negative influence of the scene context features on the storage demand features. For example, a positive influence indicates that an increase in the scene context features will lead to an increase in the storage demand features, and a negative influence indicates that an increase in the scene context features will lead to a decrease in the storage demand features.
[0067] As an implementation method, step S240 inputs the scene context features and the storage requirement features into the feature association model for collaborative analysis and processing to generate a feature association descriptor reflecting the intrinsic association between the scene context features and the storage requirement features. Specifically, the following steps S241 to S243 may be implemented:
[0068] Step S241: Obtain a historical association record set of stored data in historical emergency command and rescue scenarios, where the historical association record set includes historical scenario context features, historical storage demand features, and actual operation effect feedback information of corresponding storage nodes.
[0069] Historical emergency command and rescue scenarios are various emergency rescue events that occurred in the past, such as fire rescue, earthquake rescue, and flood rescue. In these scenarios, multiple types of emergency terminal devices continuously collect data, which is then processed through feature identification to form historical scenario context features and historical storage demand features. Historical scenario context features reflect the specific circumstances of different historical rescue scenarios. For example, in a certain earthquake rescue, scenario context features may include information such as the earthquake magnitude, the scope of the affected area, and the type of building collapse. This information helps to understand the complexity and characteristics of the rescue scenario. Historical storage demand features reflect the storage requirements of the data at that time. For example, the data volume may be related to the scale of monitoring data in the affected area, the update frequency may be related to the real-time changes in data at the rescue site, and the access priority may be determined based on the degree of reliance of rescue decisions on certain key data.
[0070] The actual operational performance feedback for a storage node records its performance while storing historical data. This includes the node's computational load. For example, when processing large amounts of real-time earthquake monitoring data, metrics like CPU utilization and memory usage can reflect whether its computing power meets demand. It also includes information on storage capacity usage, such as whether the node is experiencing insufficient storage space, and access latency for data storage and reading, to determine whether the node can respond to data access requests in a timely manner.
[0071] To obtain a collection of historically correlated records, relevant information can be extracted from the log files of the emergency command and rescue system. These log files detail the data collection, processing, and storage process for each rescue scenario, as well as the operating status of the storage nodes. This information can also be combined with data from the storage node's monitoring system, which provides real-time records of various performance indicators.
[0072] Step S242: Input the historical association record set into the feature association model for supervised learning training to generate an association weight mapping table of scene context features and storage demand features. The association weight mapping table is used to quantify the degree of influence of different scene context features on storage demand features.
[0073] The feature association model can use a neural network architecture called a multilayer perceptron (MLP). The MLP consists of an input layer, hidden layers, and an output layer. The MLP can have multiple hidden layers and can learn complex nonlinear relationships. This makes it ideal for analyzing the association between scene context features and storage requirements.
[0074] During the training phase, the historical association record set must first be preprocessed. The historical scene context features and historical storage demand features in the historical association record set have been converted into vector form through the previous steps. However, to ensure the effectiveness of training, these vectors must be normalized so that they are in the same dimensional range. For example, all vector element values must be scaled to the interval [0, 1]. The actual operating performance feedback information of the corresponding storage node needs to be converted into a numerical label. For example, if the storage node performs well, it can be marked as 1; if the performance is poor, it can be marked as 0.
[0075] The pre-processed set of historical association records is input into the MLP model for supervised learning training. The training goal of the MLP model is to minimize the error between the storage node operating performance predicted by the model and the actual feedback information. In each training iteration, the model performs forward propagation calculations based on the input historical scene context features and historical storage demand features to obtain the predicted storage node operating performance. Then, by calculating the error between the predicted value and the actual label, the backpropagation algorithm is used to adjust the model's parameter weights. Based on the gradient information of the error, the weight of each neuron in the model is updated in the direction of decreasing error. For example, for the connection weights from the input layer to the first hidden layer, their values are adjusted based on the partial derivatives of the error with respect to these weights, so that the model can more accurately reflect the correlation between the scene context features and the storage demand features in subsequent predictions.
[0076] After multiple iterations of training, the model's error converges to a smaller value or reaches a preset number of training iterations, and the training process ends. At this point, the model has learned the association patterns between scene context features and storage requirement features. Based on the trained model parameters, a weight mapping table for the association between scene context features and storage requirement features is generated.
[0077] The association weight mapping table is a two-dimensional matrix, with rows representing different scenario context features and columns representing different storage requirement features. Each element in the matrix represents the weight of influence of a scenario context feature on a storage requirement feature. For example, a matrix element value of 0.8 indicates that the corresponding scenario context feature has a strong positive influence on the storage requirement feature; a value of -0.3 indicates a moderate negative influence. The association weight mapping table quantifies the degree of influence of different scenario context features on the storage requirement feature, providing a precise quantitative basis for subsequent feature association calculations.
[0078] Step S243: Input the current scene context features and the current storage requirement features into the feature association model, perform weighted association calculation on the scene context features and the storage requirement features based on the association weight mapping table, and generate a feature association descriptor containing an association strength value and an association direction identifier. The association strength value indicates the degree of association between the scene context features and the storage requirement features, and the association direction identifier indicates the positive or negative impact of the scene context features on the storage requirement features.
[0079] After obtaining the association weight mapping table, the current scene context features and the current storage demand features are input into the trained feature association model. First, ensure that the vector representation of the current scene context features and the current storage demand features has the same dimension and scale as the feature vector in the training phase.
[0080] Based on the association weight mapping table, a weighted association calculation is performed on the scene context features and storage demand features. The specific calculation process is as follows:
[0081] For each pair of scene context features and storage requirement features, the corresponding association weight is obtained from the association weight mapping table. Each element in the scene context feature vector is then multiplied by the corresponding element in the storage requirement feature vector, multiplied by the association weight, and finally all products are summed. The absolute value of the weighted association calculation result S is the association strength value, which indicates the closeness of the association between the scene context feature and the storage requirement feature. A larger association strength value indicates a closer association between the two.
[0082] The association direction indicator is determined by the sign of the weighted association calculation result S. If S is greater than 0, the association direction indicator is positive, indicating that the scene context feature has a positive impact on the storage demand feature. That is, an increase in the scene context feature will lead to an increase in the storage demand feature. If S is less than 0, the association direction indicator is negative, indicating that the scene context feature has a negative impact on the storage demand feature. That is, an increase in the scene context feature will lead to a decrease in the storage demand feature. Combining the association strength value and the association direction indicator yields a feature association descriptor. For example, a feature association descriptor can be expressed as (S, sign(S)), where S is the association strength value and sign(S) is the association direction indicator (sign(S) takes a value of 1 for positive and -1 for negative).
[0083] Step S250: The scene context features, storage requirement features and feature association descriptions are combined and integrated to obtain an emergency data feature set.
[0084] To integrate and combine scene context features, storage requirement features, and feature association descriptors, their representation must first be unified. Scene context features and storage requirement features are already in vector form, while feature association descriptors contain an association strength value and an association direction identifier. The association strength value can be treated as a separate element, and the association direction identifier can be converted to a numeric value (e.g., 1 for positive and -1 for negative) and added to a new vector. The scene context feature vector, storage requirement feature vector, and processed feature association descriptor vector are then concatenated in sequence.
[0085] Step S300: Call the pre-built distributed storage strategy model to perform storage node adaptation analysis and processing on the emergency data feature set, and generate a distributed storage mapping scheme that matches the storage node computing power, storage space capacity and access delay characteristics. The distributed storage mapping scheme contains the correspondence information between the data and the target storage node.
[0086] The pre-built distributed storage strategy model is trained based on historical data. Its purpose is to select the most suitable storage node for data based on the emergency data feature set. The computing power, storage capacity, and access latency of a storage node are factors that influence data storage and access performance. Storage nodes with strong computing power can process data faster and are suitable for storing data with high update frequency and requiring complex computations. Storage nodes with large storage capacity can accommodate large amounts of data. Storage nodes with low access latency ensure fast data access. By invoking the distributed storage strategy model with the emergency data feature set as input, the model performs an adaptive analysis of each data item and each storage node based on its learned rules and patterns. During this analysis, the model comprehensively considers the degree of compatibility between the storage node's computing power, storage capacity, and access latency characteristics and the storage requirements of the data in the emergency data feature set. Ultimately, a distributed storage mapping solution is generated, which specifies the target storage node on which each data item should be stored, ensuring efficient data storage and access.
[0087] As an implementation method, the pre-built distributed storage policy model is constructed through the following steps S301 to S305:
[0088] Step S301: Acquire a historical rescue scene perception data set collected by multiple types of emergency terminal devices in historical emergency command and rescue scenarios. The historical rescue scene perception data set includes environmental monitoring data, equipment status data, and personnel positioning data in different historical rescue stages.
[0089] Historical rescue scene perception data is a valuable resource collected from past emergency command and rescue scenarios. Multiple types of emergency terminal devices continuously collect data at different stages of historical rescue operations. This data reflects the actual conditions at the rescue scene at different points in time.
[0090] Environmental monitoring data contains different information at different stages of historical rescue operations. During the early stages of rescue operations, such as the reconnaissance phase, environmental monitoring data may focus on the overall environmental conditions at the rescue site, including basic data such as temperature, humidity, and air quality, to provide rescuers with a preliminary understanding of the level of danger at the site. During the mid-stage implementation phase, more attention may be paid to environmental data directly related to the rescue operation, such as the extent of smoke spread at a fire site or the frequency of aftershocks at an earthquake site. During the final stages of rescue operations, environmental monitoring data may be used to assess the environmental impact of the rescue operation, such as the presence of residual pollutants.
[0091] Equipment status data records the operating status of rescue equipment at different historical rescue stages. During the reconnaissance phase, equipment status data may be primarily used to verify the proper startup and operation of the equipment, such as whether the fire truck's engine can ignite properly and whether the rescue robot's sensors are functioning properly. During the implementation phase, the focus is on the equipment's real-time operating parameters, such as the fire truck's water pump pressure and the rescue robot's power consumption, to ensure the equipment can continuously and stably support the rescue operation. During the finalization phase, equipment status data can be used to assess equipment wear and tear, providing a basis for subsequent maintenance and servicing.
[0092] Personnel location data tracks the changes in the locations of rescuers and trapped individuals throughout the rescue process. During the reconnaissance phase, personnel location data helps determine the initial deployment of rescuers and the approximate distribution of trapped individuals. During the implementation phase, real-time personnel location data helps commanders coordinate rescue operations, avoiding duplication of effort and conflicts among rescuers, and promptly identifying and rescuing trapped individuals. During the final phase, personnel location data is used to confirm the safe evacuation of all rescuers from the scene.
[0093] By obtaining a collection of historical rescue scene perception data, a rich set of training samples is provided for the pre-built distributed storage strategy model, enabling the model to learn the characteristics and storage requirements of data in different rescue scenarios and stages.
[0094] Step S302: Perform feature identification processing on the historical rescue scene perception data set to generate historical scene context features and historical storage demand features.
[0095] The process of feature identification for historical rescue scene perception data sets is similar to that for current rescue scene perception data sets. First, the environmental monitoring data, equipment status data, and personnel location data in the historical rescue scene perception data sets are format-standardized. Since different emergency terminal devices may have adopted different data collection methods and encoding rules in different historical periods, format standardization is required. For example, the temperature data collected by some early environmental sensors may be represented in hexadecimal encoding, while modern sensors may use decimal floating-point representation. Through format standardization, all temperature data are uniformly converted to the same decimal floating-point representation, eliminating differences in data types and encoding rules, and obtaining a standardized historical data set.
[0096] Next, the standardized historical data sets are subjected to scene semantic parsing to extract descriptive information related to the rescue scenario type, rescue phase, and key scene elements, generating historical scene context features. For example, in historical fire rescue scenarios, semantic parsing of temperature and smoke concentration data from environmental monitoring data, combined with the deployment of fire trucks from equipment status data and the location of rescue personnel from personnel positioning data, can determine information such as the scale of the fire, its development trend, and key rescue areas. This information constitutes the historical scene context features.
[0097] At the same time, storage attribute analysis is performed on standardized historical data sets, extracting attribute information related to data size, update frequency, and access priority to generate historical storage demand characteristics. Taking historical environmental monitoring data as an example, the data size over different time periods is counted, the update frequency is analyzed (such as the number of updates per minute, hour, or day), and the access priority is determined based on the data's importance in the historical rescue process. This results in the historical storage demand characteristics of this data.
[0098] By generating historical scenario context features and historical storage demand features, targeted feature information is provided for the pre-built distributed storage strategy model, enabling the model to better understand the characteristics and storage requirements of historical data, and then learn the adaptation relationship between data and storage nodes.
[0099] Step S303: collecting historical operating status information of each storage node in the historical distributed storage cluster, where the historical operating status information includes the historical computing load, historical storage space capacity, and historical access delay of the storage node.
[0100] The historical distributed storage cluster is a distributed storage system used in previous emergency command and rescue processes. The historical operating status information of each storage node reflects its performance in different historical periods.
[0101] The historical computing load of a storage node refers to the computing resource usage of the storage node over a historical period of time. Computational load can be measured using multiple metrics, such as CPU utilization and memory utilization. In historical rescue scenarios, the computing load of storage nodes can vary significantly at different stages. During the early stages of a rescue, during the reconnaissance phase, the CPU utilization of the storage node may be high due to the need to process and analyze large amounts of initial data. During the mid-stage implementation phase of the rescue, the memory utilization of the storage node may increase significantly as data is updated and processed in real time. By analyzing historical computing load, we can understand the computing bottlenecks of storage nodes in different situations, providing a reference for subsequent storage node adaptation.
[0102] Historical storage capacity refers to the amount of available storage space on a storage node during a historical time period. During the historical rescue process, as data is continuously stored and deleted, the storage capacity of a storage node changes dynamically. For example, in the early stages of a rescue, a storage node may need to store a large amount of initial data, causing its storage capacity to decrease rapidly. Later in the rescue process, some temporary data may be deleted, causing the storage capacity to increase accordingly. Understanding historical storage capacity changes helps determine whether a storage node can meet the data storage needs of varying data volumes.
[0103] Historical access latency refers to the delay in accessing data from a storage node over a historical period. Access latency is affected by various factors, such as network bandwidth and the read and write speed of the storage device. In historical rescue scenarios, high-priority data has higher access latency requirements. For example, in a fire rescue scenario, data on the location and spread of the fire must be accessed quickly to support real-time rescue decisions. By analyzing historical access latency, we can evaluate storage node access performance and select appropriate storage nodes for high-priority data.
[0104] Collecting historical operating status information of each storage node in a distributed storage cluster can be achieved through the storage node's monitoring system and log records. The monitoring system can record information such as the storage node's computing load, storage space capacity, and access latency in real time and save this information in log files.
[0105] Step S304: Use historical scene context features, historical storage demand features, and historical operating status information as training samples and input them into the initial model for supervised learning training. The training goal is to minimize the fitness score error between the data and the storage node.
[0106] The initial model can use the random forest algorithm. Random forest is an ensemble learning method consisting of multiple decision trees, each trained on a different subset of samples and features. The final prediction result is obtained by voting or averaging the outputs of multiple decision trees. Random forests offer high accuracy and robustness, can handle high-dimensional data and nonlinear relationships, and are suitable for solving data and storage node adaptation issues.
[0107] Combine historical scene context features, historical storage requirement features, and historical operating status information into training samples. The historical scene context features and historical storage requirement features have already been converted into vector form in the previous steps, and the historical operating status information has also been processed and converted into a numerical feature vector. Each training sample contains the feature information of a data point and the operating status information of the corresponding storage node.
[0108] During the training phase, each training sample is assigned a fitness score, which indicates how well the data matches the storage node. The fitness score can be determined based on historical storage performance. For example, if the data storage and access performance on a storage node is good, the fitness score can be set high; if the performance is poor, the score can be set low.
[0109] The training samples are input into the random forest model for supervised learning. The training process of the random forest model involves randomly extracting a certain number of samples from the training sample to form a training subset for each decision tree. Furthermore, a subset of features is randomly selected as input features for each decision tree to increase model diversity. Each decision tree is trained based on its training subset and input features to construct a decision tree model. The decision tree recursively partitions the features to form different decision rules to predict the fitness score between the data and the storage node. After all decision trees are trained, each one outputs a fitness score prediction for a new input sample. The random forest model averages or votes on the predictions from all decision trees to obtain the final fitness score prediction.
[0110] The training goal is to minimize the error between the model's predicted fitness score and the actual fitness score. By continuously adjusting the structure and parameters of the decision tree in the random forest model, such as the tree's depth and the minimum number of samples per leaf node, the model's prediction error is gradually reduced. The mean squared error (MSE) can be used as a loss function to measure prediction error.
[0111] Through multiple iterative training, the random forest model will gradually learn the adaptation relationship between data and storage nodes, improving the prediction accuracy of the model.
[0112] Step S305: The parameter weights of the initial model are adjusted through iterative optimization until the matching degree between the fitness score output by the model and the actual storage effect reaches a preset training accuracy threshold, thereby generating a pre-built distributed storage strategy model.
[0113] During the training process of a random forest model, the model's parameter weights must be continuously iteratively optimized to improve model performance. These parameter weights include the structural parameters of the decision tree (such as the tree depth and the minimum number of samples per leaf node) and the weight of each decision tree (in the ensemble voting process). This iterative optimization process can utilize the stochastic gradient descent (SGD) algorithm or its variants, such as Adagrad and Adadelta. For example, in each iteration, a mini-batch of training samples is randomly selected and the gradient of the loss function is calculated for that mini-batch. Based on the direction and magnitude of the gradient, the model's parameter weights are adjusted.
[0114] The preset training accuracy threshold is a pre-set error range. When the fitness score output by the model matches the actual storage performance to this threshold, the model training is considered to have achieved satisfactory results. After each iteration, the error between the model's fitness score on the validation set and the actual storage performance is calculated and compared with the preset training accuracy threshold. If the error is less than the preset training accuracy threshold, training is terminated; otherwise, iterative optimization continues. Through continuous iterative optimization, the predictive ability of the random forest model will gradually improve, and the match between the fitness score and the actual storage performance will gradually approach the preset training accuracy threshold. When the preset training accuracy threshold is reached, training is terminated, and the random forest model at this point becomes the pre-built distributed storage strategy model.
[0115] As an implementation manner, step S300 can be specifically implemented as the following steps S310 to S340:
[0116] Step S310: Input the scene context features and storage requirement features in the emergency data feature set into the feature input layer of the distributed storage strategy model, perform feature dimension alignment processing, and obtain a feature input vector with a unified dimensional representation.
[0117] The feature input layer of the distributed storage strategy model is the first layer that receives input data. Although the scenario context features and storage requirement features in the emergency data feature set have been converted to vector form in the previous steps, their dimensions may not be consistent. The purpose of feature dimension alignment is to ensure that the scenario context feature vector and the storage requirement feature vector have the same dimensions so that they can be correctly input into the model for analysis.
[0118] Feature dimension alignment can be done using the following methods:
[0119] 1. Feature selection: Analyze the importance of each feature in the scene context and storage requirement features, select features that have a significant impact on model predictions, and remove irrelevant or redundant features to ensure that the dimensions of the two feature vectors are consistent. Feature selection algorithms, such as those based on correlation analysis, can be used to calculate the correlation coefficient between each feature and the model output (fitness score) and select features with high correlation.
[0120] 2. Feature padding: If the dimensions of the scene context feature vector and the storage requirement feature vector are different, you can pad them with zero vectors or average vectors to make them consistent. For example, if the dimension of the scene context feature vector is m and the dimension of the storage requirement feature vector is n, and m>n, you can pad the storage requirement feature vector with mn zero elements to make its dimension the same as the scene context feature vector.
[0121] 3. Feature transformation: For features with large dimensionality differences, feature transformation methods can be used to convert them to the same dimensional space. For example, principal component analysis (PCA) can be used to reduce the dimensionality of feature vectors, converting high-dimensional feature vectors into low-dimensional feature vectors while retaining most of the important information.
[0122] After feature dimension alignment, the scene context feature vector and the storage requirement feature vector are sequentially concatenated to produce a feature input vector with a unified dimensional representation. This feature input vector contains both the scene context information and storage requirement information for the emergency data, providing a complete feature input for subsequent storage node adaptation analysis.
[0123] Step S320: Through the storage node feature extraction layer of the distributed storage strategy model, the computing power features, storage space capacity features and access delay features of all available storage nodes are extracted to generate a storage node feature set.
[0124] The storage node feature extraction layer of the distributed storage strategy model is responsible for extracting features related to computing power, storage space capacity, and access latency from the relevant information of the storage nodes.
[0125] Computing power can be measured using multiple metrics, such as the number of CPU cores, CPU frequency, memory capacity, and bandwidth. The storage node feature extraction layer can obtain these computing power-related metrics by querying the storage node's system information and monitoring data. For example, by reading the storage node's operating system files, the number of CPU cores and frequency can be obtained; monitoring software can monitor memory usage and bandwidth in real time. These metrics are normalized to the same dimensional range. For example, scaling metrics such as the number of CPU cores, frequency, memory capacity, and bandwidth to the [0, 1] range yields a computing power feature vector.
[0126] The storage capacity feature focuses on the available and total storage space of a storage node. Available and total storage space values can be obtained by querying the storage node's file system information. To facilitate comparison and analysis, the available and total storage space values are normalized. For example, the available storage space is divided by the total storage space to obtain a value representing the available space ratio, which is used as the storage capacity feature.
[0127] Access latency characteristics can be measured using network performance testing tools. For example, use the ping or traceroute command to measure the network latency between the storage node and the data access end. To obtain a more accurate access latency characteristic, perform multiple measurements over different time periods and take the average value as the final access latency characteristic value. Similarly, normalize the access latency characteristic value to ensure it falls within an appropriate dimensional range.
[0128] The computing power feature vector, storage space capacity feature, and access latency feature are combined to form a comprehensive storage node feature vector. The feature extraction process is repeated for all available storage nodes to generate a storage node feature set. This storage node feature set contains detailed feature information for each available storage node, providing a foundation for subsequent adaptation analysis.
[0129] Step S330: Through the adaptation analysis layer of the distributed storage strategy model, the feature input vector and the storage node feature set are associated and matched, and the fitness score between each data and each storage node is calculated. The fitness score is generated based on the degree of matching between the storage demand characteristics of the data and the computing power characteristics, storage space capacity characteristics and access delay characteristics of the storage node.
[0130] The distributed storage strategy model's adaptation analysis layer is the core component of the association matching process. The feature input vector includes the emergency data's scenario context and storage requirements, while the storage node feature set includes the computing power, storage capacity, and access latency characteristics of all available storage nodes.
[0131] The adaptation analysis layer calculates the compatibility score between each data and each storage node based on the degree of match between the data's storage requirements and the storage node's computing power, storage space capacity, and access latency. The specific matching process is as follows:
[0132] 1. Data Size Matching: The data storage requirement features include a data size classification indicator, and the storage node feature set includes a storage space capacity indicator. If the data size classification indicator displays a small file indicator, it indicates a small data size. The adaptation analysis layer then searches for storage nodes with remaining storage space that meets the small file storage requirements. This small file storage requirement can be determined based on experience or system settings. For example, the remaining storage space must be a certain multiple of the small file data size, assuming it is set to 2 times. This ensures that the storage node has sufficient space to store small files while avoiding excessive storage space waste. If the data size classification indicator displays a large file indicator, the adaptation analysis layer matches the node with remaining storage space that meets the large file storage requirements. Large file storage requirements are more stringent and may require not only remaining storage space that is greater than the large file data size but also take into account possible future expansion needs, such as requiring remaining storage space to be at least 3 times the large file data size. This ensures that large files are stored completely and stably on appropriate storage nodes.
[0133] 2. Update frequency matching: The data update frequency classification identifier reflects the frequency of data updates. For data marked as frequently updated, the storage node's computing power is relatively high. The adaptation analysis layer searches for nodes whose computing power meets the high-frequency processing conditions. High-frequency processing conditions typically involve indicators such as the storage node's CPU performance and memory bandwidth. For example, the storage node is required to have a large number of CPU cores, a high main frequency, and memory bandwidth that can support the rapid reading and writing of high-frequency data. Such storage nodes can promptly process frequently updated data, ensuring the real-time and accuracy of the data. If the data is marked as infrequently updated, the adaptation analysis layer matches nodes whose computing power meets the low-frequency processing conditions. Low-frequency processing conditions are relatively relaxed, with lower requirements for CPU and memory. This allows for the rational use of storage node resources and prevents high-performance storage nodes from being occupied by low-demand data.
[0134] 3. Access priority matching: The access priority classification identifier of the data reflects the importance and urgency of the data in the emergency command system. When the data is identified as an emergency access, it means that the data needs to be accessed quickly in a short period of time, and the adaptation analysis layer will match the nodes whose access delay meets the emergency response conditions. The emergency response conditions require that the access delay of the storage node is extremely low, for example, the access delay must be less than 10 milliseconds. This can be achieved by optimizing the network configuration of the storage node, using high-speed storage devices, etc. If the data is identified as a regular access, the adaptation analysis layer will select the node whose access delay meets the regular response conditions. The access delay requirements of the regular response conditions are relatively low, generally ranging from tens of milliseconds to hundreds of milliseconds. This can reduce storage costs while ensuring normal access to data.
[0135] After completing the matching of the three dimensions described above, the adaptation analysis layer calculates the matching scores between the data and the storage node in the three dimensions of storage space, computing power, and access latency based on the matching rules. The matching scores for each dimension can be quantified based on the degree of match. For example, in the storage space dimension, if the storage node's remaining storage space fully meets the data's storage requirements, the matching score can be set to 100 points; if it partially meets the requirements, a corresponding score is assigned based on the proportion of satisfaction. A similar method is used to score the computing power and access latency dimensions. The matching scores for each dimension are then assigned a corresponding weight. This weighting can be adjusted based on actual circumstances. For example, for high-priority data, the weight of the access latency dimension can be appropriately increased; for large data volumes, the weight of the storage space dimension can be increased. The matching scores for each dimension are multiplied by the corresponding weights and then weighted summed to form the compatibility score between the data and the storage node. For example, assuming the matching score for the storage space dimension is 80 points with a weight of 0.3; the matching score for the computing power dimension is 70 points with a weight of 0.3; and the matching score for the access latency dimension is 90 points with a weight of 0.4, then the adaptability score is 80 × 0.3 + 70 × 0.3 + 90 × 0.4 = 81 points.
[0136] Step S340: The distributed storage strategy model's allocation decision layer uses a greedy algorithm to allocate data based on the fitness scores. For each data item, the storage node with the highest fitness score is selected as the target storage node, generating a distributed storage mapping solution that includes the correspondence between the data item and the target storage node. After receiving the fitness scores for each data item and each storage node output by the fitness analysis layer, the distributed storage strategy model's allocation decision layer uses a greedy algorithm to allocate data storage nodes. The core idea of the greedy algorithm is to take the optimal choice under the current state at each selection step, hoping to ultimately achieve the global optimal solution. In this scenario, for each data item, the allocation decision layer directly selects the storage node with the highest fitness score as the target storage node. For example, if data A has a fitness score of 85 with storage node 1, 78 with storage node 2, and 82 with storage node 3, the allocation decision layer will select storage node 1 as the target storage node for data A. After traversing all data, the allocation decision layer will organize the correspondence between each piece of data and its corresponding target storage node, forming a distributed storage mapping scheme. This scheme clearly defines the storage location of each piece of data in the distributed storage system and provides detailed guidance for subsequent data distribution and storage operations. This greedy algorithm allocation process can, to a certain extent, ensure that data is stored on the most appropriate storage node, improving data storage efficiency and access performance, while also fully utilizing the resources of each storage node in the distributed storage system.
[0137] Step S400: Distribute each data in the rescue scene perception data set to the corresponding target storage node according to the distributed storage mapping solution, and obtain a distributed data storage cluster with a multi-node collaborative storage structure.
[0138] The distributed storage mapping scheme provides clear targets and paths for data distribution. Rescue scene perception data sets include a large amount of environmental monitoring data, equipment status data, and personnel location data. This data needs to be accurately distributed to its respective target storage nodes according to the mapping scheme. During the distribution process, data integrity and accuracy must be ensured to avoid data loss or corruption. The multi-node collaborative storage structure is a key feature of distributed data storage clusters. Multiple storage nodes work together to share data storage tasks, improving data storage reliability and performance. Different storage nodes can store different types of data based on their own characteristics and advantages, achieving rational data distribution and efficient utilization.
[0139] As an implementation manner, step S400 can be specifically implemented as the following steps S410 to S450:
[0140] Step S410: performing data sharding processing on each data in the rescue scene perception data set, dividing the large file data into multiple data sharding units, and keeping the small file data in the original data form, to obtain the data set to be stored.
[0141] Data sharding is designed to better meet the needs of distributed storage. Due to the large amount of data, direct storage of large file data may put a lot of pressure on the performance and storage space of the storage node. After splitting the large file into multiple data shard units, each shard unit can be stored independently on different storage nodes, which can improve the parallelism and flexibility of data storage. For example, a large file of 1GB can be split into 10 data shard units of 100MB. Small file data does not need to be split due to its small amount of data, and can be kept in the original data form, which can reduce the complexity of data processing. After data sharding, the data set to be stored contains the shard units of the large file data and the original form of the small file data, providing a suitable data form for subsequent data distribution and storage.
[0142] Step S420: According to the correspondence information between data and target storage nodes in the distributed storage mapping solution, a unique storage path identifier is allocated to each data to be stored. The storage path identifier includes the target storage node address and the storage location information of the data in the node.
[0143] The distributed storage mapping solution records in detail which target storage node each piece of data should be stored on. Based on this correspondence information, a unique storage path identifier is generated for each piece of data to be stored. The target storage node address is a unique identifier for the storage node in the network, such as an IP address, which allows data to be accurately transferred to the target storage node. The data's storage location information within the node specifies the specific storage location of the data on the target storage node, such as the folder path and file name in the file system. By assigning a unique storage path identifier to each piece of data to be stored, it can be ensured that the data is accurately stored in the designated location of the target storage node, facilitating subsequent data access and management.
[0144] Step S430: Distribute the data to be stored to the corresponding target storage node according to the storage path identifier through the data transmission protocol.
[0145] During the distribution process, the data to be stored is sent to the corresponding target storage node based on the target storage node address in the storage path identifier. To enhance data transmission security, data encryption technology can be employed to encrypt the data before transmission. For example, SSL / TLS encryption protocols can be used to encrypt data to prevent it from being eavesdropped or tampered with during transmission.
[0146] Step S440: Execute a data storage confirmation operation in each target storage node to verify the integrity and consistency of the data and generate a storage success flag or a storage failure flag.
[0147] Data storage confirmation is a critical step in ensuring data is stored correctly. After the target storage node receives the data to be stored, it needs to verify the integrity and consistency of the data. Data integrity refers to the absence of data loss or corruption during transmission and storage, while data consistency refers to the consistency of data content and metadata.
[0148] As an implementation manner, step S440 may be specifically implemented as the following steps S441 to S444:
[0149] Step S441: performing hash value calculation processing on the data to be stored in the target storage node to generate a real-time hash value of the data.
[0150] A hash algorithm, such as SHA-256, is used to calculate the hash value of the data to be stored in the target storage node. The SHA-256 algorithm has high security and uniqueness, ensuring that different data generates different hash values.
[0151] Step S442: extract the original hash value of the data to be stored before distribution, and compare the real-time hash value with the original hash value. If the two are consistent, it is determined that the data integrity verification is passed.
[0152] Before data distribution, the original hash value of the data to be stored is pre-calculated and saved. On the target storage node, the original hash value of the data to be stored is extracted and compared with the real-time hash value. If the real-time hash value matches the original hash value, the data has not changed during transmission and storage, and the data integrity verification is considered passed. If the two do not match, the data may have been lost or damaged, and the data integrity verification is considered failed. By comparing hash values, data integrity can be simply and effectively verified, ensuring that the data stored on the target storage node is consistent with the original data.
[0153] Step S443: Check the metadata information of the data to be stored to verify the matching between the metadata information and the data content.
[0154] Metadata for stored data is descriptive information about the data, such as the file name, file type, creation time, and modification time. Metadata for stored data is verified to ensure consistency with the data content. For example, the file name is checked to ensure it matches the actual data content and the file type is correct. Mismatches between metadata and data can lead to problems with data management and usage. Verifying the consistency of metadata with data ensures data consistency and improves data availability and reliability.
[0155] Step S444: If the metadata information matches the data content, the data consistency verification is determined to be passed; if both the integrity verification and the consistency verification are passed, a storage success flag is generated; if any one of the verifications fails, a storage failure flag is generated.
[0156] Based on the comparison results of the metadata information and data content, it is determined whether the data consistency verification has passed. If the metadata information matches the data content, the data consistency verification is considered to have passed; otherwise, the data consistency verification is considered to have failed. Combining the results of the data integrity verification and data consistency verification, if both pass, it means that the data was not lost or damaged during the transmission and storage process, and the metadata information and data content match each other, and a storage success flag is generated. If either verification fails, a storage failure flag is generated. The storage success flag and storage failure flag can be used for subsequent data management and processing, providing reliable feedback information for data storage and use.
[0157] Step S450: retransmitting the data to be stored that failed to be stored until all the data to be stored are successfully stored in the corresponding target storage nodes, thereby obtaining a distributed data storage cluster.
[0158] Data to be stored that generates a storage failure flag requires retransmission. This retransmission process redistributes the failed data to the corresponding target storage node according to the storage path flag. During the retransmission process, data integrity and consistency verification are also performed to ensure correct data storage. This retransmission and verification process is repeated until all data to be stored is successfully stored on the target storage node. After retransmission, all data is accurately stored on the target storage node, forming a distributed data storage cluster with a multi-node collaborative storage structure.
[0159] Step S500: collecting real-time operating status information of the distributed data storage cluster, dynamically optimizing and adjusting the distributed storage mapping scheme based on the real-time operating status information and the emergency data feature set, and generating an updated distributed storage mapping scheme.
[0160] During the operation of a distributed data storage cluster, the status of its storage nodes changes over time. For example, the computational load of a storage node may increase or decrease due to data read and write operations, storage space capacity may change as data is stored and deleted, and access latency may be affected by network conditions and storage device performance. Collecting real-time operational status information from the distributed data storage cluster provides timely insights into the current status of storage nodes. This real-time operational status information includes the current computational load, remaining storage space, and real-time access latency of storage nodes. Based on this real-time operational status information and the emergency data feature set, the existing distributed storage mapping scheme is dynamically optimized and adjusted. The emergency data feature set includes information such as the data's scenario context, storage requirement characteristics, and feature association descriptors, reflecting the data's characteristics and storage requirements. By combining real-time operational status information with the emergency data feature set, it is possible to determine whether the current distributed storage mapping scheme is still appropriate. If it is found that certain storage nodes are overloaded or have insufficient storage space, or if the storage requirements of certain data have changed, the distributed storage mapping scheme needs to be adjusted to ensure that data is always stored on the most appropriate storage nodes. After dynamic optimization and adjustment, an updated distributed storage mapping solution is generated, which can better adapt to the real-time operation status of the distributed data storage cluster and improve the efficiency and reliability of data storage.
[0161] As an implementation manner, step S500 can be specifically implemented as the following steps S510 to S550:
[0162] Step S510: collecting real-time operation status information of each storage node in the distributed data storage cluster, where the real-time operation status information includes the current computing load, remaining storage space size, and real-time access delay of the storage node.
[0163] The real-time operating status information of each storage node in the distributed data storage cluster is collected through the storage node's monitoring system. The current computing load of the storage node can be obtained by monitoring indicators such as CPU usage and memory usage. For example, use performance monitoring software to monitor the CPU usage of the storage node in real time to understand the usage of its computing resources. The remaining storage space size can be queried through the storage node's file system management tool to obtain the current available storage space size of the storage node. Real-time access latency can be measured using network performance testing tools, such as using the ping command or traceroute command to measure the latency when data is accessed from the storage node. By collecting this real-time operating status information, the current status of the storage node can be fully understood, providing a basis for subsequent adjustments to the storage mapping solution.
[0164] Step S520: input the real-time operation status information and the emergency data feature set into the storage status analysis model for joint analysis and processing, and identify the storage load imbalance nodes and storage space insufficient nodes in the current distributed storage mapping solution.
[0165] The storage status analysis model is a model specifically used to analyze the status of storage nodes in a distributed storage cluster. Real-time operating status information and emergency data feature sets are input into the model for joint analysis and processing. Real-time operating status information reflects the current status of the storage node, and the emergency data feature set reflects the data storage requirements and associated attributes with the scenario. Through joint analysis and processing, the model can determine whether the current distributed storage mapping scheme is reasonable. Nodes with unbalanced storage load refer to storage nodes with excessively high or low computing loads. For example, the CPU utilization of some storage nodes remains high for a long time, while the CPU utilization of other storage nodes is low, which indicates that there is an imbalance in storage load. Nodes with insufficient storage space refer to storage nodes with remaining storage space less than a certain threshold. Through analysis by the storage status analysis model, the nodes with unbalanced storage load and insufficient storage space in the current distributed storage mapping scheme are identified, providing targets for subsequent adjustments.
[0166] As an implementation manner, step S520 may be specifically implemented as the following steps S521 to S526:
[0167] Step S521: performing threshold determination processing on the current computing load of the storage nodes in the real-time operation status information, and screening out the storage nodes whose computing load exceeds a preset load threshold as storage load imbalance nodes.
[0168] The preset load threshold is a pre-set critical value for computing load, used to determine whether a storage node's computing load is excessive. A threshold determination is performed on the current computing load of storage nodes in real-time operational status information, comparing each storage node's computing load with the preset load threshold. If a storage node's computing load exceeds the preset load threshold, for example, if its CPU utilization exceeds 80%, it is identified as a node with an unbalanced storage load. By identifying nodes with unbalanced storage load, load issues within the storage cluster can be promptly identified, providing a target for subsequent adjustments.
[0169] Step S522: performing a threshold determination process on the remaining storage space of the storage nodes in the real-time operation status information, and screening out storage nodes whose remaining storage space is less than a preset space threshold as nodes with insufficient storage space.
[0170] The preset space threshold is a pre-set critical value for storage space used to determine whether a storage node has sufficient remaining storage space. A threshold determination is performed on the remaining storage space of storage nodes in real-time operational status information, comparing each storage node's remaining storage space with the preset space threshold. If a storage node's remaining storage space is less than the preset space threshold, for example, less than 10% of the total storage space, the node is identified as insufficiently spaced. By identifying insufficiently spaced nodes, space issues within the storage cluster can be promptly identified, preventing data storage failures caused by insufficient storage space.
[0171] Step S523: Perform data feature backtracking processing on the target storage data corresponding to the nodes with unbalanced storage load and insufficient storage space, and extract the scene context features and storage demand features of the corresponding data.
[0172] For nodes identified as experiencing imbalanced storage loads and insufficient storage space, further analysis is required regarding their corresponding target storage data. Data feature retrospective processing is used to extract the scenario context and storage requirement characteristics of this data from the storage system. Scenario context characteristics help understand the attributes associated with the data and the rescue scenario, while storage requirement characteristics reflect the data's storage requirements, such as data size, update frequency, and access priority. By extracting these characteristics, we can better understand the mismatch between data and storage nodes, providing a basis for subsequent adjustments.
[0173] Step S524: Analyze the matching relationship between the computing load of the node with unbalanced storage load and the corresponding data storage demand characteristics to determine whether the load is too high due to the mismatch between the data storage demand and the storage node capacity.
[0174] Perform a detailed analysis of the computing load of nodes with imbalanced storage loads and the storage requirements of the corresponding data. For example, check whether the data update frequency is too high, and the storage node's computing power cannot meet the processing requirements of this high-frequency data; or whether the data has a high access priority and requires a fast response, but the storage node's access latency is too high. By analyzing the matching relationship, determine whether the excessive load is caused by a mismatch between the data's storage requirements and the storage node's capabilities. If this is the case, adjust the data's storage location and migrate the data to a more suitable storage node.
[0175] Step S525: Analyze the matching relationship between the remaining storage space size of the node with insufficient storage space and the corresponding data storage demand characteristics to determine whether the insufficient space is caused by the mismatch between the data storage demand and the storage node space.
[0176] For nodes with insufficient storage space, analyze the relationship between their remaining storage space and the storage requirements of the corresponding data. Check whether the data volume exceeds the remaining storage space of the storage node or whether there is excessive concentration of data on a single node. If insufficient space is caused by a mismatch between the data storage requirements and the storage node space, consider migrating some data to other nodes with sufficient storage space or expanding the storage node capacity.
[0177] Step S526: Integrate the information of the nodes with unbalanced storage loads and the nodes with insufficient storage space and their corresponding mismatched data to generate a storage node-data association list that needs to be adjusted.
[0178] The information on nodes with unbalanced storage loads and insufficient storage space obtained previously is integrated with the contextual characteristics and storage requirements of the corresponding mismatched data. A list of storage node-data associations requiring adjustment is generated. This list details which storage nodes are experiencing issues and which data on these nodes requires adjustment. This list provides a clear goal and basis for subsequent data reallocation and storage mapping solution adjustments.
[0179] Step S530: re-feature extraction processing is performed on the target storage data corresponding to the nodes with unbalanced storage load and the nodes with insufficient storage space, to generate an adjusted data feature set including adjusted scene context features and adjusted storage demand features.
[0180] Since the status of the storage nodes has changed and the matching relationship between the data and the storage nodes has been analyzed, it is necessary to re-extract features of the target storage data corresponding to the nodes with unbalanced storage load and insufficient storage space. The scene context features and storage requirement features of the data are re-extracted, and the adjusted scene context features and adjusted storage requirement features are generated taking into account the current actual situation and adjustment needs. For example, if it is found that the computing load of a storage node is too high due to the storage of a large amount of frequently updated data, the update frequency classification identifier of this data can be appropriately adjusted when re-extracting features. The adjusted scene context features and adjusted storage requirement features are integrated together to form an adjusted data feature set. This set reflects the storage requirements of the adjusted data and the associated attributes with the scene, providing new feature information for subsequent storage node adaptation analysis.
[0181] Step S540: calling the distributed storage strategy model to perform adaptation analysis on the adjusted data feature set and the storage node feature sets of the remaining available storage nodes, and generating new correspondence information between the adjusted data and the remaining available storage nodes.
[0182] The generated feature set for the adjusted data and the feature set for the remaining available storage nodes are input into the distributed storage strategy model for adaptation analysis. The distributed storage strategy model rematches the adjusted data with the remaining available storage nodes based on its internally learned rules and patterns. During this matching process, the degree of compatibility between the storage nodes' computing power, storage capacity, and access latency characteristics and the storage requirements of the adjusted data is considered. Ultimately, a new correspondence between the adjusted data and the remaining available storage nodes is generated, clarifying which adjusted data should be stored on which remaining available storage nodes.
[0183] Step S550: merging and integrating the new correspondence information with the correspondence information of the unadjusted data in the original distributed storage mapping scheme to generate an updated distributed storage mapping scheme.
[0184] The new correspondence information between the adjusted data and the remaining available storage nodes is merged and integrated with the correspondence information between the unadjusted data in the original distributed storage mapping scheme. The correspondence information for the unadjusted data remains unchanged, and the new correspondence information for the adjusted data is added to the original scheme. Through this merger and integration, an updated distributed storage mapping scheme is generated. This scheme takes into account the real-time operating status of the distributed data storage cluster and the actual storage needs of the data. It can better achieve reasonable data storage and efficient management, improve the performance and reliability of the distributed storage system, and provide more stable data support for emergency command and rescue.
[0185] It is understandable that the various algorithms involved in the above-mentioned introductions of the embodiments of the present invention, such as greedy algorithms, hash algorithms, feature selection algorithms, etc., can all be learned from the relevant content in the prior art. In order to save space, they will not be expanded too much in the embodiments of the present invention. In addition, when implementing the scheme of the present invention, those skilled in the art can supplement the details according to the common knowledge in this field. For example, according to the common knowledge in this field, normalization can be used to eliminate dimensional conflicts before feature fusion, interpolation can be used to eliminate dimensional differences, and thresholds can be reasonably set based on historical data, experience or business scenario requirements. The model can be trained based on a general model training method, and the number of layers in the model structure can be set based on actual needs, the activation function can be selected, etc. The present invention will no longer provide redundant introductions to the overly detailed implementation process.
[0186] See also Figure 2 , Figure 2This is a schematic diagram of the structure of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 may be connected via a bus or other means. The processor 101 (also known as the Central Processing Unit (CPU)) is the computing and control core of the computer system, capable of parsing various instructions within the computer system and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, a mobile communication interface, etc.), which can be used to send and receive data under the control of the processor 101. The communication interface 102 may also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system for storing programs and data. It is understood that the memory 103 herein may include both the built-in memory of the computer system and, of course, the extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system, but this is not limited to this in the present invention.
[0187] In one embodiment, the processor 101 executes the artificial intelligence-based distributed storage method for emergency command and rescue data provided in the above embodiment of the present invention by running the computer program in the memory 103.
Claims
1. A distributed storage method for emergency command and rescue data based on artificial intelligence, characterized in that: The method comprises: Acquire a rescue scene perception data set collected in real time by multiple types of emergency terminal devices, wherein the rescue scene perception data set includes environmental monitoring data, equipment status data, and personnel positioning data at different rescue stages; Performing feature identification processing on the rescue scene perception data set to generate an emergency data feature set including scene context features and storage requirement features, wherein the scene context features are used to characterize the association attributes between the data and the rescue scene, and the storage requirement features are used to characterize the storage priority and access frequency attributes of the data; Call the pre-built distributed storage strategy model to perform storage node adaptation analysis on the emergency data feature set, and generate a distributed storage mapping scheme that matches the storage node computing power, storage space capacity and access delay characteristics. The distributed storage mapping scheme contains the corresponding relationship information between data and target storage nodes; specifically, it includes: inputting the scene context features and storage demand features in the emergency data feature set into the feature input layer of the distributed storage strategy model, performing feature dimension alignment processing, and obtaining a feature input vector with a unified dimensional representation; extracting the computing power features and storage space capacity features of all available storage nodes through the storage node feature extraction layer of the distributed storage strategy model. and access delay characteristics to generate a storage node feature set; through the adaptation analysis layer of the distributed storage strategy model, the feature input vector and the storage node feature set are associated and matched, and a fitness score between each data and each storage node is calculated, and the fitness score is generated based on the degree of matching between the storage demand characteristics of the data and the computing power characteristics, storage space capacity characteristics and access delay characteristics of the storage node; through the allocation decision layer of the distributed storage strategy model, a greedy algorithm allocation process is performed according to the fitness score, and the storage node with the highest fitness score is selected as the target storage node for each data, and a distributed storage mapping scheme containing the correspondence between the data and the target storage node is generated; Distributing each data in the rescue scene perception data set to a corresponding target storage node according to the distributed storage mapping scheme to obtain a distributed data storage cluster with a multi-node collaborative storage structure; Real-time operating status information of the distributed data storage cluster is collected, and the distributed storage mapping scheme is dynamically optimized and adjusted based on the real-time operating status information and the emergency data feature set to generate an updated distributed storage mapping scheme.
2. The artificial intelligence-based distributed storage method for emergency command and rescue data according to claim 1 is characterized in that: The feature identification processing of the rescue scene perception data set to generate an emergency data feature set including scene context features and storage requirement features includes: Standardizing the formats of the environmental monitoring data, equipment status data, and personnel positioning data in the rescue scene perception data set to eliminate differences in data types and encoding rules among data collected by different emergency terminal devices, thereby obtaining a standardized data set; Performing scene semantic parsing on each data in the standardized data set, extracting descriptive information related to the rescue scene type, rescue stage, and key scene elements from the data as scene context features; Performing storage attribute analysis on each data in the standardized data set, extracting attribute information related to data size, update frequency, and access priority from the data as storage demand characteristics; Inputting the scene context feature and the storage requirement feature into a feature association model for collaborative analysis and processing to generate a feature association descriptor reflecting the intrinsic association relationship between the scene context feature and the storage requirement feature; The scene context feature, the storage requirement feature and the feature association description are combined and integrated to obtain the emergency data feature set.
3. The artificial intelligence-based distributed storage method for emergency command and rescue data according to claim 2 is characterized in that: The step of performing scene semantic parsing on each data in the standardized data set and extracting descriptive information related to the rescue scene type, rescue stage, and scene key elements from the data as scene context features includes: Performing semantic keyword extraction processing on the environmental monitoring data in the standardized data set to extract environmental feature descriptors related to the rescue scene; Performing semantic keyword extraction processing on the device status data in the standardized data set, extracting attribute information related to data size, update frequency, and access priority from the data as device feature descriptors related to the rescue scenario; Performing semantic keyword extraction processing on the personnel positioning data in the standardized data set to extract personnel feature description words related to the rescue scene; Inputting the environmental feature descriptors, equipment feature descriptors, and personnel feature descriptors into a scene type classification model for classification processing to generate a scene type identifier corresponding to the rescue scene type; Inputting the environmental feature descriptors, equipment feature descriptors, and personnel feature descriptors into a rescue stage identification model for time series analysis and processing to generate a rescue stage identifier corresponding to the data collection time point; The scene type identifier, rescue stage identifier and scene key element description words are merged and integrated to obtain the scene context feature.
4. The artificial intelligence-based distributed storage method for emergency command and rescue data according to claim 2 is characterized in that: The performing of storage attribute analysis on each data in the standardized data set and extracting attribute information related to data size, update frequency, and access priority from the data as storage demand characteristics includes: Performing data volume statistics on each data in the standardized data set, and generating a corresponding data volume classification identifier based on a comparison result of the number of bytes of the data with a preset data volume threshold; Performing update frequency statistical processing on each data in the standardized data set, calculating the number of times the data is updated within a preset time window, and generating a corresponding data update frequency classification identifier based on a comparison result of the number of updates with a preset frequency threshold; Performing access record analysis on each data in the standardized data set, collecting the number of times the data was called by the emergency command system and the urgency of the calls within a historical period, and generating a corresponding data access priority classification identifier based on a comparison result between the statistical results and a preset priority threshold; The data size classification identifier, the data update frequency classification identifier and the data access priority classification identifier are combined and integrated to obtain the storage demand characteristics.
5. The artificial intelligence-based distributed storage method for emergency command and rescue data according to claim 1 is characterized in that: The adaptation analysis layer of the distributed storage strategy model performs association matching processing on the feature input vector and the storage node feature set, and calculates the fitness score between each data and each storage node, including: Extracting storage requirement features of the data in the feature input vector, including data size classification identifier, data update frequency classification identifier, and data access priority classification identifier; Extracting computing capability features, storage space capacity features, and access delay features of a single storage node from the storage node feature set; Establish matching rules between data size classification identifiers and storage space capacity characteristics. If the data is identified as a small file, match the nodes whose remaining storage space meets the small file storage conditions. If the data is identified as a large file, match the nodes whose remaining storage space meets the large file storage conditions. Establish matching rules between data update frequency classification identifiers and computing power characteristics. If the data has a high-frequency update identifier, match nodes whose computing power meets the high-frequency processing conditions. If the data has a low-frequency update identifier, match nodes whose computing power meets the low-frequency processing conditions. Establish matching rules between data access priority classification identifiers and access delay characteristics. If the data has an emergency access identifier, match nodes whose access delays meet emergency response conditions. If the data has a regular access identifier, match nodes whose access delays meet regular response conditions. The matching scores between the data and the storage node in three dimensions, namely, storage space, computing power and access delay, are calculated according to the matching rules, and the matching scores in the three dimensions are weighted and summed to generate the fitness score.
6. The artificial intelligence-based distributed storage method for emergency command and rescue data according to claim 1 is characterized in that: The collecting of real-time operating status information of the distributed data storage cluster, dynamically optimizing and adjusting the distributed storage mapping scheme based on the real-time operating status information and the emergency data feature set, and generating an updated distributed storage mapping scheme includes: Collecting real-time operating status information of each storage node in the distributed data storage cluster, the real-time operating status information including the current computing load, remaining storage space size and real-time access delay of the storage node; Inputting the real-time operation status information and the emergency data feature set into a storage status analysis model for joint analysis and processing to identify nodes with unbalanced storage loads and nodes with insufficient storage space in the current distributed storage mapping solution; Re-feature extraction is performed on the target storage data corresponding to the nodes with unbalanced storage load and insufficient storage space to generate an adjusted data feature set including adjusted scene context features and adjusted storage demand features; Calling the distributed storage strategy model to perform adaptation analysis on the adjusted data feature set and the storage node feature sets of the remaining available storage nodes, and generating new correspondence relationship information between the adjusted data and the remaining available storage nodes; The new correspondence information is combined and integrated with the correspondence information of the unadjusted data in the original distributed storage mapping scheme to generate the updated distributed storage mapping scheme.
7. The artificial intelligence-based distributed storage method for emergency command and rescue data according to claim 6 is characterized in that: The step of inputting the real-time operation status information and the emergency data feature set into a storage status analysis model for joint analysis and processing to identify nodes with unbalanced storage loads and nodes with insufficient storage space in the current distributed storage mapping solution includes: Performing threshold judgment processing on the current computing load of the storage nodes in the real-time operation status information, and screening out the storage nodes whose computing load exceeds a preset load threshold as storage load imbalance nodes; Performing a threshold determination process on the remaining storage space of the storage nodes in the real-time operation status information, and screening out storage nodes whose remaining storage space is less than a preset space threshold as nodes with insufficient storage space; Performing data feature backtracking processing on the target storage data corresponding to the nodes with unbalanced storage load and insufficient storage space, and extracting scene context features and storage demand features of the corresponding data; Analyze the matching relationship between the computing load of the node with the imbalanced storage load and the corresponding data storage demand characteristics to determine whether the excessive load is caused by a mismatch between the data storage demand and the storage node capacity; Analyze the matching relationship between the remaining storage space size of the node with insufficient storage space and the corresponding data storage demand characteristics to determine whether the insufficient space is caused by the mismatch between the data storage demand and the storage node space; The nodes with unbalanced storage loads and nodes with insufficient storage space and their corresponding mismatched data information are integrated to generate a storage node-data association list that needs to be adjusted.
8. The artificial intelligence-based distributed storage method for emergency command and rescue data according to claim 1 is characterized in that: The method of distributing each data in the rescue scene perception data set to a corresponding target storage node according to the distributed storage mapping scheme to obtain a distributed data storage cluster with a multi-node collaborative storage structure includes: Performing data sharding on each data in the rescue scene perception data set, dividing the large file data into multiple data shard units, and keeping the small file data in the original data form, to obtain a data set to be stored; According to the correspondence information between the data and the target storage node in the distributed storage mapping scheme, a unique storage path identifier is assigned to each data to be stored, wherein the storage path identifier includes the target storage node address and the storage location information of the data in the node; Distributing the data to be stored to the corresponding target storage node according to the storage path identifier through a data transmission protocol; Perform data storage confirmation operations in each target storage node to verify the integrity and consistency of the data and generate a storage success flag or a storage failure flag; The data to be stored that failed to be stored is retransmitted until all the data to be stored are successfully stored in the corresponding target storage nodes, thereby obtaining the distributed data storage cluster.
9. A computer system, characterized in that: include: a memory storing a computer program; A processor, configured to load the computer program to implement the artificial intelligence-based distributed storage method for emergency command and rescue data as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Big data-based file analysis, storage and division mode and distributed storage system
CN119861881A
Marine first aid auxiliary diagnosis method and system based on artificial intelligence
CN120021940A
Meteorological metadata storage method and system based on machine learning
CN120104579A