Real-time data distributed storage method based on double-spectrum bird detection system

By building a hierarchical distributed storage architecture and intelligent scheduling technology, the problems of insufficient storage capacity, network congestion and single-point failure in the dual-spectrum bird detection system are solved, real-time, reliable and efficient storage of data is achieved, and the overall level of data management is improved.

CN120407859AInactive Publication Date: 2025-08-01BEIJING JIRUIXIANG AVIATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510497945.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional centralized storage method has problems such as insufficient storage capacity, network congestion, single point failure and low data management efficiency in the dual-spectrum bird detection system, which affects the real-time and integrity of the data.

Method used

Adopting a hierarchical distributed storage architecture, combining wavelet transformation algorithm, software-defined network, multi-replica redundant storage and erasure coding technology, adjust the sampling frequency through ambient light and bird activities, build core, regional and edge storage nodes, realize data sharding and intelligent scheduling, establish a collaboration mechanism to ensure data reliability and real-timeness.

Benefits of technology

It realizes flexible expansion of the storage system, reduces network pressure, ensures timely transmission and storage of data, improves the real-time and integrity of data, enhances the security and query efficiency of data, and ensures the reliability and availability of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407859A_ABST
    Figure CN120407859A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage, and particularly discloses a real-time data distributed storage method based on a double-spectrum bird detection system, which comprises the following steps: S1, acquiring data by using the double-spectrum bird detection system, and adjusting the sampling frequency according to the ambient light intensity and the bird activity activeness; de-noising and calibrating the original spectral data by adopting a wavelet transform algorithm; s2, constructing a hierarchical architecture consisting of core, region and edge storage nodes; s3, calculating scores of the fragment data distributed to the nodes according to the feature fragment data, and distributing the data according to the scores; and S4, establishing a cooperation mechanism among the nodes, determining the number of copies by adopting a multi-copy redundancy storage strategy, and introducing an erasure code technology. The hierarchical distributed storage architecture enables the capacity of the storage system to be flexibly expanded by increasing storage nodes of different levels, can easily cope with the continuous increase of the data volume of the double-spectrum bird detection system, and overcomes the problem that the traditional centralized storage capacity is limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage, and in particular to a real-time data distributed storage method based on a dual-spectrum bird detection system. Background Art

[0002] The dual-spectral bird detection system combines visible and infrared light technologies to accurately detect bird activity in a variety of environments, providing critical data support for bird ecology research, nature reserve management, and airport bird repellent operations. As monitoring scale expands and accuracy requirements increase, the amount of data generated by the system is rapidly increasing. Traditional centralized storage methods have exposed numerous problems in handling this massive amount of real-time data.

[0003] The storage capacity of centralized storage servers is limited and cannot meet the demand for continuous data growth; the centralized transmission of large amounts of data can easily cause network congestion, resulting in data transmission delays or even loss, seriously affecting the real-time and integrity of the data; and centralized storage has the risk of single point failure. Once the server fails, the entire data storage system will be paralyzed, causing bird detection data to be at risk of loss or inaccessibility, greatly affecting the progress of related work. Summary of the Invention

[0004] The purpose of the present invention is to provide a distributed storage method for real-time data based on a dual-spectrum bird detection system to solve the problems of insufficient storage capacity, network congestion, single point failure and low data management efficiency in traditional centralized storage when processing real-time data of a dual-spectrum bird detection system.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] A real-time data distributed storage method based on a dual-spectrum bird detection system includes the following steps:

[0007] S1: Data is collected using a dual-spectral bird detection system, and the sampling frequency is adjusted according to the ambient light intensity and bird activity level. The raw spectral data is denoised and calibrated using a wavelet transform algorithm.

[0008] S2: Builds a hierarchical architecture consisting of core, regional, and edge storage nodes. Nodes are connected via high-speed redundant networks and software-defined networking technology is used to allocate network bandwidth.

[0009] S3: Calculate the score of the shard data assigned to the node based on the feature shard data, and distribute the data based on the score; calculate the node load after the data is distributed;

[0010] S4: Establish a collaboration mechanism between nodes, adopt a multi-copy redundant storage strategy, determine the number of copies, and introduce erasure coding technology.

[0011] As a further solution of the present invention: S1 specifically includes:

[0012] Obtain the environmental light intensity L and the bird activity activity A;

[0013] Adjust the sampling frequency through the formula, and the formula is as follows:

[0014] F = αL + βA + γ, where α, β, and γ are weight coefficients;

[0015] Obtain the original spectral data S raw , and calibrate it through the formula, and the formula is as follows:

[0016] S calibrated = kS raw + b, where S calibrated is the calibrated data, and k and b are calibration coefficients.

[0017] As a further solution of the present invention: S2 specifically includes:

[0018] Calculate the network bandwidth B allocated to the i-th node i :

[0019]

[0020] Among them, B total represents the total bandwidth, D i is the data transmission requirement of the i-th node, and n is the total number of nodes.

[0021] As a further solution of the present invention: S3 specifically includes:

[0022] Obtain the feature shard data, and the feature shard data includes time series, bird species classification, and monitoring area;

[0023] Calculate the score Score of the feature data shard d allocated to the node i di :

[0024]

[0025] Among them, P d represents the priority of the feature data shard d, S i represents the storage performance of the node i, L i represents the network latency of the node i, and Loard i represents the load of the node i;

[0026] Allocate the feature data shard d to the node with the highest score.

[0027] As a further solution of the present invention: The calculation formula of the load of the node is as follows:

[0028] Load = w1O capacity + w2O cpu + w3O bandwidth ;

[0029] Wherein, O capacity represents the storage capacity occupancy rate, O cpu represents the CPU usage rate, O bandwidth represents the network bandwidth utilization rate, w1, w2, and w3 are the weight coefficients of each index respectively, and w1 + w2 + w3 = 1.

[0030] As a further solution of the present invention: The formula for determining the number of replicas is as follows:

[0031]

[0032] Wherein, I d is the data importance, F d is the usage frequency, and θ1, θ2, and θ3 are the weight coefficients.

[0033] As a further solution of the present invention: In data preprocessing, the parameters of the wavelet transform algorithm are optimally selected according to the frequency characteristics and noise level of the data.

[0034] As a further solution of the present invention: When the scores of multiple nodes are the same, the node with the lowest network latency is preferentially selected to store the data shard.

[0035] Advantages of the present invention: The hierarchical distributed storage architecture enables the capacity of the storage system to be flexibly expanded by adding storage nodes at different levels, easily coping with the continuous growth of the data volume of the dual-spectrum bird detection system and overcoming the problem of limited capacity of traditional centralized storage. The intelligent scheduling of data sharding storage and network traffic effectively reduces the network pressure caused by centralized data transmission, avoids network congestion, ensures the timely and accurate transmission and storage of data, and improves the real-time performance and integrity of data. Through the network bandwidth allocation formula, the bandwidth can be reasonably allocated according to the real-time needs of nodes to ensure smooth data transmission. The combination of replica redundant storage and erasure code technology, as well as the cooperation mechanism between nodes, greatly improves the data security and fault tolerance. Even if multiple nodes fail, data loss can be guaranteed, ensuring the reliability and availability of bird detection data. Through the dynamic adjustment of the replica number formula, replicas can be reasonably allocated according to the importance and usage frequency of data to improve data security. The multi-layer index structure and intelligent query algorithm significantly improve the data query efficiency and can quickly respond to various query requests. At the same time, a reasonable data update and deletion management strategy ensures data consistency and efficient utilization of storage resources, improving the overall level of data management. The query hit rate formula can be used to evaluate query performance, and the storage capacity can be calculated to reasonably manage storage resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The present invention will be further described below with reference to the accompanying drawings.

[0037] Figure 1 It is a schematic structural diagram of a real-time data distributed storage method based on a dual-spectrum bird detection system according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] Please refer to Figure 1 As shown, the present invention is a real-time data distributed storage method based on a dual-spectrum bird detection system, including the following steps:

[0040] The dual-spectrum bird detection system continuously collects visible light and infrared light data at a high frequency. During the data collection process, an adaptive sampling technology is adopted to intelligently adjust the sampling frequency according to the changes in environmental light and the dynamic situation of bird activities.

[0041] Let the environmental light intensity be L, the activity level of birds be A, and the sampling frequency f be determined by the following formula:

[0042] F = αL + βA + γ;

[0043] Where α, β, and γ are pre-set weight coefficients that can be adjusted according to the actual situation.

[0044] For the collected raw data, first perform noise filtering. Use the wavelet transform algorithm to remove the noise signal in the data. This algorithm can effectively separate the noise components in the signal and retain the useful features of the data. Then, calibrate the data through the spectral calibration model. This model is trained based on known standard spectral data and can accurately correct the spectral deviation of the data to improve the accuracy of the data. Let the original spectral data be S raw , and the calibrated spectral data be S calibrated , and the calibration formula is:

[0045] S calibrated = kS raw + b;

[0046] Where k and b are calibration coefficients obtained through training with standard spectral data.

[0047] Construct a hierarchical distributed storage architecture, which consists of a core storage node, regional storage nodes, and edge storage nodes. The core storage node has powerful computing and storage capabilities, is responsible for storing key data and index information, and manages and schedules the entire storage system. The regional storage nodes are distributed in different geographical regions and are used to store bird detection data within their respective regions to reduce the load on the core storage node. The edge storage nodes are close to the data collection end and are responsible for temporarily storing the collected data, acting as a data cache and preprocessing function.

[0048] Each storage node is connected through a high-speed and redundant network to ensure the stability and efficiency of data transmission. Use software-defined network (SDN) technology to achieve intelligent scheduling of network traffic, dynamically allocate network bandwidth according to data transmission requirements, and avoid network congestion. Let the total network bandwidth be B total , the real-time data transmission requirement of the i-th node be D i , and the bandwidth B i allocated to this node is:

[0049]

[0050] Where n is the total number of storage nodes.

[0051] Slice the preprocessed data according to multi-dimensional features such as the time series of the data, bird species classification, and monitoring area. Design a data allocation algorithm based on priority and load balancing. For data with high real-time requirements, such as abnormal behavior data of birds, assign a higher priority and preferentially allocate it to nodes with good storage performance and low network latency; for regular data, allocate it according to the real-time load conditions of each node.

[0052] Let the priority of data slice d be P d , the storage performance of node i be S i , the network latency be L i , the load be Load i , then the score Score for data slice d to be allocated to node i d,n is:

[0053]

[0054] The system allocates the data slice to the node with the highest score. Monitor the load conditions of each storage node in real time, including indicators such as the occupancy rate O capacity of the storage capacity, the CPU usage rate O CPU , and the network bandwidth utilization rate O bandwidth . The formula for calculating the comprehensive load Load of the node is:

[0055] Load = w1O capacity + w2O cpu + w3O bandwidth ;

[0056] where w1, w2, and w3 are the weight coefficients of each indicator, and w1 + w2 + w3 = 1.

[0057] Establish an efficient cooperation mechanism between storage nodes to ensure data consistency and reliability through a distributed consensus algorithm. The regional storage nodes regularly synchronize data to the core storage nodes, and the edge storage nodes upload the data to the regional storage nodes in a timely manner after the data processing is completed.

[0058] Adopt a multi-copy redundant storage strategy. For important data, store copies on multiple nodes in different geographical locations. Dynamically adjust the number of copies N d according to the importance I d of the data and the usage frequency F copies :

[0059]

[0060] where θ1, θ2, and θ3 are weight coefficients, Indicates rounding up. At the same time, erasure coding technology is introduced to disperse the encoded data and store it on multiple nodes. Even if some nodes fail, the data can be restored through the erasure coding algorithm, further improving the fault tolerance of the data.

[0061] Construct a multi-layer index structure for the data stored in each node. The first layer is the global index, stored in the core storage node, which contains the key meta-information of the data and the storage location pointer; the second layer is the regional index, stored in the regional storage node, which performs a more detailed index on the data within the region; the third layer is the local index, stored in the edge storage node, used to quickly locate the locally stored data. Design an intelligent query algorithm to preferentially search in the local index or regional index according to the type and conditions of the query request. If no matching data is found, search in the global index. Adopt caching technology, set up data buffer areas in the core storage node and regional storage node, store the frequently queried data in the cache, and reduce the data reading time. Let the hit rate of the query request in the local index be H local , the hit rate in the regional index is H regional , and the hit rate in the global index is H global , then the total query hit rate H total is:

[0062] H total = H local +(1 - H local )H regional +(1 - H local )(1 - H regional )H global ;

[0063] At the same time, use distributed query optimization technology to decompose and distribute the query tasks to multiple nodes for parallel processing, improving the query efficiency.

[0064] When new data is collected and stored, the system automatically updates the corresponding index information. Adopt a combination of incremental update and batch update. For the update of a small amount of data, use the incremental update method to directly modify the data on the storage node and synchronously update the index; for the update of a large amount of data, first perform batch processing, and then update the storage node and index at one time.

[0065] For expired or no-longer-needed data, formulate a strict data deletion policy. Before deleting the data, first check the reference situation of the data to ensure that no other module is using the data. After deleting the data, update the index information in time to release the storage resources. Let the total storage capacity of the storage system be C total , the used storage capacity be C used , the capacity of the deleted data be C delete , then the remaining storage capacity C remainingNamely:

[0066] C remaining = C total - C used + C delete ;

[0067] Meanwhile, in order to ensure the traceability of data, log records are made for the deleted data, recording information such as the deletion time and reason of the data.

[0068] The above has described an embodiment of the present invention in detail, but the described content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A real-time data distributed storage method for a dual-spectrum bird detection system, characterized in that It includes the following steps: S1: Collect data using a dual-spectrum bird detection system, and adjust the sampling frequency according to the environmental light intensity and bird activity level; Denoise and calibrate the original spectral data using the wavelet transform algorithm; S2: Construct a hierarchical architecture consisting of core, regional, and edge storage nodes, connect the nodes through a high-speed redundant network, and use software-defined network technology to allocate network bandwidth; S3: Calculate the scores for allocating sharded data to nodes based on the feature sharded data, and allocate data according to the scores; Calculate the node load after data allocation; S4: Establish a cooperation mechanism among nodes, adopt a multi-copy redundant storage strategy, determine the number of copies, and introduce erasure code technology.

2. The real-time data distributed storage method of a dual-spectrum bird detection system according to claim 1, wherein, The specific content of S1 includes: Obtain the environmental light intensity L and bird activity level A; Adjust the sampling frequency through the formula, and the formula is as follows: F = αL + βA + γ, where α, β, and γ are weight coefficients; Obtain the original spectral data S raw , and calibrate it through a formula, the formula is as follows: S calibrated = kS raw + b, where S calibrated is the calibrated spectral data, and k and b are calibration coefficients.

3. A real-time data distributed storage method for a dual-spectrum bird detection system according to claim 1, characterized in that, The specific content of S2 includes: Calculate the network bandwidth B allocated to the i-th node i : Among them, B total represents the total bandwidth, D i is the data transmission requirement of the i-th node, and n is the total number of nodes.

4. A real-time data distributed storage method for a dual-spectrum bird detection system according to claim 1, characterized in that The specific content of S3 includes: Obtain the feature sharded data, and the feature sharded data includes time series, bird species classification, and monitoring area; Calculate the score Score of the shard d of the feature data assigned to node i d,i : Among them, P d represents the priority of the feature data shard d, S i represents the storage performance of node i, L i represents the network latency of node i, Loard i represents the load of node i; Allocate the feature data shard d to the node with the highest score.

5. A real-time data distributed storage method for a dual-spectrum bird detection system according to claim 4, characterized in that, The calculation formula for the load of the node is as follows: Load = w1O capacity + w2O cpu + w3O bandwidth ; Among them, O capacity represents the storage capacity occupancy rate, O cpu represents the CPU usage rate, O bandwidth represents the network bandwidth utilization rate. w1, w2, and w3 are the weight coefficients of each index respectively, and w1 + w2 + w3 = 1.

6. A real-time data distributed storage method for a dual-spectrum bird detection system according to claim 1, characterized in that, The formula for determining the number of copies is as follows: Among them, I d is the data importance, F d is the usage frequency, and θ1, θ2, and θ3 are weight coefficients.

7. A real-time data distributed storage method for a dual-spectrum bird detection system according to claim 1, characterized in that, In data preprocessing, the parameters of the wavelet transform algorithm are optimally selected according to the frequency characteristics and noise level of the data.

8. A real-time data distributed storage method for a dual-spectrum bird detection system according to claim 1, characterized in that, When the scores of multiple nodes are the same, preferentially select the node with the lowest network latency to store the data shard.