Artificial Intelligence-Based Information Management Method, System, and Device for Data Recorders

An AI-driven data management system for data recorders addresses inefficiencies by enabling real-time processing, adaptive compression, and dynamic storage optimization, improving data accuracy and system efficiency for complex environmental data handling.

CN119938673BActive Publication Date: 2025-07-15徐州意泰林电子科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510445732.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-15
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The traditional data recorder information management method lacks real-time data preprocessing and automated analysis capabilities, resulting in high labor costs, high operational difficulties, and insufficient storage management methods, resulting in waste of resources and low access efficiency.

Method used

Using an artificial intelligence-based data recorder information management method, real-time data is obtained through multi-source sensors for deep semantic analysis and filtering and noise reduction, combined with distributed storage architecture and adaptive lossless compression, the index structure is dynamically adjusted to optimize storage and retrieval.

Benefits of technology

It improves the real-time and accuracy of data collection, reduces storage space usage, enhances the stability and efficiency of the system, optimizes data access and processing speed, and reduces manual intervention and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938673B_ABST
    Figure CN119938673B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data management, and particularly to a data recorder information management method, system and device based on artificial intelligence. The method includes the following steps: obtaining real-time environmental monitoring data based on multi-source sensors of a data recorder; performing multi-period data deep semantic analysis on the real-time environmental monitoring data to obtain data indexes for each time period; predicting the evolution of environmental parameters for the filtered and noise-reduced monitoring data, and performing parameter visualization rendering to construct a visual view of the environmental prediction trend; performing adaptive lossless compression on the historical stored data, and constructing a distributed storage architecture to obtain a distributed data node storage framework; embedding the visual view of the environmental prediction trend into the distributed data node storage framework based on the data indexes for each time period to generate a real-time data embedded storage framework. The present invention realizes efficient and highly utilized information management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and in particular, to an information management method, system and device for data recorders based on artificial intelligence. Background Art

[0002] With the continuous development and application of intelligent technologies, data recorders are playing an increasingly important role in various fields. Data recorders are widely used in multiple fields such as industry, environmental monitoring, smart home, and energy management, and are responsible for the tasks of real-time collection, storage, and analysis of various key environmental and device data. Its main function is to continuously monitor and record the data of the working environment or the operating state of the device, ensure that the system can reflect various key parameters in real time, and thus provide accurate basis for subsequent analysis, decision-making, and optimization.

[0003] However, in traditional data recorder information management methods, the collection, storage, and analysis of data usually rely on static preset rules and manual operations. Although this method meets the basic requirements in some simple application scenarios, with the sharp increase in data volume and the increasing complexity of application scenarios, traditional methods have gradually exposed a series of problems. First of all, traditional data recorders often do not have the ability of real-time data preprocessing and automated analysis, resulting in the need for manual intervention and manual analysis after the data is recorded, thus increasing the labor cost and operation difficulty. Secondly, due to the complexity of environmental conditions, traditional methods have limited capabilities in data redundancy processing, fault detection, and real-time adjustment, and are unable to effectively cope with the increasing amount of monitoring data and diverse data management requirements. Traditional data storage management methods usually lack intelligent data compression, storage optimization, and dynamic adjustment mechanisms, which are prone to waste of data storage resources and low access efficiency. Therefore, a more intelligent data recorder information management method is needed. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes an information management method, system and device for data recorders based on artificial intelligence to solve at least one of the above technical problems.

[0005] To achieve the above object, the present invention provides an information management method for data recorders based on artificial intelligence, including the following steps:

[0006] Step S1: Obtain real-time environmental monitoring data based on multi-source sensors of the data recorder; perform in-depth semantic parsing of the real-time environmental monitoring data for multiple time periods, and perform intelligent index definition to obtain data indexes for each time period;

[0007] Step S2: Predict the evolution of environmental parameters for the filtered and noise-reduced monitoring data, and perform parameter visualization rendering to construct a visual view of the environmental prediction trend;

[0008] Step S3: Obtain the historical stored data of the data recorder; perform adaptive lossless compression on the historical stored data, and construct a distributed storage architecture, so as to obtain a distributed data node storage framework;

[0009] Step S4: Based on the data index of each time period, embed the environmental prediction trend visualization view into the distributed data node storage framework to generate a real-time data embedded storage framework;

[0010] Step S5: Perform data redundancy analysis and redundant node deletion processing based on the real-time data embedded storage framework to obtain a redundant node optimization framework;

[0011] Step S6: Calculate the storage load of each node in the redundant node optimization framework and perform dynamic index structure adjustment, so as to construct a dynamic index adjustment storage framework to execute the data recorder information management task.

[0012] Through the multi-source sensors equipped in the data recorder, the environmental parameters are monitored in real time, which enables the system to respond to environmental changes in a timely manner, improving the real-time and comprehensiveness of data collection. Through the in-depth semantic analysis of multi-period data, the deep information in the environmental data can be better extracted. Not only the original data is obtained, but also the spatio-temporal correlation of the data can be understood. For future analysis and decision-making, through filtering and noise reduction, the system can remove errors and interferences in the environmental monitoring data, ensuring the accuracy and reliability of the data, which makes subsequent analysis and prediction more accurate. Based on historical data and real-time data, through the evolution prediction model, the trend of environmental parameters is predicted, potential environmental changes are identified in advance, and early warning and prediction decisions are helped to be made. The adaptive lossless compression technology can reduce the occupancy of storage space and improve the storage efficiency without losing any data quality. The design of the distributed storage architecture ensures the high availability and fault tolerance of the data. Whether it is data storage, reading or backup, the stability and efficiency of the system can be guaranteed, and data loss or system downtime caused by single-point failures can be avoided. By embedding the visual view of the predicted trend into the distributed storage framework, it can ensure that the data is not only optimized in storage, but also the rapid access and processing of real-time data can be guaranteed. This embedding method improves the real-time and dynamic response capabilities of the system. The data embedding storage framework optimizes the storage method of the data, ensuring that environmental data and prediction data can be efficiently fused and stored in the distributed storage, further improving the overall performance of the system. Through redundancy analysis, redundant data nodes in the storage system can be identified and the deletion process of redundant nodes can be carried out. This not only reduces unnecessary storage consumption, but also optimizes the storage architecture, improving the efficiency of data storage and the utilization rate of system resources. By removing redundant nodes, the storage architecture becomes more streamlined, reducing the burden on the system, increasing the speed of data reading and writing, and improving the overall performance of the system. By calculating the storage load of each node, the system can make dynamic adjustments according to the load conditions of the nodes, ensuring that the storage and computing pressures of each node are within a reasonable range. This load balancing mechanism can effectively improve the stability and processing efficiency of the system. According to the load changes and data access requirements, dynamically adjusting the index structure helps to improve the efficiency of data retrieval and the adaptability of the system. Especially in the case of a large amount of data and frequent changes, dynamic adjustment can enable the system to better respond to changes. By dynamically adjusting the storage architecture and index structure, the information management of the data recorder can be better carried out, ensuring the efficient access and processing of data, and ultimately supporting the precise management and analysis of environmental monitoring data by the system.

[0013] Preferably, step S1 includes the following steps:

[0014] Step S11: Obtain real-time environmental monitoring data based on the multi-source sensors of the data recorder;

[0015] Step S12: Analyze the parameters of the real-time environmental monitoring data for normal distribution to generate a normal distribution range;

[0016] Step S13: Filter out abnormal outlier parameters based on the normal distribution range to obtain outlier-filtered monitoring data;

[0017] Step S14: Perform high-frequency filtering and noise reduction on the outlier-filtered monitoring data to obtain filtered and noise-reduced monitoring data;

[0018] Step S15: Perform in-depth semantic analysis on the filtered and noise-reduced monitoring data for multiple time periods to generate in-depth semantic features for each time period;

[0019] Step S16: Define intelligent indexes based on the in-depth semantic features for each time period to obtain data indexes for each time period.

[0020] The present invention collects environmental data in real time through multi-source sensors (such as temperature and humidity sensors, barometric pressure sensors, light sensors, etc.), which can comprehensively and meticulously reflect the environmental state, meet the monitoring requirements of multiple dimensions. The use of a data recorder ensures the acquisition of real-time data, can promptly reflect environmental changes, supports rapid response and timely decision-making. Through normalized distribution analysis, the data is converted into a standardized format, enabling unified processing of data collected by different sensors. For subsequent data analysis and comparison, normalized distribution analysis can clarify the normal range of various environmental parameters, identify the expected distribution and fluctuation range of the data, laying a foundation for anomaly detection. By comparing with the normalized distribution range, abnormal data (such as sensor failures, environmental mutations, etc.) can be effectively detected. This step can effectively remove outlier data deviating from the normal range, ensuring the accuracy of the data. After filtering out the outlier data, the remaining data is more representative and reliable, avoiding misanalysis or misjudgment caused by abnormal data. Through high-frequency filtering and noise reduction, the noise components (such as equipment noise, electromagnetic interference, etc.) in the real-time monitoring data can be removed, obtaining smoother and more accurate data. The denoised data is closer to the actual environmental state, which helps to conduct more accurate trend analysis and prediction. The filtered data provides high-quality input for further analysis and modeling, improving the sensitivity and response ability of the system to environmental changes. Through deep semantic parsing, not only can the data characteristics of each time period be obtained, but also the meaning and patterns behind these data can be deeply understood. This deep analysis can reveal the internal laws of environmental parameter changes over time. The generated data characteristics of time periods help the system identify the change trends of environmental parameters. Especially in data with a long time span, valuable information such as seasonal changes and periodic fluctuations can be effectively extracted. Through intelligent indexing based on deep semantic features, an index can be automatically generated for the data of each time period, making the retrieval and management of data more efficient. The data index enables the system to quickly locate and retrieve the required historical data or real-time data. Especially in an environment with a large amount of data, it greatly improves the efficiency of data processing and query. The intelligent index structure can automatically optimize the storage strategy according to the characteristics of the data, thereby improving the data access speed and storage space utilization rate.

[0021] Preferably, step S2 includes the following steps:

[0022] Step S21: Mine the change characteristics of environmental parameters from the filtered and noise-reduced monitoring data to generate the change trend characteristics of environmental parameters;

[0023] Step S22: Calculate the change frequency of the change trend characteristics of environmental parameters to extract the change frequency of environmental parameters;

[0024] Step S23: Identify the periodic evolution law based on the change frequency of environmental parameters to generate the periodic law of environmental change trends;

[0025] Step S24: Based on the periodic law of environmental change trends, perform environmental parameter evolution prediction on the filtered and noise-reduced monitoring data to obtain future environmental trend prediction features;

[0026] Step S25: Perform parameter visualization rendering on the future environmental trend prediction features to construct a visualized view of the environmental prediction trend.

[0027] The present invention reveals the potential laws of environmental parameter changes by mining the change features of the filtered and noise-reduced data. For example, parameters such as temperature, humidity, and air quality change with seasons, time, or other factors. Through feature mining, the system can identify these change patterns. The mined change features provide a deeper understanding of data analysis, helping to reveal the environmental change trends hidden behind the data and providing support for subsequent analysis and decision-making. Calculating the frequency of environmental parameter changes can help the system identify the periodic features of environmental changes. For example, certain parameters exhibit specific fluctuation patterns daily, weekly, or monthly. By extracting the frequency, the system can accurately describe the dynamic features of environmental changes. The extracted change frequency information can help identify the periodic laws of environmental changes, providing a basis for predicting periodic changes, which is of great significance for predicting weather changes, seasonal climate changes, etc. By identifying the periodic evolution law based on frequency, the system can identify the periodic laws of environmental parameter changes, such as the seasonal fluctuations of temperature and humidity, or the diurnal changes of air quality. Accurately grasp the long-term trend of environmental changes. The identification of periodic laws provides a scientific basis for long-term trend prediction. The system can predict the environmental change trend within a certain future time period, helping users anticipate future changes and make corresponding management and adjustments. Through the periodic law and the evolution prediction model of environmental parameters, the system can predict the future environmental change trend, which is not just a simple extension of current data but an intelligent prediction based on environmental change laws. Accurate environmental trend prediction can provide early warnings for decision-makers, making environmental management more forward-looking. For example, factors such as meteorological changes and climate fluctuations can be warned in advance to help relevant departments take preventive measures and reduce the potential impacts brought by environmental changes. Visualizing the future environmental trend prediction features can help users understand the future change trends of environmental parameters in an intuitive way. Through visualization means such as charts, curves, and heat maps, users can quickly understand the data and prediction results. The visualized view provides clear basis for decision-makers, facilitating them to make decisions according to the prediction results. For example, based on the visualized view of temperature and humidity change trends, users can easily judge the future environmental conditions and take corresponding measures.

[0028] Preferably, the specific steps of step S21 are:

[0029] The filtered and noise-reduced monitoring data includes environmental temperature parameters, environmental humidity parameters, light intensity, air pollutant monitoring data, and wind monitoring data;

[0030] Fit the temporal variation characteristics of the environmental temperature parameters and environmental humidity parameters to construct an environmental temperature change curve and an environmental humidity change curve;

[0031] Conduct an analysis of the spatial distribution of light intensity for the light intensity to obtain the spatial distribution characteristics of light intensity;

[0032] Calculate the pollutant peaks for the air pollutant monitoring data and extract the pollutant peak points;

[0033] Identify the wind direction and wind force intensity for the wind monitoring data to extract the wind force characteristics;

[0034] Mine the environmental change trends for the spatial distribution characteristics of light intensity, pollutant peak points, wind force characteristics, environmental temperature change curve, and environmental humidity change curve to obtain the environmental parameter change trend characteristics.

[0035] Through filtering and noise reduction processing of data, the present invention reduces equipment noise, interference signals and errors, ensuring the accuracy and reliability of data and laying a good foundation for subsequent analysis. By integrating multiple sensor data, an all-round and three-dimensional environmental data view is formed, providing more comprehensive and accurate information support for subsequent trend analysis, prediction, etc. Through fitting the temporal variation characteristics of temperature and humidity data, the variation curves of these two environmental parameters can be accurately depicted. The changes in temperature and humidity are key factors in environmental monitoring, and the fitted curves can visually show their trends over time. Environmental temperature and humidity usually have significant seasonal or periodic variations. Temporal fitting can help identify these patterns and provide valuable references for subsequent prediction and analysis. The spatial distribution analysis of light intensity helps to reveal the light distribution at different locations and times, which is crucial for analyzing the light conditions of the environment and evaluating the impact of light on parameters such as temperature and humidity. In fields such as agriculture, construction, and energy, the spatial distribution characteristics of light intensity are used to optimize resource allocation. For example, it helps the agricultural sector evaluate the light conditions for crop growth or helps building managers design energy-saving solutions. By calculating the peak values of pollutant data, extreme situations of pollutant concentration can be identified, such as peak pollutant values or high pollution periods, which are important indicators for judging the severity of environmental pollution. The extraction of peak points helps to promptly identify pollution events and supports the rapid response of the environmental monitoring system. For example, when the Air Quality Index (AQI) reaches a certain threshold, the warning system is triggered to remind relevant departments to take countermeasures. Wind direction and wind intensity are important factors in meteorological data. Through the analysis of wind monitoring data, key features such as the direction and intensity of the wind are extracted, comprehensively understanding the variation law and spatial distribution of wind power. The extraction of wind power characteristics has important applications in fields such as wind power generation and meteorological prediction. The changes in wind speed and wind direction directly affect energy production and the accuracy of weather forecasting. Therefore, accurately identifying wind power characteristics can provide strong support for relevant fields. By comprehensively analyzing multi-dimensional environmental parameters such as light, pollutant peaks, wind power, temperature, and humidity, the trends of environmental changes can be mined from multiple perspectives. This comprehensive analysis reveals the interaction and influence relationships between different environmental factors. This multi-factor trend mining helps to identify complex patterns of environmental changes. For example, the correlation between temperature changes and air pollution, and the interaction effect between wind power changes and humidity. Through these analyses, the system can more comprehensively understand the laws of environmental changes. By extracting the characteristics of environmental change trends, more accurate inputs are provided for the prediction model of environmental changes, helping to predict sudden changes or long-term trends in the environment in advance (such as climate change, air quality changes, etc.).

[0036] Preferably, the specific steps of step S3 are as follows:

[0037] Step S31: Obtain the historical storage data of the data recorder;

[0038] Step S32: Define a time span based on the historical storage data, and perform uniform segmentation processing according to the time span, so as to generate multiple historical storage data nodes;

[0039] Step S33: Perform adaptive lossless compression on multiple historical storage data nodes to obtain multiple lossless compression data nodes;

[0040] Step S34: Construct a distributed storage architecture for multiple lossless compression data nodes to obtain a distributed data node storage framework.

[0041] The present invention helps to split historical data into different time periods by defining a suitable time span, which is convenient for data analysis and storage management. The definition of this time span is adjusted according to actual needs, such as splitting by day, week, month or year, so as to make data management more flexible. Splitting the data evenly according to the time span can facilitate the independent processing and analysis of data in different time periods, which helps to avoid the processing difficulties caused by excessive data. Each data node is stored and accessed independently, improving the manageability of the data. Lossless compression can reduce the data volume without losing data, thus greatly saving storage space. In the case of a large amount of historical data, using an adaptive compression method can better reduce the storage cost. Different from lossy compression, lossless compression ensures the integrity and accuracy of the data. Even after compression, the data can still be completely restored, which is suitable for application scenarios that require high data precision and accuracy, such as environmental monitoring data. Through distributed storage, the data is dispersed and stored in multiple physical locations, which can improve the fault tolerance and reliability of the system. Even if a storage node fails, other nodes can still ensure the normal access to the data, avoiding the risk of single point of failure. The distributed storage architecture can make data storage more flexible and efficient. After the data is distributed on different nodes, parallel access can improve the data processing speed and access efficiency, which is suitable for the storage and fast access requirements of large-scale data. The distributed storage architecture has high scalability. According to the increase in storage requirements, new storage nodes can be added at any time to meet the increasing demand for data volume. As the data volume increases, the storage resources can be seamlessly expanded to avoid storage bottlenecks.

[0042] Preferably, the specific steps of Step S4 are as follows:

[0043] Step S41: Quantify the data sensitivity of the environmental prediction trend visualization view to obtain the sensitivity of the instant parameter view;

[0044] Step S42: Analyze the data access requirements of the sensitivity of the instant parameter view to generate the access requirement characteristics of real-time environmental monitoring data;

[0045] Step S43: Make a storage decision based on the characteristics of real-time environmental monitoring data access requirements to obtain the storage location of real-time data;

[0046] Step S44: Index and mark the environmental prediction trend visualization view according to the data index of each time period, and embed the distributed storage into the distributed data node storage framework based on the storage location of real-time data to generate a real-time data embedded storage framework.

[0047] By quantifying the data sensitivity of the visualization view of the environmental prediction trend, the present invention means that the system can evaluate the importance of each environmental monitoring parameter (such as temperature, humidity, air quality, etc.) to the overall prediction result. Identify which parameters have the greatest impact on the environmental trend in the prediction model and which parameters contribute less to the decision-making. By quantifying the data sensitivity, prioritize those environmental parameters that have a greater impact on the prediction result, which helps to determine the allocation and optimization of storage resources. For example, parameters with higher sensitivity are updated and accessed more frequently in the system, while parameters with lower sensitivity are updated periodically. By generating data access requirement characteristics, the system can dynamically adjust storage and processing strategies. For example, during periods or conditions of high demand, the system schedules more computing resources to process popular data to ensure the efficiency of real-time monitoring and warning. By deeply understanding the real-time data access requirements, the system optimizes the data retrieval and processing processes, reduces interference from irrelevant data, accelerates the response speed of key data, and ensures that the system can quickly respond to environmental changes. Frequently accessed data is stored in faster and more responsive storage locations (such as caches or SSDs), while less frequently accessed data is stored on more economical and larger-capacity storage media (such as HDDs), which reduces storage costs while ensuring performance. This storage decision based on real-time access requirements can flexibly adapt to different application scenarios. For example, in environmental prediction, data needs to be accessed or updated frequently during certain periods, while during other periods, the data access requirements will significantly decrease. The system automatically adjusts the storage strategy to optimize resource utilization. By creating indexes and marking for the data of each time period, the system can quickly locate the relevant data nodes, improving the speed and accuracy of data query. The marking of time indexes enables the data to be retrieved efficiently in chronological order, providing strong support for time series analysis. By combining the storage decision with the storage location of real-time data, the real-time data is embedded into the distributed storage framework, further enhancing the flexibility and scalability of the storage architecture. Distributed storage can ensure the high availability and fault tolerance of data and can scale horizontally to cope with future data volume growth.

[0048] Preferably, the specific steps of Step S5 are:

[0049] Step S51: Identify the user access behavior for each of the multiple lossless compression data nodes, and extract the access behavior data for each node;

[0050] Step S52: Calculate the data access frequency for the access behavior data of each node to obtain the data access frequency for each node;

[0051] Step S53: Analyze the access time range for the access behavior data of each node to generate the access time range feature for each node;

[0052] Step S54: Evaluate the historical data access requirements based on the data access frequency of each node and the access time range feature of each node, so as to obtain the access requirement evaluation value for each node;

[0053] Step S55: Compare the data redundancy of the access requirement evaluation value of each node based on a preset node access requirement threshold. When the access requirement evaluation value of the node is lower than the preset node access requirement threshold, perform redundant node deletion processing on the real-time data embedded storage framework to obtain a redundant node optimization framework.

[0054] The present invention collects the access behavior data of each node, and the system adopts different storage and processing strategies for different types of data nodes. For example, frequently accessed nodes are preferentially stored in high-efficiency storage media, while rarely accessed nodes are stored in low-cost storage media, thereby optimizing the use of storage resources. By analyzing the access behavior of each node, the system can identify redundant and inefficient data nodes, avoiding waste of resources, which lays a foundation for subsequent optimization of redundant data. Through the calculation of access frequency, the system sorts the data according to the actual access frequency. For example, frequently accessed nodes are stored on high-efficiency storage media to improve data reading speed; while for low-frequency accessed nodes, they are stored on low-cost storage media to achieve a balance between storage cost and performance. According to the access frequency, the system sets different access strategies and scheduling methods for different categories of data nodes, thereby improving the overall storage and access efficiency. By analyzing the access time range of each data node, the access pattern of each node in terms of time can be further understood. For example, some nodes are frequently accessed on weekdays but hardly accessed on weekends; some nodes are accessed continuously for 24 hours, while some nodes are only accessed during specific periods. Based on the characteristics of the access time range, the system can intelligently schedule the access and storage of data. For example, during certain high-access-frequency periods, the system preferentially provides data access rights for these nodes; while during low-access periods, energy-saving or degraded storage strategies are adopted to reduce unnecessary data access burdens. Through the evaluation of the access requirements of each node, the system can more accurately determine which data needs to be retained in high-efficiency storage media for a long time, which data needs to be transferred to low-cost storage media, and which data needs to be deleted or archived, thus reducing data redundancy and improving storage efficiency. Based on the evaluation of historical data access requirements, the system dynamically adjusts the storage strategy to ensure that the storage and access of data are more intelligent, which is particularly important for the management of large-scale environmental monitoring data, and can effectively reduce storage pressure and improve access efficiency. After deleting redundant nodes, the system can centrally store and process important data, reducing access to low-demand data, thereby improving data storage and access efficiency, which is crucial for an efficient storage architecture, especially in a big data environment, to avoid redundant data occupying precious storage resources. The deletion of redundant data not only improves storage efficiency but also significantly reduces storage costs. By cleaning up low-access-frequency data, the system can reduce its dependence on expensive storage resources and lower the overall storage cost. The process of redundant comparison based on the access requirement threshold makes data management more intelligent, and the system can automatically adjust the storage strategy according to the actual usage situation without manual intervention, which not only improves operation efficiency but also can adapt to different storage requirements, ensuring the flexibility and scalability of the system.

[0055] Preferably, the specific steps of step S6 are as follows:

[0056] Step S61: Calculate the storage load of each node in the redundant node optimization framework to obtain the storage load characteristics of each node;

[0057] Step S62: Predict the peak access time of each node based on the storage load characteristics of each node to obtain the predicted peak access time point of the node;

[0058] Step S63: Predict the future usage requirements of users based on the predicted peak access time point of the node to obtain the predicted usage requirements of each node;

[0059] Step S64: Dynamically adjust the index structure of the redundant node optimization framework according to the predicted usage requirements of each node, so as to construct a dynamically index-adjusted storage framework to execute the data recorder information management task.

[0060] The present invention calculates the storage load of each node in the redundant node optimization framework, so that the system can understand the storage pressure and load of each node in real time. The load calculation helps the system identify which nodes have a heavy storage burden and which nodes have idle storage, thereby optimizing the allocation of storage resources. By calculating the storage load characteristics, the system dynamically adjusts the storage resources according to the load conditions. For example, for nodes with higher loads, consider allocating more storage space, or adjusting the storage strategy of nodes with lower access frequency to reduce system bottlenecks. By analyzing the storage load characteristics, the system can predict the access peak time point of each node, which means that the system can predict large-scale data access events that will occur in the future (such as peak access, emergency data reading, etc.) and prepare in advance. The node access peak prediction helps to formulate resource scheduling strategies in advance. For example, when the system predicts that the access peak is about to arrive, it increases the storage and computing resources of the node in advance to ensure stability and rapid response during the peak period. By combining the node access peak prediction, The system can accurately predict future user usage needs. For example, at specific time points or conditions, some nodes will be accessed in large quantities (such as holidays or large-scale environmental changes), while some nodes will maintain a lower access frequency. The system will intelligently adjust storage resources based on future usage demand predictions. For upcoming high access demands, the system will allocate more storage and computing resources to related nodes; for nodes with lower demands, resource allocation will be reduced to maximize resource utilization. Based on the node usage demand predictions, the system can adjust the index structure in real time to adapt to different data access needs. For example, for frequently accessed nodes, a more efficient index structure, such as time series index or partition index, is used; for nodes with low access frequency, a simpler index structure is used to reduce storage overhead. By dynamically adjusting the index structure, the system reduces the search time during data access and improves the overall data processing speed. For example, frequently accessed nodes will be given a more efficient retrieval path, thereby improving the response speed and real-time performance of data processing.

[0061] In this specification, a data recorder information management system based on artificial intelligence is provided, which is used to execute the data recorder information management method based on artificial intelligence as described above, including:

[0062] A deep semantic analysis module is used to obtain real-time environmental monitoring data based on multi-source sensors of the data recorder; perform multi-period data deep semantic analysis on the real-time environmental monitoring data, and perform intelligent index definition to obtain the data index for each time period;

[0063] The data view module is used to predict the evolution of environmental parameters of the filtered noise reduction monitoring data, perform parameter visualization, and build a visualization view of the environmental prediction trend;

[0064] A node storage framework module for obtaining the historical storage data of a data recorder; performing adaptive lossless compression on the historical storage data and constructing a distributed storage architecture to obtain a distributed data node storage framework;

[0065] A storage embedding module for performing distributed storage embedding on the distributed data node storage framework based on the environmental prediction trend visualization view for each time period's data index to generate a real-time data embedded storage framework;

[0066] A redundancy optimization module for performing data redundancy analysis and redundant node deletion processing based on the real-time data embedded storage framework to obtain a redundant node optimization framework;

[0067] A dynamic framework adjustment module for calculating the storage load of each node on the redundant node optimization framework and performing dynamic index structure adjustment to construct a dynamic index adjustment storage framework for executing the data recorder information management operation.

[0068] Through the deep semantic analysis module, the system can conduct a more in-depth analysis of sensor data. Instead of merely staying at the data acquisition level, it can identify the potential relationships and patterns among the data, extract more information. Through the intelligent index definition, the system can generate efficient indexes for the data in each time period, making future data retrieval and query more efficient. Users can quickly locate the required data through time or other parameters, improving the query efficiency. Through the visualization rendering of the data, users can view the evolution trends of different environmental parameters (such as temperature, humidity, pollutants, etc.) in real time. Such graphical displays help users more intuitively understand and analyze the data, supporting faster decision-making. This module not only shows the current environmental conditions but also can predict future environmental changes through trends, helping users identify potential change trends in advance and facilitating the early response to environmental risks. Through adaptive lossless compression, the system can effectively reduce the storage occupancy of the data while ensuring the integrity of the data. This compression method does not lose data quality and is suitable for the storage of large-scale environmental monitoring data. Through the construction of a distributed storage architecture, the system can flexibly expand the storage capacity to adapt to the growing data storage requirements. In addition, the distributed architecture can improve the data access efficiency and reduce the risk of single-point failures. By combining the environmental prediction trend view with the data index and storage framework, the system can more efficiently manage the storage and retrieval of real-time data. The tight embedding of real-time data and historical data can reduce the latency during data query and improve the overall system response speed. This module provides a flexible framework for the storage of real-time data. According to the different access frequencies and storage requirements of the data, it can intelligently select the appropriate storage location to further improve the storage efficiency. Through the analysis and optimization of redundant nodes, the system can reduce unnecessary redundant data storage, effectively reducing the storage cost and resource waste. Redundancy optimization not only alleviates the burden on the storage system but also improves the utilization rate of the storage space, enabling more important data to be preferentially stored and accessed. By calculating the storage load of the nodes, the system dynamically adjusts the allocation of storage resources according to the actual needs. For nodes with higher loads, more resources are allocated to them, while for nodes with lower loads, appropriate resource release or degraded storage is performed. The dynamic index structure adjustment ensures that the data storage system can still operate efficiently under changing requirements. By providing a more efficient indexing method for high-demand nodes, the system can improve the query and data access speed.

[0069] The present invention also provides an information management device for a data recorder based on artificial intelligence, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the information management method for a data recorder based on artificial intelligence described in any one of the above are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1Schematic diagram of the step process of a data recorder information management method based on artificial intelligence according to the present invention;

[0071] Figure 2 Schematic diagram of the detailed implementation steps of step S1;

[0072] Figure 3 Schematic diagram of the detailed implementation steps of step S2;

[0073] Figure 4 Schematic diagram of the detailed implementation steps of step S3. Specific implementation manners

[0074] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0075] The embodiments of the present application provide a data recorder information management method, system and device based on artificial intelligence. The execution subjects of the data recorder information management method, system and device based on artificial intelligence include but are not limited to: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. that can be regarded as general computing nodes of the present application. The data processing platform includes but is not limited to at least one of: audio and image management systems, information management systems, and cloud data management systems.

[0076] Please refer to Figures 1 to 4 , the present invention provides a data recorder information management method based on artificial intelligence. The data recorder information management method based on artificial intelligence includes the following steps:

[0077] Step S1: Obtain real-time environmental monitoring data based on multi-source sensors of the data recorder; perform in-depth semantic analysis of the real-time environmental monitoring data in multiple time periods, and perform intelligent index definition to obtain data indexes for each time period;

[0078] Step S2: Predict the evolution of environmental parameters for the filtered and noise-reduced monitoring data, and perform parameter visualization rendering to construct a visual view of the environmental prediction trend;

[0079] Step S3: Obtain the historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture to obtain a distributed data node storage framework;

[0080] Step S4: Embed the visual view of the environmental prediction trend into the distributed data node storage framework based on the data index for each time period to generate a real-time data embedded storage framework;

[0081] Step S5: Perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain an optimized redundant node framework;

[0082] Step S6: Calculate the storage load of each node in the optimized redundant node framework and perform dynamic index structure adjustment, thereby constructing a dynamically index-adjusted storage framework to execute the data recorder information management task.

[0083] The present invention uses multi-source sensors equipped in the data recorder to monitor environmental parameters in real time, enabling the system to respond promptly to environmental changes, enhancing the real-time and comprehensiveness of data collection. Through in-depth semantic analysis of multi-period data, it can better extract the deep information in environmental data, not only obtaining the original data but also understanding the spatio-temporal correlation of the data. For future analysis and decision-making, through filtering and noise reduction, the system can remove errors and interference in environmental monitoring data, ensuring the accuracy and reliability of the data, which makes subsequent analysis and prediction more precise. Based on historical data and real-time data, through an evolution prediction model, trend prediction of environmental parameters is carried out to identify potential environmental changes in advance, helping to make early warning and prediction decisions. The adaptive lossless compression technology can reduce the storage space occupancy and improve the storage efficiency without losing any data quality. The design of the distributed storage architecture ensures the high availability and fault tolerance of the data. Whether it is data storage, reading, or backup, it can ensure the stability and efficiency of the system, avoiding data loss or system downtime caused by a single point of failure. By embedding the visual view of the predicted trend into the distributed storage framework, it can ensure that the data is not only optimized in storage but also enables fast access and processing of real-time data. This embedding method enhances the real-time and dynamic response capabilities of the system. The data embedding storage framework optimizes the storage method of the data, ensuring the efficient fusion and storage of environmental data and prediction data in the distributed storage, further enhancing the overall performance of the system. Through redundancy analysis, redundant data nodes in the storage system can be identified and the redundant nodes can be deleted. This not only reduces unnecessary storage consumption but also optimizes the storage architecture, improving the efficiency of data storage and the utilization rate of system resources. By removing redundant nodes, the storage architecture becomes more streamlined, reducing the burden on the system, increasing the speed of data reading and writing, and improving the overall performance of the system. By calculating the storage load of each node, the system can dynamically adjust according to the load situation of the node to ensure that the storage and computing pressure of each node is within a reasonable range. This load balancing mechanism can effectively enhance the stability and processing efficiency of the system. According to the load changes and data access requirements, dynamically adjusting the index structure helps to improve the efficiency of data retrieval and the adaptability of the system. Especially in the case of a large amount of data and diverse changes, dynamic adjustment enables the system to better cope with the changes. Through the dynamic adjustment of the storage architecture and the index structure, it is possible to better manage the information of the data recorder, ensuring the efficient access and processing of data, and ultimately supporting the precise management and analysis of environmental monitoring data by the system.

[0084] In the embodiment of the present invention, refer to Figure 1 , which is a schematic diagram of the step flow of a method for managing data recorder information based on artificial intelligence according to the present invention. In this example, the steps of the method for managing data recorder information based on artificial intelligence include:

[0085] Step S1: Obtain real-time environmental monitoring data based on the multi-source sensors of the data recorder; perform in-depth semantic analysis on the real-time environmental monitoring data for multiple time periods, and conduct intelligent index definition to obtain the data index for each time period.

[0086] In this embodiment, determine the environmental parameters to be monitored, such as temperature, humidity, air quality, light intensity, and sound level, etc. Select high-precision sensors to ensure the accuracy and reliability of the data. Use multiple sensors (such as DHT22, MQ135, photoresistors, etc.) for data collection, and ensure the compatibility between sensors for subsequent data integration. Configure the data recorder to connect all sensors to ensure that it can stably collect and store data. Use a microcontroller (such as Arduino or RaspberryPi) to control the sensors, write code to periodically read sensor data, set the data collection frequency, for example, record once per minute. Start the data recorder to begin obtaining real-time data from the sensors. The data should include the sensor ID, timestamp, and their corresponding environmental parameter values. Store the collected data in a local database (such as SQLite, MySQL) or a cloud database to ensure the accessibility and security of the data. Set up a regular backup mechanism to prevent data loss. Before performing in-depth semantic analysis, clean the collected data to remove missing values and outliers. Use the Pandas library in Python for data preprocessing. Perform in-depth analysis on the data for each time period, extract meaningful features, such as average value, maximum value, standard deviation, etc. Use the sliding window technique to analyze the data, set the time window (such as 1 hour, 1 day) for aggregation. According to the parsed data, define an intelligent index strategy to improve data retrieval efficiency. Generate a unique identifier as the index for the data of each time period. The index should include the timestamp, sensor ID, and the monitored parameter to facilitate quick search for the monitoring data of a specific time period. Use the index function of the database (such as INDEX in MySQL) to create an index for the stored data to accelerate query operations, ensuring that the index covers all key attributes. Store the results of parsing and index generation in the database, including the aggregated data and its index information for each time period, to ensure the integrity and consistency of the data. Record the access statistics information for each time period for subsequent analysis. Use data visualization tools (such as Matplotlib or Tableau) to display the changes in real-time monitoring data and generate a dashboard for users to view the environmental monitoring situation in real-time.

[0087] Step S2: Predict the evolution of environmental parameters for the filtered and noise-reduced monitoring data, and perform parameter visualization rendering to construct a visual view of the environmental prediction trend.

[0088] In this embodiment, environmental monitoring data, including parameters such as temperature, humidity, and air quality, are extracted from the previous processing stage to ensure the integrity and accuracy of the data. These data should include timestamps for subsequent analysis. A suitable filtering algorithm (such as Kalman filtering or moving average) is used to denoise the data, and appropriate parameters and window sizes are selected to balance data smoothing and feature retention. For example, the moving average effectively removes short-term fluctuations while maintaining long-term trends. After the filtering process, the clarity and accuracy of the data are verified by visually comparing the original data with the filtered data to ensure that the noise has been effectively reduced. According to the characteristics of the data, a suitable time series prediction model, such as SARIMA (Seasonal AutoRegressive Integrated Moving Average model) or a machine learning model (such as LSTM), is selected. These models can capture trends and seasonal variations in the data. The data is divided into a training set and a test set, usually with 70% of the data used for training and 30% for testing. The training set is used to train the model, and the prediction performance of the model is evaluated through the test set to ensure its accuracy. The trained model is used to predict future environmental parameters, such as temperature and humidity changes in the next 7 days, and the prediction results and their confidence intervals are recorded for subsequent analysis. A suitable visualization tool (such as Matplotlib or Seaborn) is selected for data visualization. According to the requirements, static charts or interactive charts are selected. The original data, filtered data, and predicted data are displayed in the same chart for users to intuitively compare. Different data types are distinguished by different colors and markers to ensure clear information transmission. The visualization effect is verified to ensure that the chart is clear and easy to understand and can effectively convey information about changes in environmental parameters, and adjustments and optimizations are made according to user feedback, such as changing the color scheme or adding annotations.

[0089] Step S3: Obtain the historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data and construct a distributed storage architecture to obtain a distributed data node storage framework;

[0090] In this embodiment, historical monitoring data is extracted from the storage system of the data recorder. These data are stored in a local database or cloud storage to ensure the integrity and losslessness of the acquired data. Using an appropriate database query language (such as SQL) or API interface, all historical data within the required time range is extracted. For example, all monitoring records for the past year are extracted for subsequent processing. After the data is extracted, ensure that its format is suitable for subsequent processing. Usually, the data is converted into a unified structured format, such as CSV or JSON, for subsequent compression and storage. Perform data type checks to ensure data type consistency to avoid errors in subsequent processing. Adopt adaptive lossless compression algorithms (such as LZ77, LZ78, or Deflate). These algorithms dynamically adjust compression parameters according to the characteristics of the data to optimize the compression effect. When selecting an algorithm, consider the characteristics of the data (such as repeatability, data type), and select the most suitable compression method to ensure compression efficiency. Compress the extracted historical data. During the compression process, monitor the compression ratio and compression time to ensure finding a balance between compression efficiency and processing time. Record the data sizes before and after compression to evaluate the compression effect. Ideally, the compression ratio should be between 2:1 and 5:1, depending on the redundancy of the data. According to the usage requirements and access patterns of the data, design a distributed storage architecture. Select a suitable distributed file system (such as HDFS, Ceph, or GlusterFS) as the storage backend to support large-scale data storage and access. Determine the data sharding rules and replica strategies to evenly distribute the storage load among multiple nodes to ensure high availability and data redundancy. Shard and store the compressed data on different nodes, with each node storing a certain amount of data shards to avoid single points of failure. Configure a load balancing strategy to ensure reasonable distribution of read and write requests for each node. Implement a monitoring mechanism to monitor the status and performance of each storage node in real time. Use monitoring tools (such as Prometheus or Grafana) to collect and display storage performance metrics. According to the monitoring data, continuously optimize the storage architecture, adjust the data distribution and replica strategies to cope with changes in access patterns and growth in storage requirements.

[0091] Step S4: Based on the data index of each time period, embed the environmental prediction trend visualization view into the distributed data node storage framework to generate a real-time data embedded storage framework;

[0092] In this embodiment, based on the environmental prediction results for each time period in the previous processing, corresponding data indexes are generated. The indexes should include timestamps, sensor IDs, and environmental parameters (such as temperature, humidity, etc.) to facilitate subsequent rapid retrieval. Indexes are created using database technologies (such as Elasticsearch or Apache Solr) to make data retrieval operations more efficient. Trend data is extracted from the output of the environmental prediction model and its change patterns are analyzed. This data should include time periods of rising, falling, or stable trends for real-time storage. These trend data are combined with the indexes to create a complete data record for each time period, facilitating subsequent storage and processing. A real-time data embedded storage framework is designed to ensure that it can support distributed storage and fast data access. A suitable distributed database (such as Cassandra, HBase, or MongoDB) is selected as the data storage backend, and a data sharding strategy is determined to evenly distribute the data across each storage node to ensure load balancing. Sharding is performed based on time periods or sensor IDs. The distributed storage nodes are configured to ensure that each node can handle data read and write requests. Appropriate hardware resources (such as CPU, memory, and storage space) are set to meet the requirements of real-time processing. A data replication strategy is adopted to ensure redundant storage of data, improving the reliability and availability of the system. The indexes and environmental prediction trend data are embedded into the distributed storage framework. By writing a data writing script, the data is written to the corresponding nodes according to the designed sharding strategy. During the writing process, the writing latency and error rate are monitored to ensure that the data can be stored quickly and accurately. A real-time data update mechanism is implemented so that when environmental parameters change, new data can be immediately embedded into the storage framework. A message queue (such as Kafka or RabbitMQ) is used to process real-time data streams to ensure data real-time. A data update frequency is set, for example, updated once a minute, to ensure that the stored data is always up-to-date. A monitoring system is established to monitor the performance of the distributed storage framework in real-time, collecting metrics such as the load, response time, and storage usage of each node to promptly discover and solve problems. Visualization tools (such as Grafana) are used to display the monitoring data to help operation and maintenance personnel quickly identify bottlenecks and faults. According to the monitoring results, the storage framework is continuously optimized, and the data sharding strategy and replication strategy are adjusted to improve data access efficiency and the overall performance of the system. The performance of the storage nodes is regularly evaluated, and node expansion is carried out when necessary to cope with the growing data demand.

[0093] Step S5: Based on the real-time data embedded storage framework, perform data redundancy analysis and redundant node deletion processing to obtain a redundant node optimized framework;

[0094] In this embodiment, redundant data refers to duplicate or redundant data records in a storage system. Select metrics for redundancy analysis, such as data duplication rate, storage occupancy rate, and access frequency. Set a threshold to determine which nodes or data records are considered redundant. For example, if the proportion of duplicate data in a certain node exceeds 50%, the data of that node can be considered redundant. Extract the data of all nodes from the real-time data embedding storage framework, organize and summarize it. Use query tools or scripts to obtain the storage data of each node to ensure data integrity. Record information such as the data volume, storage space occupancy, and access frequency of each node for subsequent analysis. By writing scripts, compare the data of each node to identify duplicate data records. Use a hash algorithm to generate a unique identifier for each data record to quickly find duplicates. Calculate the data duplication rate of each node and compare it with the previously set threshold to determine whether it is a redundant node. For example, if the data duplication rate of a certain node is 60%, then that node is marked as redundant. According to the analysis results, identify all redundant nodes, record information such as the IDs, stored data volume, and duplication rate of these nodes, and generate a redundant node report. Cross-check the list of redundant nodes with the system's access logs to evaluate its impact on system performance. According to the results of the redundancy analysis, formulate a deletion strategy for redundant nodes. Select to completely delete redundant nodes or archive them for subsequent auditing and data recovery. Determine the deletion priority, for example, give priority to deleting nodes with low access frequency and high data duplication rate. Use the deletion command of the database or the API of the storage service to delete the marked redundant nodes one by one. During the deletion process, ensure data consistency and integrity to avoid affecting the data of other normal nodes. Before the deletion operation, perform a data backup in case important data is lost due to unexpected situations. After the redundant node deletion process, monitor the performance of the storage framework in real time, evaluate the impact of the deletion operation on system load and response time. Use monitoring tools to collect performance metrics such as storage utilization and access latency. Compare the performance data before and after deletion, and analyze the improvement effect of redundant node deletion on the performance of the storage framework. According to the monitoring data and user feedback, continuously optimize the management strategy of redundant nodes. Regularly perform redundancy analysis to identify new redundant nodes to ensure that the storage framework always maintains the best state. As the data volume grows, adjust the frequency of redundancy analysis and the deletion strategy to ensure the efficient use of system resources.

[0095] Step S6: Calculate the storage load for each node in the redundant node optimization framework and perform dynamic index structure adjustment, thereby constructing a dynamic index adjustment storage framework to execute the data recorder information management task.

[0096] In this embodiment, metrics for storage load calculation are determined. These metrics include storage capacity utilization rate, data access frequency, node response time, CPU utilization rate, etc. Appropriate weights are selected to comprehensively evaluate the importance of each metric. The weights for storage capacity and access frequency are set to 0.4, and the weights for response time and CPU utilization rate are set to 0.3. Based on this, the overall load is calculated. Performance data of each node are extracted from the storage framework, including current storage usage, access logs, and system resource occupancy, etc. Monitoring tools are used to collect real-time data and organize it into an analyzable format, such as CSV or a database table. Information such as the ID, storage capacity, used space, access times, and response time of each node is recorded to provide basic data for subsequent load calculation. A calculation script is written to perform load calculation for each node, applying the previously defined combination of metrics and weights to calculate the comprehensive load index of each node. For example, the load index. The calculation results are analyzed to identify nodes with high load and nodes with low load. A threshold is set. For example, nodes with a load index exceeding 0.7 are regarded as high-load nodes. The load conditions of different nodes are compared to evaluate the load balance degree of the current storage framework and identify potential bottlenecks. Based on the results of the load calculation, it is evaluated whether the current index structure meets the requirements of data access. Considering the data access pattern, indexes that need to be optimized are identified. For example, for high-load nodes, auxiliary indexes need to be added or existing indexes need to be adjusted to improve data retrieval efficiency. Based on the evaluation, dynamic index adjustment is implemented. According to the node load situation, the index structure is dynamically created, deleted, or modified. For example, for nodes with high access frequency, time-range-based indexes are added to accelerate the access of time-series data. The index management function of the database management system is used to ensure that the adjustment process does not affect the normal operation of the system. The performance changes after adjustment are monitored to ensure the effectiveness of the index structure. The results of the dynamic index adjustment are integrated into the storage framework to construct a dynamic index adjustment storage framework, ensuring that the new framework can support real-time data updates and flexible index management. Performance testing is carried out to evaluate the performance of the new framework in the data recorder information management job, including data writing speed, query response time, and system load, etc. After the new framework runs, continuous monitoring is implemented to track the storage load and index performance in real time. According to the monitoring results, regular load calculations and index adjustments are performed to ensure the efficiency and adaptability of the framework. As the data volume grows and the access pattern changes, the index strategy is adjusted in a timely manner to optimize the performance of the storage framework and ensure that it can always meet the requirements of the information management job.

[0097] In this embodiment, refer to Figure 2 , which is a schematic diagram of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include:

[0098] Step S11: Obtain real-time environmental monitoring data based on the multi-source sensors of the data recorder;

[0099] Step S12: Perform parametric normal distribution analysis on the real-time environmental monitoring data to generate a normal distribution range;

[0100] Step S13: Filter out abnormal outlier parameters based on the normal distribution range to obtain outlier-filtered monitoring data;

[0101] Step S14: Perform high-frequency filtering and noise reduction on the outlier-filtered monitoring data to obtain filtered and noise-reduced monitoring data;

[0102] Step S15: Perform multi-period data deep semantic parsing on the filtered and noise-reduced monitoring data to generate data deep semantic features for each time period;

[0103] Step S16: Define intelligent indexes based on the data deep semantic features for each time period to obtain data indexes for each time period.

[0104] In this embodiment, a data recorder is configured to ensure that it can be connected to a variety of sensors, including temperature sensors, humidity sensors, barometric pressure sensors, PM2.5, PM10 sensors, and GPS modules, etc. These sensors are used to monitor environmental conditions and vehicle status in real time. Ensure that the sensors are correctly calibrated, and the data acquisition frequency is set to once per second to ensure the real-time and accuracy of the data. Start the data recorder to begin obtaining real-time data from each sensor. The data includes information such as real-time temperature, humidity level, barometric pressure, particulate matter concentration, and vehicle location. Record the timestamp of the data so that the time variation of the data can be tracked during subsequent analysis. Store the collected real-time monitoring data in a local storage device, and at the same time upload the data to a cloud server via a wireless network for subsequent access and analysis. Ensure the security during the data transmission process, and protect the integrity and privacy of the data through an encrypted connection. Obtain real-time environmental monitoring data from the cloud server and perform preliminary cleaning on the data to remove missing values and outliers to ensure the integrity and accuracy of the data. Organize each monitoring parameter into a data frame (DataFrame) for subsequent analysis. Use statistical analysis software (such as the SciPy library in Python) to perform a normal distribution test on each environmental monitoring parameter. Common methods include the Shapiro-Wilk test or the Kolmogorov-Smirnov test to determine whether the data conforms to a normal distribution. If the data does not conform to a normal distribution, it is processed by methods such as logarithmic transformation, Box-Cox transformation, or Z-score standardization. Once it is confirmed that the data conforms to or conforms to a normal distribution after processing, calculate the mean and standard deviation of each parameter to generate a normal distribution range. Record the range of each parameter. For example, the normal range of temperature is the values within ±2 standard deviations, which will be used for subsequent outlier detection. According to the normal distribution range generated in step S12, perform outlier detection on the real-time monitoring data. Usually, the threshold is set to the mean ±3 standard deviations, and the values outside this range are regarded as abnormal data. Use the NumPy library and Pandas library in Python to write code to quickly identify whether there are outliers for each parameter. Remove the identified outliers from the monitoring data set to form the monitoring data after outlier filtering, ensuring that only the values that conform to the normal range are retained in the data set. Record the data changes during the filtering process, such as the number of original data and the number after filtering, to ensure the transparency and traceability of data processing. Select a suitable high-frequency filtering algorithm. Common methods include moving average filtering, Kalman filtering, or low-pass filters. The choice of the appropriate method depends on the characteristics of the data and the type of noise. For example, moving average filtering is suitable for processing smooth signals, while Kalman filtering is applicable to the state estimation of dynamic systems. Input the monitoring data after outlier filtering into the selected filtering algorithm for high-frequency filtering processing, and ensure that the filtering parameters (such as the window size) are reasonably set according to the characteristics of the data. Record the changes of each parameter before and after filtering, such as the change trend of temperature.Ensure that the filtering process achieves the expected effect, save the filtered monitoring data for subsequent analysis and processing, ensure that the data format and structure are consistent with the previous ones, so as to carry out continuous analysis, divide the monitoring data into multiple time periods (such as every hour, every half hour or every minute) according to the timestamp of the data, use the time series processing function of the Pandas library to implement data resampling, set the length of each time period, ensure the rationality of the division, so as to facilitate subsequent in-depth analysis, conduct in-depth analysis of the data in each time period, extract relevant semantic features, including statistical features such as average, maximum, minimum, volatility, as well as trend and periodicity analysis, use machine learning algorithms (such as cluster analysis or principal component analysis) to mine the potential patterns and features in the data, and organize the extracted deep semantic features into structured data , and store them in the database for subsequent query and analysis, ensuring that the feature records of each time period are clear and concise. According to the extracted deep semantic features, formulate an indexing strategy to determine which features are most representative and suitable for indexing. Usually, features with high variability and information content are selected. For example, the average temperature, humidity and air quality index of each time period are selected as the main index features. A unique index value is generated for each time period. A hash algorithm or a unique identifier (UUID) is usually used to ensure the uniqueness and efficiency of the index. The index is associated with the corresponding deep semantic features for subsequent rapid retrieval and analysis. The generated index and related data features are stored in the database to ensure efficient retrieval capabilities. By building an index, data for a specific time period can be quickly found, which improves the speed and efficiency of data analysis. ,

[0105] In this embodiment, refer to Figure 3 , is a flowchart of detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include:

[0106] Step S21: mining the environmental parameter change characteristics of the filtered noise reduction monitoring data to generate environmental parameter change trend characteristics;

[0107] Step S22: Calculate the change frequency of the environmental parameter change trend characteristics to extract the environmental parameter change frequency;

[0108] Step S23: identifying periodic evolution rules based on the frequency of environmental parameter changes to generate periodic rules of environmental change trends;

[0109] Step S24: predicting the evolution of environmental parameters of the filtered noise reduction monitoring data based on the cycle law of environmental change trends, so as to obtain the prediction characteristics of future environmental trends;

[0110] Step S25: Parameter visualization rendering is performed on the future trend prediction features of the environment to construct a visualization view of the environment prediction trend.

[0111] In this embodiment, the time series analysis method, such as the sliding window method, is used to calculate the change rate and abnormal fluctuation of environmental parameters in each time period. For example, by calculating the difference between each time point and the previous time point, the change amount of each parameter is obtained, and the formula is Δx(t)=x(t)-x(t - 1). Further, statistical analysis tools (such as the Pandas and NumPy libraries in Python) are applied to extract the change characteristics of each parameter, including the mean, variance, maximum value, and minimum value, etc. By visualizing the extracted change characteristics (such as plotting a trend chart), the long-term change trend of each environmental parameter is identified. For example, the temperature shows periodic fluctuations with the change of seasons. Subsequently, a data frame is generated based on the trend characteristics to record the change trend of each parameter, including the change direction (rising, falling, or stable) and the change amplitude for subsequent analysis. Based on the change trend characteristics extracted in step S21, a threshold for feature change is defined. For example, a threshold (such as ±2°C) is set to judge the significant change of temperature. By traversing the change data of each time period, the number of changes exceeding the threshold is calculated to obtain the frequency. For example, if the temperature changes more than ±2°C 10 times in a month, the frequency is 10 times / month. The change frequency of each environmental parameter is recorded in the structured data, including information such as the parameter name, change frequency, and observation time period. The Pandas library is used to create a data frame for subsequent analysis and visualization. Statistical analysis is performed on the frequency data, such as calculating the average change frequency and standard deviation of each parameter to evaluate its change stability. The Python SciPy library is used for these statistical calculations. The Fourier transform (FFT) analysis method is used to extract the frequency components of the environmental parameter changes. By converting the time series data into the frequency domain, the main periodic components are identified. The FFT function in the Python NumPy library is used to perform the Fourier transform on the time series data to identify the main frequencies and the corresponding amplitudes. According to the results of the Fourier transform, the significant periodic components are extracted, and the lengths of these periods are determined. For example, if the main period of temperature change is 30 days, it is inferred that there is a seasonal change. The periodic characteristics of each parameter are recorded, including the period length, amplitude, and phase information. These information will be used for subsequent prediction analysis. The extracted periodic rules are stored in the database for subsequent query and analysis, ensuring that the periodic characteristics of each parameter are recorded and relevant statistical information is provided. A suitable prediction model is selected for future trend prediction. For example, time series prediction methods such as the seasonal ARIMA model (SARIMA) or the long short-term memory network (LSTM) are used. The parameters of the model are set. According to the periodic characteristics extracted in step S23, appropriate lag periods and seasonal parameters are selected. The selected prediction model is trained using the historical filtered and noise-reduced monitoring data. The data set is split into a training set and a test set to ensure the generalization ability of the model. The performance of the model is evaluated through cross-validation and prediction errors (such as RMSE).To ensure the effectiveness of the model in actual prediction, once the model training is completed, use the model to predict the next environmental parameters, record the predicted values of each parameter in the future time period, such as the temperature and humidity changes in the next 30 days, etc., and organize the prediction results into structured data for subsequent analysis and visualization. Select a suitable visualization tool (such as Matplotlib, Seaborn, or Plotly) to plot the prediction trend chart of environmental parameters. These tools can generate high-quality charts and support interactive visualization. Input the prediction results in step S24 into the visualization tool to generate the trend chart of each environmental parameter. The chart should display the time axis, predicted values, and their confidence intervals to intuitively reflect the change trend and uncertainty. For example, plot the temperature change trend in the next 30 days to show the rise and fall of temperature and the corresponding confidence intervals. Organize the generated visualization charts into a report for easy sharing with relevant teams and decision-makers. The report should include an explanation of the environmental change trend and a discussion of potential impacts. Add annotations to the visualization charts to highlight key change points and potential risks for subsequent decision-making.

[0112] In this embodiment, the specific steps of step S21 are as follows:

[0113] The filtered and noise-reduced monitoring data includes environmental temperature parameters, environmental humidity parameters, light intensity, air pollutant monitoring data, and wind monitoring data;

[0114] Fit the temporal variation characteristics of environmental temperature parameters and environmental humidity parameters to construct an environmental temperature change curve and an environmental humidity change curve;

[0115] Conduct an analysis of the spatial distribution of light intensity to obtain the spatial distribution characteristics of light intensity;

[0116] Calculate the pollutant peak values for the air pollutant monitoring data and extract the pollutant peak points;

[0117] Identify the wind direction and wind force intensity for the wind monitoring data to extract the wind force characteristics;

[0118] Mine the environmental change trend from the spatial distribution characteristics of light intensity, pollutant peak points, wind force characteristics, environmental temperature change curve, and environmental humidity change curve to obtain the environmental parameter change trend characteristics.

[0119] In this embodiment, environmental temperature and humidity parameters are extracted from the filtered and noise-reduced monitoring data, including timestamps, temperature values, and humidity values. To ensure the integrity and consistency of the data, missing values and outliers are removed. The data is organized into a time series format using the Pandas library, with the timestamp set as the index. Time series analysis methods, such as the autoregressive integrated moving average model (ARIMA) or exponential smoothing, are used to fit the temperature and humidity data. First, a stationarity test (such as the ADF test) is performed. If the data is non-stationary, differencing is carried out. The ARIMA model is fitted using the statsmodels library in Python, and appropriate model parameters (p, d, q) are selected to minimize the AIC value. Once the model fitting is complete, the fitting results are visualized using the Matplotlib library, and temperature change curves and humidity change curves are plotted. The curves should display the actual observed values and the fitted values to facilitate the identification of the model's fitting effect. Annotations are added to the chart to emphasize trends and seasonal variations. For example, the temperature reaches its peak in summer, while the humidity is higher in winter. Light intensity data, including time, location (GPS coordinates), and light intensity values, is extracted from the monitoring data. Ensure that the data has spatial reference for subsequent analysis. Geographical Information System (GIS) tools (such as QGIS or Geopandas) are used to visualize the light intensity data spatially. A heatmap of the light intensity is created to intuitively show the distribution differences of the light intensity in different regions. The mean and standard deviation of the light intensity in each region are calculated to identify regions with high and low light intensity as the focus of analysis. Combining the results of the spatial distribution analysis, light intensity distribution characteristics are extracted, such as the number, distribution density of high light intensity regions, and their relationship with the surrounding environment. These characteristics will be used for subsequent mining of environmental change trends. Monitoring data of various air pollutants (such as PM2.5, PM10, NO2, SO2, etc.) are extracted from the monitoring data. Ensure the completeness of the data and perform necessary cleaning. Using data analysis methods, the peak points of each pollutant are calculated. This is achieved through the sliding window method. The window size (such as 3 hours) is set, and the maximum value within the window is calculated. The find_peaks function in the SciPy library of Python is used to identify the peak points in the pollutant concentration curve and record their timestamps and corresponding concentration values. The extracted peak points are organized into a data frame, recording the peak time, concentration, and their positions in the time series for each pollutant. The frequency and time period of peak occurrences are analyzed to identify the high-incidence periods of pollution sources. Wind data, including wind speed, wind direction, and timestamp, is extracted from the monitoring data. Ensure the accuracy of the data and remove outliers. Polar plots of wind speed and wind direction are plotted using data visualization tools (such as Matplotlib or Seaborn) to intuitively show the distribution characteristics of the wind. The average value, maximum value of the wind speed, and the main direction of the wind direction are calculated to identify the periods with high wind speed and the main wind direction. For example, a histogram is used to show the frequency distribution of the wind direction.Record the extracted wind power characteristics in structured data, including time period, average wind speed, maximum wind speed, and main wind direction. Analyze how these characteristics affect other environmental parameters, such as temperature and humidity. Use multiple regression analysis or machine learning methods (such as decision trees or random forests) to explore the mutual influence between different environmental parameters. By constructing a model, identify which parameters have a significant impact on the environmental change trend. Use cross-validation to evaluate the stability and accuracy of the model to ensure the reliability of the results. Record the mined environmental change trend characteristics in the database to ensure the data is structured for subsequent analysis and query. Use visualization tools to generate comprehensive trend charts to display the change trends and relationships of different environmental parameters and provide intuitive analysis results.

[0120] In this embodiment, refer to Figure 4 , which is a schematic diagram of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include:

[0121] Step S31: Obtain the historical storage data of the data recorder;

[0122] Step S32: Define a time span based on the historical storage data and perform uniform segmentation processing according to the time span to generate multiple historical storage data nodes;

[0123] Step S33: Perform adaptive lossless compression on multiple historical storage data nodes to obtain multiple lossless compression data nodes;

[0124] Step S34: Construct a distributed storage architecture for multiple lossless compression data nodes to obtain a distributed data node storage framework.

[0125] In this embodiment, to ensure the normal operation and integrity of the data recorder, before acquiring data, check the storage status of the device to ensure no data loss or damage. Determine the data types, including environmental temperature, humidity, light intensity, air pollutant concentration, wind monitoring data, etc. Connect the data recorder via USB interface, Wi-Fi, or Bluetooth, etc. Use appropriate software tools (such as data management software or custom scripts) to extract historical data from the device, ensuring the consistency of the data format during the data extraction process, such as JSON, CSV, or database format. The extracted data should include timestamps for subsequent time series analysis. Store the extracted data in the local file system or database and perform data backup to prevent data loss. Use a version control system (such as Git) to manage the historical versions of the data. During the data storage process, record the time of data acquisition, device status, and data volume for subsequent tracking and analysis. Determine the time range for data analysis, such as the past year, six months, or during a specific event. Set a reasonable time span according to the data collection frequency (such as every minute, every hour). By analyzing the time distribution of the historical data, select an appropriate start time and end time to ensure that the data within the time span is complete enough. Use a programming language (such as Python) to evenly segment the historical stored data, set the segmentation time interval, such as daily, weekly, or monthly, depending on the characteristics of the data and the analysis requirements. Use the resample function of the Pandas library to resample the data at the set time interval. Select a suitable lossless compression algorithm, such as LZ77, Huffman coding, or Zstandard. These algorithms can reduce the storage size of the data without losing data. Evaluate the compression efficiency and speed of different algorithms to select the most suitable one. Compare the compression ratio and processing time of different algorithms when processing a specific data set through experiments. Use the compression libraries in Python (such as zlib or lzma) to compress each data node and store the compressed data in a new data structure, ensuring that the original size and compressed size of each compressed node are recorded for subsequent performance evaluation. Evaluate different distributed storage technologies (such as Hadoop HDFS, Apache Cassandra, Amazon S3, etc.) and select a suitable storage solution. Consider factors including data volume, access frequency, scalability, and fault tolerance. Set the architecture of the storage cluster, determine the number of nodes, storage capacity, and network configuration. Distribute the compressed data nodes to multiple storage nodes to ensure data redundancy and security. For example, adopt data sharding and replication strategies to ensure that each data node exists in multiple storage nodes. Use the API of the distributed file system (such as the HDFS API of Hadoop) to upload and manage the compressed data nodes. Design a data access interface to ensure that users and applications can efficiently access and retrieve the data stored in the distributed architecture.Provide data query functions using RESTful APIs or GraphQL interfaces, record the status and performance metrics of each storage node for monitoring and maintenance to ensure the stability and high availability of the system.

[0126] In this embodiment, step S4 includes the following steps:

[0127] Step S41: Quantify the data sensitivity of the environmental prediction trend visualization view to obtain the sensitivity of the instant parameter view;

[0128] Step S42: Analyze the data access requirements of the instant parameter view sensitivity to generate the characteristics of the real-time environmental monitoring data access requirements;

[0129] Step S43: Make a storage decision based on the characteristics of the real-time environmental monitoring data access requirements to obtain the storage location of the real-time data;

[0130] Step S44: Index and mark the environmental prediction trend visualization view according to the data index of each time period, and perform distributed storage embedding on the distributed data node storage framework based on the storage location of the real-time data to generate a real-time data embedded storage framework.

[0131] In this embodiment, select a suitable sensitivity analysis method, such as local sensitivity analysis (e.g., the first-order partial derivative method) or global sensitivity analysis (e.g., the Sobol method or variance decomposition method). Local sensitivity analysis is suitable for changes in a small range, while global sensitivity analysis is more suitable for the evaluation of the overall parameter impact. Determine the parameters to be analyzed, such as temperature, humidity, light intensity, air pollutant concentration, and wind force. Extract the predicted values and their change ranges of each parameter from the environmental prediction trend visualization view. Use the Pandas library in Python to perform statistical analysis on the historical data, calculate the mean and standard deviation of each parameter. Make a small perturbation (e.g., ±5%) for each parameter within its change range, observe the changes in the prediction results, and record the output results. Calculate the sensitivity indicators of each parameter. Usually, use the sensitivity coefficient SC = , where Y is the predicted result, X is the parameter value, and ΔY and ΔX are the change amounts. Through user research, questionnaires, or usage log analysis, collect information on users' access frequencies, access times, and access types (such as queries, analyses, or downloads) of environmental monitoring data. Use data analysis tools (such as Google Analytics or custom data analysis scripts) to analyze user behavior to identify access patterns and parameters of high-frequency requests. Based on the collected data, create a user access feature data frame to record the access frequency, access time period, and user type (such as researchers, managers, etc.) of each parameter. Calculate the average access frequency and peak access period of each parameter to identify which parameters are most concerned. Record the access requirement characteristics in a database and use data visualization tools (such as Matplotlib or Seaborn) to generate an access requirement heat map to intuitively display users' access requirements for each parameter. Ensure that the change trend of each feature is recorded for subsequent storage decisions and optimization. Combine the results of the access requirement analysis to determine which data should be stored on high-performance storage devices (such as SSDs) to meet the needs of high-frequency access. For data with a lower access frequency, choose to store it on a lower-cost device (such as HDDs) to reduce storage costs. Design a distributed storage architecture, considering data redundancy backup and load balancing. Utilize a distributed file system (such as HDFS) to achieve high availability and fast access. Set a data sharding strategy to distribute high-frequency access data across multiple nodes to improve concurrent access capabilities. Create a storage decision document to record the storage location, storage type, and storage strategy of each parameter. Ensure the traceability of the document for subsequent optimization and adjustment. According to the data indexes generated in steps S41 and S42, mark the environmental prediction trend visualization view for each time period. Ensure that each visualization element (such as curves, charts) can quickly access the corresponding data nodes. Use the index mechanism of the database (such as the indexing function of Elasticsearch or SQL databases) to accelerate data retrieval. Based on the storage location of real-time data, embed the data nodes into the distributed storage framework. Use data management tools (such as Apache Kafka or RabbitMQ) to achieve real-time transmission and processing of data streams. Ensure that the embedded storage framework can support dynamic expansion to adapt to changes in data volume and user access requirements. Test the performance of the new storage framework and access mechanism, record the access response time and system load conditions to evaluate the optimization effect. Make corresponding adjustments and optimizations according to the test results to ensure the high efficiency and stability of the system operation.

[0132] In this embodiment, step S5 includes the following steps:

[0133] Step S51: Identify the user access behavior of each data node one by one for multiple lossless compression data nodes, and extract the access behavior data of each node;

[0134] Step S52: Calculate the data access frequency for the access behavior data of each node to obtain the data access frequency of each node;

[0135] Step S53: Analyze the access time range of the access behavior data of each node to generate the access time range feature of each node;

[0136] Step S54: Evaluate the historical data access requirements based on the data access frequency of each node and the access time range feature of each node, so as to obtain the access requirement evaluation value of each node;

[0137] Step S55: Perform data redundancy comparison on the access requirement evaluation value of each node based on a preset node access requirement threshold. When the access requirement evaluation value of the node is lower than the preset node access requirement threshold, perform redundant node deletion processing on the real-time data embedding storage framework to obtain a redundant node optimization framework.

[0138] In this embodiment, user access logs are extracted from the data storage system. These logs contain user IDs, access times, accessed node identifiers, and access methods (such as reading, writing, etc.). To ensure the integrity and accuracy of the log data, data analysis tools (such as ELK Stack or Splunk) are used for centralized management and analysis of the logs. Scripts are written to analyze the log data to identify the access behaviors of each node. The Pandas library in Python is used to load the log data into a data frame for subsequent processing. The access records of each node are grouped to count the number of accesses, the accessing users, and their access methods for each node. Based on the access behavior data extracted in step S51, the access frequency of each node within a specific time range is calculated. For example, the time range is set as the past month, and the Pandas library is used to process the data to calculate the access frequency of each node. The calculation results are organized into a data frame containing information such as node ID, access frequency, and time period, ensuring that the changes in the access frequency of each node are recorded for subsequent analysis. Statistical information about the access frequency, such as mean, maximum value, and standard deviation, is added to the data frame for analyzing the access patterns. Using the access behavior data in step S51, the access timestamps of each node are extracted and their time distribution is analyzed. The access time range of each node is calculated, including the earliest and latest access times. The characteristics of the access time range of each node are recorded, including the number of accesses and access intervals (such as average access interval, maximum interval, etc.). The peak access periods (such as specific time periods of each day) of each node are counted, and a visualization chart (such as a heat map) is generated to display the access time characteristics. Based on the access frequency and the characteristics of the access time range calculated in steps S52 and S53, an evaluation model is constructed. A simple weighted model is adopted, and weights are assigned to each node according to the access frequency and the access time range. The weights are set. For example, the access frequency accounts for 70% and the access time range accounts for 30%, Demand_Score = 0.7×Frequency + 0.3×Time_Range, calculate the access requirement evaluation value for each node using the above formula, and record the results in a data frame, ensuring that the results include the node ID, access frequency, time range, and their corresponding evaluation values. Statistically analyze the evaluation values of all nodes to identify which nodes have high access requirements and which have relatively low ones. According to historical data and business requirements, set a preset threshold for node access requirements. For example, set the threshold to the average or median of the access requirement evaluation values, ensuring that the threshold is reasonable. Adjust the threshold by analyzing historical access records to ensure that it can effectively distinguish between high-frequency and low-frequency nodes. Traverse the access requirement evaluation values of each node and perform redundant comparisons. When the access requirement evaluation value of a node is lower than the preset threshold, mark it as a redundant node. Use database operations (such as DELETE statements) or storage management tools to delete the real-time data embedding of these redundant nodes to free up storage resources. Record the operation logs of the redundant node deletions and generate a report indicating the deleted nodes, reasons, and expected storage savings. Re-evaluate the performance of the storage framework to ensure that the system can still meet user access requirements after deleting the redundant nodes.

[0139] In this embodiment, step S6 includes the following steps:

[0140] Step S61: Calculate the storage load of each node for the redundant node optimization framework node by node to obtain the storage load characteristics of each node;

[0141] Step S62: Predict the access peak of each node based on the storage load characteristics of each node to obtain the predicted time point of the access peak of the node;

[0142] Step S63: Predict the future usage requirements of users based on the predicted time point of the access peak of the node to obtain the predicted usage requirements of each node;

[0143] Step S64: Dynamically adjust the index structure of the redundant node optimization framework according to the predicted usage requirements of each node, thereby constructing a dynamically indexed adjusted storage framework to perform the data recorder information management operation.

[0144] In this embodiment, the storage information of each node is extracted from the redundant node optimization framework, including used space, available space, data type, data size, etc., to ensure data integrity and avoid the impact of missing values on the calculation results. Monitoring tools (such as Prometheus or Nagios) are used to collect real-time storage load data for analysis. Each node is analyzed one by one to calculate storage load characteristics, including storage utilization rate, read / write rate, number of I / O requests, etc. Storage_Utilization = Used_Space / Total_Space × 100%. The storage load characteristics of each node (such as storage utilization rate, read / write rate, etc.) are organized into a data frame, recording the node ID and various load metrics, ensuring that all necessary information is included in the data frame for subsequent analysis. Visualization charts (such as bar charts) are generated to show the storage load conditions of each node, facilitating the identification of high-load and low-load nodes. Appropriate time series prediction models, such as ARIMA, SARIMA, or LSTM, etc., are selected. These models can predict future access patterns based on historical data. Using historical access data and storage load characteristics, a prediction model is constructed. The statsmodels library or TensorFlow / Keras in Python is used for model training. The historical data is divided into a training set and a test set. The model is trained using the training set and the accuracy of the model is verified through the test set. Appropriate performance metrics (such as MSE or MAE) are selected to evaluate the model effect. Once the model training is completed, the model is used to predict the future access conditions of each node, especially the time points of access peaks. The predicted time points of access peaks of each node are recorded in the data frame, including the node ID, the predicted peak time, and the corresponding predicted access frequency. Visualization charts are generated to show the predicted access peak conditions of each node, helping to identify upcoming high-load periods. Combining the predicted time points of access peaks of the nodes, the historical usage behavior of users is analyzed to construct a demand prediction model. Regression analysis, time series analysis, or machine learning models (such as random forest) are used for demand prediction. Historical data is used to calculate usage demand metrics, such as the number of accesses, the number of users, data request types, etc. The historical usage data is used to train the demand prediction model, setting the target variable as the future usage demand. For example, predict the usage demand of each node within the next week. Using the trained model, the usage demand of each node at the predicted time points of access peaks is predicted and the prediction results are recorded. The predicted usage demand results of each node are organized into a data frame, including the node ID, the predicted usage demand value, and the prediction time period, ensuring the clarity of the results for subsequent reference. Visualization charts (such as line charts) are generated to show the changing trends of the predicted usage demand of each node, helping to identify potential resource demand peaks. According to the predicted usage demand of the nodes, a dynamic index structure is designed to improve data access efficiency. Consider increasing the index priority of high-demand nodes.Ensure that it can respond quickly to user requests, implement dynamic indexing using data structures such as B-trees or hash tables, adjust the index structure according to the predicted results of demand, write scripts or use database management tools to update the index structure of each node according to the predicted results of usage requirements, ensure that the integrity and consistency of existing data are not affected during the update process, record each index adjustment operation, ensure that there is a complete operation log for subsequent auditing and optimization, conduct performance testing in the adjusted storage framework, monitor data access latency and system load conditions to evaluate the effect of the dynamic index structure, and further optimize according to the test results and adjust the index strategy to ensure the efficiency and stability of the system in the data recorder information management operation.

[0145] In this embodiment, a data recorder information management system based on artificial intelligence is provided for performing the data recorder information management method based on artificial intelligence as described above, including:

[0146] A deep semantic parsing module for acquiring real-time environmental monitoring data based on multi-source sensors of the data recorder; performing multi-period data deep semantic parsing on the real-time environmental monitoring data and performing intelligent index definition to obtain data indexes for each time period;

[0147] A data view module for predicting the evolution of environmental parameters for the filtered and noise-reduced monitoring data and performing parameter visualization rendering to construct a visual view of the environmental prediction trend;

[0148] A node storage framework module for acquiring historical storage data of the data recorder; performing adaptive lossless compression on the historical storage data and constructing a distributed storage architecture to obtain a distributed data node storage framework;

[0149] A storage embedding module for performing distributed storage embedding of the environmental prediction trend visual view on the distributed data node storage framework based on the data indexes for each time period to generate a real-time data embedded storage framework;

[0150] A redundancy optimization module for performing data redundancy analysis and redundant node deletion processing based on the real-time data embedded storage framework to obtain a redundant node optimization framework;

[0151] A dynamic framework adjustment module for calculating the storage load of each node for the redundant node optimization framework and performing dynamic index structure adjustment to construct a dynamic index adjustment storage framework for performing the data recorder information management operation.

[0152] Through the deep semantic analysis module, the system can conduct a more in-depth analysis of sensor data, not just staying at the data acquisition level, but being able to identify potential relationships and patterns among the data, mine more information. Through the intelligent index definition, the system can generate efficient indexes for the data in each time period, making future data retrieval and query more efficient. Users can quickly locate the required data through time or other parameters, improving the query efficiency. Through the visualization rendering of data, users can view the evolution trends of different environmental parameters (such as temperature, humidity, pollutants, etc.) in real time. Such graphical displays help users understand and analyze data more intuitively, supporting faster decision-making. This module can not only display the current environmental conditions, but also predict future environmental changes through trends, helping users identify potential change trends in advance and facilitating early response to environmental risks. Through adaptive lossless compression, the system can effectively reduce the storage occupancy of data while ensuring data integrity. This compression method does not lose data quality and is suitable for the storage of large-scale environmental monitoring data. Through the construction of a distributed storage architecture, the system can flexibly expand the storage capacity to adapt to the growing data storage requirements. In addition, the distributed architecture can improve the data access efficiency and reduce the risk of single-point failures. By combining the environmental prediction trend view with the data index and storage framework, the system can manage the storage and retrieval of real-time data more efficiently. The tight embedding of real-time data and historical data can reduce the latency during data query and improve the overall system response speed. This module provides a flexible framework for the storage of real-time data, intelligently selects appropriate storage locations according to the different access frequencies and storage requirements of the data, and further improves the storage efficiency. Through the analysis and optimization of redundant nodes, the system can reduce unnecessary redundant data storage, effectively reduce the storage cost and resource waste. Redundancy optimization not only reduces the burden on the storage system, but also improves the utilization rate of storage space, enabling more important data to be preferentially stored and accessed. By calculating the storage load of nodes, the system dynamically adjusts the allocation of storage resources according to actual needs. For nodes with higher loads, more resources are allocated to them, while for nodes with lower loads, appropriate resource release or degraded storage is carried out. Dynamic index structure adjustment ensures that the data storage system can still operate efficiently under changing requirements. By providing a more efficient indexing method for high-demand nodes, the system can improve the query and data access speed.

[0153] The present invention also provides an information management device for a data recorder based on artificial intelligence, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the information management method for the data recorder based on artificial intelligence described in any one of the above are implemented.

[0154] Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, systems and units described above refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0155] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it is stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, is embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that store program codes.

[0156] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes that fall within the meaning and scope of the equivalent elements of the application documents within the present invention.

[0157] As described above, these are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.

Claims

1. A method for managing data recorder information based on artificial intelligence, characterized in that It includes the following steps: Step S1: Obtain real-time environmental monitoring data based on the multi-source sensors of the data recorder; Perform in-depth semantic parsing of the real-time environmental monitoring data in multiple time periods, and perform intelligent index definition to obtain the data index for each time period; Step S2: Predict the evolution of environmental parameters for the filtered and noise-reduced monitoring data, and perform parameter visualization rendering to construct a visual view of the environmental prediction trend; Step S3: Obtain the historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture to obtain a distributed data node storage framework; Step S4: Embed the visual view of the environmental prediction trend into the distributed data node storage framework based on the data index for each time period to generate a real-time data embedded storage framework; Step S5: Perform data redundancy analysis and redundant node deletion processing on the real-time data embedded storage framework to obtain a redundant node optimization framework; Step S6: Calculate the storage load for each node of the redundant node optimization framework, and perform dynamic index structure adjustment to construct a dynamic index adjustment storage framework to execute the data recorder information management operation; Among them, the specific steps of Step S2 are: Step S21: Mine the change characteristics of environmental parameters for the filtered and noise-reduced monitoring data to generate environmental parameter change trend characteristics; Step S22: Calculate the change frequency of the environmental parameter change trend characteristics to extract the environmental parameter change frequency; Step S23: Identify the periodic evolution law based on the environmental parameter change frequency to generate the environmental change trend periodic law; Step S24: Predict the evolution of environmental parameters for the filtered and noise-reduced monitoring data based on the environmental change trend periodic law to obtain the environmental future trend prediction characteristics; Step S25: Perform parameter visualization rendering on the environmental future trend prediction characteristics to construct a visual view of the environmental prediction trend; Among them, the specific steps of Step S21 are: The filtered and noise-reduced monitoring data includes environmental temperature parameters, environmental humidity parameters, light intensity, air pollutant monitoring data, and wind force monitoring data; Fit the time-series change characteristics of the environmental temperature parameters and environmental humidity parameters to construct an environmental temperature change curve and an environmental humidity change curve; Perform light intensity spatial distribution analysis on the light intensity to obtain the light intensity spatial distribution characteristics; Calculate the pollutant peak value for the air pollutant monitoring data to extract the pollutant peak points; Identify the wind direction and wind force intensity for the wind force monitoring data to extract the wind force characteristics; Mine the environmental change trend of the light intensity spatial distribution characteristics, pollutant peak points, wind force characteristics, environmental temperature change curve, and environmental humidity change curve to obtain the environmental parameter change trend characteristics.

2. The method for managing data recorder information based on artificial intelligence according to claim 1, wherein The specific steps of Step S1 are: Step S11: Obtain real-time environmental monitoring data based on the multi-source sensors of the data recorder; Step S12: Perform parameter normal distribution analysis on the real-time environmental monitoring data to generate a normal distribution range; Step S13: Filter abnormal outlier parameters based on the normal distribution range to obtain outlier-filtered monitoring data; Step S14: Perform high-frequency filtering and noise reduction on the outlier-filtered monitoring data to obtain filtered and noise-reduced monitoring data; Step S15: Perform multi-period data deep semantic analysis on the filtered and noise-reduced monitoring data to generate data deep semantic features for each time period; Step S16: Define intelligent indexes based on the data deep semantic features for each time period to obtain data indexes for each time period.

3. The method for managing data recorder information based on artificial intelligence according to claim 1 is characterized in that, The specific steps of Step S3 are as follows: Step S31: Obtain the historical storage data of the data recorder; Step S32: Define a time span based on the historical storage data and perform uniform segmentation processing according to the time span to generate multiple historical storage data nodes; Step S33: Perform adaptive lossless compression on multiple historical storage data nodes to obtain multiple lossless compression data nodes; Step S34: Construct a distributed storage architecture for multiple lossless compression data nodes to obtain a distributed data node storage framework.

4. The information management method for a data recorder based on artificial intelligence according to claim 1, wherein The specific steps of Step S4 are as follows: Step S41: Quantify the data sensitivity of the environmental prediction trend visualization view to obtain the sensitivity of the instant parameter view; Step S42: Analyze the data access requirements of the sensitivity of the instant parameter view to generate real-time environmental monitoring data access requirement characteristics; Step S43: Make a storage decision based on the real-time environmental monitoring data access requirement characteristics to obtain the storage location of the real-time data; Step S44: Index and mark the environmental prediction trend visualization view according to the data indexes for each time period, and perform distributed storage embedding on the distributed data node storage framework based on the storage location of the real-time data to generate a real-time data embedded storage framework.

5. The method for managing data recorder information based on artificial intelligence according to claim 1, wherein The specific steps of Step S5 are as follows: Step S51: Identify the user access behavior of each data node among multiple lossless compression data nodes and extract the access behavior data of each node; Step S52: Calculate the data access frequency of the access behavior data of each node to obtain the data access frequency of each node; Step S53: Analyze the access time range of the access behavior data of each node to generate the access time range characteristics of each node; Step S54: Evaluate the historical data access requirements based on the data access frequency of each node and the access time range characteristics of each node to obtain the access requirement evaluation value of each node; Step S55: Compare the data redundancy of the access requirement evaluation value of each node based on a preset node access requirement threshold. When the access requirement evaluation value of the node is lower than the preset node access requirement threshold, perform redundant node deletion processing on the real-time data embedded storage framework to obtain a redundant node optimized framework.

6. The method for managing data recorder information based on artificial intelligence according to claim 1, wherein The specific steps of Step S6 are as follows: Step S61: Calculate the storage load of each node for the redundant node optimized framework to obtain the storage load characteristics of each node; Step S62: Predict the access peak of each node based on the storage load characteristics of each node to obtain the access peak prediction time point of the node; Step S63: Predict the future usage requirements of users based on the access peak prediction time point of the node to obtain the usage requirement prediction of each node; Step S64: Dynamically adjust the index structure of the redundant node optimization framework according to the usage requirement prediction of each node, so as to construct a dynamic index adjustment storage framework for performing the data recorder information management operation.

7. An information management system for a data recorder based on artificial intelligence, characterized in that, For performing the artificial intelligence-based data recorder information management method as described in claim 1, including: A deep semantic parsing module, configured to obtain real-time environmental monitoring data based on multi-source sensors of the data recorder; perform multi-period data deep semantic parsing on the real-time environmental monitoring data, and perform intelligent index definition to obtain data indexes for each time period; A data view module, configured to predict the evolution of environmental parameters for the filtered and noise-reduced monitoring data, and perform parameter visualization rendering to construct a visual view of the environmental prediction trend; A node storage framework module, configured to obtain the historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture to obtain a distributed data node storage framework; A storage embedding module, configured to perform distributed storage embedding of the visual view of the environmental prediction trend on the distributed data node storage framework based on the data indexes for each time period to generate a real-time data embedding storage framework; A redundancy optimization module, configured to perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain a redundant node optimization framework; A dynamic framework adjustment module, configured to calculate the storage load of each node for the redundant node optimization framework, and perform dynamic index structure adjustment, so as to construct a dynamic index adjustment storage framework for performing the data recorder information management operation.

8. An information management device for a data recorder based on artificial intelligence, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the artificial intelligence-based data recorder information management method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Rolling bearing fault diagnosis method based on WOA-VMD and GAT

    CN116662848A

  • Vehicle-mounted data voice tag system based on large language model

    CN117763194A