Data recorder information management method, system and device based on artificial intelligence
By introducing artificial intelligence technology into the data logger to analyze and predict environmental data in real time, the problem of low data acquisition and analysis efficiency in traditional methods is solved, and efficient and automated data management and storage are achieved.
Patent Information
- Application Number
- CN202510445732.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Traditional data recorder information management methods lack real-time data preprocessing and automated analysis capabilities, resulting in high cost and low efficiency of data acquisition and analysis, and the inability to effectively respond to complex environmental conditions and large-scale data management needs.
Using the data recorder information management method based on artificial intelligence, real-time environmental data is obtained through multi-source sensors, deep semantic analysis and intelligent index definition, real-time prediction and visualization of environmental parameters are realized, and data storage and retrieval are optimized by combining adaptive lossless compression and distributed storage architecture.
It improves the real-time and comprehensiveness of data collection, enhances the accuracy and automation of data analysis, optimizes data storage and retrieval efficiency, reduces storage costs and resource waste, and improves the stability and responsiveness of the system.
Smart Images

Figure CN119938673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to an artificial intelligence-based data recorder information management method, system and device. Background Art
[0002] With the continuous development and application of intelligent technology, the role of data loggers in various fields is becoming increasingly important. Data loggers are widely used in many fields such as industry, environmental monitoring, smart home, energy management, etc., and undertake the task of real-time collection, storage and analysis of various key environmental and equipment data. Its main function is to continuously monitor the working environment or equipment operation status and record data to ensure that the system can reflect various key parameters in real time, thereby providing an accurate basis for subsequent analysis, decision-making and optimization.
[0003] However, in the traditional data recorder information management method, data collection, storage and analysis usually rely on static preset rules and manual operations. Although this method meets basic needs in some simple application scenarios, with the sharp increase in data volume and the increasing complexity of application scenarios, the traditional method has gradually exposed a series of problems. First, traditional data recorders often do not have real-time data preprocessing and automated analysis capabilities, resulting in manual intervention and manual analysis of data after recording, which increases labor costs and operational difficulties. Secondly, due to the complexity of environmental conditions, traditional methods have limited capabilities in data redundancy processing, fault detection and real-time adjustment, and cannot effectively cope with the growing amount of monitoring data and diversified data management needs. Traditional data storage management methods usually lack intelligent data compression, storage optimization and dynamic adjustment mechanisms, which easily lead to waste of data storage resources and low access efficiency. Therefore, a more intelligent data recorder information management method is needed. Summary of the invention
[0004] In order to solve the above technical problems, the present invention proposes an artificial intelligence-based data recorder information management method, system and device to solve at least one of the above technical problems.
[0005] To achieve the above object, the present invention provides a data recorder information management method based on artificial intelligence, comprising the following steps: Step S1: Acquire real-time environmental monitoring data based on multi-source sensors of a data recorder; perform multi-period data deep semantic analysis on the real-time environmental monitoring data, and perform intelligent index definition to obtain a data index for each time period; Step S2: Predict the evolution of environmental parameters for the filtered noise reduction monitoring data, perform parameter visualization rendering, and construct a visualization view of the environmental prediction trend; Step S3: Acquire historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture, thereby obtaining a distributed data node storage framework; Step S4: embedding the environment prediction trend visualization view into the distributed data node storage framework based on the data index of each time period, and generating a real-time data embedding storage framework; Step S5: Perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain a redundant node optimization framework; Step S6: Calculate the storage load of each node in the redundant node optimization framework and adjust the dynamic index structure, thereby constructing a dynamic index adjustment storage framework to perform data recorder information management operations.
[0006] The present invention monitors environmental parameters in real time through multi-source sensors equipped with data recorders, which enables the system to respond to environmental changes in a timely manner and improve the real-time and comprehensiveness of data collection. Through deep semantic analysis of multi-period data, deep information in environmental data can be better extracted, not only to obtain raw data, but also to understand the temporal and spatial correlation of data, future analysis and decision-making, and filtering and noise reduction. The system can remove errors and interference in environmental monitoring data to ensure the accuracy and reliability of data, which makes subsequent analysis and prediction more accurate. Based on historical data and real-time data, the trend of environmental parameters is predicted through the evolution prediction model, potential environmental changes are identified in advance, and early warning and prediction decisions are helped. Adaptive lossless compression technology can reduce storage space occupancy and improve storage efficiency without losing any data quality. The design of distributed storage architecture ensures high availability and fault tolerance of data. Whether it is data storage, reading or backup, it can ensure the stability and efficiency of the system and avoid data loss or system downtime caused by single point failure. By embedding the visual view of the predicted trend into the distributed storage framework, it can ensure that the data is not only optimized in storage, but also can ensure fast access and processing of real-time data. This embedding method improves the real-time and comprehensiveness of the system. Dynamic response capability. The data embedded storage framework optimizes the data storage method, ensuring that environmental data and prediction data can be efficiently integrated and stored in distributed storage, further improving the overall performance of the system. Through redundancy analysis, it can identify redundant data nodes in the storage system and delete redundant nodes. This not only reduces unnecessary storage consumption, but also optimizes the storage architecture, improves the efficiency of data storage and the utilization of system resources. By removing redundant nodes, the storage architecture becomes more streamlined, reducing the burden on the system, improving the speed of data reading and writing, and improving the overall performance of the system. By calculating the storage load of each node, the system can dynamically adjust according to the load of the node to ensure that the storage and computing pressure of each node is within a reasonable range. This load balancing mechanism can effectively improve the stability and processing efficiency of the system. According to load changes and data access requirements, the index structure is dynamically adjusted, which helps to improve the efficiency of data retrieval and the adaptability of the system. Especially when the data volume is large and changes, dynamic adjustment can enable the system to better cope with changes. Through the dynamic adjustment of the storage architecture and index structure, data recorder information management can be better performed to ensure efficient access and processing of data, and ultimately support the system to accurately manage and analyze environmental monitoring data.
[0007] Preferably, step S1 comprises the following steps: Step S11: Acquire real-time environmental monitoring data based on multi-source sensors of the data recorder; Step S12: performing parameter normalization distribution analysis on the real-time environmental monitoring data to generate a normalization distribution range; Step S13: filtering abnormal outlier parameters based on the normalized distribution range, thereby obtaining outlier filtering monitoring data; Step S14: performing high-frequency filtering and noise reduction on the outlier filtered monitoring data to obtain filtered noise reduction monitoring data; Step S15: performing multi-time period data deep semantic analysis on the filtered noise reduction monitoring data to generate deep semantic features of the data for each time period; Step S16: Intelligent index definition is performed based on the deep semantic features of the data in each time period to obtain the data index for each time period.
[0008] The present invention collects environmental data in real time through multi-source sensors (such as temperature and humidity sensors, air pressure sensors, light sensors, etc.), which can comprehensively and meticulously reflect the environmental status and meet multi-dimensional monitoring needs. The use of data recorders ensures the acquisition of real-time data, can timely reflect environmental changes, support rapid response and timely decision-making, and converts data into a standardized format through normalized distribution analysis, so that data collected by different sensors are uniformly processed. In subsequent data analysis and comparison, normalized distribution analysis can clarify the normal range of various environmental parameters, identify the expected distribution and fluctuation range of data, and lay the foundation for anomaly detection. By comparing with the normalized distribution range, abnormal data (such as sensor failure, environmental mutation, etc.) can be effectively detected. This step can effectively remove outlier data that deviates from the normal range and ensure the accuracy of the data. After filtering out the outlier data, the remaining data is more representative and credible, avoiding erroneous analysis or misjudgment caused by abnormal data. By high-frequency filtering and noise reduction, noise components (such as equipment noise, electromagnetic interference, etc.) in real-time monitoring data can be removed to obtain smoother and more accurate data. The data after noise is closer to the actual environmental state, which helps to perform trend analysis and prediction more accurately. The filtered data provides high-quality input for further analysis and modeling, and improves the system's sensitivity and responsiveness to environmental changes. Through deep semantic analysis, not only can the data characteristics of each time period be obtained, but also the meaning and pattern behind these data can be deeply understood. This deep analysis can reveal the inherent laws of environmental parameters changing over time. The generated time period data characteristics help the system identify the changing trends of environmental parameters, especially in long-term data, and can effectively extract valuable information such as seasonal changes and cyclical fluctuations. Through intelligent indexing based on deep semantic features, indexes can be automatically generated for data in each time period, making data retrieval and management more efficient. Data indexing enables the system to quickly locate and retrieve the required historical data or real-time data, especially in a large data volume environment, which greatly improves the efficiency of data processing and query. The intelligent index structure can automatically optimize the storage strategy according to the characteristics of the data, thereby improving data access speed and storage space utilization.
[0009] Preferably, step S2 comprises the following steps: Step S21: mining the environmental parameter change characteristics of the filtered noise reduction monitoring data to generate environmental parameter change trend characteristics; Step S22: Calculate the change frequency of the environmental parameter change trend characteristics to extract the environmental parameter change frequency; Step S23: identifying periodic evolution rules based on the frequency of environmental parameter changes to generate periodic rules of environmental change trends; Step S24: predicting the evolution of environmental parameters of the filtered noise reduction monitoring data based on the cycle law of environmental change trends, so as to obtain the prediction characteristics of future environmental trends; Step S25: Parameter visualization rendering is performed on the future trend prediction features of the environment to construct a visualization view of the environment prediction trend.
[0010] The present invention reveals the potential laws of environmental parameter changes by mining the change characteristics of the filtered and denoised data. For example, parameters such as temperature, humidity, and air quality will change with seasons, time or other factors. Through feature mining, the system can identify the patterns of these changes. The mined change characteristics provide a deeper understanding for data analysis, help reveal the environmental change trends hidden behind the data, and provide support for subsequent analysis and decision-making. Calculating the frequency of environmental parameter changes can help the system identify the periodic characteristics of environmental changes, such as certain parameters that have specific fluctuation patterns every day, week or month. By extracting the frequency, the system can accurately describe the dynamic characteristics of environmental changes. The extracted change frequency information can help identify the periodic laws of environmental changes and provide a basis for the prediction of periodic changes, which is of great significance for predicting weather changes, seasonal climate changes, etc. By identifying the periodic evolution law based on frequency, the system can identify the periodic laws of environmental parameter changes, such as seasonal fluctuations in temperature and humidity, or diurnal changes in air quality. Accurately grasp the long-term trend of environmental changes. The identification of periodic laws provides a scientific basis for long-term trend prediction. The system can predict environmental change trends within a certain period of time in the future, helping users to foresee future changes and make corresponding management and adjustments. Through the evolution prediction model of periodic laws and environmental parameters, the system can predict future environmental change trends. This is not just a simple extension of current data, but an intelligent prediction based on the law of environmental change. Accurate environmental trend prediction can provide early warning for decision makers, making environmental management more forward-looking. For example, factors such as meteorological changes and climate fluctuations can provide early warnings to help relevant departments take preventive measures to reduce the potential impact of environmental changes. Visualizing the future trend prediction features of the environment can help users understand the future trend of environmental parameters in an intuitive way. Through visualization methods such as charts, curves, and heat maps, users can quickly understand data and prediction results. Visual views provide decision makers with a clear basis for making decisions based on prediction results. For example, based on the visualization of temperature and humidity change trends, users can easily judge future environmental conditions and take corresponding measures.
[0011] Preferably, the specific steps of step S21 are: The filtering and noise reduction monitoring data includes environmental temperature parameters, environmental humidity parameters, light intensity, air pollutant monitoring data and wind monitoring data; Fit the time series variation characteristics of ambient temperature parameters and ambient humidity parameters to construct ambient temperature variation curves and ambient humidity variation curves; Performing light intensity spatial distribution analysis on light intensity to obtain light intensity spatial distribution characteristics; Calculate the pollutant peak value of air pollutant monitoring data and extract the pollutant peak point; Identify wind direction and wind intensity from wind monitoring data to extract wind characteristics; The environmental change trend is mined by the spatial distribution characteristics of light intensity, peak points of pollutants, wind characteristics, ambient temperature change curve and ambient humidity change curve to obtain the trend characteristics of environmental parameter changes.
[0012] The present invention reduces equipment noise, interference signals and errors by filtering and denoising the data, ensures the accuracy and reliability of the data, and lays a good foundation for subsequent analysis. Integrate multiple sensor data to form a comprehensive and three-dimensional environmental data view, providing more comprehensive and accurate information support for subsequent trend analysis, prediction, etc. By fitting the time series change characteristics of temperature and humidity data, the change curves of these two environmental parameters can be accurately depicted. Changes in temperature and humidity are key factors in environmental monitoring, and the fitted curves can intuitively show their trends over time. Ambient temperature and humidity usually have significant seasonal or periodic changes, and time series fitting can help identify these laws and provide valuable references for subsequent predictions and analysis. Spatial distribution analysis of light intensity helps to reveal the distribution of light in different places and at different times, which is crucial for analyzing the light conditions of the environment and evaluating the impact of light on parameters such as temperature and humidity. In the fields of agriculture, construction, energy, etc., the spatial distribution characteristics of light intensity are used to optimize resource allocation, for example, to help the agricultural sector evaluate the light conditions for crop growth, or to help building managers design energy-saving solutions. By calculating the peak value of pollutant data, we can identify extreme cases of pollutant concentration, such as the peak value of pollutants or the period of high pollution, which is an important indicator for judging the severity of environmental pollution. The extraction of peak points helps to identify pollution events in a timely manner and supports the rapid response of environmental monitoring systems. For example, when the air quality index (AQI) reaches a certain threshold, the early warning system is triggered to remind relevant departments to take countermeasures. Wind direction and wind strength are important factors in meteorological data. Through the analysis of wind monitoring data, key features such as wind direction and strength are extracted. Comprehensively understand the changing laws and spatial distribution of wind power. Wind feature extraction has important applications in wind power generation, meteorological forecasting and other fields. Changes in wind speed and wind direction directly affect energy production and weather forecast accuracy. Therefore, accurate identification of wind features can provide strong support for related fields. By comprehensively analyzing multi-dimensional environmental parameters such as light, pollutant peak, wind, temperature and humidity, the trend of environmental changes can be explored from multiple angles. This comprehensive analysis reveals the interaction and influence relationship between different environmental factors. This multi-factor trend mining helps to identify complex patterns of environmental changes. For example, the correlation between temperature change and air pollution, the interactive effect between wind speed change and humidity, etc. Through these analyses, the system can more comprehensively understand the changing laws of the environment. By extracting the characteristics of environmental change trends, it can provide more accurate input for the prediction model of environmental change, helping to predict sudden changes or long-term trends in the environment (such as climate change, air quality changes, etc.) in advance.
[0013] Preferably, step S3 specifically comprises the following steps: Step S31: Acquire historical storage data of the data recorder; Step S32: defining a time span based on the historical storage data, and performing uniform segmentation processing according to the time span, thereby generating a plurality of historical storage data nodes; Step S33: performing adaptive lossless compression on multiple historical storage data nodes to obtain multiple lossless compressed data nodes; Step S34: construct a distributed storage architecture for multiple lossless compression data nodes, thereby obtaining a distributed data node storage framework.
[0014] The present invention helps to split historical data into different time periods by defining an appropriate time span, which is convenient for data analysis and storage management. The delineation of this time span can be adjusted according to actual needs, such as dividing by day, week, month or year, so that data management is more flexible. Evenly dividing the data by time span can facilitate independent processing and analysis of data in different time periods, which helps to avoid processing difficulties caused by excessive data. Each data node is independently stored and accessed, which improves the manageability of data. Lossless compression can reduce the data volume without losing data, thereby greatly saving storage space. For the case of a large amount of historical data, the use of an adaptive compression method can better reduce storage costs. Unlike lossy compression, lossless compression ensures the integrity and accuracy of the data. Even after compression, The data can still be fully restored, which is suitable for application scenarios that require high data precision and accuracy, such as environmental monitoring data. Through distributed storage, data is stored in multiple physical locations, which can improve the fault tolerance and reliability of the system. Even if a storage node fails, other nodes can still ensure normal access to the data, avoiding the risk of single point failure. The distributed storage architecture can make data storage more flexible and efficient. After the data is distributed on different nodes, it can be accessed in parallel, which improves the data processing speed and access efficiency. It is suitable for large-scale data storage and fast access needs. The distributed storage architecture has high scalability. According to the increase in storage needs, new storage nodes can be expanded at any time to meet the increasing demand for data volume. As the amount of data increases, storage resources can be seamlessly expanded to avoid storage bottlenecks.
[0015] Preferably, the specific steps of step S4 are: Step S41: quantifying the data sensitivity of the environment prediction trend visualization view to obtain the sensitivity of the real-time parameter view; Step S42: performing data access demand analysis on the instant parameter view sensitivity to generate real-time environment monitoring data access demand characteristics; Step S43: making a storage decision based on the access demand characteristics of the real-time environment monitoring data to obtain the storage location of the real-time data; Step S44: index and mark the environment prediction trend visualization view according to the data index of each time period, and embed the distributed data node storage framework into a distributed storage based on the storage location of the real-time data to generate a real-time data embedding storage framework.
[0016] The present invention quantifies data sensitivity through the visual view of the environmental forecast trend, which means that the system can evaluate the importance of each environmental monitoring parameter (such as temperature, humidity, air quality, etc.) to the overall forecast result. Identify which parameters have the greatest impact on environmental trends in the forecast model and which parameters contribute less to the decision. By quantifying data sensitivity, priority is given to those environmental parameters that have a greater impact on the forecast results, which helps to determine the allocation and optimization of storage resources. For example, parameters with higher sensitivity are updated and accessed more frequently in the system, while parameters with lower sensitivity are updated periodically. By generating data access demand characteristics, the system can dynamically adjust storage and processing strategies. For example, in time periods or conditions with higher demand, the system dispatches more computing resources to process hot data to ensure the efficiency of real-time monitoring and early warning. By deeply understanding the real-time data access needs, the system optimizes the data retrieval and processing process, reduces the interference of irrelevant data, accelerates the response speed of key data, and ensures that the system can respond quickly to environmental changes. Highly accessed data is stored in faster and more responsive storage locations (such as cache or SSD), while low-frequency accessed data is stored on more economical and larger storage media (such as HDD). This ensures performance while reducing storage costs. This storage decision based on real-time access needs can flexibly adapt to different application scenarios. For example, in environmental forecasting, data needs to be frequently accessed or updated during certain periods, while data access needs are significantly reduced during other periods. The system automatically adjusts storage strategies to optimize resource usage. By creating indexes and marking data for each time period, the system can quickly locate relevant data nodes and improve the speed and accuracy of data queries. The marking of time indexes enables data to be efficiently retrieved in chronological order, providing strong support for time series analysis. By combining data storage decisions with the storage location of real-time data, real-time data is embedded in the distributed storage framework, further improving the flexibility and scalability of the storage architecture. Distributed storage ensures high availability and fault tolerance of data, and can cope with future data growth through horizontal expansion.
[0017] Preferably, the specific steps of step S5 are: Step S51: performing user access behavior identification on multiple lossless compressed data nodes one by one, and extracting access behavior data of each node; Step S52: Calculate the data access frequency of each node's access behavior data to obtain the data access frequency of each node; Step S53: Analyze the access time range of the access behavior data of each node to generate the access time range characteristics of each node; Step S54: Perform historical data access demand evaluation based on the data access frequency of each node and the access time range characteristics of each node, thereby obtaining an access demand evaluation value of each node; Step S55: Perform data redundancy comparison on the access demand evaluation value of each node based on the preset node access demand threshold. When the access demand evaluation value of the node is lower than the preset node access demand threshold, perform redundant node deletion processing on the real-time data embedded storage framework to obtain a redundant node optimization framework.
[0018] The present invention collects the access behavior data of each node, and the system adopts different storage and processing strategies for different types of data nodes. For example, frequently accessed nodes are preferentially stored in efficient storage media, while rarely accessed nodes are stored in low-cost storage media, thereby optimizing the use of storage resources. By analyzing the access behavior of each node, the system can identify redundant and inefficient data nodes to avoid waste of resources, which lays the foundation for subsequent redundant data optimization. By calculating the access frequency, the system prioritizes the data according to the actual access frequency. For example, for frequently accessed nodes, they are stored on efficient storage media to improve the data reading speed; while for infrequently accessed nodes, they are stored on low-cost storage media to achieve a balance between storage cost and performance. According to the access frequency, the system sets different access strategies and scheduling methods for different categories of data nodes, thereby improving the overall storage and access efficiency. By analyzing the access time range of each data node, the access pattern of each node in time can be further understood. For example, some nodes are frequently accessed on weekdays, but almost never accessed on weekends; some nodes are continuously accessed 24 hours a day, while some nodes are only accessed during specific periods of time. Based on the characteristics of the access time range, the system can intelligently schedule data access. For example, during certain periods of high access frequency, the system gives priority to providing data access rights to these nodes; during low access periods, energy-saving or downgraded storage strategies are adopted to reduce unnecessary data access burdens. By evaluating the access needs of each node, the system can more accurately determine which data needs to be retained in high-efficiency storage media for a long time, which data needs to be transferred to low-cost storage media, and which data needs to be deleted or archived, thereby reducing data redundancy and improving storage efficiency. Based on the historical data access demand evaluation, the system dynamically adjusts the storage strategy to ensure that data storage and access are more intelligent, which is particularly important for the management of large-scale environmental monitoring data, and can effectively reduce storage pressure and improve access efficiency. After deleting redundant nodes, the system can centrally store and process important data, reduce access to low-demand data, and thus improve data storage and access efficiency, which is crucial for efficient storage architecture, especially in big data environments, to avoid redundant data occupying valuable storage resources. Deleting redundant data not only improves storage efficiency, but also significantly reduces storage costs. By cleaning up data with low access frequency, the system can reduce its dependence on expensive storage resources and reduce overall storage costs. The process of redundancy comparison based on access demand thresholds makes data management more intelligent. The system can automatically adjust storage strategies according to actual usage without manual intervention. This not only improves operational efficiency, but also can adapt to different storage needs and ensure the flexibility and scalability of the system.
[0019] Preferably, the specific steps of step S6 are: Step S61: Calculate the storage load of each node on the redundant node optimization framework to obtain the storage load characteristics of each node; Step S62: performing node access peak prediction on the storage load characteristics of each node to obtain the node access peak prediction time point; Step S63: predicting the future usage demand of users according to the predicted access peak time point of the node, and obtaining the usage demand prediction of each node; Step S64: dynamically adjust the index structure of the redundant node optimization framework according to the usage demand prediction of each node, thereby constructing a dynamic index adjustment storage framework to perform data recorder information management operations.
[0020] The present invention calculates the storage load of each node in the redundant node optimization framework, so that the system can understand the storage pressure and load of each node in real time. The load calculation helps the system identify which nodes have a heavy storage burden and which nodes have idle storage, thereby optimizing the allocation of storage resources. By calculating the storage load characteristics, the system dynamically adjusts the storage resources according to the load conditions. For example, for nodes with higher loads, consider allocating more storage space, or adjusting the storage strategy of nodes with lower access frequency to reduce system bottlenecks. By analyzing the storage load characteristics, the system can predict the access peak time point of each node, which means that the system can predict large-scale data access events that will occur in the future (such as peak access, emergency data reading, etc.) and prepare in advance. The node access peak prediction helps to formulate resource scheduling strategies in advance. For example, when the system predicts that the access peak is about to arrive, it increases the storage and computing resources of the node in advance to ensure stability and rapid response during the peak period. By combining the node access peak prediction, The system can accurately predict future user usage needs. For example, at specific time points or conditions, some nodes will be accessed in large quantities (such as holidays or large-scale environmental changes), while some nodes will maintain a lower access frequency. The system will intelligently adjust storage resources based on future usage demand predictions. For upcoming high access demands, the system will allocate more storage and computing resources to related nodes; for nodes with lower demands, resource allocation will be reduced to maximize resource utilization. Based on the node usage demand predictions, the system can adjust the index structure in real time to adapt to different data access needs. For example, for frequently accessed nodes, a more efficient index structure, such as time series index or partition index, is used; for nodes with low access frequency, a simpler index structure is used to reduce storage overhead. By dynamically adjusting the index structure, the system reduces the search time during data access and improves the overall data processing speed. For example, frequently accessed nodes will be given a more efficient retrieval path, thereby improving the response speed and real-time performance of data processing.
[0021] In this specification, a data recorder information management system based on artificial intelligence is provided, which is used to execute the data recorder information management method based on artificial intelligence as described above, including: A deep semantic analysis module is used to obtain real-time environmental monitoring data based on multi-source sensors of the data recorder; perform multi-period data deep semantic analysis on the real-time environmental monitoring data, and perform intelligent index definition to obtain the data index for each time period; The data view module is used to predict the evolution of environmental parameters of the filtered noise reduction monitoring data, perform parameter visualization, and build a visualization view of the environmental prediction trend; The node storage framework module is used to obtain the historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture to obtain a distributed data node storage framework; The storage embedding module is used to embed the environmental prediction trend visualization view into the distributed data node storage framework based on the data index of each time period, and generate a real-time data embedding storage framework; A redundancy optimization module is used to perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain a redundant node optimization framework; The dynamic framework adjustment module is used to calculate the storage load of each node of the redundant node optimization framework and adjust the dynamic index structure, so as to build a dynamic index adjustment storage framework to perform data recorder information management operations.
[0022] Through the deep semantic analysis module of the present invention, the system can conduct a deeper analysis of the sensor data, not just staying at the data collection level, but can identify the potential relationships and rules between the data and mine more information. Through the intelligent index definition, the system can generate an efficient index for the data of each time period, making future retrieval and query of data more efficient. Users can quickly locate the required data through time or other parameters to improve query efficiency. Through the visual rendering of data, users can view the evolution trends of different environmental parameters (such as temperature, humidity, pollutants, etc.) in real time. Such graphical display helps users understand and analyze data more intuitively and supports faster decision-making. The module not only displays the current environmental conditions, but also predicts future environmental changes through trends, helping users to identify potential change trends in advance and facilitate early response to environmental risks. Through adaptive lossless compression, the system can effectively reduce the storage occupancy of data while ensuring the integrity of the data. This compression method will not lose data quality and is suitable for the storage of large-scale environmental monitoring data. Through the construction of a distributed storage architecture, the system can flexibly expand the storage capacity to adapt to the growing data storage needs. In addition, the distributed architecture can improve Improve data access efficiency and reduce the risk of single point failure. By combining the environmental prediction trend view with the data index and storage framework, the system can manage the storage and retrieval of real-time data more efficiently. The close embedding of real-time data and historical data can reduce the delay in data query and improve the response speed of the overall system. This module provides a flexible framework for the storage of real-time data. According to the different access frequencies and storage requirements of the data, it intelligently selects the appropriate storage location to further improve the storage efficiency. Through the analysis and optimization of redundant nodes, the system can reduce unnecessary redundant data storage and effectively reduce storage costs and resource waste. Redundancy optimization not only reduces the burden on the storage system, but also improves the utilization of storage space, so that more important data can be stored and accessed first. By calculating the node storage load, the system dynamically adjusts the allocation of storage resources according to actual needs. For nodes with higher loads, more resources are allocated to them, and for nodes with lower loads, appropriate resource release or storage downgrade is performed. Dynamic index structure adjustment ensures that the data storage system can still operate efficiently under changing needs. By providing a more efficient indexing method for high-demand nodes, the system can improve the speed of query and data access.
[0023] The present invention also provides an artificial intelligence-based data recorder information management device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the artificial intelligence-based data recorder information management methods described above are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1A schematic diagram of the steps of a data recorder information management method based on artificial intelligence of the present invention; Figure 2 Detailed implementation flow chart of step S1; Figure 3 Detailed implementation flow chart of step S2; Figure 4 Detailed implementation flow chart of step S3. DETAILED DESCRIPTION
[0025] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0026] The present application example provides a data recorder information management method, system and device based on artificial intelligence. The execution subjects of the data recorder information management method, system and device based on artificial intelligence include but are not limited to: mechanical equipment, data processing platform, cloud server node, network upload device, etc. equipped with the system can be regarded as the general computing node of the present application, and the data processing platform includes but is not limited to: at least one of an audio image management system, an information management system, and a cloud data management system.
[0027] See also Figures 1 to 4 The present invention provides a data recorder information management method based on artificial intelligence, and the data recorder information management method based on artificial intelligence comprises the following steps: Step S1: Acquire real-time environmental monitoring data based on multi-source sensors of a data recorder; perform multi-time period data deep semantic analysis on the real-time environmental monitoring data, and perform intelligent index definition to obtain a data index for each time period; Step S2: Predict the evolution of environmental parameters for the filtered noise reduction monitoring data, perform parameter visualization rendering, and construct a visualization view of the environmental prediction trend; Step S3: Acquire historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture, thereby obtaining a distributed data node storage framework; Step S4: embedding the environment prediction trend visualization view into the distributed data node storage framework based on the data index of each time period, and generating a real-time data embedding storage framework; Step S5: Perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain a redundant node optimization framework; Step S6: Calculate the storage load of each node in the redundant node optimization framework and adjust the dynamic index structure, thereby constructing a dynamic index adjustment storage framework to perform data recorder information management operations.
[0028] The present invention monitors environmental parameters in real time through multi-source sensors equipped with data recorders, which enables the system to respond to environmental changes in a timely manner and improve the real-time and comprehensiveness of data collection. Through deep semantic analysis of multi-period data, deep information in environmental data can be better extracted, not only to obtain raw data, but also to understand the temporal and spatial correlation of data, future analysis and decision-making, and filtering and noise reduction. The system can remove errors and interference in environmental monitoring data to ensure the accuracy and reliability of data, which makes subsequent analysis and prediction more accurate. Based on historical data and real-time data, the trend of environmental parameters is predicted through the evolution prediction model, potential environmental changes are identified in advance, and early warning and prediction decisions are helped. Adaptive lossless compression technology can reduce storage space occupancy and improve storage efficiency without losing any data quality. The design of distributed storage architecture ensures high availability and fault tolerance of data. Whether it is data storage, reading or backup, it can ensure the stability and efficiency of the system and avoid data loss or system downtime caused by single point failure. By embedding the visual view of the predicted trend into the distributed storage framework, it can ensure that the data is not only optimized in storage, but also can ensure fast access and processing of real-time data. This embedding method improves the real-time and comprehensiveness of the system. Dynamic response capability. The data embedded storage framework optimizes the data storage method, ensuring that environmental data and prediction data can be efficiently integrated and stored in distributed storage, further improving the overall performance of the system. Through redundancy analysis, it can identify redundant data nodes in the storage system and delete redundant nodes. This not only reduces unnecessary storage consumption, but also optimizes the storage architecture, improves the efficiency of data storage and the utilization of system resources. By removing redundant nodes, the storage architecture becomes more streamlined, reduces the burden on the system, improves the speed of data reading and writing, and improves the overall performance of the system. By calculating the storage load of each node, the system can dynamically adjust according to the load of the node to ensure that the storage and computing pressure of each node is within a reasonable range. This load balancing mechanism can effectively improve the stability and processing efficiency of the system. According to load changes and data access requirements, the index structure is dynamically adjusted, which helps to improve the efficiency of data retrieval and the adaptability of the system. Especially when the data volume is large and changes, dynamic adjustment can enable the system to better cope with changes. Through the dynamic adjustment of the storage architecture and index structure, data recorder information management can be better performed to ensure efficient access and processing of data, and ultimately support the system to accurately manage and analyze environmental monitoring data.
[0029] In the embodiment of the present invention, refer to Figure 1 , is a schematic flow chart of the steps of a data recorder information management method based on artificial intelligence of the present invention. In this example, the steps of the data recorder information management method based on artificial intelligence include: Step S1: Acquire real-time environmental monitoring data based on multi-source sensors of a data recorder; perform multi-time period data deep semantic analysis on the real-time environmental monitoring data, and perform intelligent index definition to obtain a data index for each time period; In this embodiment, the environmental parameters to be monitored are determined, such as temperature, humidity, air quality, light intensity, and sound level. High-precision sensors are selected to ensure the accuracy and reliability of the data. A variety of sensors (such as DHT22, MQ135, photoresistors, etc.) are used for data collection, and the compatibility between sensors is ensured to facilitate subsequent data integration. The data logger is configured to connect all sensors to ensure that it can stably collect and store data. A microcontroller (such as Arduino or RaspberryPi) is used to control the sensor. Code is written to periodically read sensor data and set the data collection frequency, such as recording once every minute. The data logger is started to start acquiring real-time data from the sensor. The data should include the sensor ID, timestamp, and its corresponding environmental parameter value. The collected data is stored in a local database (such as SQLite, MySQL) or a cloud database to ensure data accessibility and security. A regular backup mechanism is set to prevent data loss. Before performing deep semantic analysis, the collected data is cleaned to remove missing values and outliers. Python is used to perform deep semantic analysis. n's Pandas library for data preprocessing, conduct in-depth analysis of the data in each time period, extract meaningful features such as the average, maximum value, standard deviation, etc., use sliding window technology to analyze the data, set time windows (such as 1 hour, 1 day) for aggregation, define intelligent indexing strategies based on the parsed data to improve data retrieval efficiency, generate unique identifiers as indexes for the data in each time period, and the indexes should include timestamps, sensor IDs, and monitored parameters to facilitate quick search of monitoring data for a specific period of time. Use the indexing function of the database (such as MySQL's INDEX) to create indexes for the stored data to speed up query operations and ensure that the index covers all key attributes. Store the results of parsing and indexing in the database, including aggregated data for each time period and its index information, to ensure data integrity and consistency. Record access statistics for each time period for subsequent analysis. Use data visualization tools (such as Matplotlib or Tableau) to display changes in real-time monitoring data and generate dashboards to facilitate users to view environmental monitoring conditions in real time.
[0030] Step S2: Predict the evolution of environmental parameters for the filtered noise reduction monitoring data, perform parameter visualization rendering, and construct a visualization view of the environmental prediction trend; In this embodiment, environmental monitoring data, including parameters such as temperature, humidity, and air quality, are extracted from the previous processing stage to ensure the integrity and accuracy of the data. These data should contain timestamps to facilitate subsequent analysis. A suitable filtering algorithm (such as Kalman filtering or moving average) is used to reduce noise on the data. Appropriate parameters and window sizes are selected to balance data smoothing and feature retention. For example, moving average effectively removes short-term fluctuations while maintaining long-term trends. After filtering, the clarity and accuracy of the data are verified. The original data and the filtered data are compared by visualization to ensure that the noise is effectively reduced. According to the characteristics of the data, a suitable time series prediction model is selected, such as SARIMA (Seasonal Autoregressive Integrated Moving Average Model) or a machine learning model (such as LSTM). These models can capture trends and seasonal changes in the data. The data is divided into a training set and a training set. Test set, usually 70% of the data is used for training and 30% for testing. Use the training set to train the model, and use the test set to evaluate the model's prediction performance to ensure its accuracy. Use the trained model to predict future environmental parameters, such as temperature and humidity changes in the next 7 days. Record the prediction results and their confidence intervals for subsequent analysis. Select appropriate visualization tools (such as Matplotlib or Seaborn) for data visualization. Select static charts or interactive charts as needed. Display the original data, filtered data, and predicted data in the same chart for intuitive comparison by users. Use different colors and marks to distinguish different data types to ensure clear information communication. Verify the visualization effect, ensure that the chart is clear and easy to understand, and can effectively convey information about changes in environmental parameters. Make adjustments and optimizations based on user feedback, such as changing the color scheme or adding annotations.
[0031] Step S3: Acquire historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture, thereby obtaining a distributed data node storage framework; In this embodiment, historical monitoring data is extracted from the storage system of the data recorder. These data are stored in a local database or cloud storage to ensure that the acquired data is complete and lossless. An appropriate database query language (such as SQL) or API interface is used to extract all historical data within the required time range. For example, all monitoring records in the past year are extracted for subsequent processing. After the data is extracted, its format is ensured to be suitable for subsequent processing. Usually, the data is converted into a unified structured format, such as CSV or JSON, to facilitate subsequent compression and storage. A data type check is performed to ensure the type consistency of the data to avoid errors in subsequent processing. An adaptive lossless compression algorithm (such as LZ77, LZ78 or Deflate) is used. These algorithms dynamically adjust compression parameters according to the characteristics of the data to optimize the compression effect. When selecting an algorithm, the characteristics of the data (such as repeatability and data type) are considered to select the most suitable compression method to ensure compression efficiency. The extracted historical data is compressed. During the compression process, the compression ratio and compression time are monitored to ensure that the compression efficiency and processing time are balanced. Find a balance point between compression and storage, record the data size before and after compression to evaluate the compression effect. Ideally, the compression ratio should be between 2:1 and 5:1, depending on the degree of data redundancy. Design a distributed storage architecture based on data usage requirements and access patterns, select a suitable distributed file system (such as HDFS, Ceph, or GlusterFS) as the storage backend to support large-scale data storage and access, determine data sharding rules and replication strategies to evenly distribute storage loads among multiple nodes to ensure high availability and data redundancy, shard and store the compressed data on different nodes, with each node storing a certain amount of data shards to avoid single point failures, configure load balancing strategies to ensure that read and write requests from each node are reasonably distributed, implement a monitoring mechanism to monitor the status and performance of each storage node in real time, use monitoring tools (such as Prometheus or Grafana) to collect and display storage performance indicators, and continuously optimize the storage architecture based on monitoring data, adjust data distribution and replication strategies to cope with changes in access patterns and growth in storage requirements.
[0032] Step S4: embedding the environment prediction trend visualization view into the distributed data node storage framework based on the data index of each time period, and generating a real-time data embedding storage framework; In this embodiment, based on the environmental prediction results of each time period in the previous processing, a corresponding data index is generated. The index should include timestamp, sensor ID and environmental parameters (such as temperature, humidity, etc.) to facilitate subsequent rapid retrieval. The index is created using database technology (such as Elasticsearch or ApacheSolr) to make the data retrieval operation more efficient. Trend data is extracted from the output of the environmental prediction model and its change pattern is analyzed. These data should include the rising, falling or stable time period of the trend for real-time storage. These trend data are combined with the index to create a complete data record for each time period to facilitate subsequent storage and processing. A real-time data embedding storage framework is designed to ensure that it can support distributed storage and fast data access. A suitable distributed database (such as Cassandra, HBase or MongoDB) is selected as the data storage backend. The data sharding strategy is determined to evenly distribute the data to each storage node to ensure load balancing. The data is sharded according to the time period or sensor ID. The distributed storage nodes are configured to ensure that each node can process the read and write requests of the data. Appropriate hardware resources (such as CPU, memory and storage space) are set to meet the needs of real-time processing. Requirement, adopt data copy strategy to ensure redundant storage of data, improve system reliability and availability, embed index and environmental prediction trend data into distributed storage framework, write data to corresponding nodes according to designed sharding strategy by writing data writing script, monitor write delay and error rate during writing process to ensure data can be stored quickly and accurately, implement real-time data update mechanism so that new data can be immediately embedded into storage framework when environmental parameters change, use message queue (such as Kafka or RabbitMQ) to process real-time data stream to ensure real-time data, set data update frequency, such as once a minute, to ensure that stored data is always up to date, establish monitoring system, monitor the performance of distributed storage framework in real time, collect indicators such as load, response time and storage usage of each node, so as to find and solve problems in time, use visualization tools (such as Grafana) to display monitoring data, help operation and maintenance personnel quickly identify bottlenecks and faults, continuously optimize storage framework according to monitoring results, adjust data sharding strategy and copy strategy to improve data access efficiency and overall system performance, regularly evaluate the performance of storage nodes, and expand nodes when necessary to cope with growing data demand.
[0033] Step S5: Perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain a redundant node optimization framework; In this embodiment, redundant data refers to repeated or redundant data records in the storage system. An indicator for redundancy analysis is selected, such as data repetition rate, storage occupancy rate and access frequency. A threshold is set to determine which nodes or data records are considered redundant. For example, if the duplicate data of a node accounts for more than 50%, the data of the node can be considered redundant. The data of all nodes are extracted from the real-time data embedded storage framework, sorted and summarized, and the storage data of each node is obtained using a query tool or script to ensure data integrity. The data volume, storage space occupancy and access frequency of each node are recorded for subsequent analysis. By writing a script, the data of each node is compared to identify duplicate data records. A unique identifier for each data record is generated using a hash algorithm to quickly find duplicates. The data repetition rate of each node is calculated and compared with the previously set threshold to determine whether it is a redundant node. For example, if the data repetition rate of a node is 60%, the node is marked as redundant. According to the analysis results, all redundant nodes are identified, the IDs, stored data volume and repetition rate of these nodes are recorded, and a redundant node report is generated. The list of redundant nodes is compared with the access date of the system. Cross-check the redundant nodes to evaluate their impact on system performance. According to the results of redundancy analysis, formulate a deletion strategy for redundant nodes. Choose to completely delete redundant nodes or archive them for subsequent audit and data recovery. Determine the priority of deletion, such as deleting nodes with low access frequency and high data duplication rate first. Use the database deletion command or the storage service API to delete the nodes marked as redundant one by one. During the deletion process, ensure the consistency and integrity of the data to avoid affecting the data of other normal nodes. Before the deletion operation, back up the data to prevent important data loss due to unexpected circumstances. After the redundant nodes are deleted, monitor the performance of the storage framework in real time to evaluate the impact of the deletion operation on the system load and response time. Use monitoring tools to collect performance indicators such as storage utilization and access latency. Compare the performance data before and after deletion and analyze the improvement effect of redundant node deletion on the performance of the storage framework. According to monitoring data and user feedback, continuously optimize the management strategy of redundant nodes. Regularly perform redundancy analysis to identify new redundant nodes to ensure that the storage framework always maintains the best state. As the amount of data grows, adjust the frequency of redundancy analysis and deletion strategy to ensure efficient use of system resources.
[0034] Step S6: Calculate the storage load of each node in the redundant node optimization framework and adjust the dynamic index structure, thereby constructing a dynamic index adjustment storage framework to perform data recorder information management operations.
[0035] In this embodiment, the indicators for storage load calculation are determined, including storage capacity utilization, data access frequency, node response time, and CPU utilization, etc. Appropriate weights are selected to comprehensively evaluate the importance of each indicator. The weights of storage capacity and access frequency are set to 0.4, and the weights of response time and CPU utilization are set to 0.3. The overall load is calculated in this way. The performance data of each node is extracted from the storage framework, including current storage usage, access logs, and system resource usage, etc. The real-time data is collected using a monitoring tool and organized into an analyzable format, such as a CSV or database table. The ID, storage capacity, used space, number of accesses, and response time of each node are recorded to provide basic data for subsequent load calculation. A calculation script is written to perform load calculation for each node. The previously defined combination of indicators and weights is applied to calculate the comprehensive load index of each node, such as the load index. The calculation results are analyzed to identify nodes with high load and nodes with low load. A threshold is set. For example, a node with a load index exceeding 0.7 is regarded as a high-load node. The load conditions of different nodes are compared to evaluate the load balance of the current storage framework and identify potential bottlenecks. Based on the results of the load calculation, it is evaluated whether the current index structure meets the needs of data access. Consider the data access pattern and identify the indexes that need to be optimized. For example, for high-load nodes, it is necessary to add auxiliary indexes or adjust existing indexes to improve data retrieval efficiency. Based on the evaluation, implement dynamic index adjustment and dynamically create, delete or modify the index structure according to the node load. For example, for nodes with high access frequency, add time range-based indexes to speed up the access of time series data. Use the index management function of the database management system to ensure that the adjustment process does not affect the normal operation of the system. Monitor the performance changes after adjustment to ensure the effectiveness of the index structure. Integrate the results of dynamic index adjustment into the storage framework and build a dynamic index adjustment storage framework to ensure that the new framework can support real-time data updates and flexible index management. Perform performance testing to evaluate the performance of the new framework in data logger information management operations, including data writing speed, query response time, and system load. After the new framework is running, implement continuous monitoring to track storage load and index performance in real time. According to the monitoring results, perform regular load calculation and index adjustment to ensure the efficiency and adaptability of the framework. As the amount of data increases and the access pattern changes, adjust the index strategy in a timely manner to optimize the performance of the storage framework to ensure that it can always meet the needs of information management operations.
[0036] In this embodiment, refer to Figure 2 , is a flowchart of detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Step S11: Acquire real-time environmental monitoring data based on multi-source sensors of the data recorder; Step S12: performing parameter normalization distribution analysis on the real-time environmental monitoring data to generate a normalization distribution range; Step S13: filtering abnormal outlier parameters based on the normalized distribution range, thereby obtaining outlier filtering monitoring data; Step S14: performing high-frequency filtering and noise reduction on the outlier filtered monitoring data to obtain filtered noise reduction monitoring data; Step S15: performing multi-time period data deep semantic analysis on the filtered noise reduction monitoring data to generate deep semantic features of the data for each time period; Step S16: Intelligent index definition is performed based on the deep semantic features of the data in each time period to obtain the data index for each time period.
[0037] In this embodiment, the data recorder is configured to ensure that it can be connected to a variety of sensors, including temperature sensors, humidity sensors, air pressure sensors, PM2.5, PM10 sensors and GPS modules, etc. These sensors are used to monitor environmental conditions and vehicle status in real time, ensure that the sensors are calibrated correctly, set the data collection frequency to once per second, ensure the real-time and accuracy of the data, start the data recorder, and start to obtain real-time data from each sensor. The data includes real-time temperature, humidity level, air pressure, particulate matter concentration and vehicle location information, record the timestamp of the data, so that the time change of the data can be tracked in subsequent analysis, and store the collected real-time monitoring data in a local storage device, and upload the data to a cloud service through a wireless network. The server is used for subsequent access and analysis, ensuring the security of data transmission, protecting the integrity and privacy of data through encrypted connections, obtaining real-time environmental monitoring data from the cloud server, and preliminarily cleaning the data to remove missing values and outliers to ensure the integrity and accuracy of the data. Each monitoring parameter is organized into a data frame (DataFrame) for subsequent analysis. Statistical analysis software (such as Python's SciPy library) is used to perform a normal distribution test on each environmental monitoring parameter. Commonly used methods include the Shapiro-Wilk test or the Kolmogorov-Smirnov test to determine whether the data conforms to the normal distribution. If the data does not conform to the normal distribution, logarithmic transformation, Box -Cox transformation or Z-score standardization and other methods are used for processing. Once it is confirmed that the data conforms to or conforms to the normal distribution after processing, the mean and standard deviation of each parameter are calculated to generate a normalized distribution range, and the range of each parameter is recorded. For example, the normalized range of temperature is a value within ±2 standard deviations, which will be used for subsequent anomaly detection. According to the normalized distribution range generated in step S12, outlier detection is performed on the real-time monitoring data. The threshold is usually set to ±3 standard deviations of the mean. Values outside this range are regarded as abnormal data. The code is written using Python's NumPy library and Pandas library to quickly identify whether there are outliers for each parameter, and the identified outliers are removed from the monitoring data set to form an outlier filter. monitoring data, ensure that only values that meet the normalization range are retained in the data set, record data changes during the filtering process, such as the number of original data and the number after filtering, ensure the transparency and traceability of data processing, select a suitable high-frequency filtering algorithm, common methods include moving average filtering, Kalman filtering or low-pass filtering, and the selection of a suitable method depends on the characteristics of the data and the type of noise. For example, moving average filtering is suitable for processing smooth signals, while Kalman filtering is suitable for state estimation of dynamic systems. Input the monitoring data after outlier filtering into the selected filtering algorithm for high-frequency filtering, ensure that the filtering parameters (such as window size) are reasonably set according to the characteristics of the data, and record the changes of each parameter before and after filtering, such as the temperature change trend.Ensure that the filtering process achieves the expected effect, save the filtered monitoring data for subsequent analysis and processing, ensure that the data format and structure are consistent with the previous ones, so as to carry out continuous analysis, divide the monitoring data into multiple time periods (such as every hour, every half hour or every minute) according to the timestamp of the data, use the time series processing function of the Pandas library to implement data resampling, set the length of each time period, ensure the rationality of the division, so as to facilitate subsequent in-depth analysis, conduct in-depth analysis of the data in each time period, extract relevant semantic features, including statistical features such as average, maximum, minimum, volatility, as well as trend and periodicity analysis, use machine learning algorithms (such as cluster analysis or principal component analysis) to mine the potential patterns and features in the data, and organize the extracted deep semantic features into structured data , and store them in the database for subsequent query and analysis, ensuring that the feature records of each time period are clear and concise. According to the extracted deep semantic features, formulate an indexing strategy to determine which features are most representative and suitable for indexing. Usually, features with high variability and information content are selected. For example, the average temperature, humidity and air quality index of each time period are selected as the main index features. A unique index value is generated for each time period. A hash algorithm or a unique identifier (UUID) is usually used to ensure the uniqueness and efficiency of the index. The index is associated with the corresponding deep semantic features for subsequent rapid retrieval and analysis. The generated index and related data features are stored in the database to ensure efficient retrieval capabilities. By building an index, data for a specific time period can be quickly found, which improves the speed and efficiency of data analysis. ,
[0038] In this embodiment, refer to Figure 3 , is a flowchart of detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Step S21: mining the environmental parameter change characteristics of the filtered noise reduction monitoring data to generate environmental parameter change trend characteristics; Step S22: Calculate the change frequency of the environmental parameter change trend characteristics to extract the environmental parameter change frequency; Step S23: identifying periodic evolution rules based on the frequency of environmental parameter changes to generate periodic rules of environmental change trends; Step S24: predicting the evolution of environmental parameters of the filtered noise reduction monitoring data based on the cycle law of environmental change trends, so as to obtain the prediction characteristics of future environmental trends; Step S25: Parameter visualization rendering is performed on the future trend prediction features of the environment to construct a visualization view of the environment prediction trend.
[0039] In this embodiment, a time series analysis method, such as a sliding window method, is used to calculate the rate of change and abnormal fluctuation of the environmental parameters in each time period. For example, by calculating the difference between each time point and the previous time point, the change amount of each parameter is obtained, and the formula is Δx(t)=x(t)-x(t-1). Statistical analysis tools (such as Pandas and NumPy libraries in Python) are further used to extract the change characteristics of each parameter, including mean, variance, maximum value and minimum value, etc. The extracted change characteristics are visualized (such as drawing a trend graph) to identify the long-term change trend of each environmental parameter. For example, the temperature fluctuates periodically with the change of seasons. Subsequently, a data frame is generated based on the trend characteristics to record the change characteristics of each parameter. The change trend of the number, including the change direction (increase, decrease or stability) and the change amplitude, is used for subsequent analysis. Based on the change trend feature extracted in step S21, a threshold value of feature change is defined. For example, a threshold value (such as ±2°C) is set to judge the significant change of temperature. The change data of each time period is traversed, and the number of changes exceeding the threshold is calculated to obtain the frequency. For example, if the temperature changes by more than ±2°C 10 times in a month, the frequency is 10 times / month. The change frequency of each environmental parameter is recorded in structured data, including parameter name, change frequency, observation time period and other information. The Pandas library is used to create a data frame for subsequent analysis and visualization. Statistical analysis is performed on the frequency data, such as calculation The average change frequency and standard deviation of each parameter are used to evaluate the stability of its changes. The Python SciPy library is used to perform these statistical calculations. The Fourier transform (FFT) analysis method is used to extract the frequency components of environmental parameter changes. The main periodic components are identified by converting the time series data into the frequency domain. The FFT function in the Python NumPy library is used to Fourier transform the time series data to identify the main frequencies and corresponding amplitudes. Based on the results of the Fourier transform, significant periodic components are extracted and the lengths of these cycles are determined. For example, if the main cycle of temperature change is 30 days, it is inferred that there is seasonal change. The periodic characteristics of each parameter are recorded, including the cycle length, amplitude and phase The information will be used for subsequent forecasting and analysis. The extracted periodic patterns will be stored in the database for subsequent query and analysis. The periodic characteristics of each parameter will be recorded and relevant statistical information will be provided. A suitable forecasting model will be selected for future trend forecasting. For example, a time series forecasting method such as seasonal ARIMA model (SARIMA) or long short-term memory network (LSTM) will be used to set the parameters of the model. According to the periodic characteristics extracted in step S23, a suitable lag period and seasonal parameters will be selected. The selected forecasting model will be trained using the historical filtered and denoised monitoring data. The data set will be divided into a training set and a test set to ensure the generalization ability of the model. The model performance will be evaluated through cross-validation and prediction error (such as RMSE).Ensure the effectiveness of the model in actual prediction. Once the model training is completed, use the model to predict the next environmental parameters and record the predicted values of each parameter in the future time period, such as the temperature and humidity changes in the next 30 days. Organize the prediction results into structured data for subsequent analysis and visualization. Select suitable visualization tools (such as Matplotlib, Seaborn or Plotly) to draw the prediction trend chart of the environmental parameters. These tools can generate high-quality charts and support interactive visualization. Input the prediction results in step S24 into the visualization tool to generate a trend chart for each environmental parameter. The chart should display the timeline, the predicted value and its confidence interval to intuitively reflect the change trend and uncertainty. For example, draw the temperature change trend in the next 30 days, show the rise and fall of temperature, and the corresponding confidence interval. Organize the generated visualization chart into a report for easy sharing with relevant teams and decision makers. The report should include an explanation of the environmental change trend and a discussion of the potential impact. Add annotations to the visualization chart to highlight key change points and potential risks to facilitate subsequent decision making. ,
[0040] In this embodiment, the specific steps of step S21 are: The filtering and noise reduction monitoring data includes environmental temperature parameters, environmental humidity parameters, light intensity, air pollutant monitoring data and wind monitoring data; Fit the time series variation characteristics of ambient temperature parameters and ambient humidity parameters to construct ambient temperature variation curves and ambient humidity variation curves; Performing light intensity spatial distribution analysis on light intensity to obtain light intensity spatial distribution characteristics; Calculate the pollutant peak value of air pollutant monitoring data and extract the pollutant peak point; Identify wind direction and wind intensity from wind monitoring data to extract wind characteristics; The environmental change trend is mined by the spatial distribution characteristics of light intensity, peak points of pollutants, wind characteristics, ambient temperature change curve and ambient humidity change curve to obtain the trend characteristics of environmental parameter changes.
[0041] In this embodiment, ambient temperature and humidity parameters, including timestamps, temperature values, and humidity values, are extracted from the filtered noise reduction monitoring data. Ensure the integrity and consistency of the data and remove missing values and outliers. Use the Pandas library to organize the data into a time series format and set the timestamp as the index. Use time series analysis methods, such as autoregressive moving average model (ARIMA) or exponential smoothing, to fit the temperature and humidity data. First, a stationarity test (such as ADF test) is performed, and if the data is not stable, differential processing is performed. Use Python's statsmodels library to fit the ARIMA model, and select appropriate model parameters (p, d, q) to minimize the AIC value. Once the model is fitted, use the Matplotlib library to visualize the fitting results and draw temperature change curves and humidity change curves. The curve should show the actual observed values and the fitted values to facilitate the identification of the model's fitting effect. Add annotations to the chart to emphasize trends and seasonal changes. For example, the temperature peaks in the summer, while the humidity is higher in the winter. Extract light intensity data from the monitoring data, including time, location (GPS coordinates) and light intensity values. Ensure that the data has spatial references to facilitate subsequent analysis. Use geographic information system (GIS) tools (such as QGIS or Geopandas) to spatially visualize the light intensity data. Create a heat map of light intensity to intuitively show the distribution differences of light intensity in different areas. Calculate the mean and standard deviation of light intensity in each area, and identify areas with high and low light intensity as the focus of analysis. Combined with the results of spatial distribution analysis, extract light intensity distribution characteristics, such as the number of high light intensity areas, distribution density and their relationship with the surrounding environment. These characteristics will be used for subsequent environmental change trend mining. Extract monitoring data of various air pollutants (such as PM2.5, PM10, NO2, SO2, etc.) from the monitoring data, ensure data integrity and perform necessary cleaning. Use data analysis methods to calculate the peak point of each pollutant. This is achieved through the sliding window method, set the window size (such as 3 hours), and calculate the maximum value within the window. Use the find_peaks function in Python's SciPy library to identify the peak points in the pollutant concentration curve and record their timestamps and corresponding concentration values. Organize the extracted peak points into a data frame, record the peak time, concentration and position of each pollutant in the time series. Analyze the frequency and time period of peak occurrence to identify the peak period of pollution sources. Extract wind data from the monitoring data, including wind speed, wind direction and timestamp. Ensure the accuracy of the data and remove outliers. Use data visualization tools (such as Matplotlib or Seaborn) to draw polar coordinate graphs of wind speed and wind direction to intuitively display the distribution characteristics of the wind. Calculate the average, maximum and main directions of wind speed to identify periods with higher wind speeds and main wind directions. For example, use a histogram to display the frequency distribution of wind direction.Record the extracted wind characteristics into structured data, including time period, average wind speed, maximum wind speed, and main wind direction. Analyze how these characteristics affect other environmental parameters, such as temperature and humidity. Use multiple regression analysis or machine learning methods (such as decision trees or random forests) to explore the mutual influence between different environmental parameters. By building a model, identify which parameters have a significant impact on environmental change trends. Use cross-validation to evaluate the stability and accuracy of the model to ensure the reliability of the results. Record the mined environmental change trend characteristics into a database to ensure that the data is structured for subsequent analysis and query. Use visualization tools to generate comprehensive trend charts to show the changing trends and relationships of different environmental parameters and provide intuitive analysis results.
[0042] In this embodiment, refer to Figure 4 , is a flowchart of detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include: Step S31: Acquire historical storage data of the data recorder; Step S32: defining a time span based on the historical storage data, and performing uniform segmentation processing according to the time span, thereby generating a plurality of historical storage data nodes; Step S33: performing adaptive lossless compression on multiple historical storage data nodes to obtain multiple lossless compressed data nodes; Step S34: construct a distributed storage architecture for multiple lossless compression data nodes, thereby obtaining a distributed data node storage framework.
[0043] In this embodiment, ensure that the data logger is working properly and the data is complete. Before acquiring data, check the storage status of the device to ensure that no data is lost or damaged. Determine the data type, including ambient temperature, humidity, light intensity, air pollutant concentration, and wind monitoring data. Connect the data logger through a USB interface, Wi-Fi, or Bluetooth. Use appropriate software tools (such as data management software or custom scripts) to extract historical data from the device. Ensure that the data format is consistent during the data extraction process, such as JSON, CSV, or database format. The extracted data should include a timestamp for subsequent time series analysis. Store the extracted data in a local file system or database, and back up the data to prevent data loss. Use a version control system (such as Git) to manage historical versions of data. During data storage, record the time of data acquisition, device status, and data volume for subsequent tracking and analysis. Determine the time range of data analysis, such as the past year, six months, or a specific event period. Set a reasonable time span based on the frequency of data collection (such as every minute, every hour). By analyzing the time distribution of historical data, select an appropriate start time and end time to ensure that the data within the time span is complete enough. Use a programming language (such as Python) to evenly split the historical storage data and set the time interval for the split, such as daily, weekly, or monthly, depending on the characteristics of the data and analysis requirements. Use the res ample function, resample the data at a set time interval, select a suitable lossless compression algorithm, such as LZ77, Huffman coding, or Zstandard, which can reduce the storage size of data without losing data, evaluate the compression efficiency and speed of different algorithms to select the most suitable algorithm, compare the compression ratio and processing time of different algorithms when processing specific data sets through experiments, use compression libraries in Python (such as zlib or lzma) to compress each data node, store the compressed data in a new data structure, ensure that the original size and compressed size of each compressed node are recorded for subsequent performance evaluation, and evaluate different distributed storage technologies (such as Ha doopHDFS, Apache Cassandra, Amazon S3, etc.), select the appropriate storage solution, consider factors including data volume, access frequency, scalability and fault tolerance, set the architecture of the storage cluster, determine the number of nodes, storage capacity and network configuration, distribute the compressed data nodes to multiple storage nodes, ensure data redundancy and security, for example, adopt data sharding and replication strategies to ensure that each data node can exist in multiple storage nodes, use the distributed file system API (such as Hadoop's HDFS API) to upload and manage compressed data nodes, design data access interfaces to ensure that users and applications can efficiently access and retrieve data stored in the distributed architecture,Use RESTful API or GraphQL interface to provide data query function, record the status and performance indicators of each storage node for monitoring and maintenance, and ensure the stability and high availability of the system.
[0044] In this embodiment, step S4 includes the following steps: Step S41: quantifying the data sensitivity of the environment prediction trend visualization view to obtain the sensitivity of the real-time parameter view; Step S42: performing data access demand analysis on the instant parameter view sensitivity to generate real-time environment monitoring data access demand characteristics; Step S43: making a storage decision based on the access demand characteristics of the real-time environment monitoring data to obtain the storage location of the real-time data; Step S44: index and mark the environment prediction trend visualization view according to the data index of each time period, and embed the distributed data node storage framework into a distributed storage based on the storage location of the real-time data to generate a real-time data embedding storage framework.
[0045] In this embodiment, a suitable sensitivity analysis method is selected, such as local sensitivity analysis (such as the first-order partial derivative method) or global sensitivity analysis (such as the Sobol method or the variance decomposition method). Local sensitivity analysis is suitable for changes within a small range, while global sensitivity analysis is more suitable for the evaluation of the impact of overall parameters. Determine the parameters that need to be analyzed, such as temperature, humidity, light intensity, air pollutant concentration, and wind speed. Extract the predicted values of each parameter and its range of variation from the environmental forecast trend visualization view. Use Python's Pandas library to perform statistical analysis on historical data and calculate the mean and standard deviation of each parameter. Make a small disturbance (for example, ±5%) to each parameter within its range of variation, observe the changes in the prediction results, and record the output results. Calculate the sensitivity index of each parameter, usually using the sensitivity coefficient SC= , where Y is the prediction result, X is the parameter value, and ΔY and ΔX are the changes. Collect information about the frequency, time and type of access (such as query, analysis or download) of users to environmental monitoring data through user surveys, questionnaires or usage log analysis. Use data analysis tools (such as Google Analytics or custom data analysis scripts) to analyze user behavior to identify access patterns and parameters with high frequency requests. Based on the collected data, create a user access feature data frame to record the access frequency, access time period and user type (such as researchers, managers, etc.) of each parameter. Calculate the average access frequency and access peak period of each parameter to identify which parameters are of most concern. Record the access demand features in the database and use data visualization tools (such as Matplotlib or Seaborn) to generate access demand heat maps to intuitively display users' access needs for each parameter. Make sure to record the change trend of each feature for subsequent storage decisions and optimization. Combined with the results of the access demand analysis, determine which data should be stored on high-performance storage devices (such as SSDs) to meet the needs of high-frequency access. For data with low access frequency, choose to store it on low-cost devices (such as HDDs) to reduce storage costs. Design a distributed storage architecture, taking into account data redundancy and load balancing. Use a distributed file system (such as HDFS) to achieve high availability and fast access. Set a data sharding strategy to distribute frequently accessed data on multiple nodes to improve concurrent access capabilities. Create a storage decision document to record the storage location, storage type, and storage strategy for each parameter. Ensure the traceability of the document for subsequent optimization and adjustment. Mark the environmental forecast trend visualization for each time period based on the data index generated in steps S41 and S42. Ensure that each visualization element (such as curves and charts) can quickly access the corresponding data node. Use the database indexing mechanism (such as the indexing function of Elasticsearch or SQL database) to speed up data retrieval. Embed data nodes into the distributed storage framework based on the storage location of real-time data. Use data management tools (such as Apache Kafka or RabbitMQ) to achieve real-time transmission and processing of data streams. Ensure that the embedded storage framework can support dynamic expansion to adapt to changes in data volume and user access needs. Test the performance of the new storage framework and access mechanism, record the access response time and system load, and evaluate the optimization effect. Make corresponding adjustments and optimizations based on the test results to ensure the efficiency and stability of system operation.
[0046] In this embodiment, step S5 includes the following steps: Step S51: performing user access behavior identification on multiple lossless compressed data nodes one by one, and extracting access behavior data of each node; Step S52: Calculate the data access frequency of each node's access behavior data to obtain the data access frequency of each node; Step S53: Analyze the access time range of the access behavior data of each node to generate the access time range characteristics of each node; Step S54: Perform historical data access demand evaluation based on the data access frequency of each node and the access time range characteristics of each node, thereby obtaining an access demand evaluation value of each node; Step S55: Perform data redundancy comparison on the access demand evaluation value of each node based on the preset node access demand threshold. When the access demand evaluation value of the node is lower than the preset node access demand threshold, perform redundant node deletion processing on the real-time data embedded storage framework to obtain a redundant node optimization framework.
[0047] In this embodiment, user access logs are extracted from the data storage system. These logs contain user ID, access time, node identifier of the access, and access mode (such as read, write, etc.). The integrity and accuracy of the log data are ensured. Data analysis tools (such as ELKStack or Splunk) are used to centrally manage and analyze the logs. Scripts are written to analyze the log data to identify the access behavior of each node. The log data is loaded into a data frame using Python's Pandas library for subsequent processing. The access records of each node are grouped, and the number of visits, access users, and access modes of each node are counted. Based on the access behavior data extracted in step S51, the access frequency of each node within a specific time range is calculated. For example, the time range is set to the past month. The Pandas library is used to process the data, and the access frequency of each node is calculated. The calculation results are organized into a data frame containing information such as node ID, access frequency, and time period, to ensure that each node is recorded. The access frequency changes of the points are analyzed for subsequent analysis. The statistical information of the access frequency, such as the mean, maximum value and standard deviation, is added to the data frame to analyze the access pattern. The access behavior data in step S51 is used to extract the access timestamp of each node and analyze its time distribution. The access time range of each node, including the earliest and latest access time, is calculated. The access time range characteristics of each node, including the number of visits and the access interval (such as the average access interval, the maximum interval, etc.), are recorded. The access peak time of each node (such as a specific time period of the day) is counted, and a visualization chart (such as a heat map) is generated to display the access time characteristics. Based on the access frequency and access time range characteristics calculated in steps S52 and S53, an evaluation model is constructed. A simple weighted model is used to assign weights to each node according to the access frequency and access time range. The weights are set. For example, the access frequency accounts for 70%, the access time range accounts for 30%, and Demand_Score=0.7×Frequency+0.3×Time_Range, use the above formula to calculate the access demand evaluation value of each node, and record the result in the data frame, ensure that the result contains the node ID, access frequency, time range and its corresponding evaluation value, count the evaluation values of all nodes, analyze which nodes have higher access demand and which nodes have relatively lower access demand, set the preset node access demand threshold according to historical data and business needs, for example, set the threshold to the average or median of the access demand evaluation value, ensure that the threshold is reasonable, adjust the threshold by analyzing historical access records to ensure that it can effectively distinguish high-frequency and low-frequency nodes, traverse the access demand evaluation value of each node, and perform redundancy comparison. When the access demand evaluation value of the node is lower than the preset threshold, mark it as a redundant node, use database operations (such as DELETE statements) or storage management tools to delete the real-time data embedding of these redundant nodes to release storage resources, record the operation log of redundant node deletion, and generate a report to explain the deleted nodes, reasons and expected storage savings, re-evaluate the performance of the storage framework to ensure that the system can still meet the user's access needs after deleting redundant nodes. .
[0048] In this embodiment, step S6 includes the following steps: Step S61: Calculate the storage load of each node on the redundant node optimization framework to obtain the storage load characteristics of each node; Step S62: performing node access peak prediction on the storage load characteristics of each node to obtain the node access peak prediction time point; Step S63: predicting the future usage demand of users according to the predicted access peak time point of the node, and obtaining the usage demand prediction of each node; Step S64: dynamically adjust the index structure of the redundant node optimization framework according to the usage demand prediction of each node, thereby constructing a dynamic index adjustment storage framework to perform data recorder information management operations.
[0049] In this embodiment, the storage information of each node is extracted from the redundant node optimization framework, including the used space, available space, data type and data size, etc., to ensure the integrity of the data and avoid the influence of missing values on the calculation results. A monitoring tool (such as Prometheus or Nagios) is used to collect real-time storage load data for analysis. Each node is analyzed one by one to calculate the storage load characteristics, including storage utilization, read and write rate, number of I / O requests, etc. Storage_Utilization=Used_Space / Total_Space×100%. The storage load characteristics of each node (such as storage utilization, read and write rate, etc.) are organized into a data frame and the node I is recorded. D and various load indicators, ensure that the data frame contains all necessary information for subsequent analysis, generate visual charts (such as bar charts) to show the storage load of each node, so as to identify high-load and low-load nodes, select appropriate time series prediction models, such as ARIMA, SARIMA or LSTM, which can predict future access patterns based on historical data, use historical access data and storage load characteristics to build a prediction model, use Python's statsmodels library or TensorFlow / Keras for model training, divide historical data into training set and test set, use the training set to train the model, and verify the accuracy of the model through the test set, select appropriate performance indicators ( The model effect is evaluated by using the MSE or MAE (such as MSE or MAE). Once the model training is completed, the model is used to predict the future access of each node, especially the time point of the access peak. The predicted time point of the access peak of each node is recorded in the data frame, including the node ID, the predicted peak time and the corresponding expected access frequency. A visual chart is generated to show the predicted access peak of each node to help identify the upcoming high-load period. Combined with the predicted time point of the node's access peak, the user's historical usage behavior is analyzed to build a demand forecasting model. Regression analysis, time series analysis or machine learning models (such as random forests) are used for demand forecasting. Historical data is used to calculate usage demand indicators, such as the number of visits, the number of users, and the type of data request. ,Use historical usage data to train the demand prediction model, set the target variable as the future usage demand, for example, predict the usage demand of each node in the next week, use the trained model to predict the usage demand of each node at the predicted access peak time point, and record the prediction results, organize the usage demand prediction results of each node into a data frame, including the node ID, predicted usage demand value and prediction time period, to ensure the clarity of the results for subsequent reference, generate visual charts (such as line charts), show the usage demand prediction trend of each node, help identify potential resource demand peaks, design dynamic index structures based on the node usage demand prediction to improve data access efficiency, and consider increasing the index priority of high-demand nodes,Ensure that it can respond quickly to user requests, use data structures (such as B-trees or hash tables) to implement dynamic indexing, adjust the index structure according to demand forecasting results, write scripts or use database management tools to update the index structure of each node according to the demand forecasting results, ensure that the integrity and consistency of existing data are not affected during the update process, record each index adjustment operation, ensure that there is a complete operation log for subsequent auditing and optimization, perform performance testing in the adjusted storage framework, monitor data access latency and system load to evaluate the effect of the dynamic index structure, further optimize according to the test results, and adjust the index strategy to ensure the efficiency and stability of the system in data logger information management operations. ,
[0050] In this embodiment, a data recorder information management system based on artificial intelligence is provided, which is used to execute the data recorder information management method based on artificial intelligence as described above, including: A deep semantic analysis module is used to obtain real-time environmental monitoring data based on multi-source sensors of the data recorder; perform multi-period data deep semantic analysis on the real-time environmental monitoring data, and perform intelligent index definition to obtain the data index for each time period; The data view module is used to predict the evolution of environmental parameters of the filtered noise reduction monitoring data, perform parameter visualization, and build a visualization view of the environmental prediction trend; The node storage framework module is used to obtain the historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture to obtain a distributed data node storage framework; The storage embedding module is used to embed the environmental prediction trend visualization view into the distributed data node storage framework based on the data index of each time period, and generate a real-time data embedding storage framework; A redundancy optimization module is used to perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain a redundant node optimization framework; The dynamic framework adjustment module is used to calculate the storage load of each node of the redundant node optimization framework and adjust the dynamic index structure, so as to build a dynamic index adjustment storage framework to perform data recorder information management operations.
[0051] Through the deep semantic analysis module of the present invention, the system can conduct a deeper analysis of the sensor data, not just staying at the data collection level, but can identify the potential relationships and rules between the data and mine more information. Through the intelligent index definition, the system can generate an efficient index for the data of each time period, making future retrieval and query of data more efficient. Users can quickly locate the required data through time or other parameters to improve query efficiency. Through the visual rendering of data, users can view the evolution trends of different environmental parameters (such as temperature, humidity, pollutants, etc.) in real time. Such graphical display helps users understand and analyze data more intuitively and supports faster decision-making. The module not only displays the current environmental conditions, but also predicts future environmental changes through trends, helping users to identify potential change trends in advance and facilitate early response to environmental risks. Through adaptive lossless compression, the system can effectively reduce the storage occupancy of data while ensuring the integrity of the data. This compression method will not lose data quality and is suitable for the storage of large-scale environmental monitoring data. Through the construction of a distributed storage architecture, the system can flexibly expand the storage capacity to adapt to the growing data storage needs. In addition, the distributed architecture can improve Improve data access efficiency and reduce the risk of single point failure. By combining the environmental prediction trend view with the data index and storage framework, the system can manage the storage and retrieval of real-time data more efficiently. The close embedding of real-time data and historical data can reduce the delay in data query and improve the response speed of the overall system. This module provides a flexible framework for the storage of real-time data. According to the different access frequencies and storage requirements of the data, it intelligently selects the appropriate storage location to further improve the storage efficiency. Through the analysis and optimization of redundant nodes, the system can reduce unnecessary redundant data storage and effectively reduce storage costs and resource waste. Redundancy optimization not only reduces the burden on the storage system, but also improves the utilization of storage space, so that more important data can be stored and accessed first. By calculating the node storage load, the system dynamically adjusts the allocation of storage resources according to actual needs. For nodes with higher loads, more resources are allocated to them, and for nodes with lower loads, appropriate resource release or storage downgrade is performed. Dynamic index structure adjustment ensures that the data storage system can still operate efficiently under changing needs. By providing a more efficient indexing method for high-demand nodes, the system can improve the speed of query and data access.
[0052] The present invention also provides an artificial intelligence-based data recorder information management device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the artificial intelligence-based data recorder information management methods described above are implemented.
[0053] Those skilled in the art clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0054] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it is stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or partly or all or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media for storing program codes such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks or optical disks.
[0055] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is therefore intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.
[0056] The above is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A data recorder information management method based on artificial intelligence, characterized in that: The following steps are involved: Step S1: Acquire real-time environmental monitoring data based on multi-source sensors of the data recorder; Perform multi-time period data deep semantic analysis on the real-time environmental monitoring data, and perform intelligent index definition to obtain the data index for each time period; Step S2: Predict the evolution of environmental parameters for the filtered noise reduction monitoring data, perform parameter visualization rendering, and construct a visualization view of the environmental prediction trend; Step S3: Acquire historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture, thereby obtaining a distributed data node storage framework; Step S4: embedding the environment prediction trend visualization view into the distributed data node storage framework based on the data index of each time period, and generating a real-time data embedding storage framework; Step S5: Perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain a redundant node optimization framework; Step S6: Calculate the storage load of each node in the redundant node optimization framework and adjust the dynamic index structure, thereby constructing a dynamic index adjustment storage framework to perform data recorder information management operations.
2. The data recorder information management method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of step S1 are: Step S11: Acquire real-time environmental monitoring data based on multi-source sensors of the data recorder; Step S12: performing parameter normalization distribution analysis on the real-time environmental monitoring data to generate a normalization distribution range; Step S13: filtering abnormal outlier parameters based on the normalized distribution range, thereby obtaining outlier filtering monitoring data; Step S14: performing high-frequency filtering and noise reduction on the outlier filtered monitoring data to obtain filtered noise reduction monitoring data; Step S15: performing multi-time period data deep semantic analysis on the filtered noise reduction monitoring data to generate deep semantic features of the data for each time period; Step S16: Intelligent index definition is performed based on the deep semantic features of the data in each time period to obtain the data index for each time period.
3. The data recorder information management method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of step S2 are: Step S21: mining the environmental parameter change characteristics of the filtered noise reduction monitoring data to generate environmental parameter change trend characteristics; Step S22: Calculate the change frequency of the environmental parameter change trend characteristics to extract the environmental parameter change frequency; Step S23: identifying periodic evolution rules based on the frequency of environmental parameter changes to generate periodic rules of environmental change trends; Step S24: predicting the evolution of environmental parameters of the filtered noise reduction monitoring data based on the cycle law of environmental change trends, so as to obtain the prediction characteristics of future environmental trends; Step S25: Parameter visualization rendering is performed on the future trend prediction features of the environment to construct a visualization view of the environment prediction trend.
4. The data recorder information management method based on artificial intelligence according to claim 3 is characterized in that: The specific steps of step S21 are: The filtering and noise reduction monitoring data includes environmental temperature parameters, environmental humidity parameters, light intensity, air pollutant monitoring data and wind monitoring data; Fit the time series variation characteristics of ambient temperature parameters and ambient humidity parameters to construct ambient temperature variation curves and ambient humidity variation curves; Performing light intensity spatial distribution analysis on light intensity to obtain light intensity spatial distribution characteristics; Calculate the pollutant peak value of air pollutant monitoring data and extract the pollutant peak point; Identify wind direction and wind intensity from wind monitoring data to extract wind characteristics; The environmental change trend is mined by the spatial distribution characteristics of light intensity, peak points of pollutants, wind characteristics, ambient temperature change curve and ambient humidity change curve to obtain the trend characteristics of environmental parameter changes.
5. The data recorder information management method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of step S3 are: Step S31: Acquire historical storage data of the data recorder; Step S32: defining a time span based on the historical storage data, and performing uniform segmentation processing according to the time span, thereby generating a plurality of historical storage data nodes; Step S33: performing adaptive lossless compression on multiple historical storage data nodes to obtain multiple lossless compressed data nodes; Step S34: construct a distributed storage architecture for multiple lossless compression data nodes, thereby obtaining a distributed data node storage framework.
6. The data recorder information management method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of step S4 are: Step S41: quantifying the data sensitivity of the environment prediction trend visualization view to obtain the sensitivity of the real-time parameter view; Step S42: performing data access demand analysis on the instant parameter view sensitivity to generate real-time environment monitoring data access demand characteristics; Step S43: making a storage decision based on the access demand characteristics of the real-time environment monitoring data to obtain the storage location of the real-time data; Step S44: index and mark the environment prediction trend visualization view according to the data index of each time period, and embed the distributed data node storage framework into a distributed storage based on the storage location of the real-time data to generate a real-time data embedding storage framework.
7. The data recorder information management method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of step S5 are: Step S51: performing user access behavior identification on multiple lossless compressed data nodes one by one, and extracting access behavior data of each node; Step S52: Calculate the data access frequency of each node's access behavior data to obtain the data access frequency of each node; Step S53: Analyze the access time range of the access behavior data of each node to generate the access time range characteristics of each node; Step S54: Perform historical data access demand evaluation based on the data access frequency of each node and the access time range characteristics of each node, thereby obtaining an access demand evaluation value of each node; Step S55: Perform data redundancy comparison on the access demand evaluation value of each node based on the preset node access demand threshold. When the access demand evaluation value of the node is lower than the preset node access demand threshold, perform redundant node deletion processing on the real-time data embedded storage framework to obtain a redundant node optimization framework.
8. The data recorder information management method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of step S6 are: Step S61: Calculate the storage load of each node on the redundant node optimization framework to obtain the storage load characteristics of each node; Step S62: performing node access peak prediction on the storage load characteristics of each node to obtain the node access peak prediction time point; Step S63: predicting the future usage demand of users according to the predicted access peak time point of the node, and obtaining the usage demand prediction of each node; Step S64: dynamically adjust the index structure of the redundant node optimization framework according to the usage demand prediction of each node, thereby constructing a dynamic index adjustment storage framework to perform data recorder information management operations.
9. An artificial intelligence-based data recorder information management system, characterized in that: The method for executing the data recorder information management method based on artificial intelligence as claimed in claim 1 comprises: A deep semantic analysis module is used to obtain real-time environmental monitoring data based on multi-source sensors of the data recorder; perform multi-period data deep semantic analysis on the real-time environmental monitoring data, and perform intelligent index definition to obtain the data index for each time period; The data view module is used to predict the evolution of environmental parameters of the filtered noise reduction monitoring data, perform parameter visualization, and build a visualization view of the environmental prediction trend; The node storage framework module is used to obtain the historical storage data of the data recorder; perform adaptive lossless compression on the historical storage data, and construct a distributed storage architecture to obtain a distributed data node storage framework; The storage embedding module is used to embed the environmental prediction trend visualization view into the distributed data node storage framework based on the data index of each time period, and generate a real-time data embedding storage framework; A redundancy optimization module is used to perform data redundancy analysis and redundant node deletion processing based on the real-time data embedding storage framework to obtain a redundant node optimization framework; The dynamic framework adjustment module is used to calculate the storage load of each node of the redundant node optimization framework and adjust the dynamic index structure, so as to build a dynamic index adjustment storage framework to perform data recorder information management operations.
10. An artificial intelligence-based data recorder information management device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that: When the processor executes the computer program, the steps of the data recorder information management method based on artificial intelligence described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Rolling bearing fault diagnosis method based on WOA-VMD and GAT
CN116662848A
Vehicle-mounted data voice tag system based on large language model
CN117763194A
Non-operational interactive live broadcast data intelligent analysis system
CN118413708A
Hard disk data protection method and device based on artificial intelligence
CN119646906A
Method to extract parameters from in-situ monitored signals for prognostices
US20100100337A1
Cited By
Monitoring video data processing method and device based on artificial intelligence large model
CN120166200A
A method and device for processing monitoring video data based on artificial intelligence large models
CN120166200B
Train operation data analysis method based on self-learning
CN120503852A
Geological survey data management method and system
CN120688092A
Geological survey data management method and system
CN120688092B