Automated machine learning-based workflow for time series forecasting

An automated machine learning-based workflow efficiently processes sensor data to train predictive models for timely and accurate forecasting, addressing inefficiencies in managing complex systems by reducing human error and optimizing resource allocation.

JP2026062536APending Publication Date: 2026-04-09FUJITSU LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing methods for managing complex systems like data centers and HVAC equipment struggle with efficiently utilizing sensor data for predictive maintenance and optimization, leading to inefficient resource allocation and potential downtime due to manual data processing and lack of accurate forecasting.

Method used

An automated machine learning-based workflow that receives sensor data, stores it in a relational database, determines cutoff records, and trains predictive models using a threshold-based approach to ensure timely and accurate forecasting.

Benefits of technology

This approach reduces human error, optimizes resource utilization, and enables proactive maintenance by efficiently processing large amounts of sensor data, ensuring predictive models remain accurate over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062536000001_ABST
    Figure 2026062536000001_ABST
Patent Text Reader

Abstract

A workflow for time series forecasting is provided. [Solution] In one embodiment, a workflow for time series forecasting can be performed based on automated machine learning. Sensor data for measurement parameters is received from multiple sensors installed in the built environment, and the received sensor data is stored in a relational database table. Cutoff records are determined for the prediction model for the measurement parameters, associated with previous training checkpoints. Records are determined that include new records whose timestamps occur after the measurement timestamp of the cutoff record. The size of the determined records is compared to a threshold size, and a training dataset is prepared. Based on this comparison, the prediction model is trained on the training dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments described in this disclosure relate to an automated machine learning-based workflow for time series prediction.

Background Art

[0002] Maintaining, managing, and optimizing complex systems with multiple interconnected components, such as data centers, industrial facilities, and large-scale Heating, Ventilation, and Air Conditioning (HVAC) equipment, presents significant challenges for operators and facility managers. These systems generate vast amounts of sensor data regarding environmental conditions, equipment performance, and energy consumption. However, it has been found that effectively utilizing this data to predict maintenance needs, prevent failures, and optimize operations is difficult with traditional approaches.

[0003] Conventional methods often rely on fixed maintenance schedules or reactive responses to equipment failures, leading to inefficient resource allocation and potential downtime. Some facilities have implemented basic monitoring systems, but they typically lack the ability to accurately predict future maintenance requirements or provide actionable insights. Additionally, the vast amount and complexity of sensor data collected from modern facilities can overwhelm human operators and make it difficult to identify meaningful patterns or trends without advanced analysis tools. There is a clear need for a more sophisticated data-driven approach that can leverage the abundant available information to actively manage complex systems and improve overall operational efficiency.

[0004] The subject matter claimed in this disclosure is not limited to embodiments that solve problems or operate only in the environments such as those described above. Rather, this background art description is provided only to illustrate an example of the technical area in which the embodiments described in this disclosure may be implemented. [Overview of the project]

[0005] According to one aspect of one embodiment, the method may include a set of operations that includes receiving sensor data for a measurement parameter from one of several sensors installed in a construction environment. The set of operations may further include storing the received sensor data as new records in a relational database table to determine a cutoff record associated with a previous training checkpoint of a predictive model for the measurement parameter. The set of operations may further include determining from the table a record containing a new record in which each measurement timestamp defined in the table occurs after the measurement timestamp of the cutoff record. The size of the determined record may be compared to a threshold size, a training dataset may be prepared based on the determined record, and a predictive model may be trained on the training dataset based on the comparison.

[0006] The objectives and advantages of the embodiments will be realized and achieved by at least the elements, mechanisms and combinations specifically indicated in the claims.

[0007] Both the above general description and the following detailed description are given as examples and are illustrative only; they do not limit the claimed invention. [Brief explanation of the drawing]

[0008] Embodiments will be described and explained in more specific and detail through the use of the attached drawings, including the following figures. [Figure 1] This diagram illustrates an example environment for an automated, machine learning-based workflow for time series forecasting. [Figure 2] This block diagram shows an exemplary system for an automated, machine learning-based workflow for time series forecasting. [Figure 3A]This figure, along with Figure 3B, shows a flowchart illustrating an example method for training and time series forecasting using an automated machine learning-based workflow. [Figure 3B] This figure, along with Figure 3A, shows a flowchart illustrating an example method for training and time series forecasting using an automated machine learning-based workflow. [Figure 4A] Figure 4B shows an example user interface that displays predicted values ​​along with suggestions on the user interface. [Figure 4B] This figure, along with Figure 4A, shows an example user interface that displays predicted values ​​along with suggestions on the user interface. [Figure 5] This diagram shows an example flowchart of an automated, machine learning-based workflow for time series forecasting.

[0009] These all follow at least one embodiment described in this disclosure. [Modes for carrying out the invention]

[0010] Some embodiments described in this disclosure may relate to methods and systems for automated machine learning-based workflows for time series forecasting. In this disclosure, sensor data may be received from one of several sensors installed in a constructed environment. The sensors may include temperature sensors, humidity sensors, barometric pressure sensors, and air quality sensors, etc. The constructed environment may be a data center, an HVAC system, or a similar facility. The received sensor data may be stored as new records in a relational database table. Each of the sensor data may become a new record in the relational database, enabling organized storage and easy searching for analysis. The database table may function as a structured repository for all incoming sensor data. A cutoff record may be determined, which may be associated with a previous training checkpoint of the predictive model for the measurement parameters. From the table, records may be determined that contain new records whose respective measurement timestamps occur after the measurement timestamp of the cutoff record. The size of the determined records may be compared to a threshold size. A training dataset may be prepared based on the determined records. Based on the comparison of the record size to the threshold size, the predictive model may be trained on the training dataset. If the record size exceeds the threshold size, the predictive model may be trained on the training dataset. If the record size is below a threshold size, input for a predictive model can be prepared based on the received sensor data.

[0011] The technical field of automated machine learning-based workflows for time series forecasting can be improved by configuring a system to forecast the workflow with respect to time series data. The system can receive sensor data about measurement parameters from sensors installed in the build environment and store this sensor data as new records in a relational database table. The system can determine cutoff records associated with previous training checkpoints of the forecasting model. Records containing new records with measurement timestamps occurring after the timestamp of the cutoff record can be determined from the table. To prepare a training dataset, the size of these records can be compared to a threshold size. Based on this comparison, a forecasting model can be trained on this training dataset.

[0012] This approach may offer several advantages: 1. Automated data processing: The system can efficiently manage large amounts of sensor data, reducing manual labor and the potential for human error. 2. Adaptive model training: Predictive models can be updated based on new data, ensuring that their predictions remain accurate over time. 3. Efficient resource utilization: By comparing the record size to a threshold, the system can optimize when to retrain the model and balance computational resources with predictive accuracy. 4. Improved decision-making: Predictive capabilities can enable proactive maintenance and optimization of the built environment.

[0013] Traditional methods for detecting and managing environmental parameters in a built environment may involve sending data to an IoT server-based cloud system. These systems can allow operators or administrators to monitor environmental conditions and receive alerts for unusual situations. However, managing data from IoT systems can present several challenges: 1. Human error: Manual data entry or monitoring is prone to errors and can potentially lead to inaccurate data and costly mistakes. 2. Time-intensive processing: Manually processing large amounts of data can reduce efficiency and productivity. 3. Data format mismatch: IoT devices generate data in various formats, complicating manual integration and standardization and increasing the likelihood of errors. 4. Real-time processing requirements: IoT data may need to be processed in real time to be useful, and manual management can struggle to achieve this.

[0014] This disclosure may address these challenges by providing an automated, machine learning-based approach to time series forecasting. This approach can enable more efficient, accurate, and timely processing of sensor data, leading to improved management and optimization of the built environment.

[0015] Embodiments of this disclosure will be described with reference to the accompanying drawings.

[0016] Figure 1 is a diagram illustrating an exemplary environment relating to an automated machine learning-based workflow for time series forecasting, configured according to at least one embodiment described in this disclosure. Referring to Figure 1, Environment 100 is shown. Environment 100 may include a system 102, a build environment 104, a remote server 106, a relational database 108, a communication network 110, and a display device 112.

[0017] System 102 may include preferred logic, circuitry, and interfaces that can be configured to receive sensor data about measurement parameters from multiple sensors 126A-126N installed in the construction environment 104. System 102 may determine cutoff records associated with previous training checkpoints of the prediction model 206. System 102 may further determine records containing new records whose respective measurement timestamps occur after the measurement timestamp of the cutoff record. Records containing new records may be stored in a table in the relational database 108. System 102 may also compare the sizes of the determined records and prepare a training dataset. Based on this comparison, System 102 may train the prediction model 206 on the training dataset. The prediction model 206 may be trained on the training dataset based on the determination that the size of the records exceeds a threshold size.

[0018] System 102 may further prepare inputs for the predictive model 206 based on the received sensor data and the determination that the record size is below a threshold size. System 102 uses the predictive model 206 to predict future values ​​of measurement parameters based on the prepared inputs. To provide suggestions for equipment in the built environment 104 (e.g., multiple sensors 126A-126N), System 102 may query the knowledge database 214 using the predicted values. These suggestions may include actions such as repairing, servicing, maintaining, or replacing the equipment. Furthermore, the suggestions may indicate whether the predicted values ​​correspond to a failure state of the equipment.

[0019] In one embodiment, system 102 may control display device 112 associated with an administrator 128 of built environment 104. This enables predicted values and corresponding proposals to be displayed for the administrator's reference. To prepare a training data set, system 102 may extract feature columns from table records. These columns may include measurement parameters and their corresponding measurement timestamps. Their values are sorted based on the measurement timestamps, and the median interval between such timestamps may be calculated. Next, an aggregation query may be executed to generate an aggregated feature column. System 102 may identify missing measurement timestamps within this feature column and determine a sliding window size. Finally, using the sliding window size, a training data set can be obtained from the aggregated feature column.

[0020] The built environment 104 can be, for example, a data center, a heating, ventilation, and air conditioning (HVAC) system, a manufacturing building, a warehouse, a distribution center, or a power plant, or similar facilities. The data center can be considered an exemplary built environment (as shown in FIG. 1). For example, as shown in the illustration, the built environment 104 as a data center may include a set of racks 116A - 116N, surveillance cameras 118, access monitoring devices 120, cooling components 122 - 124, and a plurality of sensors 126A - 126N. The set of racks 116A - 116N may house servers, switches, routers, power distribution units, patch panels, and cooling components 122 - 124 within the built environment 104. The built environment 104, remote server 106, and display device 112 may be communicatively coupled to each other via communication network 110.

[0021] In some embodiments, system 102 may include a built environment 104. In some other embodiments, system 102 may be separately disposed outside the built environment 104. In these or other embodiments, system 102 may represent various built environments such as, for example, a data center, a heating, ventilation, and air conditioning (HVAC) system, a manufacturing building, a warehouse, a distribution center, or a power plant, or similar facilities.

[0022] The plurality of sensors 126A - 126N may include, for example, temperature sensors, humidity sensors, pressure sensors, air quality index sensors, and power consumption monitoring sensors. In an example embodiment, the built environment 104 can incorporate a storage system, network equipment, a power system, and cable management components. The set of racks 116A - 116N may include a plurality of servers ranging from traditional rack - mount servers to blade servers. The storage system may include hard drives, solid - state drives, and storage area networks (SANs). The network equipment may include routers, switches, and firewalls for managing data traffic and ensuring secure communication between devices, including sensors 126A - 126N. The plurality of sensors 126A - 126N may be placed within the built environment 104 based on specific requirements.

[0023] The access monitoring device 120, also known as an access control system, can be a security device used to regulate and monitor entry and exit points in the built environment 104. The access monitoring device 120 can be designed to ensure that only authorized individuals can access restricted areas such as, for example, an HVAC system, a data center, or a confidential area within a facility. The cooling components 122 - 124 can use various cooling systems such as, for example, air cooling, containment of hot aisles and cold aisles, in - row and in - rack cooling, chilled water systems, liquid cooling, and free cooling to manage the heat generated by servers

[0024] The remote server 106 may include logic, interfaces, and / or code configured to store sensor data received from at least one of a plurality of sensors 126A-126N as relational data 114. The remote server 106 may be configured to retrieve sensor data stored as new records in a table in the relational database 108 from the relational database 108. In one embodiment, the remote server 106 may be configured remotely or located within the system 102. In at least one embodiment, the remote server 106 may be implemented as a plurality of distributed cloud-based resources using some techniques well known to those skilled in the art. In certain embodiments, the functionality of the remote server 106 may be incorporated into the system 102 in whole or at least in part without departing from the scope of this disclosure.

[0025] The relational database 108 may be stored or cached on a device such as a remote server 106 or system 102. The relational database 108 may receive sensor data from system 102 or multiple sensors 126A-126N and provide queried data based on received queries. The received sensor data may be stored in the relational database 108 in the form of tables. The relational database 108 may be hosted on multiple servers located in the same or different locations. The operation of the relational database 108 may be performed using hardware including a processor, a microprocessor (for example, to perform or control the execution of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some other examples, the relational database 108 may be implemented using software.

[0026] The communication network 110 may include various communication media through which system 102 can communicate with a remote server 106 or a device storing sensor data. Examples of the communication network 110 may include, but are not limited to, the Internet, cloud networks, Wireless Fidelity (Wi-Fi®) networks, personal area networks (PANs), local area networks (LANs), cellular networks (e.g., Long-Term Evolution (or 4G) networks and 5G cellular networks), satellite networks (e.g., networks of low-Earth orbit satellites), and / or metropolitan area networks (MANs). Various devices within environment 100 may connect to the communication network 110 using various wired and wireless communication protocols, including TCP / IP, UDP, HTTP, FTP, ZigBee®, EDGE, IEEE 802.11, Li-Fi, IEEE 802.16, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocols, and Bluetooth®.

[0027] The display device 112 may include logic, circuitry, and interfaces configured to display predicted values ​​and suggestions, optionally in a graphical format. Examples of the display device 112 may include cathode ray tube (CRT) monitors, liquid crystal displays (LCDs), light-emitting diode (LED) displays, plasma displays, electronic paper displays (E-paper), touchscreen displays, and head-mounted displays. The display device 112 may include preferred logic, circuitry, and interfaces that can be configured to display predicted values ​​together with suggestions. The display device 112 may further display predicted values ​​in a graphical format for evaluation. Examples of the display device 112 may include, but are not limited to, cathode ray tube (CRT) monitors, liquid crystal displays (LCDs), light-emitting diode (LED) displays, plasma displays, electronic paper displays (E-paper), touchscreen displays, and head-mounted displays.

[0028] During operation, system 102 can receive sensor data about measurement parameters from multiple sensors 126A-126N installed in the construction environment 104. Measurement parameters may include, for example, temperature values ​​received from a temperature sensor and an air quality index received from an air quality sensor. System 102 may store the received sensor data as new records in a table in the relational database 108. Checks may be performed at regular intervals or after a defined period. As part of these checks, system 102 may determine cutoff records associated with previous training checkpoints of the prediction model 206 for the measurement parameters. A cutoff record may be defined as the last record used in previous training of the prediction model 206. For example, if the prediction model 206 was last trained using temperature data up to a specific date timestamp '12:00:00', this timestamp may be considered the cutoff record. All temperature values ​​recorded up to this timestamp, including this timestamp, were used in previous training of the prediction model 206. From the table, records containing new records whose respective measurement timestamps occur after the measurement timestamp of the cutoff record can be determined. The size of such records can be compared to a threshold size, and a training dataset can be prepared based on this comparison. If the size of the records exceeds the threshold, a predictive model 206 can be trained on the training dataset. The predictive model 206 can be trained using an automated ML tool 212. An automated machine learning tool (AutoML) 212 may refer to a system 102 or process that automates tasks involved in applying machine learning to real-world problems. Using the AutoML tool 212, the predictive model 206 can be automatically developed and refined based on sensor data from the table. The predictive model 206 can be a machine learning model configured to predict future values ​​or trends based on historical data.The predictive model 206 and the AutoML tool 212 can be further illustrated in Figures 2, 3A, and 3B.

[0029] In another embodiment, if the record size is below a threshold size, the system 102 can prepare input for the predictive model 206 based on the received sensor data.

[0030] In some embodiments, system 102 can prepare a training dataset by extracting feature columns from records in a table. Feature columns may store the values ​​of measurement parameters and their respective measurement timestamps. The values ​​in the feature columns can be sorted based on the measurement timestamps, and the median interval between timestamps can be calculated. Aggregation queries can be executed on the sorted feature columns to generate aggregated feature columns. Preparing the training dataset may include filling in missing values ​​corresponding to missing measurement timestamps in the aggregated feature columns and determining the sliding window size for the aggregated feature columns. Further details regarding the preparation of the training dataset and the execution of aggregate queries are provided in Figures 3A and 3B.

[0031] The predictive model 206 can be applied to the prepared input to generate predicted values ​​for measurement parameters for future timestamps. Based on these predicted values, the knowledge database 214 can be queried to determine suggestions for equipment installed in the build environment 104 (e.g., servers located in rack 116). The suggestions may include actions to perform repair, maintenance, inspection, or replacement of the equipment. The suggestions may indicate whether the predicted values ​​correspond to a failure state of the equipment or whether the equipment requires assistance. The display device 112 may be controlled to display the predicted values ​​along with the suggestions. The display device 112 may be associated with the administrator 128 of the build environment 104.

[0032] Queries from the knowledge database 214 may include determining historical values ​​that match the predicted values. Proposals corresponding to the matching historical values ​​may be retrieved from multiple proposals. The knowledge database 214 may include states containing historical values ​​of measurement parameters and associated timestamps. States may include decision information containing proposals corresponding to the historical values. States may also include source types associated with the states and decision information.

[0033] Without departing from the scope of this disclosure, modifications, additions, or omissions may be made to Figure 1. For example, the environment 100 may include more or fewer elements than those illustrated and described in this disclosure. For example, in some embodiments, the environment 100 may include the system 102 but not the relational database 108. Also, in some embodiments, each function of the relational database 108 may be incorporated into the system 102 without departing from the scope of this disclosure.

[0034] For each type of sensor data, the operations described above may be repeated to train other predictive models for that type of sensor data. For example, temperature prediction based on sensor data received from a temperature sensor, pressure prediction based on pressure values ​​received from a pressure sensor, and footstep prediction based on received access data.

[0035] Figure 2 is a block diagram illustrating an exemplary system 102 for an automated machine learning-based workflow for time series forecasting, configured according to at least one embodiment described in this disclosure. Figure 2 will be described in relation to elements from Figure 1. Referring to Figure 2, a block diagram 200 of system 102 is shown. System 102 may include a network processor 202, memory 204, a forecasting model 206, an I / O device 208, and a network interface 210. The I / O device 208 may include a display device 112. The memory 204 may include relational data 114 and an automated machine learning (AutoML) tool 212.

[0036] Processor 202 may include preferred logic, circuitry, and / or interfaces that can be configured to execute program instructions associated with various operations performed by system 102. Processor 202 may include any preferred dedicated or general-purpose computer, computing entity, or processing device, including various computer hardware or software modules, that can be configured to execute instructions stored on any applicable computer-readable medium. For example, processor 202 may include a microprocessor, microcontroller, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other digital or analog circuitry configured to interpret and / or execute program instructions and / or process data. Although shown as a single processor in Figure 2, processor 202 may include any number of processors configured to individually or collectively perform or direct any number of operations of system 102, as described in this disclosure. Furthermore, one or more of these processors may reside on one or more different systems, such as different remote servers 106.

[0037] In some embodiments, the processor 202 may be configured to interpret and / or execute program instructions and / or process data stored in memory 204. In some embodiments, the processor 202 may fetch program instructions from the predictive model 206 and load them into memory 204. After the program instructions are loaded into memory 204, the processor 202 may execute them. Some examples of the processor 202 may be a graphical processing unit (GPU), a central processing unit (CPU), a reduced instruction set computer (RISC) processor, an application-specific integrated circuit (ASIC) processor, a composite instruction set computer (CISC) processor, a coprocessor, and / or a combination thereof.

[0038] Memory 204 may include preferred logic, circuitry, and / or interfaces that can be configured to store program instructions executable by the processor 202. In certain embodiments, memory 204 may be configured to store information such as sensor data, a table of records including new records with their respective timestamps, and predicted values, but is not limited to these. Memory 204 may include a computer-readable storage medium for carrying or storing computer-executable instructions or data structures. Such a computer-readable storage medium may include any available medium that can be accessed by a general-purpose or dedicated computer, such as the processor 202.

[0039] Such computer-readable storage media may include, but are not limited to, tangible or non-temporary computer-readable storage media, including random access memory (RAM), read-only memory (ROM), electrically erasable and programmable read-only memory (EEPROM), compact disk read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory devices (e.g., solid-state memory devices), or other storage media that can be used to carry or store specific program code or data structures in the form of computer-executable instructions and that can be accessed by a general-purpose or dedicated computer. Combinations of the above may also be included in the scope of computer-readable storage media. Computer-executable instructions may include, for example, instructions and data configured to cause the processor 202 to execute specific processes or sets of processes related to system 102.

[0040] Memory 204 can further store automated ML tools 212. In some embodiments, automated ML tools 212 may be located outside of memory 204. Automated machine learning tools (AutoML) 212 may refer to a system 102 or process that automates tasks involved in applying machine learning to real-world problems. AutoML may encompass a variety of techniques and algorithms designed to automatically select, configure, and optimize machine learning models (e.g., predictive models 206) without extensive manual intervention.

[0041] In some embodiments, the AutoML tool 212 may perform tasks such as data preprocessing, feature selection, model selection, hyperparameter tuning, and model evaluation. The AutoML tool 212 may iterate through multiple combinations of algorithms, feature engineering techniques, and hyperparameters to identify a model that performs best for a given dataset and problem.

[0042] The AutoML tool 212 can efficiently explore the space of possible models and configurations by utilizing techniques such as Bayesian optimization, genetic algorithms, or neural architecture search. In some cases, the AutoML tool 212 can handle tasks such as imputing missing values, coding categorical variables, and scaling numerical features.

[0043] The AutoML tool 212 can start with input data (e.g., records from a table) and proceed through various stages, including data cleaning, feature engineering, model selection, hyperparameter optimization, and ensemble creation. Its output can be a fully trained predictive model 206 ready for deployment, along with performance metrics and explanations of the model's decision-making process.

[0044] In the context of predictive maintenance for data centers or industrial facilities, the AutoML tool 212 can be used to automatically develop and refine a predictive model 206 based on sensor data from a table.

[0045] The predictive model 206 can be a machine learning model configured to predict future values ​​or trends based on historical data. In some embodiments, the predictive model 206 may use various algorithms and techniques to analyze patterns and relationships in the input data and generate predictions. The predictive model 206 may be trained on a dataset containing historical measurements and corresponding results to learn underlying patterns and correlations.

[0046] In some implementations, the prediction model 206 may be a time series prediction model, such as an autoregressive integrated moving average (ARIMA) model, which can capture trends and seasonality in the data. Alternatively, the model may be based on more advanced techniques, such as a long short-term memory (LSTM) neural network, which can learn long-term dependencies in sequential data.

[0047] The predictive model 206 can take various forms depending on the specific requirements of the application. For example, in a data center environment, the model may be used to predict future power consumption based on historical usage patterns, environmental conditions, and scheduled workloads. In an industrial environment, the model may predict the probability of equipment failure by analyzing sensor data, such as vibration, temperature, and pressure readings.

[0048] Model predictions can be used to generate actionable insights for system operators. For example, predictive model 206 may predict, based on the performance metrics of a particular HVAC component, that it is likely to fail within the following month. This prediction can then be used to schedule preventative maintenance, potentially avoiding costly downtime and extending the lifespan of the equipment.

[0049] In some cases, the predictive model 206 may be part of a larger predictive maintenance system (e.g., system 102). System 102 may combine the output of the predictive model with other data sources and expertise to provide comprehensive recommendations for system 102 optimization and maintenance scheduling.

[0050] In some embodiments, the prediction model 206 may be defined by its hyperparameters, such as the number of weights, cost function, input size, and number of layers. During training, the parameters of the ML model may be tuned, and the weights may be updated to move toward the global minimum of the cost function of the ML model. After several epochs of training on feature information in the training dataset, the ML model may be trained to output prediction / regression results for a set of inputs. The prediction results may show values ​​for each input of the set of inputs (multiple measurement parameters).

[0051] The ML model may include electronic data that can be implemented, for example, as a software component of an application executable on the sensor data of system 102. The ML model may rely on libraries, external scripts, or other logic / instructions for execution by a processing device, for example, a processor 202. The ML model may include code and routines configured to enable a computing device, for example, a processor 202, to perform one or more operations, such as predicting the values ​​of measurement parameters of multiple sensors 126A-126N installed in the build environment 104. In addition or alternatively, the ML model may be implemented using hardware including a processor, a microprocessor (for example, to perform or control one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the ML model may be implemented using a combination of hardware and software.

[0052] The I / O device 208 may include suitable logic, circuitry, interfaces, and / or code that can be configured to receive user input. The I / O device 208 may further be configured to provide output in response to user input. The I / O device 208 may include various input and output devices that can be configured to communicate with the processor 202 and other components such as a network interface 210. Examples of input devices include, but are not limited to, a touchscreen, keyboard, mouse, joystick, and / or microphone. Examples of output devices include, but are not limited to, a display device 112 and a speaker. The I / O device 208 may be configured within or outside of the system 102.

[0053] The network interface 210 may communicate with a network via wireless communication, such as the Internet, an intranet, and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). Wireless communication may use any of several communication standards, protocols, and technologies, such as GSM® (Global System for Mobile Communications), EDGE (Enhanced Data GSM Environment), Wideband Code Division Multiple Access (W-CDMA), LTE (Long Term Evolution), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth®, Wi-Fi (Wireless Fidelity) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), VoIP (Voice over Internet Protocol), Li-Fi (light fidelity), or Wi-MAX.

[0054] In certain embodiments, system 102 may include a display device 112, a remote server 106, and a relational database 108. Modifications, additions, or omissions can be made to system 102 without departing from the scope of this disclosure. For example, in some embodiments, system 102 may include any number of other components not expressly shown or described. The predictive model is described in detail in Figures 3A, 3B, 4, and 5.

[0055] Figures 3A and 3B together show flowcharts illustrating an example method for training and time series forecasting of an automated machine learning-based workflow according to one embodiment of the present disclosure. Figures 3A and 3B may be described in relation to elements from Figures 1 and 2. Referring to Figures 3A and 3B, an execution flow 300 is shown. An exemplary execution flow 300 may include a set of operations that can be performed by one or more components of Figure 1, such as system 102. These operations may include receiving sensor data, storing sensor data in relational database 108, determining cutoff records, determining records, comparing record sizes, preparing a training dataset, and training a predictive model 206. System 102 can perform a set of operations for an automated machine learning-based workflow for time series forecasting.

[0056] At 302, the operation of receiving sensor data may be performed. System 102 may be configured to receive sensor data about measurement parameters from at least one of several sensors 126A-126N installed in the built environment 104. Examples of sensor data may include, but are not limited to, temperature values, air quality values, power consumption values, and air pressure values. The built environment 104 may be a data center, an HVAC system, a manufacturing facility, a warehouse and distribution center, a power plant, or other similar facility. A data center may house computer systems and related components such as telecommunications and storage systems. A data center may be designed to ensure the continuous operation of information technology (IT) services using mechanisms such as climate control, backup power, and security systems. An HVAC system may be essential for maintaining indoor air quality and thermal comfort within a building. An HVAC system may use several sensors 126A-126N and a sensor monitoring device (e.g., system 102) to regulate temperature, humidity, and airflow to ensure a comfortable and safe environment. Warehouses and distribution centers can be used to store goods and manage inventory. These facilities can be designed for the efficient movement and storage of products and often feature high ceilings, wide aisles, and advanced logistics systems for tracking and managing inventory. Warehouses and distribution centers may use multiple sensors 126A-126N to monitor operations.

[0057] At 304, the operation of storing sensor data can be performed. Sensor data can be stored as new records in a table in the relational database 108. The relational database 108 can query and manipulate data based on sensor data input. The relational database 108 can provide the ability to efficiently store and query large amounts of structured data. Each sensor data can correspond to a table in the relational database 108. Each sensor data collected by a sensor can correspond to a table column, and each table row can consist of sensor data (measurement parameters) collected by the sensor (e.g., 126A) at a specific point in time (i.e., a certain measurement timestamp). A construction environment 104 with multiple sensors 126A-126N can continuously collect sensor data from multiple sensors 126A-126N installed throughout the construction environment 104 and monitor the received sensor data.

[0058] In 306, an operation to determine the cutoff record for the measurement parameter may be performed. Sensor data collected from at least one of the multiple sensors 126A-126N may be checked against the cutoff record associated with a previous training checkpoint of the predictive model 206 for the measurement parameter. For example, if the measurement parameter is 'temperature value', the 'temperature value' may be received every 5 seconds. If a 'temperature value' is received at the measurement timestamp '00:10:10', the next 'temperature value' may be received at '00:10:15'.

[0059] A cutoff record can be defined as the last record used in a previous training of the prediction model 206. For example, if the prediction model 206 was last trained using temperature data up to the timestamp '12:00:00' on a particular day, this timestamp can be considered the cutoff record. All temperature values ​​recorded up to this timestamp, including this timestamp, were used in the previous training of the prediction model 206.

[0060] In the previous example, if the current timestamp is '14:30:00', system 102 can determine that all temperature values ​​recorded between '12:00:00' and '14:30:00' are new records not used in previous training of the prediction model 206. These new records may be considered in the next step of the process, for example, to determine whether there is enough new data to justify retraining the model.

[0061] The cutoff record may also be defined in terms of the number of records. For example, if the predictive model 206 was last trained using 1,000,000 temperature readings, the 1,000,000th record may be considered the cutoff record. Any temperature readings collected after this record are considered new data that was not used in the previous training of the predictive model 206.

[0062] In one example, we can consider a cutoff record with a count of '50,00,000', and the 'temperature value' occurring at the next count can be considered a measurement timestamp occurring after the cutoff record and can be considered in the next step (see 310). The 'temperature value' occurring at the cutoff record (i.e., before the count '50,00,000') can be considered a previous training checkpoint. The parameter timestamp after the previous training checkpoint can be considered. In another example, if we consider 'temperature value' as a measurement parameter, a timestamp of '10:00:00' on a particular day or month may be considered a cutoff record, and the 'temperature value' occurring after the timestamp '10:00:00' can be considered in the next step (see 310). The values ​​specified by the measurement parameter and measurement timestamp can vary depending on various scenarios (e.g., the amount of records, the type of system 102 or build environment 104).

[0063] This approach allows system 102 to efficiently identify new data for potential use when updating the predictive model 206, ensuring that the model remains up-to-date and accurate as new sensor data is continuously collected from the build environment.

[0064] In step 308, it is possible to determine which records in the table contain new records whose measurement timestamps occur after the measurement timestamp of the cutoff record. Specifically, system 102 may compare the measurement timestamp of each record in the table with the timestamp of the cutoff record. Any record with a timestamp later than the timestamp of the cutoff record can be considered a new record and included in the determination. This process can effectively filter out all data used in previous training checkpoints of the predictive model 206, focusing only on newly collected sensor data. For example, if the cutoff record has a timestamp of '2023-07-30 12:00:00', system 102 may identify all records in the table with timestamps later than '2023-07-30 12:00:00'. Such records may represent new sensor data collected since the last model training checkpoint.

[0065] In step 310, it may be further determined whether the record size (determined in step 308) exceeds the threshold size. If the record size exceeds the threshold size, control may proceed to step 312. If the record size is less than the threshold size, control may proceed to step 326. Here, records that occurred after the measurement timestamp of the cutoff record may be considered in the next step 312. The operations from 312 to 320 may be performed for data preprocessing, as described here.

[0066] In step 312, feature column extraction and timestamp sorting may be performed.

[0067] The processor 202 may extract feature columns from the table records that store the values ​​of measurement parameters and their respective measurement timestamps. Values ​​from the feature columns and their corresponding measurement timestamps may be extracted and sorted based on the timestamps. Feature columns may contain specific values ​​of interest within the dataset (e.g., temperature, pressure, humidity). The table in the relational data 114 may associate measurement timestamps with measurement parameters. Measurement timestamps may record when the parameter values ​​occurred. Feature column values ​​may be sorted by timestamp. In some embodiments, measurement timestamps may be converted to a 'datetime' format for appropriate sorting. After sorting, the median interval between measurement timestamps may be calculated. The median interval can be determined using the time difference between consecutive timestamps.

[0068] In 314, after sorting the feature column, an aggregation query can be executed to generate an aggregated feature column. The aggregation query can extract the calculated timestamps, specify such timestamps as 'aggregated intervals', calculate the average value of each interval, and group the results by interval. An example of an aggregation query is given below: SELECT CAST(unixepoch(timestamp) / median_interval AS INTEGER) AS aggregated_interval, avg(feature) AS avg_feature FROM approachd_data GROUP BY aggregated_interval;

[0069] Aggregation queries can perform calculations on feature column values ​​to generate aggregated values. In one embodiment, executing an aggregation query may involve several steps of processing feature column data. In some embodiments, the processor 202 may be configured to determine multiple aggregate intervals by splitting each measurement timestamp by a calculated interval median. This splitting operation can help normalize the timestamps and group them into consistent intervals. The processor 202 may then select from the feature column a set of values ​​corresponding to each of the unique aggregate intervals of the multiple aggregate intervals. This selection process may include identifying all data points that fall within each normalized interval. After selecting the relevant values, the processor 202 may calculate the mean of the set of values ​​for each aggregate interval. This averaging operation can help smooth short-term fluctuations in the data and highlight long-term trends in the data. Finally, the processor 202 may group the mean values ​​based on the multiple aggregate intervals. The resulting aggregate feature column may contain the mean values ​​corresponding to each of the unique aggregate intervals of the multiple aggregate intervals. This grouping process can help organize the processed data into a format that can be easily used for further analysis or model training.

[0070] In some cases, the aggregation query may be customized or modified based on the specific requirements of the predictive model 206 or the nature of the sensor data being processed. For example, instead of using the mean, the query may use other aggregation functions, such as the median, maximum, or minimum, depending on the characteristics of the measurement parameter being analyzed.

[0071] At 316, missing measurement timestamps within the aggregated feature columns can be determined. The processor 202 can identify the missing timestamps and estimate such missing timestamps using a suitable interpolation method, such as linear interpolation. Linear interpolation can estimate an unknown value between two known consecutive feature column values.

[0072] In step 318, missing values ​​corresponding to missing measurement timestamps in the aggregated feature column can be filled in. Processor 202 can fill in all missing values ​​using interpolation. For example, if the measurement timestamps are '00:10:05' and '00:10:15' and the recording interval is 5 seconds, the missing values ​​between these known timestamps can be interpolated.

[0073] At 320, the sliding window size determination operation may be performed. Processor 202 may be configured to determine the sliding window size for the aggregated feature columns. The sliding window size may be defined as having a length of 2 or more (i.e., n ≥ 2, where n is an even number). Feature column values ​​that have a first set of data points that fit within the window may be considered. The window can move two data points at a time (or by a defined step size), processing each new set of data points as the window slides across the dataset.

[0074] Sliding windows can be used to create features that capture time patterns. In time series forecasting, sliding windows can create lag features by taking the value of a variable from a previous point in time and including such a value as a feature in the model at the current point in time. For example, to forecast warehouse access data for the next day, access data from the past seven days may be required. Sales data from the last seven days may be referenced as a lag feature.

[0075] For example, in real-time systems such as fraud detection or recommendation engines, sliding windows can help maintain up-to-date features by continuously aggregating recent data. For temperature sensor readings received at specific intervals (e.g., every 2 seconds), a sliding window of size 2 could be used to calculate the average temperature every two readings, and update the average when a new reading arrives and the old reading is removed.

[0076] Once the sliding window size is determined, the dataset can be split into appropriate ratios (e.g., 80% training set and 20% test set). A splitting index may be used to divide the window into training and test sets.

[0077] At 322, the operation to acquire the training dataset may be performed. Processor 202 may be configured to acquire the training dataset from aggregated feature columns based on the sliding window size. For example, the training dataset may contain 80% of the rows in chronological order, while the test dataset may consist of the remaining 20%. The predictive model 206 may be trained using the training dataset as input.

[0078] At 324, the operation of training the predictive model 206 may be performed. Processor 202 may be configured to train the predictive model 206 based on a comparison of the size of the determined record with a threshold size. Specifically, the predictive model 206 can be trained on the training dataset if the size of the record exceeds the threshold size. The training dataset may be prepared based on the determined records. Records are determined for measurement timestamps that occurred after the measurement timestamp of the cutoff record, and the size of the determined record may be compared with the threshold size.

[0079] In at least one embodiment, the predictive model 206 may be trained using automated machine learning operations. The automated machine learning operations for training the predictive model 206 (performed using an Automated Machine Learning Tool (AutoML) 212) may include several steps and techniques tailored specifically for time series forecasting. In some embodiments, such operations may begin with automated feature engineering, in which relevant features are extracted from time series data. This may include, for example, creating lag features, rolling statistics, and seasonal indicators.

[0080] Next, the operation can proceed to algorithm selection, where various time series forecasting models can be evaluated, such as ARIMA, Prophet, or advanced deep learning models like LSTM or transformer-based architectures. In some cases, the automated ML tool 212 can test multiple algorithms in parallel and compare their performance on training data. In addition, hyperparameter tuning may be performed automatically, and the system 102 may explore different combinations of model parameters to optimize performance. This may involve techniques such as grid search, random search, or Bayesian optimization. The automated ML operation can also handle data preprocessing tasks, such as handling missing values, detecting and removing outliers, and normalizing or scaling data as required by different algorithms. Furthermore, to ensure robust model evaluation and selection, cross-validation techniques specific to time series data, such as time series cross-validation or rolling window validation, may be used. In some implementations, the automated ML operation may incorporate an ensemble method that combines predictions from multiple models to improve overall prediction accuracy. Throughout the training of the predictive model 206, the automated ML operation can continuously monitor and evaluate model performance and, if necessary, implement an early stopping mechanism to prevent overfitting.

[0081] At 326, the operation to obtain predicted values ​​may be performed. Processor 202 may be configured to obtain predicted values ​​from the trained predictive model 206. If the determined record size is below the threshold size, predicted values ​​may be obtained from the trained predictive model 206. If the determined record size is above the threshold size, the predictive model 206 may be trained on the dataset.

[0082] Once the trained prediction model 206 is obtained, predictions can be extracted by performing a prediction procedure. The future time and features to be predicted can be supplied as input to the prediction model 206, and the prediction model 220 can output predicted values ​​for the features at a given time.

[0083] At 328, an operation to query the knowledge database 214 may be performed. The processor 202 may be configured to query the knowledge database 214 based on predicted values ​​to determine suggestions about the equipment installed in the build environment 104. The knowledge database 214 may include tables having elements such as status, initialization details, environmental sensor status, entries added by the administrator 128, and other relevant information.

[0084] At 330, an operation may be performed to determine historical values ​​that match the predicted values. The processor 202 may be configured to identify historical values ​​corresponding to the predicted values. Historical values ​​may be derived from sensor data collected before training the prediction model 206. These historical values ​​may be pre-stored in the knowledge database 214 for comparison by the administrator 128. In some embodiments, historical values ​​may include data collected during maintenance of equipment in the build environment 104.

[0085] In step 332, an operation may be performed to retrieve a suggestion corresponding to the identified historical value. The processor 202 may be configured to retrieve a suggestion corresponding to the matching historical value. The suggestion may be obtained from the knowledge database 214 (in step 334). The knowledge database 214 may include tables having various elements. These elements may include state 334A, decision information 334B, and source type 334C. The knowledge database 214 may include multiple elements without limitation.

[0086] In 334A, the state may include multiple historical values ​​of a measurement parameter and multiple timestamps associated with each of those historical values. For example, the state may indicate that the temperature has been rising over the past week and the most recent temperature is above 30 degrees Celsius, or that the number of 0.3 micron particles has exceeded 10,000 over the past 24 hours. The historical values ​​may be represented as text containing the names of the measurement parameters, such as temperature, humidity, or the number of 10.0 micron particles. The values ​​may be numerical values ​​associated with the historical values ​​(for example, for a temperature of 30 degrees Celsius, the value may be 30.0). The period or timestamp may be represented as a numerical value describing the elapsed time.

[0087] In 334B, decision information may include multiple proposals corresponding to multiple historical values. For example, data received during HVAC maintenance, repair, or replacement actions may be considered decision information regarding predicted values.

[0088] In 334C, the source type may be associated with state and decision information. For example, the source type may store an identification (ID) value indicating whether the source (table row) was derived from documentation (e.g., manuals and regulations), from the operation of the build environment 104 during maintenance, or from the expertise and experience of the build environment 104 administrator 128.

[0089] In one embodiment, when the system 102 is initialized for the first time, the initial state stored in the knowledge database 214 may include information derived from manuals and regulations. The system 102 can record sensor data from the build environment 104 while the build environment 104 is undergoing maintenance, repair, replacement, or similar activities. The administrator 128 of the build environment 104 can add entries to the knowledge database 214 based on their experience and expertise. These entries may include status, decision information, and source type. The system 102 can check and retrieve suggestions from the knowledge database 214 corresponding to measurement parameters and feature values. Based on the entries in the knowledge database 214, the system 102 may determine whether any action needs to be taken.

[0090] The knowledge database 214 may further include decision information having multiple suggestions corresponding to multiple historical values. Furthermore, the knowledge database 214 may include source types associated with the state and decision information.

[0091] Figures 4A and 4B together illustrate an exemplary electronic user interface (UI) 402 of a display device 112 that shows predicted values ​​along with the proposal, according to one embodiment of the present disclosure. Figures 4A and 4B will be described in relation to elements from Figures 1, 2, 3A, and 3B. The exemplary electronic UI 402 that shows the prediction method shown in the exemplary environment 400 may be implemented by any suitable system, apparatus, or device, such as the system example 102 in Figure 1 or the processor 202 in Figure 2. The measurement parameters may include one or more parameters for an automated machine learning-based workflow for time series forecasting, without departing from the scope of the present disclosure.

[0092] The electronic UI 402 includes various UI elements, such as a forecasting module 404, suggestions corresponding to measurement parameters 406, and measurement parameters 408. The electronic UI 402 can be associated with the administrator 128 of the build environment 104 to display the predicted values ​​along with the suggestions 406. The forecasting module 404 can display suggestions 406 that include predicted values ​​corresponding to equipment failure conditions. The administrator 128 can select the predicted measurement parameters (in 404A). Based on the administrator's selection, suggestions 406 can be displayed along with the predicted values. The suggestions 406 corresponding to the measurement parameters may include multiple options for the administrator's selection, such as a select action 406A and an update suggestion 406C. The select action 406A (UI element) may include a dropdown menu 406B. The dropdown menu 406B for action selection 406A may include options such as repair equipment A and maintenance of equipment B. The action may be determined based on the predictive model 206. The suggestion update 406C (UI element) may include a dropdown menu 406D. As shown in the figure, for example, the suggestion update 406C may include options such as equipment A repaired and maintenance done for equipment B. The actions and suggestions displayed in the dropdown menu may include additional UI elements. The update action or suggestion 406E (UI element) may be used to add or remove multiple actions or suggestions to the dropdown menu 406B or 406D. The selected actions and suggestions can be viewed in detail by downloading a report.

[0093] The electronic UI 402 may include option 406F for downloading or sharing a detailed report. The report may include details of actions to be taken and historical measurements. The report may include updated suggestions 406C based on the actions taken. The updated suggestions 406C may include another option for setting reminders to provide UI elements for maintenance of specific equipment or similar actions. The administrator 128 may set reminders or processors 202 based on the predictive model 206 or historical measurements.

[0094] Furthermore, the display device 112 may display measurement parameters (e.g., historical measurements) 408, such as parameter 1, parameter 2, ..., parameter n, on a UI element. The administrator 128 can view historical values ​​of measurement parameters received from multiple sensors 126A-126N (stored in the relational database 108). The electronic UI 402 may display various datasets relating to measurement parameters along with timestamps. For example, value 1-timestamp 1 (408A, 408B, 408C), value 2-timestamp 2 (408A, 408B, 408C, ...), value n-timestamp n (408n). In one embodiment, the administrator 128 may include multiple parameters using the 'add more parameters' 408D prompt. This disclosure may also be applicable to other changes, deletions, or additions to the electronic UI 402 without departing from the scope of this disclosure.

[0095] Referring to Figure 4B, graph 404D shows predicted values. The graph includes exemplary values ​​for measurement parameters such as power consumption, air quality index, and temperature. Administrator 128 can select the parameters to be predicted using 404C. Based on the selected measurement parameters, graph 404D may be shown for easy evaluation.

[0096] The electronic UI 402 may display predicted values ​​and may provide suggestions about equipment installed in the build environment 104 (e.g., servers and access doors). The electronic UI 402 may further include prompts for selecting measurement parameters (historical measurements) 404E. The selected measurement parameters 404E may be used to select measurement parameters and display them as a graph showing predicted values. In one embodiment, the graph may display predicted values ​​or values ​​for the measurement parameters. The administrator 128 may select measurement parameters to compare various measurement parameters. The administrator 128 may upload 404F a file or document containing the measurement parameters and measurement timestamps along with the predicted values. Note that the electronic UI 402 is provided only as an exemplary implementation of the display device 112 in Figure 1 and should not be construed as limiting the scope of this disclosure. This disclosure may also be applicable to other changes, deletions, or additions to the display device 112 without departing from the scope of this disclosure.

[0097] Figure 5 shows an example flowchart relating to an automated machine learning-based workflow for time series forecasting according to one embodiment of the present disclosure. Figure 5 will be described in relation to elements from Figures 1, 2, 3A, 3B, 4A, and 4B. Referring to Figure 5, an exemplary flow 500 is shown. The method shown in exemplary environment 500 can be performed by any preferred system, apparatus, or device, such as by system example 102 in Figure 1 or processor 202 in Figure 2. Although shown in separate blocks, steps and operations relating to one or more blocks of flow 500 can be divided into further blocks, combined into fewer blocks, or deleted depending on the particular embodiment. An operation may begin at 502 and proceed to 504.

[0098] In 504, sensor data regarding measurement parameters can be received from one of the multiple sensors 126A-126N installed in the construction environment 104. The sensor data may include data received from at least one of the multiple sensors 126A-126N, such as temperature / humidity sensors, pressure sensors, and air quality sensors. The multiple sensors 126A-126N can collect environmental data within the construction environment 104. The construction environment 104 can be a data center, an HVAC system, or a similar facility. Once the measurement parameters are received from the multiple sensors 126A-126N, they can be stored in the relational database 108. The relational database 108 organizes the data into tables. Each table contains rows and columns, where rows can represent individual records (e.g., measurement parameters) and columns can represent attributes of the records (values ​​related to the measurement parameters).

[0099] At 506, the received sensor data may be stored as a new record in a table in the relational database 108. Sensor data from at least one of the multiple sensors 126A-126N, such as temperature, humidity, and pressure sensors, may be collected from the construction environment 104. The relational database 108 may include a table, e.g., a 'sensor table', that stores information such as temperature and humidity along with the locations of the multiple sensors 126A-126N. The relational database 108 may also include a table, e.g., a 'sensor data table', that stores data collected from the multiple sensors 126A-126N, including a timestamp (e.g., measurement timestamp). The received sensor data may then be stored in the relational database 108. The 'sensor data table' is stored in the relational database 108, and the 'sensor data table' may have entries for each record, including time and measurement parameter values. By querying the relational database 108, specific data can be retrieved from it. For example, the administrator 128 can retrieve the temperature record from space 1. The relational database 108 can discover and provide information by looking at related tables.

[0100] In step 508, a cutoff record may be determined that is associated with a previous training checkpoint of the predictive model 206 for the measured parameters. The cutoff record can be determined based on the dataset used to train the predictive model 206. Consider an example where sensor data received between timestamps 01:00:00 and 12:00:00 may be used to train the predictive model 206. The cutoff record can be considered to be at 12:00:00 as a previous training checkpoint of the predictive model 206 for the measured parameters. The cutoff record may also be for different ranges (e.g., days, months, or years), and the predictive model 206 may change based on those different ranges.

[0101] In 510, records containing new records that occurred after the cutoff record's measurement timestamp, as defined in the table, may be determined. Records that occurred after the cutoff record's measurement timestamp may be determined as defined in the table.

[0102] In 512, the size of the above records can be compared to a threshold size. The threshold size can be the time set to train the prediction model 206. In one embodiment, the threshold can be the number of records received during the set period. For example, records received after a cutoff record can be compared to the threshold size, and the threshold size can be '500000' records. If the size of the determined records exceeds '500000' records, those records can be considered the threshold point and may be considered to have a threshold size of '500000'.

[0103] In 514, a training dataset may be prepared based on the above records. The training dataset may be prepared by determining, upon comparison, that the determined records exceed a threshold size. Preparing the training dataset may involve several steps, such as extracting feature columns, sorting values, calculating the median interval, executing aggregate queries, determining missing timestamps, filling in missing values, determining a sliding window, and obtaining the training dataset. The operations of extracting feature columns and sorting their respective timestamps may be performed.

[0104] Feature columns can store the values ​​of measurement parameters. The values ​​in the feature column and their respective measurement timestamps can be extracted and sorted based on their respective measurement timestamps. Measurement timestamps can record the time when the measurement parameter was measured. Values ​​within the feature column can be sorted based on their timestamps. In one embodiment, measurement timestamps can be converted to a 'datetime' format to ensure sorting. Once the values ​​in the feature column are sorted, the median interval between each measurement timestamp can be calculated. The time difference between consecutive timestamps can be determined to calculate the median difference. Aggregation queries can be performed on the feature column to generate aggregated feature columns. An aggregation query may include extracting timestamps divided by the median interval, naming the extracted timestamps 'aggregated intervals', calculating the average value of the 'aggregated intervals', and grouping the average values ​​by the 'aggregated intervals'. An example of an aggregation query is given below: SELECT CAST(unixepoch(timestamp) / median_interval AS INTEGER) AS aggregated_interval, avg(feature) AS avg_feature FROM approachd_data GROUP BY aggregated_interval;

[0105] Aggregation queries can be used to perform calculations on the values ​​of feature columns to generate aggregated values. Furthermore, processor 202 can be configured to determine missing measurement timestamps. Processor 202 is configured to determine missing measurement timestamps in the aggregated feature columns. Missing measurement timestamps can be identified and filled in using linear interpolation, and sliding window size determination can be performed. Processor 202 can be configured to determine the sliding window size for the aggregated feature columns. A sliding window may require defining its size. The sliding window size can be thought of as a length of 2 or greater than 2 (e.g., n≧2), where n is an even number. Sliding windows can be used to create features that capture temporal patterns. In real-time systems such as fraud detection or recommendation engines, sliding windows help maintain up-to-date features by continuously aggregating recent data. Once the sliding window size is determined, an index about the sliding window may be divided into 80% training and 20% testing. The index can be the point at which the dataset is split into a training set and a test set. This splitting index can be used to divide the window into a training set and a test set. Based on the sliding window size, the training dataset can be obtained from the aggregated feature column. The training dataset may consider a splitting sliding window (e.g., 80% training and 20% testing) for obtaining the training dataset.

[0106] In 516, a predictive model 206 may be trained on the training dataset based on comparison. A predictive model 206 may be trained based on a comparison of the determined record size with a threshold size. The training dataset can be prepared based on the determined records. Records whose measurement timestamp occurs after the measurement timestamp of the cutoff record can be determined, and the determined record size can be compared with a threshold size. A predictive model 206 can be trained on the training dataset based on the determination that the record size is greater than the threshold size. A predictive model 206 may be trained using automated machine learning operations. Predictions can be obtained from the trained predictive model 206. In one embodiment, when the measurement timestamp of a record occurs before the cutoff, a predictive model 206 can be obtained from the trained predictive model 206. When the measurement timestamp of a record occurs after the cutoff, a predictive model 206 can be trained on the dataset. Also, if the determined record size is less than the threshold size, a predictive model 206 can be obtained from the trained predictive model 206. If the determined record size is greater than the threshold size, the predictive model can be trained on the dataset.

[0107] The display device 112 having the electronic UI 402 is provided merely as an exemplary implementation of the display device 112 in Figure 1 and should not be construed as limiting the scope of this disclosure. This disclosure may also be applicable to other modifications, deletions, or additions to the electronic UI 402 without departing from the scope of this disclosure.

[0108] The embodiments described herein can be used in a number of application areas, such as heating, ventilation and air conditioning (HVAC), manufacturing buildings, warehouses and distribution centers, and power plants.

[0109] Various embodiments of this disclosure may provide one or more non-temporary computer-readable storage media configured to store instructions causing a system (e.g., system 102) to perform an action in response to being executed. Such action may include receiving sensor data about a measurement parameter from one of several sensors 126A-126N installed in a construction environment 104, and storing the received sensor data as a new record in a table of a relational database 108. Such action may further include determining a cutoff record associated with a previous training checkpoint of a predictive model 206 for the measurement parameter, and determining a record containing a new record where each measurement timestamp defined in the table occurs after the measurement timestamp of the cutoff record. Such action may further include comparing the size of the determined record to a threshold size, preparing a training dataset based on the determined record, and training the predictive model 206 on the training dataset.

[0110] As described above, the embodiments described in this disclosure may include the use of a dedicated or general-purpose computer (e.g., processor 202 in Figure 2) including various computer hardware or software modules, as will be described in more detail later. Also, as described above, the embodiments described in this disclosure may be implemented using a computer-readable medium (e.g., memory 204 or relational data 114 or automated ML tool 212 in Figure 2) that carries or stores computer executable instructions or data structures.

[0111] When used in this disclosure, the terms “module” or “component” may refer to a specific hardware implementation configured to perform the operation of that module or component, and / or a software object or software routine stored in and / or executed by general-purpose hardware of a computing system (e.g., computer-readable media, processing devices, etc.). In some embodiments, the various components, modules, engines, and services described in this disclosure may be implemented as multiple objects or processes running on system 102 (e.g., as separate threads). While some of the systems 102 and methods described in this disclosure are generally described as being implemented in software (stored in and / or executed by general-purpose hardware), specific hardware implementations, or combined implementations of software and specific hardware, are also possible and intended. In this description, “computing entity” may be any system 102 as previously defined in this disclosure, or any module or combination of modules running on system 102.

[0112] Furthermore, if there is an intended number of specific claims to be introduced, such intention will be explicitly stated in the claim; if there is no such statement, such intention does not exist. For example, for the sake of understanding, the claims attached below may include the use of the prefatory phrases “at least one” and “one or more” to introduce a claim. However, the use of such phrases should not be interpreted as meaning that the introduction of a claim by the indefinite article “a” or “an” limits any particular claim containing the claim introduced in this way to only one embodiment containing such a claim (for example, “a” and / or “an” should be interpreted as meaning “at least one” or “one or more”). The same applies to the use of the definite article used to introduce a claim.

[0113] Furthermore, any disjunct terms or phrases presenting two or more different terms, whether in the description of an embodiment, the claims, or the drawings, should be understood as intended to include one of those terms, either of those terms, or both. For example, the phrase “A or B” should be understood to include the possibilities of “A” or “B” or “A and B.”

[0114] All examples and conditional statements contained herein are intended for educational purposes to help readers understand the disclosure and concepts provided by the inventors of this application to advance the technology, and should be construed as not limiting to the examples and conditions specifically described herein. While embodiments of this disclosure have been described in detail, these embodiments can be modified, substituted, and altered in various ways without departing from the spirit and scope of this disclosure.

[0115] In relation to the above explanation, the following additional information is disclosed. (Note 1) A method that is performed by at least one processor, Receiving sensor data about measurement parameters from one of the multiple sensors installed in the built environment, The received sensor data is stored as a new record in a relational database table. Determining the cutoff record associated with the previous training checkpoint of the predictive model for the aforementioned measurement parameters, From the aforementioned table, determine the record containing a new record in which each measurement timestamp defined in the table occurs after the measurement timestamp of the cutoff record, The determined record size is compared with the threshold size, Prepare a training dataset based on the record determined above, Based on the above comparison, the predictive model is trained on the training dataset, A method of having. (Note 2) The construction environment is the method described in Note 1, which includes either a data center or an industrial facility. (Note 3) The method according to Note 1, wherein the prediction model is trained on the training dataset based on the determination that the size of the record exceeds the threshold size. (Note 4) Based on the received sensor data and the determination that the size of the record is below the threshold size, prepare the input for the prediction model, Applying the prediction model to the prepared input to generate predicted values ​​for the measurement parameters for future timestamps, Based on the aforementioned predicted values, the knowledge database is queried to determine suggestions for equipment installed in the aforementioned construction environment, The method described in Appendix 1, further comprising the above. (Note 5) The method described in Note 4, wherein the proposal includes actions to carry out the repair or maintenance of the equipment, the inspection of the equipment, or the replacement of the equipment. (Note 6) The above proposal is the method described in Note 4, which indicates whether the predicted value corresponds to the failure state of the equipment. (Appendix 7) The method of Appendix 4, further comprising controlling a display device associated with the administrator of the construction environment to display the predicted values ​​together with the proposal. (Note 8) The aforementioned knowledge database is A state including a plurality of historical values ​​of the measurement parameter and a plurality of timestamps associated with each of the plurality of historical values, Decision information including multiple proposals corresponding to the multiple historical values, and The state and the source type associated with the determination information, The method described in Appendix 4, having the characteristics of the method described in Appendix 4. (Note 9) The query to the aforementioned knowledge database is From the aforementioned plurality of historical values, a historical value that matches the predicted value is determined. Select the proposal corresponding to the historical value from the above multiple proposals. The method described in Appendix 8, which includes the following: (Note 10) Preparing the aforementioned training dataset is From the records of the table, a feature column is extracted that stores the values ​​of the measurement parameters and their respective measurement timestamps. The values ​​in the feature column are sorted based on each of the aforementioned measurement timestamps. The median interval between each of the aforementioned measurement timestamps is calculated. After the sorting, an aggregation query is executed on the feature column to generate an aggregated feature column. Determine the missing measurement timestamps within the aggregated feature column. The missing values ​​corresponding to the missing measurement timestamps in the aggregated feature columns are filled in. Determine the sliding window size for the aggregated feature column. Based on the sliding window size, the training dataset is obtained from the aggregated feature columns. The method described in Appendix 1, which includes the following: (Note 11) The execution of the aggregate query is as follows: The respective measurement timestamps are divided by the calculated interval median to determine multiple aggregate intervals. From the feature column, select a set of values ​​corresponding to each unique aggregation interval of the plurality of aggregation intervals. Calculate the average value of the set of the aforementioned values, The average values ​​are grouped based on the plurality of aggregation intervals, and the aggregated feature column includes the average values ​​corresponding to each unique aggregation interval of the plurality of aggregation intervals. The method described in Appendix 10, which includes the following: (Note 12) The predictive model is trained using automated machine learning operations, as described in Note 1. (Note 13) One or more non-temporary computer-readable storage media configured to store instructions, wherein the instructions cause the system to perform an action in response to being executed, and such action Receiving sensor data about measurement parameters from one of the multiple sensors installed in the built environment, The received sensor data is stored as a new record in a relational database table. Determining the cutoff record associated with the previous training checkpoint of the predictive model for the aforementioned measurement parameters, From the aforementioned table, determine the record containing a new record in which each measurement timestamp defined in the table occurs after the measurement timestamp of the cutoff record, The determined record size is compared with the threshold size, Prepare a training dataset based on the record determined above, Based on the above comparison, the predictive model is trained on the training dataset, One or more non-temporary computer-readable storage media having [a certain characteristic]. (Note 14) The prediction model is trained on the training dataset based on the determination that the size of the record exceeds the threshold size, on one or more non-temporary computer-readable storage media as described in Note 13. (Note 15) The above operation further, Based on the received sensor data and the determination that the size of the record is below the threshold size, prepare the input for the prediction model. Applying the prediction model to the prepared input to generate predicted values ​​for the measurement parameters for future timestamps, Based on the aforementioned predicted values, the knowledge database is queried to determine suggestions for equipment installed in the aforementioned construction environment, One or more non-temporary computer-readable storage media as described in Appendix 13, having the following characteristics: (Note 16) The operation further comprises controlling a display device associated with the administrator of the build environment to display the predicted values ​​together with the proposal, one or more non-temporary computer-readable storage media as described in Note 15. (Note 17) The aforementioned knowledge database is A state including a plurality of historical values ​​of the measurement parameter and a plurality of timestamps associated with each of the plurality of historical values, Decision information including multiple proposals corresponding to the multiple historical values, and The state and the source type associated with the determination information, One or more non-temporary computer-readable storage media as described in Appendix 15, having the following characteristics: (Note 18) The preparation of the training dataset as described above is From the records of the table, a feature column is extracted that stores the values ​​of the measurement parameters and their respective measurement timestamps. The values ​​in the feature column are sorted based on each of the aforementioned measurement timestamps. The median interval between each of the aforementioned measurement timestamps is calculated. After the sorting, an aggregation query is executed on the feature column to generate an aggregated feature column. Determine the missing measurement timestamps within the aggregated feature column. The missing values ​​corresponding to the missing measurement timestamps in the aggregated feature columns are filled in. Determine the sliding window size for the aggregated feature column. Based on the sliding window size, the training dataset is obtained from the aggregated feature columns. One or more non-temporary computer-readable storage media as described in Appendix 13, having the following characteristics: (Note 19) The execution of the aggregate query is as follows: The respective measurement timestamps are divided by the calculated interval median to determine multiple aggregate intervals. From the feature column, select a set of values ​​corresponding to each unique aggregation interval of the plurality of aggregation intervals. Calculate the average value of the set of the aforementioned values, The average values ​​are grouped based on the plurality of aggregation intervals, and the aggregated feature column includes the average values ​​corresponding to each unique aggregation interval of the plurality of aggregation intervals. One or more non-temporary computer-readable storage media as described in Appendix 18, having the following characteristics: (Note 20) Memory storing instructions, and A processor coupled to the aforementioned memory, which executes the aforementioned instructions, Receiving sensor data about measurement parameters from one of the multiple sensors installed in the built environment, The received sensor data is stored as a new record in a relational database table. Determining the cutoff record associated with the previous training checkpoint of the predictive model for the aforementioned measurement parameters, From the aforementioned table, determine the record containing a new record in which each measurement timestamp defined in the table occurs after the measurement timestamp of the cutoff record, The determined record size is compared with the threshold size, Prepare a training dataset based on the record determined above, Based on the above comparison, the predictive model is trained on the training dataset, A processor that executes a process having A system that has

Claims

1. A method that is executed by at least one processor, Receiving sensor data about measurement parameters from one of the multiple sensors installed in the built environment, The received sensor data is stored as a new record in a relational database table. Determining the cutoff record associated with the previous training checkpoint of the predictive model for the aforementioned measurement parameters, From the aforementioned table, determine the record containing a new record in which each measurement timestamp defined in the table occurs after the measurement timestamp of the cutoff record, The determined record size is compared with the threshold size, Prepare a training dataset based on the record determined above, Based on the above comparison, the predictive model is trained on the training dataset, A method of having.

2. The method according to claim 1, wherein the construction environment includes either a data center or an industrial facility.

3. The method according to claim 1, wherein the predictive model is trained on the training dataset based on the determination that the size of the record exceeds the threshold size.

4. Based on the received sensor data and the determination that the size of the record is below the threshold size, prepare the input for the prediction model. Applying the prediction model to the prepared input to generate predicted values ​​for the measurement parameters for future timestamps, Based on the aforementioned predicted values, the knowledge database is queried to determine suggestions for equipment installed in the aforementioned construction environment, The method according to claim 1, further comprising:

5. The method according to claim 4, wherein the proposal includes actions for performing repair or maintenance of the equipment, maintenance and inspection of the equipment, or replacement of the equipment.

6. The above proposal is the method according to claim 4, which indicates whether the predicted value corresponds to a failure state of the equipment.

7. The method according to claim 4, further comprising controlling a display device associated with the administrator of the construction environment to display the predicted values ​​together with the proposal.

8. The aforementioned knowledge database is A state including a plurality of historical values ​​of the measurement parameter and a plurality of timestamps associated with each of the plurality of historical values, Decision information including multiple proposals corresponding to the multiple historical values, and The state and the source type associated with the determination information, The method according to claim 4, having the following characteristics.

9. The querying of the aforementioned knowledge database is From the aforementioned plurality of historical values, a historical value that matches the predicted value is determined. Select the proposal corresponding to the historical value from the above multiple proposals. The method according to claim 8, wherein the above is achieved.

10. Preparing the aforementioned training dataset is From the records of the table, a feature column is extracted that stores the values ​​of the measurement parameters and their respective measurement timestamps. The values ​​in the feature column are sorted based on each of the aforementioned measurement timestamps. The median interval between each of the aforementioned measurement timestamps is calculated. After the sorting, an aggregation query is executed on the feature column to generate an aggregated feature column. Determine the missing measurement timestamps within the aggregated feature column. The missing values ​​corresponding to the missing measurement timestamps in the aggregated feature columns are filled in. Determine the sliding window size for the aggregated feature column. Based on the sliding window size, the training dataset is obtained from the aggregated feature columns. The method according to claim 1, wherein the method is as follows:

11. The execution of the aggregate query is as follows: The respective measurement timestamps are divided by the calculated interval median to determine multiple aggregate intervals. From the feature column, select a set of values ​​corresponding to each unique aggregation interval of the plurality of aggregation intervals. Calculate the average value of the set of the aforementioned values, The average values ​​are grouped based on the plurality of aggregation intervals, and the aggregated feature column includes the average values ​​corresponding to each unique aggregation interval of the plurality of aggregation intervals. The method according to claim 10, wherein the above is achieved.

12. The method according to claim 1, wherein the predictive model is trained using automated machine learning operations.

13. One or more non-temporary computer-readable storage media configured to store instructions, wherein the instructions cause the system to perform an action in response to being executed, and such action Receiving sensor data about measurement parameters from one of the multiple sensors installed in the built environment, The received sensor data is stored as a new record in a relational database table. Determining the cutoff record associated with the previous training checkpoint of the predictive model for the aforementioned measurement parameters, From the aforementioned table, determine the record containing a new record in which each measurement timestamp defined in the table occurs after the measurement timestamp of the cutoff record, The determined record size is compared with the threshold size, Prepare a training dataset based on the record determined above, Based on the above comparison, the predictive model is trained on the training dataset, One or more non-temporary computer-readable storage media having [a certain characteristic].

14. The memory that stores the instructions, and A processor coupled to the aforementioned memory, which executes the aforementioned instructions, Receiving sensor data about measurement parameters from one of the multiple sensors installed in the built environment, The received sensor data is stored as a new record in a relational database table. Determining the cutoff record associated with the previous training checkpoint of the predictive model for the aforementioned measurement parameters, From the aforementioned table, determine the record containing a new record in which each measurement timestamp defined in the table occurs after the measurement timestamp of the cutoff record, The determined record size is compared with the threshold size, Prepare a training dataset based on the record determined above, Based on the above comparison, the predictive model is trained on the training dataset, A processor that executes a process having A system that has