Systems and methods for assessing and repairing data using data quality metrics
By employing configurable pipelines with data quality indicators and aggregation processing components in building automation systems, sensor data quality issues are automatically addressed, resolving problems of missing sensor data and outliers, thereby improving the system's data accuracy and operational efficiency.
Patent Information
- Application Number
- CN202180051359.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-21
- Filing Date
- 2021-07-26
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2041-07-26
AI Technical Summary
Data quality issues generated by sensors in building automation systems, including missing data, misaligned timestamps, and outliers, lead to inaccurate and inefficient system operation, requiring extensive manual intervention for data filtering and processing.
It employs Data Quality Indicators (DQI) and Data Quality Aggregation (DQA) processing components to perform data quality analysis and remediation through configurable pipelines and adapters, including dynamically building configurable pipelines, performing data quality checks, interpolation and aggregation methods, and using the DQ core library and feature engine for automated data quality assessment and remediation.
It enables automated data quality analysis and repair, reduces manual intervention, improves data accuracy and system operation efficiency, and lowers maintenance costs.
Smart Images

Figure CN116097628B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the analysis and repair of physical sensor data in building control systems and other systems. Background Technology
[0002] Building automation systems encompass a variety of systems that facilitate the monitoring and control of various aspects of building operations. These include security systems, fire safety systems, lighting systems, and HVAC systems. The components of a building automation system are widely distributed throughout the facility. For example, an HVAC system may include temperature sensors and ventilation damper controllers, as well as other components virtually located in every area of the facility. These building automation systems typically have one or more centralized control stations from which system data can be monitored and various aspects of system operation can be controlled and / or monitored.
[0003] To allow for the monitoring and control of distributed control system components, building automation systems typically employ multi-level communication networks to transmit operational and / or alarm information between operating elements such as sensors and actuators and a central control station. An example of a building automation system controller is the DXR controller, available from Siemens Industries, Inc. (“Siemens”), Building Technologies, Buffalo Grove, Illinois. In this system, several control stations connected via Ethernet or other types of networks can be distributed across one or more building locations, each with the ability to monitor and control system operations.
[0004] To ensure the proper operation of building automation or control systems, it is crucial that the data generated by physical sensors and other devices accurately reflect the state of a specific building automation or control system or subsystem (such as the variable air volume subsystem of an HVAC system). Raw data with sufficient data quality (DQ) is required for building control processing and related data science algorithms, such as fault detection and machine learning. Due to hardware and software limitations, the raw data associated with specific sensors, devices, or subsystems captured for specific building control processing often contains problems, including but not limited to missing data, misaligned timestamps, and outlier results. Such problems typically require skilled technicians, engineers, and / or knowledgeable data scientists to spend significant time manually filtering the raw data to identify and resolve these issues, resulting in substantial cost and time commitment. Therefore, improved systems are needed. Summary of the Invention
[0005] This disclosure describes systems and methods for using data quality metrics to assess and repair data, particularly for building automation systems.
[0006] According to one embodiment, a method performed by a data processing system includes receiving input data representing the operation of physical devices of a building automation system. The method includes receiving a configuration file defining data quality (DQ) processing to be performed on the input data. The method includes the data processing system dynamically constructing a configurable pipeline based on the configuration file, the pipeline including one or more data quality indicators (DQI) or data quality aggregation (DQA) processing components from a DQ core library. The method includes the data processing system performing DQ processing on the input data, including executing each DQI or DQA processing component included in the pipeline. The method includes generating one or more DQ results based on the DQ processing. The method includes the data processing system returning one or more DQ results. This method can be performed by a data processing system or controller that is part of or communicates with a building automation system.
[0007] In various embodiments, the method is executed with a software architecture including a pipeline generator that constructs a configurable pipeline. In various embodiments, the method is executed with a software architecture including multiple adapters configured to transform data for use by one or more DQI or DQA processing components. In various embodiments, DQ processing includes generating DQ flags based on domain knowledge using fixed and fuzzy logic. In various embodiments, DQ processing includes performing energy meter overflow checks. In various embodiments, DQ processing includes performing a summarization method processing of DQ data from sensor points based on different time-domain aggregation patterns. In various embodiments, the configuration file includes the definition of a configurable pipeline, multiple DQI and DQA processing components to be executed in series and / or parallel, and connections between the multiple DQI and DQA processing components. In various embodiments, the pipeline includes a DQA processing component using one of a weighted average of DQA, the maximum DQI of DQA, the time horizon averaging of DQA, and the time horizon maximum of DQA. In various embodiments, the configuration file includes an identifier of the pattern to be applied, an identifier of the aggregation method, and an identifier of the interpolation method, and the pattern includes data quality metrics with associated weights.
[0008] In various embodiments, dynamically building a configurable pipeline includes initializing the configurable pipeline and reading patterns associated with a configuration file. In such embodiments, dynamically building a configurable pipeline includes selectively adding at least one quality check method to the configurable pipeline based on patterns and configuration files, wherein the quality check method is from a DQI processing component of the DQ core library. In such embodiments, dynamically building a configurable pipeline includes selectively adding at least one interpolation method to the configurable pipeline based on patterns and configuration files. In such embodiments, dynamically building a configurable pipeline includes selectively adding a flag assignment method to the configurable pipeline based on patterns and configuration files. In such embodiments, dynamically building a configurable pipeline includes selectively adding an aggregation method to the configurable pipeline based on patterns and configuration files, wherein the aggregation method is from a DQA processing component of the DQ core library. In such embodiments, dynamically building a configurable pipeline includes storing the configurable pipeline.
[0009] Disclosed embodiments include a building automation system comprising a plurality of sensors and at least one data processing system configured to process input data collected from the operation of at least one of the plurality of sensors and perform the processing as described herein. Disclosed embodiments also include a non-transitory machine-readable medium encoded with executable instructions that, when executed, cause at least one processor in the building automation system to perform the processing as described herein.
[0010] The foregoing has provided a fairly broad overview of some features and technical advantages of this disclosure, enabling those skilled in the art to better understand the following detailed descriptions. Additional features and advantages of this disclosure that form the subject matter of the claims will be described below. Those skilled in the art will understand that they can readily use the disclosed concepts and specific embodiments as the basis for modifying or designing other structures to achieve the same purpose of this disclosure. Those skilled in the art will also recognize that such equivalent constructions, in their broadest forms, do not depart from the spirit and scope of this disclosure.
[0011] Before proceeding with the detailed embodiments below, it may be advantageous to clarify the definitions of certain words or phrases used throughout this patent document: the terms "include" and "comprise," and their derivatives, mean to include without limitation; the term "or" is inclusive, meaning and / or; the phrases "associated with" and "associated therewith," and their derivatives, can mean to include, be included in, interconnected with, contain, be contained, connected to or connected with, coupled to or coupled with, communicate with, cooperate with, interleave, juxtapose, approach, bind to or bind with, have its characteristics, etc.; and the term "controller" means any device, system, or part thereof that controls at least one operation, whether such device is implemented in hardware, firmware, software, or some combination of at least two of these. It should be noted that the functionality associated with any particular controller can be centralized or distributed, whether local or remote. Various definitions of certain words and phrases are provided in this patent document, and those skilled in the art will understand that such definitions apply in many (if not most) instances to the prior and future use of such defined words and phrases. While some terms may include multiple embodiments, the appended claims expressly limit these terms to the specific embodiments. Attached Figure Description
[0012] To gain a more complete understanding of this disclosure and its advantages, reference is now made to the following description in conjunction with the accompanying drawings, wherein the same numerals denote the same objects, and wherein:
[0013] Figure 1 A block diagram of a building automation system according to this disclosure is shown, wherein data quality of heating, ventilation and air conditioning (HVAC) systems or other systems can be improved;
[0014] Figure 2 The following is shown in accordance with this disclosure: Figure 1 A detail of a field panel;
[0015] Figure 3 The following is shown in accordance with this disclosure: Figure 1 Details of a field controller;
[0016] Figure 4 Examples of elements of a software architecture that can be used to implement exposed processing are shown;
[0017] Figure 5A A non-limiting example of a pattern used in conjunction with a configuration file according to the disclosed embodiments is shown;
[0018] Figure 5BA non-limiting example of a configuration file according to the disclosed embodiments is shown;
[0019] Figure 6 An example of data quality aggregation processing according to a disclosed embodiment is shown;
[0020] Figure 7 An example of out-of-scope DQI functionality according to a disclosed embodiment is shown;
[0021] Figure 8 An example of single-schedule quantity calculation using sensor data with monotonic holders is shown according to a disclosed embodiment;
[0022] Figure 9 , Figure 10 , Figure 11 and Figure 12 An example of processing according to the disclosed embodiments is shown;
[0023] Figure 13A and Figure 13B Examples of DQ results according to the disclosed embodiments are shown; and
[0024] Figure 14 A block diagram of a data processing system that can be implemented in various embodiments is shown. Detailed Implementation
[0025] The accompanying drawings and various embodiments used to describe the principles of this disclosure in this patent document are for illustrative purposes only and should not be construed as limiting the scope of this disclosure in any way. Those skilled in the art will understand that the principles of this disclosure can be implemented in any suitably arranged device. Many of the innovative teachings of this application will be described with reference to exemplary, non-limiting embodiments.
[0026] The building automation system (BAS) disclosed herein can help operate the system efficiently in a space to save energy in an automated operating mode. The BAS continuously assesses environmental conditions and energy use in the space and can determine when the space is being operated most efficiently and indicate this to the user. Similarly, the BAS can identify and indicate when the system is operating inefficiently, such as because occupants exceed room controls associated with a specific space in the building due to personal preferences or due to drastic changes in weather conditions. The BAS can automatically or at user input adjust control settings to make the system operate efficiently again.
[0027] For a Building Automation System (BAS) to operate correctly, it collects data from numerous sensors and other devices located in rooms within the building and its systems, such as those that are part of heating, cooling, or ventilation equipment, or otherwise distributed throughout the building and its systems. Low-quality data, such as missing data points or data points that do not accurately reflect the condition of the building or system, can lead to incorrect or inefficient operation. Note that, as used herein, “sensor” is intended to include any physical device that collects data that can be processed as described herein, whether it is an electricity meter, water meter, thermostat, airflow sensor, or other device.
[0028] The disclosed embodiments include systems and methods for automatically analyzing the quality of data being processed and correcting the data to ensure the proper operation of the BAS.
[0029] Figure 1 A block diagram of a building automation system 100 that can implement the disclosed embodiments is shown. The building automation system 100 is an environmental control system configured to control at least one of a plurality of environmental parameters (such as temperature, humidity, lighting, etc.) within a building. For example, in a particular embodiment, the building automation system 100 may include a DXR controller, such as acting as a field controller or panel controller in the building automation system, which allows setting and / or changing various controls of the system. While a brief description of the building automation system 100 is provided below, it will be understood that the building automation system 100 described herein is merely one example of a particular form or configuration of a building automation system, and the system 100 may be implemented in any other suitable manner without departing from the scope of this disclosure.
[0030] For the illustrated embodiment, the building automation system 100 includes a field controller 102, a reporting server 104, multiple client stations 106a-c, multiple field panels 108a-b, multiple field controllers 110a-e, and multiple field devices 112a-d. Although three client stations 106, two field panels 108, five field controllers 110, and four field devices 112 are shown, it will be understood that the system 100 may include any suitable number of these components 106, 108, 110, and 112 based on the specific configuration of a particular building.
[0031] A field controller 102, which may include a computer or a general-purpose processor, is configured to provide overall control and monitoring of the building automation system 100. The field controller 102 can operate as a data server capable of exchanging data with various components of the system 100. Thus, the field controller 102 can allow for control and monitoring by a computer located on the field controller 102 or another management computer. Figure 1 Various applications (not shown in the image) run on the system to access system data.
[0032] For example, the field controller 102 can communicate with other management computers, Internet gateways, or other gateways via a management level network (MLN) 120 to connect to other external devices and to an additional network manager (which in turn can connect to more subsystems via an additional low-level data network). The field controller 102 can use the MLN 120 to exchange system data with other components on the MLN 120, such as a report server 104 and one or more client stations 106. The report server 104 can be configured to generate reports on various aspects of the system 100. Each client station 106 can be configured to communicate with the system 100 to receive information from and / or provide modifications to the system 100 in any appropriate manner. The MLN 120 may include Ethernet or a similar wired network and may employ TCP / IP, BACnet, and / or other protocols that support high-speed data communication.
[0033] The field controller 102 can also be configured to accept modifications and / or other inputs from a user. This can be achieved via the user interface of the field controller 102 or any other user interface that can be configured to communicate with the field controller 102 via any suitable network or connection. The user interface may include a keyboard, touchscreen, mouse, or other interface components. The field controller 102 is particularly configured to influence or change the operational data of the field panel 108 and other components of the system 100. The field controller 102 can use the building network (BLN) 122 to exchange system data with other elements on the BLN 122, such as the field panel 108.
[0034] Each field panel 108 may include a general-purpose processor and is configured to provide control over one or more corresponding field controllers 110 using data and / or instructions from the field controller 102. While the field controller 102 is typically used to modify one or more of the various components of the building automation system 100, the field panel 108 is also capable of providing certain modifications to one or more parameters of the system 100. Each field panel 108 may use a field-level network (FLN) 124 to exchange system data with other elements on the FLN 124, such as a subset of the field controllers 110 coupled to the field panel 108.
[0035] Each field controller 110 may include a general-purpose processor and may correspond to one of several localized standard building automation subsystems, such as a building space temperature control subsystem, a lighting control subsystem, etc. For a particular embodiment, the field controller 110 may include a DXR controller available from Siemens. However, it should be understood that the field controller 110 may include any other suitable type of controller without departing from the scope of the invention.
[0036] To perform control of its respective subsystem, each field controller 110 may be coupled to one or more field devices 112. Each field controller 110 is configured to provide control of its one or more corresponding field devices 112 using data and / or instructions from its corresponding field panel 108. In some embodiments, some field controllers 110 may control their subsystems based on sensed conditions and desired setpoint conditions. In these embodiments, these field controllers 110 may be configured to control the operation of one or more field devices 112 in an attempt to bring the sensed conditions to the desired setpoint conditions. Note that in system 100, information from the field devices 112 may be shared between the field controllers 110, field panels 108, field controllers 102, and / or any other elements on or connected to system 100.
[0037] To facilitate information sharing among subsystems, subsystems can be grouped into FLN 124. For example, subsystems corresponding to field controllers 110a and 110b can be coupled to field panel 108a to form FLN 124a. Each FLN 124 may include a low-level data network that can employ any suitable proprietary or open protocol.
[0038] Each field device 112 can be configured to measure, monitor, and / or control various parameters of the building automation system 100. Examples of field devices 112 include lights, thermostats, temperature sensors, lighting sensors, fans, damper actuators, heaters, coolers, alarms, HVAC equipment, louver controllers and sensors, and many other types of field devices. Field devices 112 are capable of receiving and / or sending signals to the field controller 110, field panel 108, and / or field controller 102 of the building automation system 100. Therefore, the building automation system 100 can control various aspects of building operation by controlling and monitoring the field devices 112. Specifically, each or any of the field devices 112 can generate data that is processed as described herein.
[0039] like Figure 1 As shown, any field panel 108, such as field panel 108a, can be directly coupled to one or more field devices 112, such as field devices 112c and 112d. For this type of embodiment, field panel 108a can be configured to provide direct control of field devices 112c and 112d, rather than control via one of field controllers 110a and 110b. Therefore, for this embodiment, the functionality of field controller 110 for one or more specific subsystems can be provided by field panel 108, eliminating the need for field controller 110.
[0040] Figure 2Details of a field panel 108 according to this disclosure are shown. For this particular embodiment, the field panel 108 includes a processor 202, a memory 204, an input / output (I / O) module 206, a communication module 208, a user interface 210, and a power module 212. The memory 204 includes any suitable data memory capable of storing data, such as instruction 220 and a database 222. It should be understood that the field panel 108 can be implemented in any other suitable manner without departing from the scope of this disclosure.
[0041] Processor 202 is configured to operate field panel 108. Therefore, processor 202 can be coupled to other components 204, 206, 208, 210, and 212 of field panel 108. Processor 202 can be configured to execute program instructions or programmed software or firmware, such as BAS application software 230, stored in instructions 220 in memory 204. In addition to storing instructions 220, memory 204 can also store other data for use by system 100 in database 222, such as various records and configuration files, graphical views, and / or other information. For example, memory 204 can store DQ software architecture 402, which performs data quality processing as described herein, as described in more detail below.
[0042] Execution of BAS application 230 by processor 202 may result in control signals being sent to any field device 112 that may be coupled to field panel 108 via I / O module 206 of field panel 108. Execution of BAS application 230 may also result in processor 202 receiving status signals and / or other data signals from field devices 112 coupled to field panel 108, storing the associated data in memory 204, and processing the data as described herein. In one embodiment, BAS application 230 may be provided by or implemented with a DXR controller commercially available from Siemens Industries. However, it should be understood that BAS application 230 may include any other suitable BAS control software.
[0043] I / O module 206 may include one or more input / output circuits configured to communicate directly with field device 112. Therefore, in some embodiments, I / O module 206 includes analog input circuitry for receiving analog signals and analog output circuitry for providing analog signals.
[0044] Communication module 208 is configured to provide communication with field controller 102, other field panels 108, and other components on BLN 122. Communication module 208 is also configured to provide communication to field controller 110 and other components on FLN 124 associated with field panel 108. Therefore, communication module 208 may include a first port coupled to BLN 122 and a second port coupled to FLN 124. Each port may include RS-485 standard port circuitry or other suitable port circuitry.
[0045] Field panel 108 can be locally accessed via interactive user interface 210. A user can control the collection of data from field device 112 through user interface 210. User interface 210 for field panel 108 may include devices for displaying data and receiving input data. These devices may be permanently fixed to field panel 108 or may be portable and movable. In some embodiments, user interface 210 may include an LCD screen and a keypad. User interface 210 can be configured to both change and display information about field panel 108 (such as status information and / or modifications to field panel 108) and the field panel 108 itself.
[0046] The power module 212 can be configured to supply power to the components of the field panel 108. The power module 212 can operate with standard 120-volt AC power, other AC voltages, or DC power supplied by one or more batteries.
[0047] Figure 3 Details of a field controller 110 according to this disclosure are shown. For this particular embodiment, the field controller 110 includes a processor 302, a memory 304, an input / output (I / O) module 306, a communication module 308, and a power supply module 312. In some embodiments, the field controller 110 may also include a user interface (…). Figure 3 (Not shown in the image), the user interface is configured to change and / or display information about the field controller 110. The memory 304 includes any suitable data storage device capable of storing data, such as instructions 320 and a database 322. It should be understood that the field controller 110 can be implemented in any other suitable manner without departing from the scope of this disclosure. In some embodiments, the field controller 110 may be located in a room of a building or in a room adjacent to a building, in which the temperature or another environmental parameter associated with the subsystem can be controlled by the field controller 110.
[0048] Processor 302 is configured to operate field controller 110. Therefore, processor 302 can be coupled to other components 304, 306, 308, and 312 of field controller 110. Processor 302 can be configured to execute program instructions or programmed software or firmware, such as subsystem application software 330, stored in instructions 320 in memory 304. For a particular example, subsystem application 330 may include a temperature control application configured to control and process data from all components of the temperature control subsystem, such as temperature sensors, damper actuators, fans, and various other field devices. In addition to storing instructions 320, memory 304 may also store other data for use by the subsystem, such as various configuration files and / or other information, in database 322. For example, memory 304 may store DQ software architecture 402, which performs data quality processing as described herein, as described in more detail below.
[0049] Execution of subsystem application 330 by processor 302 may cause control signals to be sent to any field device 112 that may be coupled to field controller 110 via I / O module 306. Execution of subsystem application 330 may also cause processor 302 to receive status signals and / or other data signals from field device 112 coupled to field controller 110 and store the associated data in memory 304.
[0050] I / O module 306 may include one or more input / output circuits configured to communicate directly with field device 112. Therefore, in some embodiments, I / O module 306 includes analog input circuitry for receiving analog signals and analog output circuitry for providing analog signals.
[0051] Communication module 308 is configured to provide communication with other components (such as other field controllers 110) on the field panel 108 corresponding to field controller 110 and on FLN 124. Therefore, communication module 308 may include a port coupling to FLN 124. This port may include RS-485 standard port circuitry or other suitable port circuitry.
[0052] The power module 312 can be configured to supply power to the components of the field controller 110. The power module 312 can operate with standard 120-volt AC power, other AC voltages, or DC power supplied by one or more batteries.
[0053] Manual inspection of BAS data for data quality (DQ) analysis is practically impossible in many implementations. Theoretically, engineers can perform this inspection using graphing software (such as...). Open the data file in a spreadsheet program and plot the data curve. Peaks or outliers in the curve may indicate a DQ problem and suggest the need for further action. However, finding such problems in volume data generated by a BAS system is daunting or impossible. Building a cluster is characterized by a large number of sensors: customers may have several reporting groups. Each reporting group includes a large number of HVAC units. Each HVAC unit contains multiple sensors. Finally, to examine all the data, engineers will need to... 4 Plotting curves on the order of 106 sensor readings is not something that can actually be done manually.
[0054] The disclosed embodiments include systems and methods for automatically analyzing BAS data to perform data quality checks, identify defective or low-quality data points, repair data, aggregate data, identify problematic system equipment and processes, and otherwise analyze and process BAS data to enable system analysis and maintenance.
[0055] These processes offer significant technical improvements over existing attempts to manually identify problems and enable data quality analysis that was previously impossible or impractical.
[0056] The disclosed embodiments include systems and methods for performing automated DQ checks and data cleaning.
[0057] Figure 4 An example of a DQ software architecture 402 that can be used to implement the disclosed processing in a data processing system 400 is shown. The data processing system 400 may be, for example, an implementation of a field controller data processing system 102, a client station 106, a report server 104, or another client or server data processing system or controller configured to operate as disclosed herein. In various embodiments, the DQ software architecture 402 may be used in the memory 204 of a field panel 108 communicating with a BAS application 230, or in the memory 304 of a field controller 110 communicating with a subsystem application 330, for a local implementation of the data quality processing and system described in detail herein. The DQ software architecture 402 described herein is exemplary and not restrictive; particular implementations may use alternative architectural components to perform similar functionality, may use different names to refer to various components, may combine or partition various operations differently relative to different components, or may otherwise use different logical structures to perform the processing as described herein, and the scope of this disclosure is intended to cover such variations. For example, in some implementations, components such as the DQ core library 410, the configurable feature engine 406, the analytics application 416, and the configurable pipeline 422 can be implemented together as the DQ core application.
[0058] DQ software architecture 402 is an example of an architecture for processing Data Quality Indicators (DQI) and Data Quality Aggregation (DQA) data from input data such as sensor or instrument data. The DQI disclosed herein provides a quantification of data quality with a unified and flexible metric. The DQA aggregated DQI disclosed herein allows users to “zoom in” and “zoom out” along temporal and spatial ranges, enabling systems and users to quickly identify data with quality issues. In some cases, DQI methods produce DQI values for given input data (e.g., one week of time-series data from a specific energy instrument) and / or results from quality checking methods / algorithms of feature engine 406; in other cases, DQA methods create DQA values composed of or derived from a number of calculated DQIs to give a comprehensive view of multiple devices, buildings, facilities, or other multiple input datasets. Note that the terms “method” and “algorithm” can be used to describe one or more specific processing components performed or executed by a data processing system.
[0059] DQ software architecture 402 may include a configurable feature engine (FE) 406 to apply or compute different features defined by configuration file 408. In various embodiments, the FE includes features (or methods) that perform data quality algorithms or metrics. FE methods may be added along with DQI and DQA processing to a configurable pipeline as described herein. In various embodiments, FE features are applied to raw data, such as time-series data, that can be received or retrieved by the system. DQI and DQA processing are used to evaluate and aggregate the processed data.
[0060] Using a time-series input, the FE406 can compute various statistical features, including mean, maximum, minimum, and other features, as described in more detail below. The output of the FE406 can include scalars or vectors defined in the corresponding "feature" plugins within the FE406. For example, the "mean" feature is computed by a plugin within the FE406. In various embodiments, the FE406 provides a set of standard features and can allow the integration or linking of third-party features. The FE406 can process and extract data from a single signal / sensor simultaneously, or it can process and extract data from multiple signals / sensors simultaneously. The feature engine 406 not only allows DQI but also features from third-party components. For example, other components can add features as labels on top of sensor data, such as references. Figure 13A As described and shown.
[0061] The FE406 can also perform basic quality checks, such as zero-value checks, outlier checks to identify abnormal data points, boundary checks to identify data exceeding user-defined or internally estimated boundaries (i.e., out of range), and energy meter overflow checks. Energy meter overflow checks can also be implemented as a flag assignment module or flag engine that can detect energy meter overflow events and flag the sensor data, as described below. Other discrete quality checks that can be performed by the feature engine 406 can include identifying empty / missing data points, identifying incomplete or incorrectly formatted data points, identifying data that does not conform to user-defined or calculated time frequencies, and identifying energy consumption meter-specific data problems, such as negative energy consumption and meter overflow.
[0062] The DQ software architecture 402 may include a DQ core library 410, which includes configurable Data Quality Indicators (DQI) and Data Quality Aggregation (DQA) computation components. Each of these components can perform specific DQI or DQA processing on the corresponding input data. The use of these configurable components will be discussed in more detail below. The DQ core library 410 may include configurable components such as machine learning (ML), artificial intelligence (AI), statistical, and rule-based analytics components.
[0063] The DQ core library 410 includes multiple data quality inspection components that will be executed as part of a reconfigurable pipeline as described below. Each DQI or DQA inspection may have a defined function or operation that can process each distinct data point available in the input sensor data 412 and generate, for each such distinct data point available in the input sensor data 412, an output value such as binary (0 or 1) or normalized (between 0 and 1). Inspections may include, but are not limited to, identifying empty / missing data points, identifying incomplete or incorrectly formatted data points, identifying data exceeding user-defined or internally estimated boundaries (i.e., out-of-range data), identifying data that does not conform to a user-defined or calculated time frequency, identifying outlier data points, and identifying energy consumption meter-specific data problems, such as negative energy consumption and meter overflow.
[0064] The DQ core library 410 may include a "summarization method" process that can be added to pipeline 422. This summarization method can generate DQ data for one or more sensor points in any sequential combination defined in configuration file 408. The summarization method can also analyze the DQ data of sensor points based on different time-domain aggregation patterns, such as daily, weekly, or monthly DQI.
[0065] DQA processing can aggregate DQI information, as described in more detail below.
[0066] DQ software architecture 402 can receive, load, store, and otherwise process sensor data 412. Sensor data 412 can be “real-time” sensor data received in a BAS system as described herein or from another system, can be stored data previously received from such a system, or can be other sensor data such as test data or calibration data. In various embodiments, sensor data 412 can be accessed by any or all other elements of DQ software architecture 402.
[0067] Sensor data 412 may include or be associated with DQ metadata 414 (such as tags, parameters, or other information that defines or describes sensor data 412). DQ metadata is a tag on the raw sensor data.
[0068] DQ software architecture 402 may include one or more analytics applications (APPs) 416, each of which may execute one or more analytics algorithms / processes in a configurable pipeline sequence 422 as described herein. Each analytics algorithm may be, but is not required to be, a DQI or DQA process derived from features of DQ core library 410 or FE 406. The system for analytics applications may include configurable pipeline sequences 422 and is selected to be included in the pipeline based on configuration file 408.
[0069] The DQ software architecture 402 may include one or more adapters 418. Adapters 418 may be used to convert or adapt data between various computing components in the DQ core library 410, APP 416, sensor data 412 and associated DQ metadata 414 as needed, and otherwise perform the processing described herein as needed.
[0070] The DQ software architecture 402, executing on the data processing system 400, can take sensor data 412 as input and generate one or more DQIs. Each DQI, whether generated by individual APPs 416 in the configurable pipeline 422 or as the final result of the configurable pipeline 422, can be generated as a normalized index within [0, 1], where 0 is the best quality and 1 is the worst data quality. In other implementations, the DQI definition may be reversed (e.g., 1 represents the best quality) or different value ranges may be used. The DQ software architecture 402 uses adjustable weights to aggregate several DQIs into a DQA. Configurable parameters such as DQI definitions, weight values, and DQA definitions can be stored in a configuration file 408. As described herein, DQIs can be aggregated into DQAs. The contents, generation, and use of the configuration file 408 are described in more detail below.
[0071] According to the disclosed embodiments, the pipeline generator 420 of the DQ software architecture 402 can build a configurable pipeline 422 in the analytics application 416 to perform DQI and DQA as needed for a specific purpose, such as a combination of data quality inspection, aggregation, indexing, and interpolation processing. Exemplary processes for generating the configurable pipeline 422 using configuration file 408 are described in detail below.
[0072] For example, configuration file 408 may include definitions of processing pipelines for the DQI and DQA processing components to be executed, such as those stored in the DQ core library 410. Configurable pipelines may include multiple DQI and DQA processing components executed in series and / or in parallel, depending on user needs. Users can specify the connections between algorithms and processing components and define processes that run in parallel or are executed serially; user specifications can be stored in configuration file 408. Note that while the exemplary configurable pipeline 422 is shown only as having two DQ algorithms / processes, a given pipeline 422 generated by pipeline generator 420 may have any number of DQ algorithms / processes, and any of these can be executed in series or in parallel with other algorithms / processes that can be defined by the pipeline configuration. Configurable pipeline 422 allows analysts to modify DQI processing and aggregation without changing the code.
[0073] To accommodate multiple configurable algorithms running in parallel with various possible data sensor data sources 412, the precise operations taken by the DQ software architecture 402 can be dynamically determined at runtime by the pipeline generator 420 and can be influenced by the order in which data enters. The configurable pipeline 422 can be generated based on the data context, metadata, and defined client requirements or other factors that can be defined in configuration file 408 or other configuration files or parameters. In this way, the pipeline generator 420 can dynamically generate the data quality analysis processing pipeline 422, which may be unique to the usage of the data being analyzed.
[0074] Pipeline generator 420 can be combined with numerous DQ algorithms / processes for a specific sensor point. For example, for point 1, pipeline 422 may include a gap checker, a frequency checker, etc. For point 2, pipeline 422 may include an outlier checker, a monotonicity checker, etc. Unlike other analysis tools, pipeline 422 is not hard-coded by data scientists or other developers who are unaware of the pipeline configuration during the development phase. This configuration can be defined by application engineers or users and captured in configuration file 408 after the developers have completed their code. Based on configuration file 408 for a specific implementation, pipeline generator 420 can build pipeline 422 at runtime, as described in further detail herein.
[0075] Configuration file 408 can specify information such as feature plugins DQI and DQA, connections between DQ feedback components to be executed, and parameters for each individual component to be executed. Configuration file 408 is flexible, allowing the source code of DQ components, features, etc., to be modified to build and execute the desired processing pipeline 422. Configuration file 408 can specify the association between sensor / input data and applicable DQI, DQA, or other processes. Configuration file 408 can include default pipeline configurations for specific types of data or specific data sources and can be edited as described herein. Furthermore, in the processes described herein, specific operations included in the configuration file can be excluded based on user input during or prior to execution.
[0076] Unlike existing methods that require reprogramming the source code to perform DQ analysis functions for different input data, configuration file 408, as described in this article, allows application engineers to work independently without adjusting the DQ core source code for individual applications (such as different buildings, manufacturing systems, or other specific input data sources).
[0077] More broadly, the disclosed embodiments enable individuals with different roles to jointly define data analysis processes. The end users of the disclosed data quality system (or the software implementing the processes described herein) are typically the "owners" responsible for the operation of the data being analyzed. For example, these users could be those individuals responsible for creating reports and analyses using BAS data, but this could be anyone involved in maintaining the quality and integrity of that data.
[0078] An application engineer can refer to the person responsible for setting up such a system to analyze data from a specific building or facility, and can be an individual who specifies the scope of quality checks to be performed for that building through a configuration file. Application engineers can also make changes to the configuration as needed, such as at the request of users or when new capabilities are added to the system.
[0079] A data scientist can be an individual responsible for developing data quality methods (such as quality checks, aggregation, and interpolation) implemented within a configurable pipeline. Data scientists can update systems as needed with new data quality methods, which can be executed as part of the pipeline.
[0080] For example, in an exemplary non-limiting use of the disclosed embodiments, an HVAC application engineer may use a text editor, a graphical editor, or other interface to create or modify configuration file 408. The application engineer may use configuration file 408 to connect different algorithms, such as those programmed by a data scientist and added to the DQ core library 410. The data scientist is responsible for algorithm development and is unaware of the hardware configuration of a specific data source. The DQ core developers may also be responsible for adding features, including DQI, DQA, etc., to the DQ core library 410 or the feature engine 406.
[0081] DQ software architecture 402 may include a flag engine 424. Flag engine 424 can be used to generate data quality flags based on domain knowledge using fixed and fuzzy logic. These flags can then be used as part of data quality analysis processing in pipeline 422 to perform DQI and DQA functions. This domain knowledge is specific to the context in which the disclosed functions operate, such as in a BAS system, manufacturing process control system, or other system. Because the definition of “bad” data points within the scope of data quality analysis is highly dependent on the system being analyzed, the classification of data will be determined by contextual data and user input. In the disclosed embodiments, all DQI / DQA components and processes use the same flag and value format, which is converted by an adapter as needed to enable general DQI processing and aggregation by the DQA process.
[0082] Simple Boolean logic (true / false) can be used for the most basic problems of quality, and the disclosed embodiments can also use the DQI described herein to measure feature importance (domain knowledge-based configuration) and use fuzzy logic to assign appropriate data quality labels. In this case, fuzzy logic refers to the absence of sharp boundaries between data quality (I / O) states and the assignment of values between 1 and 0. The definition of this non-binary boundary can be implemented using a sigmoid function, as described in more detail herein.
[0083] The flag engine 424 can be implemented as a plug-in to the pipeline generator 420. The flag engine 424 can modify sensor data, for example, by adding "flags" as columns in a sensor value table in a database. Flags can be binary, values, or strings as labels to indicate the DQ of a sensor time series.
[0084] The processing and architecture described in this paper enable collaborative workflows among users, application engineers, and data scientists or other developers. For example, developers can develop core features, while data scientists can design various plugins and DQI / DQA components from the aforementioned libraries. Application engineers or other users can then interactively define configuration files to produce configurable pipelines. Users can then execute pipelines for DQI / DQA results.
[0085] When pipeline 422 is executed, each or any processing component checks the generated relevant DQI or DQA. If the DQA is less than a threshold, or too many poor DQIs are generated, the component may refuse to execute (i.e., abort the processing pipeline) or may be configured to run another pre-processing function to improve data quality before continuing pipeline processing. For example, if the system needs to replace too many missing values through interpolation or other means, the system may determine that the model represented by the input data is invalid and therefore abort the execution of some or all operations in the configurable pipeline.
[0086] Configurable pipelines can also include runtime-defined aggregation and interpolation processes, and components for such processes can be stored in the DQ core library 410 and referenced by configuration file 408. Interpolation and aggregation steps can be determined for a given user-required runtime, specifically as can be defined in configuration file 408.
[0087] Interpolation can include, but is not limited to, basic forward value filling, linear interpolation, and polynomial interpolation.
[0088] Aggregation can include summary statistics such as counts and means, as well as data quality aggregation based on daily / weekly / monthly data quality metrics and computational data quality aggregation. Disclosed embodiments can combine different DQIs with learned weighting methods and can compute boundaries or derivatives from raw sensor data input.
[0089] Figure 5A A non-limiting example of pattern 502 used in conjunction with configuration file 514 according to a disclosed embodiment is shown. In this example, "quality pattern" 502 defines a base pattern 504 that defines elements such as weighting attributes for zero values, negative values, date formats, and outliers. Base pattern 504 also defines whether aggregation should be performed ("true") and the type of interpolation to be used ("linear"). Pattern 502 also defines a zero-weighted pattern 506 that defines elements such as weighting attributes for zero values (greater than zero in base pattern 504), negative values, and date formats. Zero-weighted pattern 506 also defines how frequently aggregation should be performed ("weekly") and the type of interpolation to be used ("linear"). Of course, any other number or different elements may be included in such pattern 502, and this simplified example is not limited to.
[0090] Figure 5BA non-limiting example of configuration file 512 according to a disclosed embodiment is shown, illustrated in YAML Ain't Markup Language (YAML) format, which can be used in conjunction with or incorporated into pattern 502. In various implementations, configuration file 512 may be implemented in Extensible Markup Language (XML) format, YAML format, JavaScript Object Notation (JSON) format, or another markup language or format. Since configuration file 512 references and may include pattern 502, configuration file 512 and pattern 502 together can be used as configuration file 408 as described herein.
[0091] In this simplified example, configuration file 512 includes the definition of input data 514 (shown as meter IDs 101, 102, 103, and 104) and the schema to be applied to each input data source. In this example, "basic_schema" 504 is used for meter IDs 101 and 102, while "zero-weighted schema" 506 is used for meter ID 104. Custom schema 516 is used for meter ID 103.
[0092] like Figure 5A and Figure 5B As can be seen in the simplified example, configuration file 408 (such as configuration file 512 together with mode 502) defines such elements of the configurable pipeline as input data and sources, data quality metric processing and corresponding weighting to be used, data quality aggregation processing to be used, and interpolation processing to be used.
[0093] Configuration file 408 may include nodes that can adopt a specific naming convention used in the calling application or the target application.
[0094] Figure 6 An example of a data quality aggregation (DQA) process 602 according to a disclosed embodiment is shown. In this example, the DQA process 602 receives sensor data 612 from multiple instruments. In other cases, the sensor data may come from... Figures 1 to 3 Any device shown or described herein. DQA processing 602 also receives data quality configuration data 608, which may (but does not have to) be in the form of a configuration file 408. Details on how to use a configuration file to define a specific DQA or DQI process and appropriate input data are described above.
[0095] DQA processing 602 can then (when the system performs DQ analysis in other ways) aggregate any generated DQ data into logical groups, for example, based on specifications in the object or based on DQ configuration data 608. In this example, individual meter / sensor data is first aggregated into report groups. Note that a single meter can be part of multiple report groups. DQI / DQA data can be further aggregated, such as by combining multiple report groups into build groups, or by aggregating report groups to be processed by a given algorithm. Furthermore, all DQI / DQA data can be further aggregated to reflect the DQ of the entire facility or campus. DQA definitions can be reused as needed.
[0096] In addition to the aggregated DQI / DQA data itself, DQA processing 602 can also output metadata 614 for data cleaning or repair processing. Metadata 614 can be combined with the raw sensor data, such as in metadata 414. Metadata 614 can include any contextual data required or useful for certain DQ decisions; for example, metadata 614 can include minimum / maximum boundaries for specific values when this is known and it is expected that this will be flagged as a problem.
[0097] The disclosed embodiments also enable novel DQ analysis processing to be stored in the DQ core library 410. Specifically, domain-specific DQI computation methods can use normalized DQI for each type of dataset. For example, when out of range DQI is defined below, a sigmoid function is used to transform values between [-∞, ∞] and [0, 1].
[0098] The disclosed embodiments can handle out-of-range DQI. When the input value exceeds the boundary, the out-of-range DQI is close to 1. When the input sensor value is s and its normal range is s∈[a, b], the out-of-range DQI is m1, as defined below:
[0099] m1=f1(x)=sigmoid(x1)+sigmoid(x2)
[0100]
[0101] The original signal / data s is scaled and shifted to produce the value x:
[0102]
[0103]
[0104] Where, scalar k is
[0105]
[0106] Figure 7An example of such an out-of-range DQI function is shown, where a = 0 and b = 2.
[0107] The disclosed embodiments can handle gap DQI. If there are more gaps in the sample data, the gap DQI approaches 1. The gap DQI can be calculated as:
[0108]
[0109] Among them, G i (s) is the i-th gap time of the sensor value s, and T is the total measurement time.
[0110] The disclosed embodiments can handle monotonic DQI. Some instrument or sensor data should increase monotonically. To quantify the non-monotonic level of the raw sensor data, the system can use the M3 metric, which represents the ratio of the area under the raw sensor data to the area under the monotonic sensor data.
[0111] Figure 8 An example of single-schedule calculation using sensor data with monotonic "holders" is shown to illustrate a non-limiting example of detected data quality problems. In this figure, solid lines represent sensor data, and dotted / dashed lines represent holder data. Such single-schedule calculations can be useful, for example, in the context of an accumulator meter indicating data quality problems with monotonically increasing failures. Dashed lines indicate the expected monotonic increase, and deviations from this represent the problem with DQI's measurement of severity in normalized terms when such DQI processing is added to a configurable pipeline as described herein. Such functions can be used to generate DQI values / exponents indicating the severity of data deviations as values between 0 and 1.
[0112] The system can use the holder function h(s[i]) and can perform processing according to the following exemplary pseudocode for the monotonic holder function h(s):
[0113]
[0114]
[0115] The system can then use the m3 metric, defined as:
[0116]
[0117] Among them, s m It is the minimum value of the sensor reading 's'. For example... Figure 9 As shown, h(s[i]) is always increasing. The metric m3 is the ratio of the area under curve s[i] to the area under line h(s[i]).
[0118] The disclosed embodiments support several processes for data quality aggregation, which aggregate data quality metrics together. Since each DQI is a metric between 0 and 1 in the various embodiments, the system can combine them with different DQA methods.
[0119] The disclosed embodiments may use a weighted average of DQA, for example, to calculate the DQ of a set of devices in a spatial range. For example, the system may use:
[0120]
[0121] Among them, w i ∈[0,1] is the weighting factor of the i-th DQI, and the time index is k. This cluster can be used as another DQI for further clustering.
[0122] The disclosed embodiments can use the maximum DQI for DQA. When some DQ problems are critical, the system can aggregate based on the maximum DQI:
[0123] m[k]=max i (m i [k])
[0124] The disclosed embodiments can use time-range averaging (downsampling) of DQA to help determine whether any DQ problems exist within a given time period. The system can also aggregate DQI based on different time-domain samples. DQA processing can take an aggregated sample for every M samples, such as:
[0125] m[Mk]=m[k]↓M
[0126] For convenience, define
[0127] m M [k] = m[Mk]
[0128] Make
[0129]
[0130] Among them, w i It's the weight. Then, m M [k] can be used as another DQI for further aggregation.
[0131] The disclosed embodiments can handle the maximum value within a time range as a substitute for or modification of the time range average. The maximum value within the time range represents another downsampling method, such as:
[0132]
[0133] Using these DQA processes, the system can aggregate DQI metrics layer by layer. From the user's perspective, the user can zoom in and out of sensor data in the spatial or temporal domain. Utilizing DQI and DQA, the user can... Figure 4 The sensor data represents the relevant metrics selected from a large number of sensors.
[0134] The DQI discussed in this article can be, for example, derived from... Figure 4 The components in the feature engine 406 are computed by processes defined in the DQ core library 410, or by other defined processes or components.
[0135] Figure 9 The processing according to the disclosed embodiments is illustrated and can be performed, for example, by a data processing system in a building automation system as described herein, such as a report server 104, field controller 102, client station 106, or other systems or controllers that can connect to a data source within the BAS 100 to generate a corresponding pipeline to implement the data quality processing techniques as described herein. In other implementations, such processing can be performed by a separate data processing system or a system using data generated by or received from the BAS. In other embodiments, such processing can be performed by a data processing system operating on data generated by or received from some other processing control system. For simplicity of description, the term "system" below refers to the data processing system that performs the processing described in any of these implementations. Figure 9 The processing may include any other processing, components or other features described herein, or a combination thereof.
[0136] The system can receive input data to be processed for data quality metrics and / or data quality aggregation (902). As used herein, “receive” can include loading from a storage device, receiving from another device or processing, receiving via interaction with a user, or otherwise, and specifically can include receiving device data from one or more sensors or other devices in a building automation system. In some implementations, the input data may include or be associated with DQ metadata that is also received at this time; in other cases, DQ metadata is not received at this time but is created, modified, or appended later as described herein. The input data may correspond to data from a single device or from multiple devices.
[0137] The system can receive a configuration file (904) that defines the DQ processing to be performed on the input data. Details and characteristics of various implementations of such configuration files have been described above. The system can receive configuration files in a variety of ways, including receiving user input to do so, receiving one or more configuration files from a dedicated email address, folder, or storage address, etc. The flexibility of receiving configuration files to build configurable pipelines is useful, for example, because application engineers can independently develop configuration files according to patterns defined as described herein and forward such configuration files to the system to prompt for “on-demand” construction of the corresponding pipeline for DQ processing of the input data. The following... Figure 10 The exemplary processing for generating such configuration files is described in the context of [the relevant context].
[0138] Note that in some cases, the nature of the input data received at 902 can determine which profile is received at 904. For example, if sensor data from a variable air volume (VAV) unit is received, the system may load or otherwise receive a profile for analyzing the VAV unit. In other cases, the opposite may be true, and a profile may be received at 904 before input data is received at 902. In these cases, appropriate input data may be loaded or otherwise received based on what is specified in the profile. In either case, the profile may define the DQI or DQA associated with the identified input data source and how they are processed in a configurable pipeline.
[0139] In some embodiments, receiving input data at 902 and / or receiving a configuration file at 904 may be initiated or executed under the control of the calling application or the client application. That is, the calling application may specify the input data and / or configuration file to be used in order to perform the processing disclosed herein, and the results of these processes may then be returned to the calling application.
[0140] The system can dynamically build configurable pipelines based on configuration files (906). As described herein, this can include building pipelines from one or more DQI or DQA processing components stored in the DQ core library, and the pipelines can include such processing components connected in parallel and / or in series with each other. In various embodiments, the pipeline is built dynamically at runtime, contrary to a pre-programmed sequence of operations. The following is further elaborated upon... Figure 11 The exemplary processes used to generate such configurable pipelines are described in the context of [the relevant context].
[0141] The system can perform DQ processing on the input data (908). This can include executing each DQI or DQA processing component included in the pipeline in the order defined in the pipeline. It can also include performing other DQ processing that is not necessarily part of the pipeline, such as DQ processing or checks performed by the feature engine described herein. It can also include generating data quality tags based on domain knowledge by the tagging engine. It can also include transforming the data through one or more adapters as needed for each DQ processing.
[0142] DQ processing may include repairing or otherwise processing the data, including performing functions such as normalizing the data, performing interpolation to augment the data, or replacing missing data points, as well as other processing that is not strictly “data-quality” processing but is still useful in performing the DQ processing described herein.
[0143] Based on DQ processing, the system can generate one or more DQ results corresponding to the input data (910). DQ results may include DQIs corresponding to the input data, DQIs corresponding to specific data points in the input data, DQIs corresponding to the device that generated the specific input data, etc. DQ results can identify low-quality or missing data points, including any outliers or other features described herein. DQ results may include aggregated data generated from the input data using DQA processing as described herein. Compared to time series or historical trends, DQ results may include data markers and other indicators of specific and aggregated DQIs.
[0144] The system can return DQ results (912). "Returning" can include storing the DQ results in a storage device, displaying the DQ results to a device used in the user interface, transferring the DQ results to another device, or processing them. In various embodiments, the DQ results can be returned to the calling application for further processing, such as data cleaning, fault analysis in a physical system represented by the input data, filtering, visualization, etc.
[0145] Generating and returning DQ results may include parsing and storing data structures such as the expanded sensor data table 1300 and DQ aggregation table 1330 described below. Returning DQ results may also include storing, displaying, or transmitting indicators that low-quality data points can indicate correct data from the device, indicating anomalies reflected by the data points, such as excessively high temperatures recorded by a thermostat device. In such cases, "low-quality" data points may indicate a failure or misconfiguration of other devices in the system that generated the input data, and the DQ results may indicate such problems. Similarly, DQ results (whether as individual DQIs or from aggregated data) may indicate one or more devices or portions of the system that generated the input data as a problem area.
[0146] In some embodiments, the format of DQ results and / or input data may be JSON in a Representational State Transfer (REST) application programming interface (API), but the input or output may be implemented using any API or other format that the calling application or client application may require.
[0147] Figure 10 Processes according to disclosed embodiments are illustrated, which can be performed, for example, by a data processing system in a building automation system as described herein, such as via user interaction with report server 104, field controller 102, client workstation 106, or other systems or controllers to generate configuration files as disclosed herein. In other implementations, such processes may be performed by a separate data processing system or by a system using data generated by or received from a BAS. In other embodiments, such processes may be performed via user interaction with a data processing system that operates on data generated by or received from some other processing control system. For simplicity of description, the term "system" below refers to the data processing system that performs the processes described in any of these implementations. Figure 10 The processing may include any other processing, components or other features described herein, or a combination thereof.
[0148] The system or user can determine whether one or more existing patterns can be used for desired processing on a particular building (1002). As part of this processing, the system may, for example, identify any patterns using specific input data, DQA processing, DQI processing, or other factors disclosed herein. "Building" refers to the building automation system and its sensors and other hardware within the physical structure. Data generated by the BAS will be processed as disclosed herein.
[0149] If no existing pattern exists that can be used for the required processing ("No"), the system may interact with the user to receive or create a new pattern for the building (1004). Pattern 502 above is an example of a pattern that can be created.
[0150] Once a schema is created in 1004, or if it already exists in 1002 ("yes"), the system can interact with the user to receive or create a new profile for the building (1006). Profile 512 above is an example of such a profile that can be created.
[0151] The system then adds the pattern to the configuration file (1008). A combined configuration file and pattern, or a configuration file that references and incorporates patterns, can serve as the aforementioned configuration file 408. This configuration file defines the DQI, DQA, and other processing that will be performed on the input data.
[0152] The system can interact with the user to define the input data for the configuration file (1010). In this example, this can be done by adding the meter ID of the meter that generates the input data to the configuration file.
[0153] The system stores the completed configuration (1012). For example, the system can store configuration file 408 in the configuration directory under the building ID, which identifies the building that is being analyzed and generating input data.
[0154] Complete configuration file settings (1014).
[0155] Figure 11 The disclosed embodiments illustrate processes that can be performed, for example, by a data processing system in a building automation system as described herein, such as report server 104, field controller 102, client station 106, or other systems or controllers, to generate a configurable pipeline as disclosed herein. In other implementations, such processes may be performed by a separate data processing system or by a system using data generated by or received from a BAS. In other embodiments, such processes may be performed by a data processing system operating on data generated by or received from some other processing control system. For simplicity of description, the term "system" below refers to the data processing system that performs the processes described in any of these implementations. Figure 11 The processing may include any other processing, components or other features described herein, or a combination thereof.
[0156] The system can receive an input (1102) indicating that a pipeline should be established. This input can be user input, input received from another device, process, or application, etc. The input can indicate a building, instrument, device, or other data source to be analyzed. For the purposes of this example, the input indication corresponds to, for example... Figure 10 The configuration file 408 created in the process.
[0157] The system receives a building configuration file (1104), such as configuration file 1106 stored in an external shared location. Configuration file 1106 may be, for example, configuration file 408. This may include retrieving configuration file 1106 from a configuration directory associated with the building's building ID. Preferably, and as described above, the configuration file is stored in an external storage location or a shared configuration drive, such that the configuration file is independently editable and not hardcoded into the DQ core library.
[0158] The system reads the relevant quality mode 1110 (1108) of the instrument defined in the configuration file. As described herein, each configuration file can define the instrument or other device to be read, along with the associated mode, such as:
[0159] meter:
[0160] -Instrument_id: 101
[0161] Quality_Mode: Basic_Mode
[0162] System initialization or instantiation of empty pipeline (1112).
[0163] The system reads the relevant quality checks to be performed as defined in the configuration file (1114). Figure 5A and Figure 5B In the example, quality checks can be read from Quality_Mode → Basic_Mode → Metrics.
[0164] If there are quality checks to be performed ("Yes" at 1114), then for each quality check, the system determines whether the user has specified that a given quality check should be excluded (1116). If yes ("Yes"), the system determines whether there are other quality checks to be performed (1118), and if yes ("Yes"), returns to 1116 to proceed to the next quality check. If no ("No"), the system proceeds to 1126 because all quality checks have been processed or excluded.
[0165] If a given quality check is not excluded at 1116 (“No”), the system reads the quality check method 1122 (1112) to be used as defined in the configuration file. In various embodiments, each of these quality checks is part of the feature engine component 406 of the DQ software architecture 402. In various embodiments, these quality checks may be DQI processing from the DQ core library 410. Quality checks may be a combination of FE processing and DQI processing. Quality checks may be discrete steps for detecting data quality problems. Each type of check (e.g., outlier detection) may have multiple algorithms / methods for performing the check, as specified in the configuration file.
[0166] The system adds the given quality check 1122 to the pipeline (1124), such as configurable pipeline 422. The system then returns to 1118 above to determine whether any additional quality checks should be performed.
[0167] Once all required and unexcluded quality checks have been added to the pipeline, the system determines whether the configuration file specifies that interpolation should be used on the input data (1126).
[0168] If interpolation is to be performed ("Yes" at 1126), the system determines whether the user has specified that interpolation should be excluded (1128). If yes ("Yes"), the system moves to 1136 because interpolation will not be performed.
[0169] If no ("No" at 1128), the system reads the interpolation method / algorithm to be executed as defined in the configuration file, 1132 (1130). Figure 5A and Figure 5B In the example, the quality check can be read from Quality_Mode → Basic_Mode → Interpolation → Method. The system adds the interpolation method / algorithm 1132 to the pipeline (1134), such as the configurable pipeline 422, and moves it to 1136.
[0170] The system receives flag assignment logic (1136). In some embodiments, the flag assignment logic may include default flag assignment logic stored in the DQ core library 410. In some embodiments, flag assignment logic may be received from the calling application or the client application. In some embodiments, flag assignment logic may be specified through specific quality checks. The flag assignment logic specifies rules such as thresholds, outliers, and other values used to flag specific data.
[0171] The system determines whether the user has specified that flag assignment should be excluded (1138). If yes, the system moves to 1146 because flag assignment will not be performed.
[0172] If no ("No" at 1138), the system reads the flag assignment logic method / algorithm 1142 (1140) to be used from the configuration file. Figure 5A and Figure 5B In the example, quality checks can be read from Quality_Mode → Basic_Mode → Flag Logic → Method. The system adds the flag assignment logic method / algorithm 1142 to the pipeline (1144), such as configurable pipeline 422, and moves it to 1146.
[0173] The system determines whether the configuration file specifies that data quality aggregation should be performed (1146). If no (“No”), the pipeline completes, and the completed pipeline can return to the entire process (1158), such as Figure 9 906 in the middle.
[0174] If data quality aggregation is to be performed ("Yes" at 1146), the system determines whether the user has specified that aggregation should be excluded (1148). If yes ("Yes"), the pipeline completes, and the completed pipeline can return to the entire process (1158), such as... Figure 9 906 in the code is because the aggregation will not be performed.
[0175] If no ("No" at 1148), the system reads the DQI 1152 and aggregation method 1156 (1150) to be used for data quality aggregation from the configuration file and / or from the quality checks. In various embodiments, each quality check has an associated DQI, which is used by aggregation as a weighting mechanism, such as... Figure 5A and Figure 5B A simplified example is shown. For instance, the DQI value is used as part of the index calculation for a given data quality fault type. Figure 5A and Figure 5B In the example, the cluster can be read from either Quality_Mode → Cluster or Quality_Mode → Basic_Mode → Cluster.
[0176] The system adds an aggregation method 1156 with an appropriate DQI 1152 to the pipeline (1154), such as a configurable pipeline 422. At this point, the pipeline completes, and the completed pipeline can return to the overall processing (1158), such as... Figure 9 906 in the middle.
[0177] Figure 12 An example of processing according to a disclosed embodiment is shown, which can be performed, for example, by a data processing system in the building automation system described herein, such as a report server 104, field controller 102, client station 106, or other system or controller, illustrating the overall processing and interaction within the DQ software architecture 402. In other implementations, such processing may be performed by a separate data processing system or by a system using data generated by or received from the BAS. In other embodiments, such processing may be performed by a data processing system operating on data generated by or received from some other processing control system. For simplicity of description, the term "system" below refers to the data processing system that performs the processing described in any of these implementations. "System" may include client application 1204, DQ software architecture 1206, pipeline 1208, feature engine / DQ core library 1210, and configuration catalog 1212, all of which interact with one or more users 1202. Figure 12 The processing may include any other processing, components or other features described herein, or a combination thereof.
[0178] Figure 12 The example assumes that some basic configurations as described in this article have already occurred. For example, data scientists or other individuals have created DQI and DQI algorithms, as well as methods that contain them, along with any DQ feature algorithms and methods that contain them, and have appropriately stored them in library 1210, such as Figure 4 The example uses FE 406 and DQ core library 410. Similarly, application engineers or other individuals have created relevant configuration files, such as by using... Figure 10The processes described herein are processed and stored in configuration directory 1212. Configuration directory 1212 can be implemented in any memory or storage device described herein, and is preferably externally accessible so that configuration files can be created and edited without recoding the DQ software architecture 1206 itself.
[0179] In this example, the system receives a request to perform DQ analysis processing (1214) via client application 1204, which receives such a request from user 1202. The system may request user input (1216) via client application 1204, such as the building, equipment, or facility to which DQ analysis processing will be performed, or other appropriate user input. The system may also receive user input from user 1202 via client application 1204 (1218).
[0180] The system can then initiate DQ analysis setup (1220) by, for example, calling client application 1204 of DQ software architecture 1206. Initiating DQ analysis setup can include any relative information, in this example, the building ID that identifies the building to be analyzed.
[0181] The system can retrieve configuration files (1222, 1224) used for DQ analysis. In this example, the DQ software architecture 1206 sends a request (1222) to the configuration directory 1212 specifying a building ID. In response, the DQ software architecture 1206 receives the corresponding configuration file (1224) from the configuration directory 1212.
[0182] Based on the configuration file, the system retrieves DQ methods / algorithms (1226, 1228), which may include any DQ checks, FE features, DQI processing, or DQA processing that can be specified in the configuration file as described herein. In this example, the DQ software architecture 1206 sends a request for a DQ method / algorithm (1226) to the FE / DQ core library 1210 and, in response, receives the requested DQ method / algorithm (1228).
[0183] The system creates pipelines (1230, 1232) based on the configuration file and using the retrieved DQ method / algorithm. This can be, for example, based on... Figure 11 The process is executed in the process. In this example, the DQ software architecture 1206 performs a pipeline creation process (1230) to produce an instance (1232) of the configured pipeline 1208.
[0184] The system can indicate that DQ analysis setup is complete (1234). This can be done, for example, by notifying the client application 1204 that DQ analysis setup is complete (1234) via DQ software architecture 1206.
[0185] The system can initialize DQ analysis processing (1236). This can be achieved, for example, by having a client application 1204 send the building / sensor data to be analyzed along with instructions to perform the analysis to the DQ software architecture 1206 (1236).
[0186] The system can perform DQ analysis processing (1238, 1240). In this example, the DQ software architecture 1206 can use building / sensor data (1238) to execute pipeline 1208 as a response to receive results (1240). During pipeline execution, each process added to the pipeline as described herein is executed to generate a DQ result. As described herein, in various cases, the processes added to the pipeline may be executed sequentially, simultaneously, or in an order different from that indicated in the configuration file. As mentioned above, the system can abort pipeline processing under predetermined conditions, such as if the DQA is less than a threshold or too many poor DQIs are generated.
[0187] The system can return DQ results (1242, 1244). In this example, DQ software architecture 1206 can return the DQ result to client application 1204 (1242), and client application 1204 can then return the DQ result to user 1202.
[0188] Figure 13A and Figure 13B Examples of DQ results are shown, such as an expanded sensor data table that combines sensor data 412 and DQ metadata 414 with DQ aggregated values, which can be generated by the processing described herein.
[0189] Figure 13A The DQ results are shown as an expanded sensor data table 1300, serving as an example of DQ aggregation for calculating the DQ index. In this example, the expanded sensor data table 1300 includes raw sensor data such as the time 1302 for each sensor reading, the meter ID 1304, and the actual sensor value 1306 for each time. In this example, and for each sensor reading, the "expanded" DQ data includes a marker 1308 for zero values, a missing value 1310, and an outlier measurement 1312. In this example, the DQ data includes a DQI indexed by time index 1318, reflecting the DQI calculated for each time-indexed sensor reading. In this example, the DQ data also includes a DQI indexed by metric 1314, which sums the zero value 1308, the missing value 1310, and the outlier measurement 1312, along with a DQI weight 1316, applied to each DQI indexed by metric 1314.
[0190] In this example, the DQ data also includes a DQ aggregate 1320 for the meter, which is calculated as the sum of the time index 1318 values of the individual DQIs. Examples of such calculations as described above use an average over the DQA, for example, to calculate the DQ of a particular device over a set of time index values. For example, the system could use:
[0191]
[0192] Among them, w i ∈[0,1] is the weighting factor for the i-th DQI, and the time index is k. In this example, the column "DQI (by time index)" is equivalent to m[k], w i The value is the DQI weight 1316, and the individual mi[k] value is the DQI index for each metric 1314.
[0193] The expanded sensor data sheet 1300 can be the DQ result of the processing described herein, and can be continuously or repeatedly updated during pipeline processing such that it reflects a "snapshot" of the data being processed by the pipeline at any point in time.
[0194] Figure 13B The DQ results are shown as a DQ aggregation table 1330. The DQ aggregation table allows users to immediately see DQ issues, such as those on a building, facility, or multiple buildings or facilities. This non-limiting example shows DQA results for multiple sensors / meters in multiple buildings, sorted by priority. The data in this example includes building 1332, meter ID 1334 in each building, meter type 1336 for each meter, location 1338 for each meter in building 1332, DQA value 1340 for each meter, and priority 1342 for each meter. The BAS disclosed herein can sort DQ results, such as by DQA value, and assign priorities to the sorting, indicating the relative severity of DQ issues related to data from each meter, sensor, or other device being evaluated. In this particular example, the DQ aggregation table 1330 sorts the DQ aggregations of individual meters, but similar sorting can be performed on aggregated DQ values for meter collections (such as all meters on a floor of a building, all meters of a specific type in a given building, all meters of a specific type in all buildings of a facility, etc.). Note that in this example, the DQA value of 1340 for meter 101 at "Zhongfu Plaza" is as follows: Figure 13A The DQ aggregation 1320 shown illustrates that the DQ aggregation table 1330 can reflect the aggregation of multiple extended sensor data tables 1300. The DQ aggregation table can show the aggregation, sorting, and prioritization of DQI calculated for individual data points on a specific instrument for any number of buildings and facilities.
[0195] Figure 14A block diagram of a data processing system 1400 that can be implemented in various embodiments is shown. The data processing system 1400 is... Figure 1 An implementation of the field controller data processing system 102 in the middle and Figure 4 An example of the implementation of the data processing system 400 in the example.
[0196] Data processing system 1400 includes a processor 1402 connected to a secondary cache / bridge 1404, which in turn is connected to a local system bus 1406. The local system bus 1406 may be, for example, a Peripheral Component Interconnect (PCI) architecture bus. In the depicted example, main memory 1408 and a graphics adapter 1410 are also connected to the local system bus 1406. The graphics adapter 1410 may be connected to a display 1411.
[0197] Other peripheral devices, such as LAN / WAN / wireless (e.g., WiFi) adapter 1412, can also be connected to the local system bus 1406. An expansion bus interface 1414 connects the local system bus 1406 to the input / output (I / O) bus 1416. The I / O bus 1416 connects to a keyboard / mouse adapter 1418, a disk controller 1420, and an I / O adapter 1422. The disk controller 1420 can be connected to a storage device 1426, which can be any suitable machine-usable or machine-readable storage medium, including but not limited to non-volatile hard-coded media such as read-only memory (ROM) or erasable, electrically programmable read-only memory (EEPROM), magnetic tape storage devices, and user-recordable media such as floppy disks, hard disk drives, and optical disc read-only memory (CD-ROM) or digital universal disc (DVD), as well as other known optical, electrical, or magnetic storage devices.
[0198] Storage device 1426 may store any program code or data used to perform the processes disclosed herein or to perform building automation tasks. In a particular embodiment, storage device 1426 may include elements such as input data 1452, libraries 1454, configuration files 1456, and other data 1458, as well as a stored copy of BAS application 1428. Other data 1458 may include the software architecture, any of its elements, or any other data, programs, code, tables, or other information or data discussed above.
[0199] In the example shown, audio adapter 1424 is also connected to I / O bus 1416, and a speaker (not shown) can be connected to this audio adapter to play sound. Keyboard / mouse adapter 1418 provides connectivity for pointing devices (not shown), such as a mouse, trackball, track indicator, etc. In some embodiments, data processing system 1400 may be implemented, for example, as a touchscreen device, such as a tablet computer or touchscreen panel. In these embodiments, the elements of keyboard / mouse adapter 1418 may be integrated with display 1411.
[0200] In various embodiments of this disclosure, the data processing system 1400 may be implemented as a workstation or field controller 102, wherein all or part of the BAS application 1428, stored in memory 1408, is configured to perform the processing as described herein and is generally used as a BAS as described herein. For example, processor 1402 executes program code of BAS application 1428 to generate a graphical user interface 1430 displayed on display 1411. In various embodiments of this disclosure, the graphical user interface 1430 provides a user with an interface to view information about and control one or more devices, objects, and / or points associated with the management system 100. The graphical user interface 1430 also provides a customizable interface to present information and control in an intuitive and user-modifiable manner.
[0201] Those skilled in the art will understand that Figure 14 The hardware depicted may vary for a particular implementation. For example, other peripheral devices such as optical disc drives may be used in addition to or in place of the depicted hardware. The examples depicted are provided for illustrative purposes only and are not intended to imply any architectural limitations with respect to this disclosure.
[0202] With appropriate modifications, one of various commercial operating systems can be used, such as Microsoft Windows, a product of Microsoft Corporation located in Redmond, Washington. TM This version. Operating systems can be modified or created based on this disclosure, for example, to enable object discovery and the generation of hierarchical structures of discovered objects.
[0203] LAN / WAN / WiFi adapter 1412 can, for example, connect to network 1432, such as Figure 1 MLN 120 in the context of this. As further explained below, network 1432 can be any combination of public or private data processing system networks or networks known to those skilled in the art, including the Internet. Data processing system 1400 can communicate with one or more computers via network 1432, which are not part of data processing system 1400 but can be implemented, for example, as a separate data processing system 1400.
[0204] Of course, those skilled in the art will recognize that, unless the order of operations is specifically indicated or required, some steps in the above process may be omitted, performed simultaneously or sequentially, or performed in a different order.
[0205] Those skilled in the art will recognize that, for simplicity and clarity, this document does not depict or describe the full structure and operation of all data processing systems suitable for use with this disclosure. Rather, only data processing systems that are unique to this disclosure or necessary for understanding this disclosure are depicted and described. The remaining structure and operation of the systems used herein may conform to any of the various current implementations and practices known in the art.
[0206] It is important to note that although this disclosure is described in the context of a full-function system, those skilled in the art will understand that at least a portion of the mechanisms of this disclosure can be distributed in the form of instructions contained in any of a variety of machine-usable, computer-usable, or computer-readable media, and this disclosure applies equally regardless of the specific type of instruction or signal-bearing medium or storage medium used to actually perform the distribution. Examples of machine-usable / readable or computer-usable / readable media include: non-volatile, hard-coded media such as read-only memory (ROM) or erasable, electrically programmable read-only memory (EEPROM), and user-recordable media such as floppy disks, hard disk drives, and optical disc read-only memory (CD-ROM) or digital universal disc (DVD).
[0207] Although exemplary embodiments of the present disclosure have been described in detail, those skilled in the art will understand that various changes, substitutions, variations, and modifications may be made to the disclosure without departing from the spirit and scope of the broadest form of the disclosure.
[0208] The descriptions in this application should not be construed as implying that any particular element, step, or function is a fundamental element that must be included in the scope of the claims: the scope of the patent subject matter is defined only by the granted claims. Furthermore, none of these claims are intended to invoke 35 USC §112(f) unless the exact phrase “means for” is followed by a participle.
Claims
1. A method in a building automation system, the method being executed by a data processing system and comprising: The data processing system receives input data representing the operation of the physical devices of the building automation system; The data processing system receives a configuration file that defines the data quality processing to be performed on the input data; The data processing system dynamically constructs a configurable pipeline based on the configuration file. The configurable pipeline includes one or more data quality metrics or data quality aggregation processing components from the data quality core library. The data processing system performs data quality processing on the input data, including executing each data quality metric or data quality aggregation processing component included in the configurable pipeline; The data processing system generates one or more data quality results based on the data quality processing. as well as The data processing system returns one or more data quality results.
2. The method according to claim 1, wherein, The method is executed with a software architecture that includes a pipeline generator for constructing the configurable pipeline.
3. The method according to claim 1, wherein, The method is executed in a software architecture that includes multiple adapters configured to transform data for use by the one or more data quality metrics or data quality aggregation processing components.
4. The method according to claim 1, wherein, The data quality processing includes generating data quality labels based on domain knowledge using fixed and fuzzy logic.
5. The method according to claim 1, wherein, The data quality processing includes performing energy meter overflow checks.
6. The method according to claim 1, wherein, The data quality processing includes summarizing data from sensor points based on different time-domain aggregation patterns.
7. The method according to claim 1, wherein, The configuration file includes the definition of the configurable pipeline, multiple data quality metrics and data quality aggregation processing components to be executed in series and / or in parallel, and the connections between the multiple data quality metrics and data quality aggregation processing components.
8. The method according to claim 1, wherein, The configurable pipeline includes a data quality aggregation processing component that uses one of the following: a weighted average of data quality aggregations, the maximum data quality metric of data quality aggregations, the time-range average of data quality aggregations, and the time-range maximum value of data quality aggregations.
9. The method according to claim 1, wherein, The configuration file includes identifiers of the patterns to be applied, clustering methods, and interpolation methods, and the patterns include data quality metrics with associated weights.
10. The method according to claim 1, wherein, Dynamically constructing the configurable pipeline includes: Initialize the configurable pipeline; Read the pattern associated with the configuration file; Based on the mode and the configuration file, at least one quality check method is selectively added to the configurable pipeline, wherein the quality check method is a data quality indicator processing component from the data quality core library; Based on the pattern and the configuration file, at least one interpolation method may be selectively added to the configurable pipeline; Based on the pattern and the configuration file, flag assignment methods are selectively added to the configurable pipeline; Based on the pattern and the configuration file, aggregation methods are selectively added to the configurable pipeline, wherein the aggregation methods are data quality aggregation processing components from the data quality core library; and Store the configurable pipeline.
11. A building automation system comprising a plurality of sensors and at least one data processing system, the data processing system being configured to process input data collected from the operation of at least one of the plurality of sensors, wherein, The building automation system is configured as follows: Receive input data representing the operation of the physical devices of the building automation system; Receive a configuration file that defines the data quality processing to be performed on the input data; The data processing system dynamically constructs a configurable pipeline based on the configuration file. The configurable pipeline includes one or more data quality metrics or data quality aggregation processing components from the data quality core library. The data processing system performs data quality processing on the input data, including executing each data quality metric or data quality aggregation processing component included in the configurable pipeline; One or more data quality results are generated based on the aforementioned data quality processing; as well as Return the one or more data quality results.
12. The building automation system according to claim 11, wherein, The building automation system is also configured to execute with a software architecture that includes a pipeline generator for building the configurable pipeline.
13. The building automation system according to claim 11, wherein, The building automation system is also configured to execute with a software architecture that includes multiple adapters configured to transform data for use by the one or more data quality metrics or data quality aggregation processing components.
14. The building automation system according to claim 11, wherein, The data quality processing includes generating data quality labels based on domain knowledge using fixed and fuzzy logic.
15. The building automation system according to claim 11, wherein, The data quality processing includes performing energy meter overflow checks or performing data quality data aggregation methods based on different time-domain aggregation patterns to analyze sensor points.
16. The building automation system according to claim 11, wherein, The configuration file includes the definition of the configurable pipeline, multiple data quality metrics and data quality aggregation processing components to be executed in series and / or in parallel, and the connections between the multiple data quality metrics and data quality aggregation processing components.
17. The building automation system according to claim 11, wherein, The configurable pipeline includes a data quality aggregation processing component that uses one of the following: a weighted average of data quality aggregations, the maximum data quality metric of data quality aggregations, the time-range average of data quality aggregations, and the time-range maximum value of data quality aggregations.
18. The building automation system according to claim 11, wherein, The configuration file includes identifiers of the patterns to be applied, clustering methods, and interpolation methods, and the patterns include data quality metrics with associated weights.
19. The building automation system according to claim 11, wherein, To dynamically construct the configurable pipeline, the data processing system is configured as follows: Initialize the configurable pipeline; Read the pattern associated with the configuration file; Based on the mode and the configuration file, at least one quality check method is selectively added to the configurable pipeline, wherein the quality check method is a data quality indicator processing component from the data quality core library; Based on the pattern and the configuration file, at least one interpolation method may be selectively added to the configurable pipeline; Based on the pattern and the configuration file, flag assignment methods are selectively added to the configurable pipeline; Based on the pattern and the configuration file, aggregation methods are selectively added to the configurable pipeline, wherein the aggregation methods are data quality aggregation processing components from the data quality core library; and Store the configurable pipeline.
Citation Information
Patent Citations
Data pipeline architecture for cloud processing of structured and unstructured data
EP3182284A1
Pipeline template configuration in a data processing system
WO2019243787A1