Collaborative systems and methods for validating device failure models in an analytics crowd-sourcing environment
Patent Information
- Application Number
- CN202180036142.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-11
- Filing Date
- 2021-05-17
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2041-05-17
AI Technical Summary
因此,使用现成商用软件或开放源选项的当前分析建模在特征方面缺乏通用性,并且几乎不具有集成来自多个源的模型和输入的任何手段
Smart Images

Figure CN115699042B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the validation of analytical models. More specifically, this invention relates to a system for logging in and validating analytical models in a crowdsourcing environment. Background Technology
[0002] For industrial equipment manufacturers and organizations investing in cost-effective equipment operation, the demanding task of monitoring the health of their associated equipment operations is undertaken using a variety of software applications. Examples of equipment can include construction machinery such as trucks, cranes, earthmoving vehicles, mining vehicles, backhoe excavators, material handling equipment, transportation equipment, oil and refining equipment, and farming equipment. Equipment maintenance has become highly complex, generating vast amounts of data, some aspects of which can be repeatedly extracted, stored, modified, and analyzed by different applications. For example, a first application might perform a health check on a group of equipment in a fleet. This health check could be based on several factors, such as machine sensor data, diagnostic or warning content, maintenance and repair history, application segmentation and severity, machine age, and service time. Machine health checks or assessments can be performed through applications with end-to-end solutions, where assessment technicians or users first collect and store data from various data sources, then extract, modify, and analyze that data, and finally store, view, and evaluate the analysis and related outputs.
[0003] It is also understandable that the same data is often generated and used by many other applications executed by different users. In the current example, each application expects a specific end-to-end solution unrelated to the evaluation, but rather generates independent end-to-end solutions. For example, a site project manager, a material supplier, and an equipment fleet manager may each have access to the same data, but with separate end-to-end solutions.
[0004] A strategy where data is reproduced or stored multiple times and used in various end-to-end solutions is typical; however, it consumes valuable computing and human resources when storing, accessing, retrieving, and modifying the same or identical data. Furthermore, because data is currently reproduced for different end-to-end solutions, there is a possibility of data loss, as no end-to-end solution benefits from the output generated by operations performed by other end-to-end solutions, since analytical results, modified data, and knowledge generated from each application are typically stored independently and separately. There is a need for modular applications capable of extracting data of interest from a common location, performing analytical operations on the extracted data, and storing the analytical models and associated analytical results in a common location accessible to other users.
[0005] Current business condition monitoring software typically offers end-to-end solutions, and some even provide APIs for integration with externally developed analytical models. However, using such software is not cost-effective because often only a portion of the software code is dedicated to developing and validating new analytical models. For example, if a user decides to create a new analytical model, using current commercial software, the user will end up generating the majority of new code for each analytical model. As currently represented, commercial analytical software wastes valuable computational resources because most of the code is replaced when a new analytical model is generated. Furthermore, if only a small amount of code is actually used to create new analytical models, the cost-effective use of data is inefficient when most of the new code is generated each time. Other commercial and open-source options may have features available for developing and validating analytical models, but they typically require significant integration work by data scientists or software engineers at high cost and often result in less general-purpose software options.
[0006] Current processes for validating analytical models typically involve close collaboration between data scientists, condition monitoring consultants (CMAs), and non-data scientists—who may be condition monitoring experts with extensive product and industry knowledge. Consequently, there are few, if any, effective and cost-efficient means of validating analytical models in real-time where data scientists and non-data scientists can constructively interact with CMAs to generate analytical models.
[0007] Because it involves the role of data scientists in developing analytical models, these individuals typically use commercial software, specially developed internal software, or open-source web applications such as Jupyter Notebooks to develop analytical models. Furthermore, crowdsourcing for generating analytical models is uncommon for data scientists, and there is often a lack of commercial solutions that allow data scientists to interact with crowdsourcing in the development of analytical models. Therefore, there is a need to develop an easy-to-use application where multiple data scientists can easily migrate models and, in coordination with the CMA, reference benchmark data in real time to validate these migrated models.
[0008] Because it involves the generality of inputs in the current modeling process of data scientists, segmented analytical modeling is ineffective, as these processes lack the ability to share valuable insights with and benefit from the insights of other analytical model developers.
[0009] Furthermore, such integration capabilities as those found in current platforms based on Python models are not readily available in commercial tools, although some commercial software offers limited means, such as APIs, for integrating with externally developed Python models. Even with this limited functionality, having an API is not cost-effective.
[0010] U.S. Patent Application No. 2019,0265,971 (“'971 Patent Application”), filed August 29, 2019 by Behzadi et al., discloses a method for data processing and enterprise applications. This application describes a general-purpose software application development platform for enterprise applications, in which applications are generated by first integrating data from various sources using different data types or data structures, and then using a model-driven architecture to generate abstract representations of the data and their correlations. The '971 Patent Application suggests that the abstract representations of algorithms, data, and data correlations enable the use of one or more of multiple algorithms to process data without needing to know the structure, logic, or interrelationships of the multiple algorithms or different data types or data structures. A supply network risk analysis model is described for generating recommendations and options for management teams to mitigate high-risk areas in the supply chain and improve supplier planning and supplier portfolio management to create appropriate redundancy, backup, and recovery options when needed. Therefore, current analytical modeling using off-the-shelf commercial software or open-source options lacks versatility in terms of features and has virtually no means of integrating models and inputs from multiple sources. Furthermore, current analytical modeling lacks a complete set of inputs because there are almost no means to integrate crowdsourced analytical inputs with traditional engineering and data science inputs, and no inputs from non-scientists. Summary of the Invention
[0011] A computer-implemented system for dynamically creating and validating predictive analytics models to transform data into actionable insights, the system comprising: an analytics server communicatively connected to: (a) a sensor configured to acquire real-time data output from an electrical system, and (b) a terminal configured to display a set of markers for analytics server operations; the analytics server comprising: an event recognition circuit for identifying events; a marker circuit for selectively markering at least one time-series data region (data region) in which identified events occur based on analytics expertise; a decision engine for comparing the data region in which identified events are marked with data regions in which identified events are not marked with data regions in which identified events are not marked; an analytics modeling engine for constructing a predictive analytics model embodying classifications generated by selective markers based on analytics expertise; a terminal communicatively connected to the analytics server and configured to display visual notations generated by executing the predictive analytics model; and a feedback engine for validating the predictive model based on feedback from at least one domain expert.
[0012] A method for dynamically creating and validating predictive analytics models to transform data into actionable insights, the method comprising: identifying events; selectively labeling at least one time-series data region (data region) where the identified events occur based on analytical expertise; comparing the labeled data region of the identified events with unlabeled data regions of the identified events; constructing a predictive analytics model based on analytical expertise that embodies classifications generated through selective labeling; displaying visual notations generated by executing the predictive analytics model; and validating the predictive model based on feedback from at least one domain expert.
[0013] These and other features, aspects, and embodiments of the invention are described below in the section entitled "Detailed Description". Attached Figure Description
[0014] Figure 1 An exemplary interface 100 showing data view features according to an embodiment of the disclosed invention is shown.
[0015] Figure 2A , 2B The 2C and 2D diagrams illustrate the interface where users can search by filtering data based on customer name, site, machine, and anomalies, respectively.
[0016] Figure 3-4 An exemplary interface is shown, allowing users to filter data of interest by selecting the machine and specific channel of interest.
[0017] Figure 5 The illustration shows an example of an add-channel function according to one embodiment of the disclosed invention.
[0018] Figure 6 An exemplary interface for data visualization of two data channels is shown.
[0019] Figure 7 An interface 700 according to an embodiment of the present invention is depicted, whereby a user can filter data of interest by selecting specific exceptions.
[0020] Figure 8 An exemplary interface 800 is provided, in which a user can view machines with selected anomalies and edit the default channel of the selected anomaly.
[0021] Figure 9 Examples of recently reported selection anomalies are provided.
[0022] Figure 10 The illustration shows a method by which a user selects a machine from a list of machines that have been reported as having selected anomalies, according to an embodiment of the disclosed invention.
[0023] Figure 11This indicates the load default channel option according to an embodiment of the present invention.
[0024] Figure 12 An example of the default channel list used for the selected exception is shown.
[0025] Figure 13 An exemplary visual representation of the selected data channel is shown.
[0026] Figure 14 Indicates an anomaly in the timeline according to one embodiment of the disclosed invention.
[0027] Figure 15 An anomaly in the timeline according to one embodiment of the disclosed invention is depicted.
[0028] Figure 16 Various data visualization options for users are shown according to an embodiment of the present invention.
[0029] Figure 17 This represents an example of a save view option in response to the user's selection of the save current view option.
[0030] Figure 18 The illustration shows a way in which a user can switch to a previously saved view according to an embodiment of the present invention.
[0031] Figure 19 Various properties of a previously saved view according to an embodiment of the present invention are shown.
[0032] Figure 20 The illustration shows an exemplary interface for creating exported channels.
[0033] Figure 21 An interface is depicted according to one embodiment of the disclosed invention, allowing users to create, edit, save, or publish exported channels.
[0034] Figure 22 An exemplary interface is shown that allows users to modify and test the algorithm used to derive new channels according to a preferred embodiment of the disclosed invention.
[0035] Figure 23 The interface displayed allows users to initiate the fleet creation process.
[0036] Figure 24 An instructive fleet settings interface that allows users to configure new fleet settings.
[0037] Figure 25 A list of machines to be added to a fleet according to one embodiment of the present invention is provided.
[0038] Figure 26An exemplary interface is shown where a newly created fleet is added to the fleet settings interface.
[0039] Figure 27 An interface is shown that allows a user to select data operations from various data operations according to an embodiment of the present invention.
[0040] Figure 28 It indicates an exemplary interface through which users can provide parameters for generating histograms.
[0041] Figure 29 An exemplary histogram generated by one embodiment of the disclosed system is shown.
[0042] Figure 30 The illustration shows a system diagram of various components according to one embodiment of the disclosed system.
[0043] Figure 31 The illustration shows a flowchart of a process according to an embodiment of the disclosed system. Detailed Implementation
[0044] Why Python?
[0045] Python is a widely used programming language by engineers and data scientists for developing high-value, advanced analytics models. The current login and validation process for typical analytics models is time-consuming and tedious, especially when validation must involve stakeholders who may be dealer employees and typically have backgrounds dissimilar to the data scientists generating the advanced analytics models, such as condition monitoring consultants and CMAs.
[0046] Many senior data scientists use Jupyter notebooks to develop their models. Jupyter notebooks are free-form text-editing web applications that can execute Python or R syntax and libraries. In other words, all tasks, including data search, data ingestion, data preparation, algorithm development, model execution, and results viewing, typically require decoding using Python syntax. Jupyter is a widely used analytics tool among senior data scientists. However, non-data scientists accepted as community members are envisioned to participate in the development and validation of analytical models. Therefore, it is understandable that Jupyter is not an easy tool to use for non-data scientist community members such as CMAs, engineers, and subject matter experts.
[0047] Why crowdsourcing?
[0048] The publicly available system integrates various data sources, data views, and rule-based failure models used in the analytics crowdsourcing environment for validation. The publicly available system supports instantaneous visualization of time-series data with few or no lag anomalies. The publicly available system provides dynamic visualization of time-series data with anomalies in real-time. Community members have the ability to create and experiment with analytical models in real time. Experiments can, for example, involve selecting data to be processed by the newly created analytical model, and viewing the data visualization interface generated by the newly created model after processing the selected data. The newly created analytical model can use exported channels, such as combined... Figure 21 It was discussed in detail.
[0049] like Figure 23 As shown, the system considers community input, that is, input from non-data scientist community members in addition to input from the data scientist community. Community members will have the following capabilities: storing and sharing templates used to create exported channels, storing and sharing model execution results, and via methods such as... Figure 23 The data visualizations shown are used for comparative analysis.
[0050] By effectively utilizing experts from (1) the crowdsourcing community, (2) the broader community of data scientists, (3) condition monitoring consultants (CMAs), and dealer community experts derived from years of hands-on use of the system, robust analytical models are generated from the publicly available system. Interestingly, each of these communities brings different insights to the system and together they form a reliable knowledge base.
[0051] Therefore, collaborative Internet applications have been proposed to leverage the insights of multiple communities (crowdsourcing, data scientists, condition monitoring consultants, and resellers). These communities can then contribute their expertise and domain knowledge to enhance the analytical capabilities of publicly available systems. For example, resellers typically have a deep understanding of a) the products they sell and b) the customers for those products.
[0052] This type of knowledge is considered a competitive advantage. Without feedback from resellers, data scientists may not know what data clients are looking for or what anomalies they want to see in data visualizations. Publicly available systems can streamline the analytics development and deployment process to dynamically produce analytical models.
[0053] Why collaborate:
[0054] Public systems provide community members with the ability to validate and modify publicly available analytical models. Collaboratively and iteratively, community members can increase the quantity, accuracy, effectiveness, and impact of flawed models. For example, data scientist communities, non-data scientist communities, and crowdsourcing communities can access, view, edit, test, and distribute any published analytical model within days rather than months.
[0055] The ease of access leads to an increased number of available failure models, which in turn can largely aid machine learning. Similarly, the publicly available system is expected to provide community members with free access to any released failure model, thus avoiding the high costs associated with commercial software packages that currently lack the robustness and generality of publicly available systems. Furthermore, the publicly available system enables a more efficient validation process and greater involvement of CMAsd in validating models. This increased efficiency and model accuracy result in higher recommendation outputs.
[0056] refer to Figure 1 An exemplary interface 100 according to an embodiment of the disclosed invention is shown. The term "user" is used in this application to refer to a data scientist or a member of the non-data scientist community (e.g., a CMA) who will use the system according to the invention. In response to a user selecting the industry option 102 to specify the industry for which the user is seeking data analysis, an industry dropdown list 104 is populated on interface 100. By selecting the view option 106, the user can view... Figure 2A The interface 200 is shown below. By selecting the creation option 108, data scientists can create, fine-tune, and backtest analytical models. A single user interface can simultaneously display the data visualization 1 and the Python code executed to generate the data visualization. Similarly, by selecting the deployment option 110, data scientists can create deployment configurations that identify which cloud server to use, how long to run it, and how frequently. By selecting the management option 112, the management assistant can gain access for authentication and data / feature rights.
[0057] Now go to Figure 2A The diagram shows a display interface 200 displaying four global filters 210, 220, 230, and 240. Users can view data of interest after filtering by following these steps: (1) selecting a specific customer name in the customer name selection area 210, (2) selecting a location in the location selection area 220, (3) selecting a machine in the machine selection area 230, or (4) selecting an anomaly from the anomaly selection area 240. When a user searches by customer name, the system searches a first database to retrieve information related to a specific customer and displays telemetry data using a first global filter, such as the customer name.
[0058] Users can provide a customer name by selecting it from the drop-down list in Customer Name 202, or they can manually enter the customer name in the Customer Name Input Area 210. In response to the user's selection of a customer name, the Site List 220, Machine List 230, and Error List 240 associated with the selected customer are populated on the interface 200.
[0059] Now go to Figure 2B The diagram illustrates interface 200, where users can filter data for searching by site name. Users can provide a site name by entering it in the site name input area 220 or by selecting it from the drop-down list in site name 204. Figure 2B The interface 200 indicates several site names, which are examples and not limitations. These site names are BMA Peak Decline, Ravensworth, and Glendell MARC.
[0060] Alternatively, such as Figure 2C As shown, users can search for data by filtering it by machine name on interface 200. The machine name can be entered into the machine name input area 230, or selected from the machine name drop-down list 206. Figure 2C As shown, the exemplary machine name drop-down list 206 lists several machine names, but is not limited to SSP00448, SSP00169, and SSP00170 as examples.
[0061] Similarly, users can filter data by exception name, such as Figure 2D The interface 200 is shown in the diagram. Users can provide an exception name by entering the desired exception name in the exception input area 240 or by selecting an exception name from the exception name drop-down list 208. Several exemplary exception names are shown in the list of exception names 208, but are not limited to, brake pressure v2, left wheel speed, and brake accumulator pressure, as examples.
[0062] Users can filter data of interest by selecting the machine and specific channels they are interested in, such as Figure 3-4 As shown, however, when the channel is a signal from a machine (time series data), anomalies can be considered for alarms. A user interface is provided to display the content by overlaying the anomaly on top of the time series data, and this user interface provides an important purpose for users to perform verification.
[0063] refer to Figure 3 The diagram illustrates an interface 300 according to one embodiment of the disclosed invention, indicating the selection of machine 330 and the selection of channel 302 for the selected machine 330. For example, in response to a user selecting machine SSP000337 in machine selection area 330, a list of channels associated with machine SSP000337 is displayed in channel selection area 302. In other words, by selecting the engine speed channel in channel selection area 302, time-series data of SSP000337 regarding the engine speed channel can be displayed on interface 300.
[0064] According to one embodiment of the disclosed invention, in order to filter data from a desired channel of a desired machine, a user can perform the following steps: (1) identify the machine of interest displayed in the machine selection area, (2) identify the channel of interest in the channel selection area, and (3) drag the selected channel onto the visualization area, and (4) cause the data from the selected channel for the selected machine to be visually displayed in the data visualization area. Such phenomena are shown... Figure 4 (Chinese) Reference Figure 4 Interface 400 indicates in response to the user-generated exemplary interface: (1) select machine “SSP000337” in machine selection area 430 ((2) select “engine speed” channel in channel selection area 402, and (3) drag the selected “engine speed” channel in data visualization area 404) so that data from the “engine speed” channel is displayed in machine vision display 408 in data visualization area 404 for SSP000337.
[0065] For the visual display 408, in the data display area 404, the x-axis, indicated by reference numeral 412, represents parameter xxx, while the y-axis, indicated by reference numeral 406, represents parameter yyy. After the user drags the selected channel in the data visualization area 404, the data from the selected channel is graphically represented in the area indicated by reference numeral 408 for the parameters selected on the X-axis 412 and Y-axis 406, where the X-axis indicates time and the Y-axis represents the channel name, such as engine speed measured in revolutions per minute (RPM). The user can zoom in, zoom out, or change the focus on a specific data range using the range selection tool 410.
[0066] Now go to Figure 5 Interface 500, as shown, illustrates an add-channel function according to an embodiment of the disclosed invention. Advantageously, if a user is interested in viewing a graphical representation comparing data generated by a previously selected channel with data generated by an additional / other (not yet selected) channel on interface 500, the user can select the add-channel option 504. When the add-channel option is selected, the user can again select another channel of interest in the channel selection area 502 and drag the newly selected channel to an unoccupied portion of the visualization area 508 to generate... Figure 6 The interface shown is 600.
[0067] exist Figure 6In the example interface 600, a segmented data visualization area 608 is illustrated, indicating a data visualization. The first portion 608A of the data visualization area 608 contains a data visualization representation of the previously selected "Left Rear Brake Fluid Temperature" channel, while the second portion 608B of the data visualization area 608 contains a data visualization representation of the newly selected "Right Rear Brake Fluid Temperature" channel. It is understood that users can easily adjust the height of the visualization bars, and users can cut off selected data areas using the selector tools on the timeline.
[0068] Alternatively, users can filter data of interest by selecting specific anomalies, in which case users can... Figure 7 On the interface 700 shown, in the exception selection area 740, users can select the exception of interest from the exception drop-down list. After selecting the desired exception, the user can choose the exception handler tool 702 to generate the exception. Figure 8 The interface shown is 800.
[0069] Now go to Figure 8 On interface 800, two options are presented to the user. The user selects option 810, which displays an explanation of "which machine has an anomaly". Figure 9 The interface shown is 900. Alternatively, by selecting option 820, the user can change the default channel list for selected exceptions.
[0070] Now for reference Figure 9 The diagram illustrates an interface 900 displaying a list of machines that have previously exhibited selected anomalies (908). Users can filter the list of machines displayed on interface 900 within specific time spans. For example, users can select options 902, 904, and 906 respectively to view the dates in the past year, past month, and past week for each machine in the list that reported anomalies. This feature helps identify the precise time the alarm was displayed, the duration of the alarm, and the time when corrective action was taken to resolve the issue that initially triggered the alarm.
[0071] Alternatively, in another embodiment, however in Figure 9As not shown, users can also filter data by the frequency with which machines have reported selected anomalies. For example, users can select a list of machines to enumerate those reporting anomalies once a year, once a month, or once a week. This feature is advantageous because users can contrast machines that require immediate attention (such as those reporting anomalies weekly) with those that do not require immediate attention (such as those reporting anomalies annually). Once verified by CMA, the alert or anomaly data becomes “foundational facts.” Foundational facts are labeled data and can be defined as data that must be available before data scientists can train supervised machine learning models. A large amount of foundational fact data can potentially lead to more accurate machine learning models.
[0072] Analysis iterations can be performed during the model validation process to prevent “edge scenarios” and generate optimal thresholds based on statistical and / or engineering experience and product / industry knowledge. Edge scenarios can be defined as situations approaching acceptable design constraints that may be selectively eliminated by the user. Initially, the analysis model may not use precise formulas to quickly detect the occurrence of specific failures; however, it can be iteratively fine-tuned by members of both the data scientist and non-data scientist communities.
[0073] Such iterative approaches attempt to identify various use cases from input from subject matter experts (such as members of the data scientist community). In this context, a use case, in its simplest form, can be a list of associations, typically defining interactions between associations reflecting the state of an equipment system to achieve an objective. In systems engineering, use cases can be used at a higher level than in software engineering, often representing tasks or stakeholder objectives. Detailed requirements can then be captured in a systems modeling language (SysML) or as contractual statements.
[0074] Furthermore, iterative methods adjust the analytical model's algorithm based on feedback received from the non-data scientist community, where feedback includes the results of those cases. Close collaboration between data scientists and subject matter experts will achieve optimal alerting standards for working best with applicable assets and increasing user throughput.
[0075] It should be understood that users can select a specific machine from the list of machines displayed on interface 900, such as Figure 10 As shown. Interface 1000 displays a list of machines that reported the selected anomaly, in which the user has selected machine "SSP00337" indicated by reference numeral 1002.
[0076] Refer again Figure 10 This shows the user selecting a machine from a list of machines that have reported selection errors. In response to the user selecting machine 1002 on interface 1000, the following is displayed: Figure 11 The interface shown is 1100.
[0077] refer to Figure 11 Each exception has a filtered list of channels associated with it, and by selecting the Load Default Channels option 1102 on interface 1100, the user can view the default channel list for the selected exception, as shown in Figure 1200.
[0078] Now for reference Figure 12 The interface 1200 displays a default list of channels for selection, and the user can drag and drop data channels of interest from the default channel list 1202 onto the visualization area 1204 to generate a visual representation of the selected data channels, such as... Figure 13 The interface shown in 1300 is shown in the image.
[0079] Now for reference Figure 13 Interface 1300 is shown, in which, after selecting two channels indicated by reference numerals 1302 and 1304, the user drags the selected channels 1302 and 1304 onto the data visualization area 1306 to generate a data visualization representation 1308. Additionally, a timeline 1310 is shown on interface 1300, which demonstrates how users can view anomalies identified for the selected machine as represented by icons on the timeline 1310. Figure 14 The date line 1310 is shown in more detail on interface 1400.
[0080] Figure 14 An anomaly is depicted on a timeline according to one embodiment of the disclosed invention. By scrolling left and right on timeline 1408, the user can see anomalies identified by anomaly icons 1402. Therefore, when such anomalies are displayed... Figure 4 When an anomaly occurs, timeline 1408 visually displays an anomaly icon 1402 and the corresponding date 1406. The user can select the anomaly icon 1402, causing the selector tool 1404 to quickly focus on and indicate it on the interface 1400. The selector tool 1404 displays data before and after the time interval indicated by the anomaly icon 1402 when the anomaly occurred. This event occurs... Figure 15 As shown in the image.
[0081] Figure 15 An anomaly in the timeline according to one embodiment of the disclosed invention is depicted. Now turn to Figure 15 Interface 1500 displays a data visualization from August 24, 30 to August 26, 30, dated October 5. This visualization can be used to identify data trends over a specific time span indicated by reference number 1502. The system's ability to not only visualize data but also narrow it down to specific anomalies is particularly useful in debugging the causes of anomalous machine behavior.
[0082] Furthermore, users can provide feedback on specific anomalies by labeling them as good, bad, or even anomalous. An anomaly can be characterized as good when it alerts to events important to the user. Conversely, an anomaly can be categorized as bad when it alerts to events unimportant to the user. Additionally, users can add comments reflecting why an anomaly is good or bad, and even suggest steps to take to improve bad anomalies.
[0083] Once the user selects the machine and / or anomaly before choosing a channel to display the data of interest, the combination of these parameters defines the view. In this case, on interface 1500, the user can select a specific asset 1510, and then select a specific fault model 1520 from the drop-down menu in the upper right corner. The user can then select a specific anomaly 1530 from the macro timeline view 1540, which displays several green dots. Each green dot represents an example when an anomaly is reported. For channels, the user can select the channel for that fault model by clicking the illuminate button 1550, or the user can create a view template. Optionally, the user can add additional channels as needed. Customer needs often change over time and with the application. However, for the same application, users may repeatedly generate the same view to access the same / similar information by creating, saving, and subsequently accessing the saved view.
[0084] refer to Figure 16 This illustrates various data visualization options for users according to an embodiment of the present invention. To save users the effort of creating the same view, the disclosed system... Figure 16 The interface 1600 provides several functions, such as saving the current view, my saved views, and histograms. Users can select option 1602 to highlight a given exception.
[0085] By selecting the "Save Current View" option 1606 from the drop-down menu 1604. In other words, the user can save the newly created current view, as shown on interface 1600, for future use. Figure 17 The interface 1700 shown is used. Users can select various options in the drop-down menu 1604 to perform other operations, such as displaying graphic labels, displaying selected data files, view data in Google Maps, and views in 3D motion.
[0086] Now go to Figure 17The user selects the "Save Current View" option 1606 on interface 1600 to generate the "Save View" window 1702 on interface 1700. The user provides a name for the view to be saved in the view name input area 1704. A list of selected channels is populated in window 1702, and the user can then delete any listed channels by selecting option 1706. Similarly, the user can add a new desired channel by selecting the "Add New Channel" option 1708.
[0087] The code for the saved view is displayed in code display area 1710. As described in more detail below, users can fine-tune view parameters by changing the code in code display area 1710. After providing the desired view name, desired channel, and desired code modifications, users can choose save option 1712 to save the newly created or modified view for later use. It is worth noting that unless users want to create entirely new views, typically, users are interested in making small but significant changes to the code for customization based on their application. It should be understood that users do not need to know computer languages such as Python, R, SAS, or SQL to make code changes.
[0088] Alternatively, such as Figure 18 As shown, users can switch to a previously saved view via selection interface 1800, displaying the saved view options 1802. Now refer to Figure 18 Using a previously saved view 1802, the user can choose to view one of the previously saved views. In response to the user's selection 1802, a drop-down list 1804 of all previously selected views is displayed. The user can select an option from the views. When an option is selected from the drop-down list of previously saved views 1804, details about the selected view are displayed to the user. This process... Figure 19 As shown in the image.
[0089] Reference Figure 19 This explains the properties of views previously saved by the user. Interface 1900 displays a view selection area 1902, which shows views previously saved by the user. When the user selects a specific view from the views listed in the view selection area 1902, various properties related to the selected and previously saved view are displayed on interface 1900, such as the view name 1904, the list of channels selected for the saved view 1906, and the software code of the saved view 1908. The user can choose the delete option 1910 to delete the saved view, or choose option 1912 to load the selected view. This feature is particularly valuable because it allows for personalized or customized view templates, where users can save new views and modify existing ones. Users can easily upload a view by selecting it from a drop-down list, rather than having to repeatedly create the view from scratch.
[0090] Figure 20 This demonstrates the user's ability to create exported channels. Exported channels are computations or datasets using original channels from the original dataset, created by the user, and can be reused by the same user or other users in the same or different applications. In response to the user selecting the Export Channels option 2002 on interface 2000, the display of interface 2100 is initiated. Alternatively, the user can select the Create Fleet option 2004 to generate the following combination... Figure 24 The discussion interface.
[0091] Now go to Figure 21 This allows users to create, edit, save, or publish exported channels, displaying a list 2106 of existing channels on interface 2100. When a user selects a specific channel, a set of channel attributes is displayed on interface 2100, including but not limited to channel name 2104, channel description 2108, channel input 2110, and delete input option 2112. This allows users to delete specific inputs, add new input options 2114, and the formula input area 2116. By modifying the code in the formula input area 2116, users can create new exported channels.
[0092] Then, the user can save the newly created exported channel by naming it with the first name. Similarly, the user can publish the exported channel. Before publishing, the newly created exported channel can be modified, for example, by further modifying the algorithm / formula, and it can be saved again with the first name. However, if the user wants to modify the published exported channel, the user can modify the published exported channel and save the modified exported channel as an exported channel, where the first name is different from the second name. In other words, the revised exported channel is a new revision of the existing exported channel.
[0093] By selecting the View Revision History option 2118 from any selected exported channel, users can view the changes made and who made them, but if needed, users can abort the update and revert to a previous version. Understandably, if a user wishes to create a new exported channel, they can select the "New" option 2102 and provide the aforementioned set of channel attributes, such as channel name, description, inputs, and formulas.
[0094] Reference Figure 22 An exemplary interface 2200 is shown, whereby after entering formula 2202, the user can test the entered code by selecting the test option 2204. It should be understood that the user does not need to know any specific programming language syntax to modify or provide formulas to modify the exported channel or generate a new exported channel.
[0095] After decoding the new algorithm, the user can test the newly written exported channel on a specific machine via test option 2204. Upon successful testing of the newly exported channel, a success notification 2206 is displayed on interface 2200. After successfully testing the exported channel, the user can save the newly written algorithm for personal use by selecting the "Add to My List" option 2210.
[0096] Alternatively, users can save and publish the algorithm by selecting save and publish option 2208 to immediately add the newly exported channel to the user's channel list. If the newly exported channel is saved as a dedicated channel, it is then added to the user's channel list.
[0097] Similarly, if a newly exported channel is saved as a public channel, it is added to the user's channel list and the channel lists of other users. After a newly exported channel is published, other users can easily add it to their respective channel lists. This saves valuable computing resources and improves computing efficiency, as one user's insights are easily communicated to other users, thus saving time and increasing productivity.
[0098] It is worth noting that the ability for users to modify existing formulas and provide new ones advantageously offers non-programmers the opportunity to change formulas to meet the needs of user-customized applications. Therefore, non-programmer users can use modified versions of the analysis tool, configured to meet the customized needs of their applications, without having to write the analysis tool's code from scratch.
[0099] refer to Figure 23 The description of the display interface 2300 will allow the user to initiate the fleet creation process. When the user selects the create fleet option 2302, the following can be used: Figure 3 The interface shown in step 2400 is used to start the fleet creation process.
[0100] Now go to Figure 24 The display interface 2400 shows that, to create a fleet, the user selects the "Create New Fleet" option 2402 on interface 2400. The user can then specify a fleet name in the fleet name input area 2404 and provide a description in the fleet description input area 2406. The user then selects the "Edit Machine List" option 2408 on interface 2400 to fill in the information. Figure 25 The interface shown is 2500.
[0101] Now describing Figure 25With machines displayed in machine display area 2502, the user can add any or all of the listed machines to the fleet the user is trying to create. The user can choose to add all or some of the machines displayed in machine display area 2502 to a selected machine display area 2506, and finally add the selected machines to the fleet created by the user.
[0102] To add a specific machine to a new fleet, a user can simply select the machine of interest and drag and drop it from zone 2502 to zone 2506. Alternatively, the user can add all machines displayed in zone 2502 to the new fleet by selecting option 2504. Conversely, the user can deselect previously selected machines by selecting option 2508. After selecting the machines to add to the new fleet, the user can select option 2510 to abort adding the selected machines to the new fleet. Alternatively, the user can select option 2512 to add the selected machines to the new fleet. If the user selects option 2512, the fleet settings interface 2600 is displayed.
[0103] Now for reference Figure 26 Interface 2600 is shown, which displays a fleet selection area 2602 showing the newly created fleet "BMA Peak Reduction". Machines previously added to the newly created fleet are displayed in the selected machine display area 2604. After creating the fleet and selecting machines, the user can delete the newly created fleet by selecting the delete option 2606.
[0104] Alternatively, users can cancel fleet creation by selecting option 2608, or save option 2610 to finally create a new fleet. Once created, the fleet is visible to the user in the fleet list. Similar to the viewing operation described above, users can select, edit, or delete previously created fleets.
[0105] Data visualization is important because it provides a way to validate a given analytical model and generate visual evidence to share with clients. Histograms are one way analytical tools visualize data for users.
[0106] Go to Figure 27 Interface 2700 displays a drop-down list 2702, which allows the user to select various data operations for the fleet, such as obtaining a histogram of the selected data and viewing the saved histogram. When the user selects "obtain histogram" for the fleet option 2704, as shown... Figure 28 The interface shown is 2800.
[0107] Now for reference Figure 28Interface 2800 is shown, where the user can provide parameters for generating a histogram, such as name, description, time span, bins, channels, etc. Alternatively, the user can select a group of machines to generate a histogram via option 2802. Alternatively, the user can select a group of machines for histogram generation via option 2804. Therefore, interface 2800 allows the user to visualize the histogram of the selected fleet or selected data. The user can abort histogram creation by selecting option 2806 to delete the histogram. Alternatively, the user can generate the histogram by selecting option 2808 to create the histogram. After creating the desired histogram, the user can view the histogram by selecting option 2810 to view the plot.
[0108] Turn now Figure 29 Interface 2900 is displayed, which shows an exemplary histogram generated by an embodiment of the disclosed system.
[0109] Now let's discuss Figure 30 The diagram illustrates a system diagram of various components according to one embodiment of the disclosed system. Data 3004 is typically received from many sources, such as equipment operating in different locations, and may also be received from external sources, such as dealer databases or externally created annotation data.
[0110] Once received, the data is sent to at least two locations, boxes 3002 and 3014. In other words, the data is sent to (or via) a set of approved analytics models 3002. Additionally, the data is also sent to box 3014, where the non-data scientist community can build new analytics models based on (1) the insights they develop while using the equipment (2) to meet the needs of their respective applications. In box 3016, the non-data scientist community can test and validate the newly built analytics models. Once testing is complete and if the newly created analytics model is approved for production use at box 3018, the newly created analytics model is stored in my SQL database 3012.
[0111] Alternatively, if the newly constructed analytics model is not approved for production at block 3018, the process moves to block 3014 after the non-data scientist analytics model is recycled into the non-data scientist community pool of models that are not yet perfected. Crowdsourcing community members can work together to try to overcome the shortcomings in these models. Collaboration is a significant benefit of crowdsourcing communities. While each community member may not have answers to all questions, collectively, the community may have or propose answers to most questions.
[0112] Typically, in box 3016, logs created during the testing and validation of a newly built analytics model can be stored in my SQL database 3012. If a non-data scientist community member decides to retain the newly created analytics model for personal use only, the newly approved analytics model is stored in MySQL database 3012. However, if a non-data scientist community member agrees to release the newly approved analytics model, then in box 3002, in addition to being stored in MySQL database 3012, the newly approved analytics model is released and made available to community members and merged into the approved models database.
[0113] All previously created analytical models are tested and approved in a stored or approved model database 3002. Received data is typically processed by at least one approved model; more specifically, at box 3006, code associated with at least one approved model is executed to generate a data visualization. Based on the results of the code execution, a log is created in box 3008 recording any anomalies or other similar logs encountered. ACM database 3010 stores the logs generated in box 3008, indicating anomalies that occurred during the execution of the approved analytical model. Similarly, ACM database 3010 may also store the data visualizations generated in box 3006. ACM database 3010 is configured with all data relevant to an application according to one embodiment of the disclosed invention.
[0114] Therefore, a computer-implemented system for dynamically creating and validating predictive analytics models to transform data into actionable insights is disclosed. The system includes: an analytics server communicatively connected to: a sensor configured to acquire real-time data output from an electrical system; and a terminal configured to display a set of markers indicating the operations of the analytics server.
[0115] The analytics server includes: an event recognition circuit for recognizing the occurrence of events; and a labeling circuit for selectively labeling at least one time-series data region (data area) where the identified events have occurred, based on analytical expertise. Additionally, the analytics server includes a decision engine for comparing the data regions where the identified events have been labeled with data regions where the identified events have not been labeled.
[0116] The analytical modeling engine, also part of the analytical engine, is used to build predictive analytical models that embody classifications generated by selective labeling, based on analytical expertise. Furthermore, endpoints are communicatively connected to the analytical server and configured to display visual notations generated by executing the predictive analytical models. Additionally, a feedback engine, also part of the publicly available system, is used to validate the predictive models based on feedback from at least one domain expert.
[0117] Industrial applicability
[0118] Now for reference Figure 31 A flowchart of a process according to one embodiment of the system is disclosed. Process 3100 begins when a data visualization request is received at box 3102 from a member of the data scientist community or at box 3106 from a non-data scientist community member. A data visualization request may typically contain three sub-parts: first, identifying the input data to be visualized, such as the channel of the data to be visualized; second, identifying anomalies in the data to be captured; and third, identifying anomalies in the data to be captured.
[0119] Combination Figures 2A-2D This describes the process of identifying the customer, site, machine, or anomaly in the data to be visualized. The second part of a data visualization request is identifying how the input data will be transformed. This phenomenon is combined with... Figure 21 Let's discuss this. The third part of the data visualization request defines how the transformed data should be displayed, for example, in... Figure 3-5 As discussed in the article.
[0120] In box 3104, the requested data is accessed via a data storage or data warehouse that stores various types of data, such as data related to various industries, including but not limited to energy, manufacturing, aviation, automotive, chemical, pharmaceutical, telecommunications, retail, insurance, healthcare, financial services, public sector and other industries.
[0121] Although the disclosed invention is described in relation to the visualization of machine performance data, those skilled in the art will understand that the disclosed invention can be applied to many other fields.
[0122] In box 3108, the analyzer or analysis server can determine the source of the data visualization request. If the process determines in box 3110 that the analysis model has been presented by a data scientist in a computer programming language and has been approved through testing, the process can move to box 3114 to determine whether the analysis model needs to be modified.
[0123] In box 3118, if it is determined that the requested modification is minor, the process moves to box 3120 to change the code in the programming language, thereby implementing the desired modification in the analysis model. Afterward, the process moves to box 3122 to test and verify the modified code before executing it in box 3124. Then, after visualizing the data in box 3126, the process exits in box 3128.
[0124] Alternatively, if in box 3130 it is determined that the requested modification is major, the process moves to box 3132 to initiate the software development project and testing. Once the software is developed, the process moves to box 3134 to execute the newly developed analysis code. After visualizing the data in box 3126, the process exits in box 3128 based on the newly developed software model.
[0125] However, if at box 3112 the process determines that the analysis model has been presented by a non-data scientist, written in a non-computer programming language, and approved through testing, the process moves to box 3136 to modify the code used for the analysis model via a user interface according to an embodiment of the disclosed invention.
[0126] Next, at box 3138, the process tests and validates the code modified via the user interface, continues executing the modified code to visualize the data, and optionally saves the modified code at box 3140 for later use. Additionally, after prompting non-data scientist community members to optionally publish the code at box 3142, the process exits at box 3128. Publishing the analytical model helps other community members, who can iteratively modify and improve the code to ultimately enhance the quality of the resulting analytical model.
[0127] The publicly available system offers numerous advantages. Some exemplary advantageous features of the publicly available system include, but are not limited to: a free-form Python syntax editor that provides high-performance integrated visualization, exported channel functionality, data creation functionality, feedback loop capabilities, etc. Furthermore, the publicly available system advantageously combines three separate modular functionalities: data management, analytical modeling, and displaying a user interface within a single integrated application.
[0128] Publicly available systems can enhance analytical capabilities and increase deployment efficiency, creating a competitive advantage. Publicly available systems can provide failure models much faster. Typically, developing an analytical model of average complexity takes several months. As mentioned above, using publicly available systems allows users to create, test, and deploy failure models instantly.
[0129] Because non-data scientist community users have the ability to directly modify only 10% of the algorithm that constitutes the code written by the data scientist community, valuable computing and human resources are saved; otherwise, these resources would be exhausted to modify the remaining 90% of the code. Therefore, the publicly available system provides an improvement to the computer itself.
[0130] All community members, data scientists, and non-data scientists are likely to benefit from the newly released analytics model because they are likely to (1) easily use the released analytics model as is, (2) make minor modifications to the released analytics model to generate new analytics models based on the released analytics model to meet the needs of their respective clients, and (3) build other analytics models based on / on the released analytics model, or build new analytics models based on insights gained from the released analytics model.
[0131] Similarly, crowdsourcing communities share the insights of community members (typically subject matter experts or end-user system experts) who publish analytical models, and then apply those insights and knowledge to future analytical models generated through crowdsourcing. This iterative process of creating and sharing increasingly accurate analytical models is truly beneficial and valuable to all parties involved. Such evolving analytical models can also perform excellent predictive analytics and prevent future damage.
[0132] Publicly available systems lead to more accurate predictive analytics because published models are supported by model validation evidence in the form of anomaly logs. The ability to instantaneously share histograms gives community members the opportunity to verify the claimed accuracy of published or soon-to-be-published analytical models.
[0133] Because public systems leverage pathways between the data scientist community and non-data scientist communities, they lead to an increase in the capacity of analytical models. The exchange of information and insights among community members can generate guidance, requests, and subsequently, additional fault models.
[0134] It should be understood that the foregoing description provides examples of the disclosed systems and techniques. However, it is conceivable that other implementations of the invention may differ in detail from the foregoing examples. All references to the invention or examples thereof are intended to refer to the specific examples discussed at that point and are not intended to imply any limitation on the scope of the invention more generally. All language used to distinguish and derogatoryly describe certain features is intended to lack preference for those features, but unless otherwise specified, does not completely exclude them from the scope of the invention.
[0135] Unless otherwise stated herein, the description of the range of values herein is intended only as a shorthand for individually referring to each individual value falling within that range, and each individual value is incorporated into the specification as if it were individually described herein. All methods described herein may be performed in any suitable order unless otherwise stated herein or clearly contradicted by the context.
[0136] In the context of describing the invention (particularly in the context of the appended claims), the use of the terms “a” and “an,” as well as “the” and “at least one” and similar indicators, should be interpreted to cover both the singular and the plural, unless otherwise stated herein or clearly contradicted by the context. The use of the term “at least one” following a list of one or more items (e.g., “at least one of A and B”) should be interpreted to mean one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise stated herein or clearly contradicted by the context.
[0137] Therefore, this invention includes all modifications and equivalents of the subject matter described in the appended claims as permitted by applicable law. Furthermore, any combination of the foregoing elements in all possible variations is covered by this invention unless otherwise stated herein or clearly contradicted by the context.
Claims
1. A computer-implemented system for dynamically creating and validating predictive analytics models to transform data (3004) into actionable insights, the system comprising: The analysis server is communicatively connected to: (a) a sensor configured to acquire real-time data (3004) output from an electrical system, and (b) a terminal configured to display a set of markers indicating the operation of the analysis server; The analysis server includes: Event recognition circuitry, used to identify events; A labeling circuit for selectively labeling at least one region (2502) of time series data (3004) based on analytical expertise, wherein the identified event occurs in the region (2502) of the time series data (3004); A decision engine is used to compare the data (3004) region (2502) where the identified event is tagged with the data (3004) region (2502) where the identified event is not tagged. An analytical modeling engine for building predictive analytical models that reflect classifications generated by selective labeling, based on analytical expertise; A terminal, communicatively connected to the analysis server and configured to display visual symbols generated by executing the predictive analysis model; and A feedback engine is used to validate the predictive analytics model based on feedback from at least one domain expert. The predictive analytics model described herein evolves iteratively and collaboratively based on feedback from communities that include two or more of the following: dealer communities, condition monitoring consultant communities, data scientist communities, and crowdsourcing communities. The construction of the predictive analysis model includes: Determine the source of the data visualization request, wherein the source is a member of the reseller community of the reseller community or a member of the data scientist community of the data scientist community, and When the data visualization request originates from a dealer community member, a new exported channel is created to selectively modify the algorithm corresponding to the predictive analytics model, which has been previously tested and approved. This allows dealer community members to change the algorithm corresponding to the predictive analytics model to meet the needs of their dealer business without writing code for the entire predictive analytics model. The algorithm corresponding to the predictive analysis model is written in a non-computer programming language.
2. The system according to claim 1, wherein the predictive analytics model is generated via a machine (1002) learning algorithm.
3. The system of claim 1, wherein the predictive analytics model becomes a basic fact once validated by the community.
4. The system of claim 1, wherein the iterative evolution of the predictive analytics model prevents edge scenarios and generates an optimal threshold range for at least one parameter monitored by the predictive analytics model.
5. The system of claim 2, wherein the collaborative evolution of the predictive analytics model is based on at least one of statistical data (3004), engineering experience, product knowledge, and industry knowledge to generate the optimal threshold range for the predictive analytics model.
6. The system according to claim 1, further comprising: A non-data scientist interface, enabling non-data scientists to search for specific assets or groups of assets to investigate at least one alert and provide feedback for new analytical models without having to code in a computer language; and The data scientist user interface allows data scientists to backtest the new analytical model for various use cases by validating the analytical results overlaid on machine data.
7. The system of claim 1, wherein the analysis server is configured to determine whether the feedback used to validate the predictive analytics model is written in a programming language or a non-programming language, and The analytics server is configured to modify a portion of the predictive analytics model's code by directly changing the code in a first case where the feedback is written in the programming language, and to modify a portion of the predictive analytics model's code via a user interface in a second case where the feedback is written in the non-programming language.
8. A method for dynamically creating and validating predictive analytics models to transform data (3004) into actionable insights, the method comprising: Identify events; Based on analytical expertise, at least one region (2502) of time series data (3004) is selectively labeled, wherein the identified event occurs; The data (3004) region (2502) where the identified event is tagged is compared with the data (3004) region (2502) where the identified event is not tagged; Based on analytical expertise, a predictive analytics model is constructed that includes classifications generated from selective labeling; Display the visual symbols generated by executing the predictive analytics model; and The predictive analytics model is validated based on feedback from at least one domain expert. The predictive analytics model described herein evolves iteratively and collaboratively based on feedback from communities that include two or more of the following: dealer communities, condition monitoring consultant communities, data scientist communities, and crowdsourcing communities. The construction of the predictive analysis model includes: Determine the source of the data visualization request, wherein the source is a member of the reseller community of the reseller community or a member of the data scientist community of the data scientist community, and When the data visualization request originates from a dealer community member, a new exported channel is created to selectively modify the algorithm corresponding to the predictive analytics model, which has been previously tested and approved. This allows dealer community members to change the algorithm corresponding to the predictive analytics model to meet the needs of their dealer business without writing code for the entire predictive analytics model. The algorithm corresponding to the predictive analysis model is written in a non-computer programming language.
9. The method of claim 8, wherein the predictive analytics model is generated via a machine (1002) learning algorithm.
10. The method of claim 8, wherein the predictive analytics model becomes a basic fact once validated by the community.
11. The method of claim 8, wherein the iterative evolution of the predictive analytics model prevents edge scenarios and generates an optimal threshold range for at least one parameter monitored by the predictive analytics model.
12. The method of claim 9, wherein the collaborative evolution of the predictive analytics model is based on at least one of statistical data (3004), engineering experience, product knowledge, and industry knowledge to generate the optimal threshold range for the predictive analytics model.
13. The method of claim 8, further comprising: Provides an interface for non-data scientists, allowing them to search for specific assets or groups of assets to investigate at least one alert and provide feedback on new analytical models without having to code in a computer language; and It provides a user interface for data scientists, allowing them to backtest the new analytical model for various use cases by validating the analytical results overlaid on machine data.
14. The method of claim 8, further comprising modifying a portion of the code of the predictive analytics model by directly changing the code in the first case where the feedback is written in a programming language, and modifying a portion of the code of the predictive analytics model by means of a user interface in the second case where the feedback is written in a non-programming language.
Citation Information
Patent Citations
System and method for managing a fleet of remote assets
US20110208567A9
System And Method For Storing Equipment Management Operations Data
US20140095554A1
Declarative debriefing for predictive pipeline
US20200151588A1