Metadata inheritance for data assets
A hierarchical model for data objects addresses the challenge of metadata propagation across diverse data sources, facilitating secure and efficient metadata inheritance for improved data visualization and business understanding.
Patent Information
- Application Number
- JP2023574495
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-01-28
- Filing Date
- 2022-05-26
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-05-26
AI Technical Summary
Organizations face challenges in effectively utilizing vast collections of data due to the difficulty in propagating metadata from diverse data sources to visualization creators or users, making it hard to improve business practices and understanding of business activities.
A hierarchical model is generated to represent dependencies between data objects, allowing metadata to be propagated by traversing the model based on queries, including security policies and timeout conditions to ensure secure and efficient metadata inheritance.
Enables secure and efficient propagation of metadata across diverse data sources, enhancing data visualization and understanding of business activities.
Smart Images

Figure 0007807468000001 
Figure 0007807468000002 
Figure 0007807468000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application is a utility patent application based on previously filed U.S. Provisional Patent Application No. 63 / 195,568, filed June 1, 2021, the benefit of the filing date of which is hereby claimed pursuant to 35 U.S.C. §119(e) and is further incorporated by reference in its entirety.
[0002] FIELD OF THE INVENTION The present invention relates generally to data visualization, and more particularly, but not exclusively, to managing data associated with objects included in a visualization. [Background technology]
[0003] Organizations are generating and collecting increasingly large amounts of data. This data may be associated with different parts of the organization, such as consumer activity, manufacturing activity, customer service, server logs, etc. For various reasons, it may be inconvenient for such organizations to effectively utilize their vast collections of data. In some cases, the volume of data may make it difficult to effectively utilize the collected data to improve business practices. Therefore, in some cases, organizations may use various applications or tools to generate visualizations based on some or all of their data. Using visualizations to represent data may enable organizations to improve their understanding of business activity, sales, customer information, employee information, key performance indicators, etc. In some cases, sophisticated visualizations may incorporate or otherwise rely on data from various sources within the organization, including different databases. In some cases, there may be many different visualizations that may rely on these diverse or different data sources. Often, the data defined in different data sources may include metadata, including descriptions, labels, tags, etc. Data source creators / designers may associate metadata with various data to convey information that may be of interest to the visualization creator or visualization user. However, the data objects accessed by a visualization creator or visualization user may be logically separate from the data in the data source. Thus, it may be difficult to determine when or whether metadata can be propagated to the data objects used by the visualization creator or visualization user. Thus, it is with respect to these and other considerations that the present invention has been made. [Brief explanation of the drawings]
[0004] Non-limiting and non-exhaustive embodiments of the present invention are described with reference to the following drawings, in which like reference numerals refer to like parts throughout the various views unless otherwise specified. For a better understanding of the described innovations, reference is made to the following detailed description of various embodiments, which should be read in conjunction with the accompanying drawings. [Figure 1] 1 illustrates an exemplary system environment in which various embodiments may be implemented. [Figure 2] 1 illustrates a schematic embodiment of a client computer. [Figure 3] 1 illustrates a schematic embodiment of a network computer. [Figure 4] FIG. 1 illustrates a logical architecture of a system for metadata inheritance for data assets, according to one or more of various embodiments. [Figure 5] 1 illustrates a logical representation of a portion of a system for metadata inheritance for data assets, according to one or more of various embodiments. [Figure 6] 1 illustrates a logical schematic diagram of a portion of a system illustrating dependencies within a data model in accordance with one or more of various embodiments. [Figure 7] 1 illustrates a logical schematic diagram of a portion of a dependency hierarchy showing at least some of the logical nodes in the dependency hierarchy, according to one or more of various embodiments. [Figure 8] 1 illustrates a logical schematic diagram of a data object including metadata information in accordance with one or more of various embodiments. [Figure 9] 1 depicts a simplified flowchart of a process for metadata inheritance for data assets, according to one or more of various embodiments. [Figure 10] 1 illustrates a flowchart of a process for metadata inheritance for a data asset, according to one or more of various embodiments. [Figure 11] 1 illustrates a flowchart of a process for metadata inheritance for a data asset, according to one or more of various embodiments. [Figure 12]1 illustrates a flowchart of a process for metadata inheritance for a data asset, according to one or more of various embodiments. [Figure 13] 1 illustrates a flowchart of a process for information security associated with metadata inheritance for data assets, according to one or more of various embodiments. [Figure 14] 1 illustrates a logical representation of a query for metadata inheritance for a data asset, according to one or more of various embodiments. [Figure 15] 1 illustrates a logical representation of a query for metadata inheritance for a data asset, according to one or more of various embodiments. [Figure 16] 1 illustrates a logical representation of a metadata query for metadata inheritance according to one or more of various embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0005] Various embodiments will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific exemplary embodiments in which the invention may be practiced. However, embodiments may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the embodiments to those skilled in the art. Among other things, the various embodiments may be methods, systems, media, or devices. Accordingly, the various embodiments may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Therefore, the following detailed description is not to be construed in a limiting sense.
[0006] Throughout the specification and claims, the following terms shall take the meanings expressly associated therewith unless the context clearly dictates otherwise. As used herein, the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may. Additionally, as used herein, the phrase "in another embodiment" does not necessarily refer to different embodiments, although it may. Thus, as described below, various embodiments can be readily combined without departing from the scope or spirit of the invention.
[0007] Additionally, as used herein, the term "or" is an inclusive "or" operator and is equivalent to the term "and / or" unless the context clearly dictates otherwise. The term "based on" is not exclusive and allows for based on additional unlisted factors unless the context clearly dictates otherwise. Additionally, throughout this specification, the meanings of "a," "an," and "the" include plural references. The meaning of "in" includes "in" and "on."
[0008] With respect to the exemplary embodiments, the following terms are also used herein in accordance with their corresponding meanings, unless the context clearly dictates otherwise:
[0009] As used herein, the term “engine” refers to logic embodied in hardware or software instructions, which may be written in a programming language such as C, C++, Objective-C, COBOL, Java™, Kotlin, PHP, Perl, JavaScript, Ruby, VBScript, C#, or other Microsoft .NET™ languages. An engine may be compiled into an executable program or written in an interpreted programming language. Software engines may be called by other engines or by themselves. An engine as described herein refers to one or more logical modules that may be merged with other engines or applications or divided into sub-engines. An engine may be stored in a non-transitory computer-readable medium or computer storage device and stored on and executed by one or more general-purpose computers, thus creating a special-purpose computer configured to provide the engine. Also, in some embodiments, one or more portions of an engine may be a hardware device, ASIC, FPGA, or the like, that performs one or more operations in support of or as part of the engine.
[0010] As used herein, the term "data model" refers to one or more data structures that represent one or more entities associated with data collected or maintained by an organization. Data models are typically configured to model various operations or activities associated with an organization. In some cases, data models are configured to provide or facilitate various data-focused operations, such as efficient storage, querying, indexing, retrieval, and updating. In general, data models can be configured to provide features related to data manipulation or data management rather than to provide an easy-to-understand presentation or visualization of data.
[0011] As used herein, the term "data object" refers to one or more entities or data structures that comprise a data model. In some cases, a data object may be considered part of a data model. A data object may represent a class or type of item, such as a database, a data source, a table, a workbook, a visualization, a workflow, etc.
[0012] As used herein, the term "data object class" or "object class" refers to one or more entities or data structures that represent a class, kind, or type of data object.
[0013] As used herein, the term "display model" refers to one or more data structures that represent one or more representations of a data model that may be suitable for use in a visualization displayed on one or more hardware displays. A display model may define styling or user interface functionality that may be made available to users other than the creator.
[0014] As used herein, the term "dependency hierarchy" refers to one or more data structures that represent a specialized model for representing lineage information of a corresponding data model. A dependency hierarchy often includes field nodes for storing / representing values of data object attributes, computation nodes representing applied data transformations that may change the semantic meaning of the data, or flow nodes representing data transformations that do not substantially change the semantic meaning of the transformed data. Edges between nodes represent that data or information from nodes higher in the dependency hierarchy can be propagated to nodes further down in the dependency hierarchy. Data dependencies between data objects in a data model can be traced through the dependency hierarchy.
[0015] As used herein, the term "metadata query" refers to query information that represents a request for metadata information for one or more data objects. In some cases, the metadata query can be resolved using a dependency hierarchy that corresponds to the data model.
[0016] As used herein, the term "display object" refers to one or more data structures that comprise a display model. In some cases, a display object may be considered part of a display model. A display object may represent an individual instance of an item or an entire class or type of item that can be displayed in a visualization. In some embodiments, a display object may be considered or referred to as a view because it provides a view of a portion of a data model.
[0017] As used herein, the term "anchor field" refers to a field node in a dependency hierarchy selected as a starting point for determining metadata inheritance. A metadata query can define an anchor field that is often the focus of the metadata query.
[0018] As used herein, the term "panel" refers to an area within a graphical user interface (GUI) that has a defined geometry (e.g., x, y, z dimensions) within the GUI. A panel may be positioned to display information to a user or to host one or more interactive controls. The geometry or style associated with a panel may be defined using configuration information, including dynamic rules. Also, in some cases, a user may perform actions on one or more panels, such as moving, showing, hiding, resizing, and reordering.
[0019] As used herein, the term "configuration information" refers to information that may include rule-based policies, pattern matching, scripts (e.g., computer-readable instructions), etc., that may be provided from a variety of sources, including configuration files, databases, user input, built-in defaults, etc., or combinations thereof.
[0020] The following briefly describes embodiments of the invention in order to provide a basic understanding of some aspects of the invention. This brief description is not intended as an extensive overview. It is not intended to identify key or critical elements or to delineate or otherwise narrow the scope. Its purpose is merely to present some concepts in a simplified form as a prelude to the more detailed description that is presented later.
[0021] Briefly, various embodiments relate to managing data using a network computer. In one or more of the various embodiments, a hierarchical model can be generated that includes one or more edges that represent dependencies between one or more field nodes, one or more computational nodes, or one or more flow nodes.
[0022] In one or more of the various embodiments, in response to a query to determine one or more values of metadata associated with an anchor field that is a field node in the hierarchical model, further operations are performed including traversing the hierarchical model upward from the anchor field based on the query and the hierarchical model, wherein performing the further operations includes collecting one or more values of metadata corresponding to the visited field nodes in response to visiting one or more field nodes in the hierarchical model, the traversal being terminated based on the type of query; terminating traversal of the hierarchical model associated with the visited computational node in response to visiting a computational node; and terminating traversal of the hierarchical model associated with the visited flow node in response to visiting a flow node that depends on two or more other nodes in the hierarchical model.
[0023] In one or more of various embodiments, a response to the query may be provided that includes one or more collected values of the inheritable metadata for the anchor field.
[0024] In one or more of various embodiments, terminating the traversal based on the type of query may include terminating the traversal of the hierarchical model in response to the query being of a first query type, in response to visiting a first ancestor field node of the anchor field that may be associated with a value of the metadata such that the value of the metadata collected from the first ancestor field node may be provided as metadata for the anchor field.
[0025] In one or more of various embodiments, terminating the traversal based on the type of query may include performing further operations in response to the query being a second query type, the operations including visiting each field node in the hierarchical model that is an ancestor of the anchor field, such that visiting an intervening computational node or an intervening multi-input flow node terminates traversal of the hierarchical model, and collecting one or more values of metadata for each visited field node that is an ancestor of the anchor field, such that the one or more values of metadata may be sorted based on one or more dependencies in the hierarchical model that may correspond to the anchor field and one or more visited ancestor field nodes.
[0026] In one or more of various embodiments, traversing the hierarchical model may include determining one or more security policies associated with the hierarchical model based on the query and the client that submitted the query, comparing the one or more visited field nodes and the client to the one or more security policies, determining one or more restricted field nodes in the hierarchical model based on the comparison such that the one or more security policies exclude the client from accessing information associated with the one or more restricted field nodes, and excluding one or more values of metadata that may be associated with the one or more restricted field nodes from the response to the query such that one or more of the identifiers or data types associated with the one or more restricted field nodes may be included in the response to the query.
[0027] In one or more of various embodiments, one or more of the timeout value or node visitation limit value may be determined based on one or more of the query type or query. In some embodiments, in response to a time providing a response to the query exceeding the timeout value, the traversal of the hierarchical model is terminated and a partial response to the query including one or more collected values of the metadata is provided. Alternatively, in response to the number of visited nodes in the hierarchical model exceeding the node visitation limit value, the traversal of the hierarchical model is terminated and a partial response including one or more collected values of the metadata is provided.
[0028] In one or more of various embodiments, the type of query may be determined based on information provided with the query. In some embodiments, one or more conditions for terminating the traversal may be determined based on the type of query. In some embodiments, response information to include in a response to a query may be determined based on the type of query, such that the response information may include one or more of: one or more values of the metadata, one or more aggregate values based on the one or more values of the metadata, or the performance of one or more other actions defined by the type of query.
[0029] In one or more of various embodiments, in response to the anchor field being associated with a metadata value, the metadata value associated with the anchor field may be included in the response to the query, and in some embodiments, in response to the metadata value not being present in the anchor field, the metadata value associated with the ancestor field node in the hierarchical model closest to the anchor field is provided in the response to the query.
[0030] Operating environment shown in the diagram 1 illustrates components of one embodiment of an environment in which embodiments of the present invention may be practiced. Not all components are required to practice the present invention, and variations in the arrangement and type of components may be made without departing from the spirit or scope of the present invention. As illustrated, system 100 of FIG. 1 includes a local area network (LAN) / wide area network (WAN) 110, a wireless network 108, client computers 102-105, a visualization server computer 116, and the like.
[0031] At least one embodiment of client computers 102-105 is described in more detail below in connection with FIG. 2. In one embodiment, at least some of client computers 102-105 can operate over one or more wired or wireless networks, such as network 108 or 110. In general, client computers 102-105 can include virtually any computer capable of communicating over a network to send and receive information, perform various online activities, perform offline operations, and the like. In one embodiment, one or more of client computers 102-105 can be configured to operate within a business or other entity to perform various services for the business or other entity. For example, client computers 102-105 may be configured to operate as a web server, a firewall, a client application, a media player, a mobile phone, a game console, a desktop computer, and the like. However, client computers 102-105 are not limited to these services and may also be used in connection with end-user computing in other embodiments, for example. It should be appreciated that more or fewer client computers (as shown in FIG. 1) may be included in a system as described herein, and thus the embodiments are not limited by the number or type of client computers used.
[0032] Computers capable of operating as client computers 102 can include computers that typically connect using wired or wireless communication media, such as personal computers, multiprocessor systems, microprocessor-based or programmable electronic devices, network PCs, and the like. In some embodiments, client computers 102-105 can include virtually any portable computer capable of connecting to and receiving information from another computer, such as laptop computers 103, mobile computers 104, and tablet computers 105. However, portable computers are not so limited and may include other portable computers, such as cellular phones, display pagers, radio frequency (RF) devices, infrared (IR) devices, personal digital assistants (PDAs), handheld computers, wearable computers, and integrated devices that combine one or more of the preceding computers. Accordingly, client computers 102-105 typically range widely in terms of capabilities and features. Furthermore, client computers 102-105 can access a variety of computing applications, including browsers or other web-based applications.
[0033] A web-enabled client computer may include a browser application configured to send requests and receive responses over the web. The browser application may be configured to receive and display graphics, text, multimedia, and the like using virtually any web-based language. In one embodiment, the browser application is enabled to display and send messages using JavaScript, HyperText Markup Language (HTML), eXtensible Markup Language (XML), JavaScript Object Notation (JSON), Cascading Style Sheets (CSS), and the like, or a combination thereof. In one embodiment, a user of the client computer may use the browser application to perform various activities over a network (online). However, other applications may also be used to perform various online activities.
[0034] The client computers 102-105 may also include at least one other client application configured to receive or send content between other computers. The client application may include functionality for sending or receiving content, etc. The client application may further provide information that identifies itself, including type, functionality, name, etc. In one embodiment, the client computers 102-105 may uniquely identify themselves via any of a variety of mechanisms, including an Internet Protocol (IP) address, a telephone number, a mobile identification number (MIN), an electronic serial number (ESN), a client certificate, or other device identifier. Such information may be provided in one or more network packets transmitted between other client computers, the visualization server computer 116, or other computers, etc.
[0035] The client computers 102-105 may further be configured to include a client application that allows an end user to log in to an end user account that may be managed by another computer, such as the visualization server computer 116. Such an end user account may be configured to allow the end user to manage one or more online activities, including, in one non-limiting example, project management, software development, systems administration, configuration management, search activity, social networking activity, browsing various websites, communicating with other users, etc. The client computers may also be configured to allow the user to view reports, interactive user interfaces, or results provided by the visualization server computer 116.
[0036] Wireless network 108 is configured to couple client computers 103-105 and their components to network 110. Wireless network 108 may include any of a variety of wireless sub-networks that may be further overlaid, such as standalone ad-hoc networks, to provide infrastructure-oriented connectivity for client computers 103-105. Such sub-networks may include mesh networks, wireless local area network (WLAN) networks, cellular networks, etc. In one embodiment, the system may include two or more wireless networks.
[0037] The wireless network 108 may further include an autonomous system of terminals, gateways, routers, etc., connected by wireless links, etc. These connectors may be configured to move freely and randomly and organize themselves arbitrarily, such that the topology of the wireless network 108 may change rapidly.
[0038] The wireless network 108 may further employ multiple access technologies, including second (2G), third (3G), fourth (4G), and fifth (5G) generation wireless access for cellular systems, WLAN, wireless router (WR) mesh, etc. Access technologies such as 2G, 3G, 4G, 5G, and future access networks may enable wide-area coverage for mobile computers, such as the client computers 103-105, with various degrees of mobility. In one non-limiting example, the wireless network 108 may enable wireless connectivity via wireless network access such as Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wideband Code Division Multiple Access (WCDMA), High Speed Downlink Packet Access (HSDPA), Long Term Evolution (LTE), etc. In essence, wireless network 108 can include virtually any wireless communication mechanism by which information can travel between client computers 103-105 and another computer, network, cloud-based network, cloud instance, etc.
[0039] Network 110 is configured to couple network computers with other computers, including visualization server computer 116, client computer 102, and client computers 103-105, such as via wireless network 108. Network 110 can use any form of computer-readable medium for communicating information from one electronic device to another. Network 110 can also include a local area network (LAN), a wide area network (WAN), as well as the Internet, direct connections via universal serial bus (USB) ports, Ethernet ports, etc., other forms of computer-readable media, or any combination thereof. In an interconnected set of LANs, including those based on different architectures and protocols, routers act as links between LANs, allowing messages to be sent to one another. Furthermore, while communication links within a LAN typically include twisted wire pairs or coaxial cable, communication links between networks may utilize other carrier mechanisms, including analog telephone lines, fully or partially dedicated digital lines including T1, T2, T3, and T4, or wireless links including, for example, E-carrier, Integrated Services Digital Network (ISDN), Digital Subscriber Line (DSL), satellite links, or other communication links known to those skilled in the art. Furthermore, communication links may use any of a variety of digital signaling technologies, including, but not limited to, DS-0, DS-1, DS-2, DS-3, DS-4, OC-3, OC-12, OC-48, etc. Furthermore, remote computers and other associated electronic devices may be remotely connected to either the LAN or WAN via modems and temporary telephone links. In one embodiment, network 110 may be configured to transport information in Internet Protocol (IP).
[0040] Additionally, communication media typically embodies computer-readable instructions, data structures, program modules, or other transport mechanisms and includes any information non-transitory or transitory distribution media. By way of example, communication media includes wired media such as twisted pair, coaxial cable, fiber optics, and wave guides, as well as other wired and wireless media such as acoustic, RF, infrared, and other wireless media.
[0041] Additionally, one embodiment of the visualization server computer 116 is described in more detail below in connection with FIG. 3. While FIG. 1 depicts the visualization server computer 116 as a single computer, innovations or embodiments are not so limited. For example, one or more functions of the visualization server computer 116 or the like may be distributed across one or more separate network computers. Moreover, in one or more embodiments, the visualization server computer 116 may be implemented using multiple network computers. Furthermore, in one or more of various embodiments, the visualization server computer 116 or the like may be implemented using one or more cloud instances in one or more cloud networks. Thus, these innovations and embodiments should not be construed as limited to a single environment or other configuration, and other architectures are contemplated.
[0042] Exemplary Client Computer 2 illustrates an embodiment of a client computer 200 that may include more or fewer components than those shown. Client computer 200 may represent, for example, one or more embodiments of the mobile computer or client computer illustrated in FIG.
[0043] The client computer 200 may include a processor 202 in communication with a memory 204 via a bus 228. The client computer 200 may also include a power supply 230, a network interface 232, an audio interface 256, a display 250, a keypad 252, an illuminator 254, a video interface 242, an input / output interface 238, a haptic interface 264, a global positioning system (GPS) receiver 258, an open-air gesture interface 260, a temperature interface 262, a camera 240, a projector 246, a pointing device interface 266, and processor-readable permanent storage 234, processor-readable removable storage 236. The client computer 200 may optionally communicate with a base station (not shown) or directly with another computer. Also, in one embodiment, although not shown, a gyroscope may be used within the client computer 200 to measure or maintain the orientation of the client computer 200.
[0044] The power supply 230 can provide power to the client computer 200. Power can be provided using a rechargeable or non-rechargeable battery. Power can also be provided by an external power source, such as an AC adapter or a powered docking cradle, which supplements or recharges the battery.
[0045] The network interface 232 includes circuitry for coupling the client computer 200 to one or more networks and is configured for use with one or more communications protocols and technologies, including, but not limited to, protocols and technologies implementing any portion of the OSI model for mobile communications (GSM), CDMA, time division multiple access (TDMA), UDP, TCP / IP, SMS, MMS, GPRS, WAP, UWB, WiMax, SIP / RTP, GPRS, EDGE, WCDMA, LTE, UMTS, OFDM, CDMA2000, EV-DO, HSDPA, or any of a variety of other wireless communications protocols. The network interface 232 is sometimes known as a transceiver, a transceiver device, or a network interface card (NIC).
[0046] Audio interface 256 may be configured to generate and receive audio signals, such as the sound of a human voice. For example, audio interface 256 may be coupled to a speaker and microphone (not shown) to enable communication with others or to generate audio acknowledgments for some actions. The microphone in audio interface 256 may also be used for input to or control of client computer 200, such as using voice recognition, detecting touch based on sound, etc.
[0047] Display 250 may be a liquid crystal display (LCD), gas plasma, electronic ink, light emitting diode (LED), organic LED (OLED), or any other type of light reflective or light transmissive display that can be used with a computer. Display 250 may also include a touch interface 244 configured to receive input from an object such as a stylus or the finger of a human hand, and may sense touch or gestures using resistive, capacitive, surface acoustic wave (SAW), infrared, radar, or other technologies.
[0048] Projector 246 may be a remote handheld projector or an integrated projector capable of projecting an image onto a remote wall or any other reflective object, such as a remote screen.
[0049] Video interface 242 can be configured to capture video images, such as still photographs, video segments, infrared video, etc. For example, video interface 242 may be coupled to a digital video camera, a webcam, etc. Video interface 242 can include a lens, an image sensor, and other electronics. The image sensor can include a complementary metal-oxide semiconductor (CMOS) integrated circuit, a charge-coupled device (CCD), or any other integrated circuit for sensing light.
[0050] Keypad 252 may comprise any input device configured to receive input from a user. For example, keypad 252 may include a touch-sensitive numeric dial or a keyboard. Keypad 252 may also include command buttons associated with selecting and sending images.
[0051] Illuminator 254 may provide a status indication or may provide light. Illuminator 254 may remain active for a specific period of time or in response to an event message. For example, when illuminator 254 is active, it backlights the buttons on keypad 252 and may remain on while the client computer is powered on. Illuminator 254 may also backlight these buttons in various patterns when certain actions are performed, such as dialing another client computer. Illuminator 254 may also cause a light source located within a transparent or translucent case of the client computer to illuminate in response to an action.
[0052] Additionally, client computer 200 may also include a hardware security module (HSM) 268 to provide additional tamper-resistant protection for generating, storing, or using security / cryptographic information, such as keys, digital certificates, passwords, passphrases, two-factor authentication information, etc. In some embodiments, the hardware security module may be employed to support one or more public key infrastructure (PKI) standards, to generate, manage, or store key pairs, etc. In some embodiments, HSM 268 may be a standalone computer, while in other cases, HSM 268 may be configured as a hardware card that can be added to the client computer.
[0053] Client computer 200 may also include input / output interface 238 for communicating with external peripherals or other computers, such as other client computers and network computers. Peripheral devices may include audio headsets, virtual reality headsets, display screen glasses, remote speaker systems, remote speaker and microphone systems, etc. Input / output interface 238 may utilize one or more technologies, such as Universal Serial Bus (USB), infrared, WiFi, WiMax, Bluetooth™, etc.
[0054] The input / output interface 238 may also include one or more sensors for determining geographic location information (e.g., GPS), monitoring power conditions (e.g., voltage sensors, current sensors, frequency sensors, etc.), monitoring weather (e.g., thermostats, barometers, anemometers, humidity detectors, precipitation gauges, etc.), etc. The sensors may be one or more hardware sensors that collect or measure data external to the client computer 200.
[0055] The haptic interface 264 may be configured to provide tactile feedback to the user of the client computer. For example, the haptic interface 264 may be used to vibrate the client computer 200 in a particular manner when another user of the computer is on the phone. The temperature interface 262 may be used to provide a temperature measurement input or a temperature change output to the user of the client computer 200. The open-air gesture interface 260 may detect physical gestures of the user of the client computer 200 by using, for example, a single or stereo video camera, radar, a gyro sensor in a computer held or worn by the user, or the like. The camera 240 may be used to track the physical eye movements of the user of the client computer 200.
[0056] The GPS transceiver 258 can determine the physical coordinates of the client computer 200 on the surface of the Earth, which typically outputs a location as latitude and longitude values. The GPS transceiver 258 can also further determine the physical location of the client computer 200 on the surface of the Earth using other geographic positioning mechanisms, including, but not limited to, triangulation, Assisted GPS (AGPS), Enhanced Observed Time Difference (E-OTD), Cell Identifier (CI), Service Area Identifier (SAI), Extended Timing Advance (ETA), Base Station Subsystem (BSS), etc. It is understood that under different conditions, the GPS transceiver 258 can determine the physical location of the client computer 200. However, in one or more embodiments, the client computer 200, via other components, may provide other information that can be used to determine the physical location of the client computer, including, for example, a Medium Access Control (MAC) address, an IP address, etc.
[0057] In at least one of various embodiments, applications such as operating system 206, client display engine 222, other client apps 224, and web browser 226 may be configured to use geographic location information to select one or more location-specific features, such as a time zone, language, currency, or calendar format. The location-specific features may be used in documents, visualizations, display objects, display models, action objects, user interfaces, reports, and internal processes or databases. In at least one of various embodiments, the geographic location information used to select the location information may be provided by GPS 258. Also, in some embodiments, the geographic location information may include information provided using one or more geographic location protocols over a network, such as wireless network 108 or network 111.
[0058] The human interface components may be peripheral devices physically separate from the client computer 200, allowing for remote input or output to the client computer 200. For example, information routed as described herein through a human interface component such as the display 250 or keyboard 252 may instead be routed via the network interface 232 to an appropriate remotely located human interface component. Examples of human interface peripheral components that may be remote include, but are not limited to, audio devices, pointing devices, keypads, displays, cameras, projectors, etc. These peripheral components may communicate via pico-networks such as Bluetooth™, Zigbee™, etc. One non-limiting example of a client computer having such peripheral human interface components is a wearable computer that may include a remote pico-projector along with one or more cameras that remotely communicate with the separately located client computer to sense a user's gestures toward a portion of an image projected by the pico-projector onto a reflective surface such as a wall or the user's hand.
[0059] The client computer may include a web browser application 226 configured to receive and send web pages, web-based messages, graphics, text, multimedia, etc. The browser application of the client computer may use virtually any programming language, including Wireless Application Protocol messages (WAP), etc. In one or more embodiments, the browser application may use Handheld Device Markup Language (HDML), Wireless Markup Language (WML), WMLScript, JavaScript, Standard Generalized Markup Language (SGML), HyperText Markup Language (HTML), Extensible Markup Language (XML), HTML5, etc.
[0060] The memory 204 may include RAM, ROM, or other types of memory. The memory 204 represents an example of a computer-readable storage medium (device) for storing information such as computer-readable instructions, data structures, program modules, or other data. The memory 204 may store a BIOS 208 for controlling the low-level operation of the client computer 200. The memory may also store an operating system 206 for controlling the operation of the client computer 200. It will be appreciated that this component may include a general-purpose operating system, such as a version of UNIX or Linux™, or a dedicated client computer communications operating system, such as the Windows Phone™, Android™, or IOS operating systems. The operating system may include or interface with a Java virtual machine module that enables control of hardware components or operating system operation via Java application programs.
[0061] Memory 204 may further include one or more data storage devices 210 that can be utilized by client computer 200 to store, among other things, applications 220 or other data. For example, data storage device 210 may also be used to store information describing various capabilities of client computer 200. The information may then be provided to another device or computer in any of a variety of ways, including as part of a header during a communication, on request, etc. Data storage device 210 may also be used to store social networking information, including address books, friend lists, aliases, user profile information, etc. Data storage device 210 may further include program code, data, algorithms, etc. used by a processor, such as processor 202, to execute and perform operations. In one embodiment, at least a portion of data storage device 210 may also be stored in another component of client computer 200, including, but not limited to, non-transitory processor-readable removable storage device 236, processor-readable persistent storage device 234, or external to the client computer.
[0062] Applications 220 may include computer-executable instructions that, when executed by client computer 200, send, receive, or process instructions and data. Applications 220 may include, for example, a client display engine 222, other client applications 224, a web browser 226, etc. Client computers may be configured to exchange communications with the visualization server computer, such as queries, searches, messages, notification messages, event messages, alerts, performance metrics, log data, API calls, and combinations thereof.
[0063] Other examples of application programs include calendars, search programs, email client applications, IM applications, SMS applications, Voice over Internet Protocol (VOIP) applications, contact managers, task managers, transcoders, database programs, word processing programs, security applications, spreadsheet programs, games, search programs, and the like.
[0064] Additionally, in one or more embodiments (not shown), client computer 200 may include, instead of a CPU, an embedded logic hardware device such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable array logic (PAL), or the like, or a combination thereof. The embedded logic hardware device may directly execute its embedded logic to perform operations. Also, in one or more embodiments (not shown), client computer 200 may include, instead of a CPU, one or more hardware microcontrollers. In one or more embodiments, the one or more microcontrollers may directly execute their own embedded logic to perform operations and access their own internal memory and their own external input / output interfaces (e.g., hardware pins or wireless transceivers) to perform operations such as a system on a chip (SOC).
[0065] Exemplary Network Computer 3 illustrates one embodiment of a network computer 300 that may be included in a system implementing one or more of the various embodiments. The network computer 300 may include more or fewer components than those shown in FIG. 3. However, the components shown are sufficient to disclose exemplary embodiments for implementing these innovations. The network computer 300 may represent, for example, one or more of the visualization server computers 116 of FIG. 1.
[0066] A network computer such as network computer 300 may include a processor 302 that can communicate with memory 304 via bus 328. In some embodiments, processor 302 may be comprised of one or more hardware processors or one or more processor cores. In some cases, one or more of the one or more processors may be specialized processors designed to perform one or more specialized operations, such as those described herein. Network computer 300 also includes a power supply 330, a network interface 332, an audio interface 356, a display 350, a keyboard 352, an input / output interface 338, a processor-readable persistent storage 334, and a processor-readable removable storage 336. Power supply 330 provides power to network computer 300.
[0067] Network interface 332 includes circuitry for coupling network computer 300 to one or more networks and is configured for use with one or more communications protocols and technologies, including, but not limited to, protocols and technologies implementing the Open Systems Interconnection Model (OSI model), Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), User Datagram Protocol (UDP), Transmission Control Protocol / Internet Protocol (TCP / IP), Short Message Service (SMS), Multimedia Messaging Service (MMS), General Packet Radio Service (GPRS), WAP, Ultra Wideband (UWB), IEEE 802.16 Worldwide Interoperability for Microwave Access (WiMax), Session Initiation Protocol / Real-time Transport Protocol (SIP / RTP), or any of a variety of other wired and wireless communications protocols. Network interface 332 is sometimes known as a transceiver, a transceiver device, or a network interface card (NIC). Network computer 300 may optionally communicate with a base station (not shown) or directly with another computer.
[0068] Audio interface 356 is configured to generate and receive audio signals, such as the sound of a human voice. For example, audio interface 356 may be coupled to a speaker and microphone (not shown) to enable communication with others or to generate audio acknowledgments for some actions. The microphone in audio interface 356 may also be used for input to or control of network computer 300, for example, using voice recognition.
[0069] Display 350 may be a liquid crystal display (LCD), gas plasma, electronic ink, light emitting diode (LED), organic LED (OLED), or any other type of light reflective or light transmissive display that can be used with a computer. In some embodiments, display 350 may be a handheld projector or picoprojector that can project an image onto a wall or other object.
[0070] Network computer 300 may also include an input / output interface 338 for communicating with external devices or computers not shown in Figure 3. Input / output interface 338 may utilize one or more wired or wireless communication technologies, such as USB™, Firewire™, WiFi, WiMax, Thunderbolt™, infrared, Bluetooth™, Zigbee™, serial port, parallel port, etc.
[0071] Input / output interface 338 may also include one or more sensors for determining geographic location information (e.g., GPS), monitoring power conditions (e.g., voltage sensors, current sensors, frequency sensors, etc.), monitoring weather (e.g., thermostat, barometer, anemometer, humidity detector, precipitation gauge, etc.), etc. The sensors may be one or more hardware sensors that collect or measure data external to network computer 300. Human interface components may be physically separate from network computer 300, allowing for remote input or output to network computer 300. For example, information routed as described herein through a human interface component such as display 350 or keyboard 352 may instead be routed via network interface 332 to an appropriate human interface component located elsewhere on the network. Human interface components include any component that enables a computer to receive input from or send output to a human user of the computer. Thus, pointing devices such as a mouse, stylus, trackball, etc. may communicate via pointing device interface 358 to receive user input.
[0072] The GPS transceiver 340 can determine the physical coordinates of the network computer 300 on the surface of the Earth, typically outputting the location as latitude and longitude values. The GPS transceiver 340 can also further determine the physical location of the network computer 300 on the surface of the Earth using other geographic positioning mechanisms, including, but not limited to, triangulation, Assisted GPS (AGPS), Enhanced Observed Time Difference (E-OTD), Cell Identifier (CI), Service Area Identifier (SAI), Extended Timing Advance (ETA), Base Station Subsystem (BSS), etc. It is understood that under different conditions, the GPS transceiver 340 can determine the physical location of the network computer 300. However, in one or more embodiments, the network computer 300, via other components, may provide other information that can be used to determine the physical location of the client computer, including, for example, a Medium Access Control (MAC) address, an IP address, etc.
[0073] In at least one of various embodiments, applications such as operating system 306, data management engine 322, display engine 324, lineage engine 326, and web services 329 may be configured to use geographic location information to select one or more location-specific features, such as a time zone, language, currency, currency format, or calendar format. The location-specific features may be used in documents, file systems, user interfaces, reports, display objects, display models, visualizations, and internal processes or databases. In at least one of various embodiments, the geographic location information used to select the location information may be provided by GPS 340. Also, in some embodiments, the geographic location information may include information provided using one or more geographic location protocols over a network, such as wireless network 108 or network 111.
[0074] Memory 304 may include random access memory (RAM), read-only memory (ROM), or other types of memory. Memory 304 represents an example of a computer-readable storage medium (device) for storing information such as computer-readable instructions, data structures, program modules, or other data. Memory 304 stores a basic input / output system (BIOS) 308 for controlling low-level operation of network computer 300. Memory also stores an operating system 306 for controlling the operation of network computer 300. It will be appreciated that this component may include a general-purpose operating system, such as a version of UNIX or Linux®, or a dedicated operating system, such as Microsoft Corporation's Windows® operating system or Apple Corporation's OSX® operating system. The operating system may include or interface with one or more virtual machine modules, such as a Java virtual machine module, which enables control of hardware components or operating system operation via Java application programs. Other runtime environments may also be included.
[0075] Memory 304 may further include one or more data storage devices 310 that may be utilized by network computer 300 to store, among other things, applications 320 or other data. For example, data storage device 310 may also be used to store information describing various capabilities of network computer 300. The information may then be provided to another device or computer in any of a variety of ways, including as part of a header during a communication, on demand, etc. Data storage device 310 may also be used to store social networking information, including address books, friend lists, aliases, user profile information, etc. Data storage device 310 may further include program code, data, algorithms, etc. used by a processor, such as processor 302, to execute and perform operations, such as those described below. In one embodiment, at least a portion of data storage device 310 may also be stored in another component of network computer 300, including, but not limited to, non-transitory media in processor-readable removable storage device 336, processor-readable non-transitory storage device 334, or any other computer-readable storage device within or external to network computer 300. Data storage 310 may include, for example, data models 314, display models 316, source data 318, etc. Data models 314 may store files, documents, versions, properties, metadata, data structures, etc. that represent one or more portions of one or more data models. Display models 316 may store display models. Source data 318 may represent memory used to store databases or other data sources that contribute to the underlying data for data models, display models, etc.
[0076] Applications 320 may include computer-executable instructions that, when executed by network computer 300, send, receive, or process messages (e.g., SMS, multimedia messaging service (MMS), instant messages (IM), email, or other messages), audio, video, and enable communication with another user of another mobile computer. Other examples of application programs include calendars, search programs, email client applications, IM applications, SMS applications, Voice over Internet Protocol (VOIP) applications, contact managers, task managers, transcoders, database programs, word processing programs, security applications, spreadsheet programs, games, search programs, and the like. Applications 320 may include a data management engine 322, a display engine 324, a lineage engine 326, web services 329, and the like, which may be configured to perform operations for embodiments described below. In one or more of various embodiments, one or more applications may be implemented as a module or component of another application. Further, in one or more of various embodiments, an application may be implemented as an operating system extension, module, plug-in, or the like.
[0077] Additionally, in one or more of various embodiments, the data management engine 322, the display engine 324, the lineage engine 326, the web services 329, etc., can operate in a cloud-based computing environment. In one or more of various embodiments, these applications, including the management platform, etc., can be implemented within virtual machines or virtual servers that can be managed in the cloud-based computing environment. In this regard, in one or more of various embodiments, applications can flow from one physical network computer to another within the cloud-based environment, depending on performance and scaling considerations that are automatically managed by the cloud computing environment. Similarly, in one or more of various embodiments, virtual machines or virtual servers dedicated to the data management engine 322, the display engine 324, the web services 329, etc. can be automatically provisioned and decommissioned.
[0078] Also, in one or more of various embodiments, the data management engine 322, display engine 324, lineage engine 326, web services 329, etc. may be located within virtual servers operating in a cloud-based computing environment rather than being tied to one or more particular physical network computers.
[0079] Additionally, network computer 300 may also include a hardware security module (HSM) 360 that provides additional tamper-resistant safeguards for generating, storing, or using security / cryptographic information, such as keys, digital certificates, passwords, passphrases, two-factor authentication information, etc. In some embodiments, the hardware security module may be used to support one or more standard public key infrastructure (PKI) systems, and may be used to generate, manage, or store key pairs, etc. In some embodiments, HSM 360 may be a standalone network computer; in other cases, HSM 360 may be configured as a hardware card that can be installed in the network computer.
[0080] Additionally, in one or more embodiments (not shown), the network computer 300 may include, instead of a CPU, an embedded logic hardware device such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable array logic (PAL), or the like, or a combination thereof. The embedded logic hardware device may directly execute its embedded logic to perform operations. Also, in one or more embodiments (not shown), the network computer may include, instead of a CPU, one or more hardware microcontrollers. In one or more embodiments, the one or more microcontrollers may directly execute their embedded logic to perform operations and access their own internal memory and their own external input / output interfaces (e.g., hardware pins or wireless transceivers) to perform operations such as a system on a chip (SOC).
[0081] Exemplary Logical System Architecture 4 illustrates a logical architecture of a system 400 for metadata inheritance for data assets, according to one or more of various embodiments. In one or more of various embodiments, the system 400 can include various components, such as a data model 402, which can include a variety of data objects ranging from one or more database objects to one or more visualizations. In this example, the data model 402 includes a database object 404, a database object 406, a table object 408, a table object 410, a table object 412, a workflow object 414, a data source object 416, a data source object 418, a workbook object 420, a sheet object 422, and a sheet object 424.
[0082] In one or more of various embodiments, a visualization server computer, such as visualization server computer 116, may be configured to use a data model, such as data model 402, that represents information that may be used to generate a visualization. In some embodiments, the data model may also be used to manage other parties in the visualization system, including users, authors, and the like.
[0083] In this example, data model 402 may have one or more root-level data objects, such as data object 404 and data object 406. Data object 404 and data object 406 represent databases that may be sources of information that drive the data model. For example, data object 404 may represent an SQL RDBMS associated with a portion of an organization, while data object 406 may represent an API gateway to another information provider or other database.
[0084] In one or more of various embodiments, data objects 408, 410, 412, etc. represent tables or table-like objects that may be provided by one or more databases. At this level of the data model, data objects may be considered to wrap or otherwise closely model entities provided from databases. Thus, in some embodiments, properties or attributes of tables or database objects may closely reflect their native representation, including attribute names, data types, table names, column names, etc. For example, a data manager may “import” a database or table into the data model such that the imported object retains some or all of the features or attributes available in the native format. In some cases, in some embodiments, one or more imported data objects may include metadata information that may also be imported.
[0085] In one or more of various embodiments, before the imported table objects can be used in a visualization, a data curator may need to perform or have one or more operations performed to prepare the visualization or the information available to a visualization creator. In this example, extract / transform / load (ETL) object 414 represents an ETL process that performs some processing on the information in table objects 410 and 412 before it is available for use in a visualization.
[0086] In one or more of various embodiments, a data source object, such as data source 416 or data source 418, represents data objects that may be available to a visualization author for incorporation into a visualization or other display model. In some embodiments, a data source object may provide a data curator with controls to manage or shape the information from a database (e.g., database 404 or database 406) that may be made available to a visualization or visualization author. For example, one or more tables in database 404 may contain sensitive information that an organization wants to exclude from a visualization. Thus, in some embodiments, by selecting mapping attributes from a table object to a data source object, a data curator can control how data from the underlying database is exposed. In some embodiments, a data curator can select specific columns or attributes from a table object to include in the data source. Also, in some embodiments, attribute names (e.g., column names) in a table object may be mapped to different names in the data source. For example, a table column named customer_identifier in a table object may be mapped to an attributed table column named "account number" in a data source. Additionally, in some embodiments, other transformations of mapping may be performed, such as data type conversion, aggregation, filtering, joining, etc. In some embodiments, extensive or complex transformations may be encapsulated in ETL objects, etc., although simpler or more general transformations may be possible without the use of separate ETL objects.
[0087] In one or more of various embodiments, the edge 448 represents a mapping from the table object to the data source. In this example, the edge 448 may represent one or more data structures that map one or more attributes (e.g., columns) of the table object 408 to the data source 416. Thus, in some embodiments, the edge 448 provides or is associated with one or more mapping rules or instructions that define what information from the table object 408 is available in the data source 416, as well as how the information from the table object 408 appears to a visualization creator.
[0088] In one or more of various embodiments, workbook object 420 represents a data object that may be associated with one or more user-level data objects, such as sheet object 422 or sheet object 424. In some embodiments, a visualization author may design a workbook, such as workbook object 420, based on information provided by one or more data sources, such as data source 416 or data source 418. In some embodiments, a visualization author may design a workbook that includes one or more sheets (e.g., sheet object 422 or sheet object 424). In some embodiments, a sheet object may include one or more visualizations, etc.
[0089] In one or more of various embodiments, sheet object 422 or sheet object 424 may represent some or all of the information that may be provided to a visualization engine or the like that provides one or more interactive visual applications or reports that may be utilized by a user. In this example, sheet object 422 or sheet object 424 may be considered to include or reference one or more of data, metadata, data structures, etc. that may be used to render one or more visualizations of information that may be provided by one or more databases. In some embodiments, a sheet may be arranged to include one or more display models, styling information, text descriptions, narrative information, stylized graphics, links to other sheets, etc.
[0090] Thus, in some embodiments, a sheet may be accessible to a user, such as user 426 or user 428. The content or visualizations in a sheet may depend on its design and the information on which it is based (e.g., information from database 404 or database 406). Typically, a sheet or included visualizations may depend on one or more attributes, columns, etc. from one or more databases. Similarly, in some embodiments, dependencies that may be associated with a database may propagate through other data objects, such as tables, data sources, workbooks, etc. In some cases, other data objects inserted between the sheet and its underlying database may introduce additional dependencies that may propagate to the sheet or visualization.
[0091] In one or more of various embodiments, it may be advantageous to propagate associated or eligible metadata that may be associated with a data object to one or more dependent data objects. However, in some cases, intervening data objects or data transformations may cause metadata associated with an ancestor object to become meaningless or confusing with respect to one or more descendant / dependent data objects or data object attributes. Thus, in some embodiments, the lineage engine may be configured to use various strategies (in the form of query types) to determine whether metadata may be eligible to be propagated to a dependent object.
[0092] 5 shows a logical representation of a portion of a system 500 for metadata inheritance of data assets, according to one or more of various embodiments. In this example, for some embodiments, data model 502 may be considered similar to data model 402 described above. However, this example illustrates how some of the data objects within the data model may be related or dependent.
[0093] In this example, data object 506 represents a data object that may have associated metadata information. Thus, in this example, some or all of the metadata associated with data object 506 may be eligible for propagation to various dependent data objects, such as data object 508, data object 510, data object 512, and data object 514. Also in this example, user 516 or user 518 may be a user that depends on or owns data object 512 or data object 514.
[0094] 6 illustrates a logical schematic diagram of a portion of a system 600 showing dependencies in a data model, according to one or more of various embodiments. In this example, system 600 includes data object 602, data object 604, data object 606, data object 608, etc.
[0095] As mentioned above, in one or more of the various embodiments, data objects within a data model may depend on other data objects within the same data model. In this example, data object 604, data object 606, and data object 608 depend on data object 602. Thus, in this example, data object 608 depends on data object 606, and so on. Note that while all data objects in this example are part of the same dependency tree with data object 602 as its root, there may be other data objects within the same data model that have different or separate dependencies.
[0096] In one or more of various embodiments, a data object in a data model may depend on one or more attributes of the other data objects and, therefore, may depend on one or more other data objects. In FIG. 6 , the lines connecting columns in a data object to other columns in other data objects represent dependencies between data objects at the attribute level. Thus, in this example, data object 602 has five columns that can be considered five attributes. In this example, the line connecting data object 602 to data object 604 represents data object 604's dependency on four attributes of data object 602. Similarly, in this example, data object 606 directly depends on three attributes of data object 604 and indirectly depends on three attributes of data object 602. Also, similarly, in this example, data object 608 directly depends on one attribute from data object 606 and indirectly depends on one attribute from data object 606, data object 604, and data object 602. Thus, in this example, all data objects depend on one or more of their parents, but dependencies between data objects may be based on fewer than all attributes. For example, line 618 represents that data object 608 depends on one attribute of data object 606 .
[0097] In some embodiments, dependencies between attributes, or data objects in general, may depend on one or more functions, filters, transformations, etc. that may be applied to attribute values as they are passed to descendent data objects. For example, a table data object may include a timestamp attribute stored as a Unix epoch timestamp. However, in this example, the dependent data source may have an attribute labeled Date that expects a conventional date representation rather than a Unix epoch timestamp. Thus, in this example, the Date attribute of the dependent data source may be associated with a conversion operation that converts the Unix epoch timestamp value provided by the parent table into a conventional date value that meets the requirements of the data source.
[0098] 7 illustrates a logical schematic diagram of a portion of a dependency hierarchy 700 illustrating at least a portion of the logical nodes in the dependency hierarchy, according to one or more of various embodiments. As described above, dependencies of data objects in a data model can be represented as a dependency hierarchy based on the attributes of the related data objects that contribute to the dependency. Accordingly, attributes of data objects represented in a dependency hierarchy may be referred to as fields to distinguish them from general attributes. In some embodiments, the dependency hierarchy may include one or more nodes, such as a field node, a flow node, a computation node, etc.
[0099] In one or more of various embodiments, a field node (e.g., a field) can represent a value of an attribute of a data object. For example, in some embodiments, a field corresponding to an attribute, such as a column, from a table from a source / originating database can provide the original value of the field. Meanwhile, a dependent field node can be considered to hold a value propagated from one or more nodes that may be higher in the dependency hierarchy.
[0100] Additionally, in some embodiments, one or more nodes in a dependency hierarchy may represent a computation (e.g., a computation node) that may be applied to a field value before it can be propagated to descendant fields. In some cases, in some embodiments, a computation node may generate a value from one or more field values that may be new or have semantic meaning from the input fields. For example, in some embodiments, a computation node configured to calculate the difference between two other field values to generate a new value for another field may be considered a computation node. Thus, in some embodiments, a computation node may be configured by a user or visualization creator to perform any computation on values from one or more input fields to generate values that can be propagated to other fields.
[0101] Also, in some embodiments, a flow node can represent a transformation operation that can be applied to a field value before it can be propagated to other fields. In one or more of various embodiments, a flow node can be used to declare a transformation operation that can be performed on a field value without changing the meaning of the field value. For example, flow nodes corresponding to operations such as text formatting, date formatting, truncation / rounding of numbers, etc. can be considered flow nodes. Also, although not shown here, a flow node may be configured to receive input values from two or more field nodes.
[0102] In this example, in some embodiments, dependency hierarchy 700 includes field node 702, field node 704, field node 706, field node 708, field node 710, field node 712, field node 714, field node 716, field node 718, and field node 720. Also in this example, dependency hierarchy 700 includes compute node 714 and flow node 720.
[0103] Thus, in some embodiments, the lineage engine may be configured to propagate the value of field node 702 to field node 706, field node 708, field node 710, and field node 712 without changing the value. In contrast, in this example, the lineage engine may be configured to provide values from field node 702 and field node 704 as inputs to computation node 714. Thus, in this example, in some embodiments, computation node 714 may propagate a new or modified value to field node 716. In some embodiments, the particular computation performed by computation node 714 may be considered a valid computation or transformation that may deviate from the semantic meaning of field node 704 or field node 702. Further, in this example, the value of field node 716 may be propagated to field node 718. Also in this example, in some embodiments, the value from field node 704 may be provided to flow node 720. Thus, in some embodiments, flow node 720 may be configured to perform one or more transformations on the value from field node 704. As noted above, the transformations performed by the flow nodes may be considered to leave intact the semantic meaning of the value of field node 704. Finally, in this example, the value produced by flow node 720 may be propagated to field node 722.
[0104] In one or more of various embodiments, the data management engine or lineage engine may be configured to automatically enforce rules or configurations defined in a data model or dependency model. Thus, in some embodiments, a user or other client can use one or more fields in a data object or visualization.
[0105] Those skilled in the art will appreciate that a production dependency hierarchy may include many more field nodes, compute nodes, flow nodes, etc. than are shown here. However, those skilled in the art will appreciate that dependency hierarchy 700 is at least sufficient to disclose the innovations contained herein.
[0106] Also, in some cases, for brevity or clarity, a field node, a computation node, a flow node, etc. may be referred to as a field, a computation, or a flow in the context of a dependency hierarchy.
[0107] Furthermore, in some embodiments, dependency hierarchies may be an inherent part of a data model. For example, rather than generating separate data structures for a data model and a dependency hierarchy, in some embodiments, a lineage engine may be configured to generate a data model that includes dependency hierarchy information. However, for purposes of brevity or clarity, dependency hierarchies may be described herein as being separate from their corresponding data models.
[0108] FIG. 8 illustrates a logical schematic diagram of a data object 800 including metadata information according to one or more of various embodiments. As described herein, in some embodiments, a data model may be configured to include or represent various data objects. In some embodiments, one or more data objects may be based on one or more original data sources, such as a database, a spreadsheet, a file, an archive, etc. Further, in some embodiments, a data object may include one or more composite objects that include data objects or data attributes from other data objects. For example, in some embodiments, a data source, workbook, worksheet, etc. may be a data object that may include one or more attributes from other data objects. As described above, a data object or data object attribute may be associated with metadata. For example, a database may be configured to provide database objects, such as tables, having one or more columns, where each table is named and each column is named. Thus, in this example, table names and column names are often used as data object names and data object attribute names in the data model. However, in some cases, in some embodiments, a database object may be associated with metadata that does not represent the values of the data object attributes but can provide insight into the meaning, context, or purpose of the database object.
[0109] In this example, in some embodiments, data object attribute 800 may be considered to represent one or more data structures that may be configured to represent a data object attribute in a data model. In this example, data object attribute 800 may include one or more inherent features, represented here by features 802. In this example, in some embodiments, the features represent various characteristics of attribute 800 that may be necessary or preferable for representing the attribute, such as identity / label, value, data type, etc. Those skilled in the art will understand that an attribute may include more or fewer inherent features depending on the kind of data object or the type of attribute. However, those skilled in the art will understand that feature 802 is intended to logically represent information that may be necessary to represent the value of a data object attribute. In this example, the label, value, and type represent inherent characteristics of data object 800. Thus, the label, value, and type may be considered necessary functionality that allows the attribute to be used in visualizations, reports, user interfaces, etc.
[0110] In contrast, in this example, metadata 804 represents metadata associated with data object attribute 800. In this example, description metadata is expressed as more fully explaining the purpose of data object attribute 800. Also in this example, metadata notes include a description describing the application / schema version that introduced this data object attribute. Clearly, in this example, feature 802 enables a display engine or the like to effectively display the value associated with the data object attribute. However, metadata 804 allows a user or visualization creator to gain an improved understanding of the context, purpose, or meaning regarding the purpose of the data object attribute. Thus, in this example, a display engine may enable a user or visualization creator to view the metadata information to help inform designing visualizations, reports, user interfaces, etc. that may include data object attribute 800.
[0111] In some embodiments, the metadata may include one or more of tags, data quality warnings, personally identifiable information (PII) detection (e.g., is this a field containing sensitive information), data freshness, hierarchical hints / information (e.g., distance between various field nodes and compute nodes, etc.), authentication / validation information (e.g., is the field trusted?), etc.
[0112] Furthermore, those skilled in the art will appreciate that a data object or data object attribute may support a wide variety of different types of metadata depending on the type of data, the source of the data, etc. Thus, in some embodiments, the lineage engine may be configured to process or recognize different types of metadata using rules, instructions, libraries, etc. provided via configuration information. Thus, in some embodiments, as new metadata types may be introduced for various data objects or data object attributes, the lineage engine may use the configuration information to take into account the new metadata types or other local requirements or circumstances.
[0113] Generalized behavior 9-13 depict generalized operations for metadata inheritance of data assets, according to one or more of various embodiments. In one or more of various embodiments, processes 900, 1000, 1100, 1200, and 1300 described in conjunction with FIGS. 9-13 may be implemented by or implemented on one or more processors on a single network computer, such as network computer 300 of FIG. 3. In other embodiments, these processes, or portions thereof, may be implemented by or on multiple network computers, such as network computer 300 of FIG. 3. In still other embodiments, these processes, or portions thereof, may be implemented by or implemented on one or more virtualized computers, such as a virtualized computer in a cloud-based environment. However, embodiments are not so limited, and various combinations of network computers, client computers, and the like may be utilized. Furthermore, in one or more of various embodiments, the processes described in conjunction with FIGS. 9-13 may be used for metadata inheritance for data assets in accordance with at least one of various embodiments or architectures, such as those described in conjunction with FIGS. 4-8. Additionally, in one or more of various embodiments, some or all of the operations performed by processes 900, 1000, 1100, 1200, and 1300 may be performed in part by a data management engine 322, a display engine 324, or a lineage engine 326 performed on one or more processors of one or more networked computers.
[0114] 9 illustrates an overview flowchart of a process 900 for metadata inheritance for data assets, according to one or more of various embodiments. After a start block, in one or more of various embodiments, a data model can be provided to a lineage engine at start block 902. As described above, a data management engine, display engine, etc. can be configured to generate a data model that can be used by a visualization creator or data curator to create data objects that can be associated with various data model layers or data object types within the data model.
[0115] At decision block 904, in one or more of various embodiments, if a metadata query can be provided to the lineage engine, control may flow to block 906. Otherwise, control may loop back to decision block 904. In one or more of various embodiments, the lineage engine may be configured to integrate with various clients or client applications that can provide metadata queries. For example, if a client application, such as a visualization authoring system, wishes to display metadata for a data object to a visualization creator, the client application may provide a metadata query to the lineage engine to provide eligible metadata, if any.
[0116] Also, in some embodiments, the metadata query information may vary depending on the type of query. Furthermore, in some embodiments, the lineage engine may be configured to support metadata queries using a variety of well-known or customized query languages, such as SQL, GraphQL, JSON, XML, Javascript, regular expressions, etc., or combinations thereof.
[0117] At block 906, in one or more of various embodiments, the lineage engine may be configured to determine a dependency hierarchy based on the data model. In one or more of various embodiments, the lineage engine may be configured to generate a dependency hierarchy for a data model on-demand or in advance. For example, in some embodiments, the lineage engine may be configured to automatically generate a dependency hierarchy corresponding to a data model when the lineage engine is launched or associated with the data model. In some embodiments, the dependency hierarchy may be inherently part of the data model. However, for brevity or clarity, dependency hierarchies are described herein separately from their corresponding data models.
[0118] At block 908, in one or more of various embodiments, the lineage engine may be configured to traverse the dependency hierarchy based on the query. In one or more of various embodiments, the metadata query information may include information identifying one or more fields or portions of the dependency hierarchy that may be relevant to the metadata query. For example, if the metadata query requests metadata for a field, the metadata query identifies this field as an anchor field for the metadata query.
[0119] At block 910, in one or more of various embodiments, the lineage engine can be configured to determine metadata information based on traversing the dependency hierarchy. In one or more of various embodiments, the metadata query type and the query itself can inform the lineage engine how to determine eligible metadata that responds to the query. In some embodiments, determining eligible metadata information (e.g., inheritable metadata) can include traversing or otherwise evaluating nodes in the dependency hierarchy. For brevity or clarity, operations related to evaluating relationships or nodes represented by the dependency hierarchy may be referred to as traversing the dependency hierarchy.
[0120] At block 912, in one or more of various embodiments, the lineage engine can be configured to provide metadata information in the query results. In one or more of various embodiments, different metadata queries or metadata query types can generate different types of results. Some metadata query types can provide metadata information based on the first qualified metadata (defined by the query) determined from the dependency hierarchy. In contrast, in some embodiments, other metadata queries or metadata query types can provide metadata for two or more fields. For example, a query for chained metadata can provide qualifications.
[0121] Control may then be returned to the calling process in one or more of various embodiments.
[0122] FIG. 10 illustrates a flowchart of a process 1000 for metadata inheritance for a data asset, according to one or more of various embodiments. After a start block, in one or more of various embodiments, at start block 1002, a metadata query may be provided to a lineage engine. In one or more of various embodiments, the lineage engine may be configured to provide one or more APIs that allow a client application to provide one or more metadata queries. In one or more of various embodiments, the metadata query information may be provided as one or more parameters via the one or more APIs. In some embodiments, the metadata query information may be provided as a JSON object that the lineage engine can parse to determine the metadata query information.
[0123] In one or more of the various embodiments, one or more APIs may allow a client to declare a query type, one or more anchor fields in a dependency hierarchy, etc. In some embodiments, the lineage engine may be configured to provide an API that allows a client to provide additional parameters, such as filters, security / authorization credentials, etc.
[0124] At block 1004, in one or more of various embodiments, the lineage engine may be configured to parse the metadata query to determine the type of metadata query. In one or more of various embodiments, the lineage engine may be configured to support one or more metadata query types. Thus, in some embodiments, the metadata query information may include direct indicators (e.g., labels), hints, etc. that enable the lineage engine to determine the query type.
[0125] In one or more of various embodiments, the lineage engine may be configured to use instructions, rules, libraries, etc. provided via configuration information to identify different query types. Thus, in some embodiments, the lineage engine may be configured to support the addition of new or different query types provided via configuration information.
[0126] At block 1006, in one or more of the various embodiments, the lineage engine may be configured to parse the metadata query to determine one or more anchor fields in the dependency hierarchy.
[0127] In one or more of various embodiments, an anchor field may be a field in a dependency hierarchy to which a metadata query may be directed. In one or more of various embodiments, the lineage engine may be configured to initiate a search for requested metadata in the anchor field. In some embodiments, the lineage engine may be configured to support metadata queries with one or more anchor fields. Thus, in one or more of various embodiments, a query that includes two or more anchor fields may execute a query for each anchor field included in the query.
[0128] In some embodiments, the anchor field identifier may be passed to the lineage engine via a query API that may be different from the query itself.
[0129] In one or more of the various embodiments, the lineage engine may be configured to traverse the dependency hierarchy based on the query type at block 1008. As described above, the lineage engine may be configured to generate the dependency hierarchy based on the relationships of the data objects in the data model.
[0130] In one or more of various embodiments, a lineage engine may be configured to respond to metadata queries on one or more anchor fields. Thus, in some embodiments, the lineage engine may be configured to traverse a dependency hierarchy to identify inheritable metadata values. Note that in some embodiments, an anchor field may have its own version of the requested metadata. Thus, in some cases, depending on the query type, the lineage engine may omit traversing the dependency hierarchy because the requested metadata value may be associated with the anchor field. In contrast, if an anchor field is not associated with the requested metadata, the lineage engine may traverse the dependency hierarchy to determine whether there are inheritable metadata values that can respond to the query.
[0131] In one or more of various embodiments, the lineage engine may be configured to perform one or more actions depending on the query type. In one or more of various embodiments, the lineage engine may be configured to determine the one or more actions based on rules, instructions, libraries, etc. provided by configuration information to take into account local requirements or circumstances.
[0132] For example, in some embodiments, some query types may be configured to provide responses that may be limited to the first metadata value that can answer the question. Also, for example, in some embodiments, other query types may be directed to more information, such as a report that identifies each field in a dependency hierarchy that may be associated with an inheritable metadata value. For example, in some embodiments, a query of type FIRST may return the first valid value of metadata determined by traversal of the dependency hierarchy. Similarly, in some embodiments, a query of type CHAIN may return requested metadata values for the entire portion of the dependency hierarchy visited in the traversal.
[0133] At block 1010, in one or more of various embodiments, the lineage engine can be configured to determine metadata information based on traversal of the dependency hierarchy and the query type. As briefly discussed above, the type of information can vary depending on the query type and the query information. Some queries can return a metadata value for one field, while other metadata queries can return metadata values for one or more fields. Similarly, in some embodiments, some queries may return an aggregate value, a count, true / false (e.g., exists), etc. Also, in some embodiments, if the requested query cannot be resolved, the lineage engine can be configured to report a null value or an error. In some cases, the lineage engine may be unable to obtain the metadata value for a field because the inheritance rules associated with the query may not provide a valid answer. For example, in some embodiments, if a field depends on two fields, the lineage engine may be prevented from automatically determining how to propagate the metadata. Thus, in this example, a null value may be reported. However, in this example, some query types may provide each alternative result rather than being unable to provide one result.
[0134] Additionally, in some embodiments, one or more data objects or corresponding fields in a dependency hierarchy may be associated with one or more permissions or privilege designations. Accordingly, in some embodiments, the lineage engine may be configured to verify that a user or client submitting a metadata query may be authorized to view or access the metadata or fields associated with the inheritable metadata. In some embodiments, the lineage engine may be configured to apply one or more information security policies depending on the metadata query, metadata query type, data model / object, etc. associated with the query results. In some embodiments, the lineage engine may be configured to use rules, instructions, etc. provided by configuration information to determine permission evaluations or user permissions associated with the metadata query results.
[0135] Control may then be returned to the calling process in one or more of various embodiments.
[0136] FIG. 11 illustrates a flowchart of a process 1100 for metadata inheritance for a data asset, according to one or more of various embodiments. After a start block, in one or more of various embodiments, at start block 1102, a metadata query for a first metadata value may be provided to a lineage engine. As described above, in some embodiments, the lineage engine may be configured to support various query types. In one or more of various embodiments, the available / allowed query types and rules, instructions, grammar, etc. for declaring queries for a particular query type may vary depending on the query type. For example, in some embodiments, the lineage engine may be configured to allow a client application to provide a name or symbol indicating the query type. In other embodiments, the lineage engine may be configured to infer the query type from the query information. For example, if the metadata query may be provided using SQL or GraphQL, the lineage engine may be configured to parse the SQL or GraphQL to determine the query type.
[0137] Thus, in some embodiments, the lineage engine can be configured to determine whether a metadata query is requesting the lineage engine to determine a first metadata value for one or more anchor fields. For example, in some embodiments, the query information can include a value such as “FIRST,” indicating that the client is requesting the first eligible metadata value, if any, for the anchor field. In this example, the FIRST query type can be configured to return the first eligible / valid value for the specified metadata. For example, if the query is requesting the value of the metadata “description” for the anchor field, the lineage engine can execute a query to find the first value of “description” for the anchor object based on the dependency hierarchy. Thus, if the anchor field has description metadata associated with it, that description information can satisfy the query. In contrast, if the anchor field does not have description metadata, the lineage engine can be configured to traverse the dependency hierarchy to determine whether there is inheritable metadata associated with other fields that can satisfy the query.
[0138] At block 1104, in one or more of various embodiments, the lineage engine may be configured to determine an anchor field based on the query. In one or more of various embodiments, the anchor field may be considered the primary subject of the query. Thus, in some embodiments, the query information may explicitly define the anchor field. In some embodiments, if multiple anchor fields may be included in the query information for a FIRST query type, the lineage engine may be configured to find the first eligible metadata value for each anchor field.
[0139] At block 1106, in one or more of various embodiments, the lineage engine may be configured to visit each next-higher node in the dependency hierarchy. In one or more of various embodiments, the lineage engine may be configured to traverse the dependency hierarchy starting from the anchor field. In one or more of various embodiments, the lineage engine may be configured to continue traversing until the query reaches a result or generates an error.
[0140] In some embodiments, the lineage engine may begin traversing the dependency hierarchy at one or more anchor fields declared in the metadata query.
[0141] At decision block 1108, in one or more of various embodiments, if the visited node may be a computational node, control may flow to block 1120. Otherwise, control may flow to decision block 1110. As noted above, in some embodiments, a field node in a dependency hierarchy may depend on one or more computational nodes. As noted above, a computational node may represent one or more computations (e.g., calculations) that may be performed on one or more other field values to generate a new or modified value.
[0142] In some cases, operations performed on a computational node may produce values with semantic meanings that may differ from the semantic meanings of the source / input fields. Thus, in some embodiments, the meaning or context of metadata associated with a field supplied to a computational node may not be relevant to the value produced by the computational node. Thus, in some embodiments, it may be disadvantageous to propagate such metadata to fields / nodes further down the dependency hierarchy, as the meaning of the metadata may no longer match the field value.
[0143] Thus, in some embodiments, a query for a first metadata value may not be answered if the computational node is encountered during traversal of the dependency hierarchy before other eligible metadata is discovered.
[0144] At decision block 1110, in one or more of various embodiments, if the visited node may be a flow node, control may flow to decision block 1112. Otherwise, control may proceed to block 1120. As mentioned above, in some embodiments, flow nodes may be similar to computation nodes in that they may represent transformations to field values that may be performed on values and passed on to descendant fields. However, in contrast to computation nodes, transformations associated with flow nodes may be assumed to preserve the semantic meaning of the transformed fields. Thus, in some embodiments, visiting a flow node in a traversal of a dependency hierarchy may not terminate a query as computation nodes do.
[0145] At decision block 1112, in one or more of various embodiments, if the visited flow node can have one input, control may flow to decision block 1114. Otherwise, control may proceed to block 1120. As described above, in some embodiments, a flow node may be configured to have multiple input nodes. In some embodiments, if a flow node has two or more input nodes, the lineage engine may be unable to determine which metadata to propagate up to the anchor field because it cannot select among one or more fields. Thus, in some embodiments, visiting a flow node with two or more input nodes may terminate the traversal and execution of the query.
[0146] At decision block 1114, in one or more of various embodiments, if the visited node has metadata information that matches the query, control may flow to flow block 1118. Otherwise, control may proceed to decision block 1116. For example, if the query requests the first "description" metadata of an anchor field, the query may be satisfied if the visited field node has a value of "description." In contrast, in some embodiments, if the visited field does not have a value for the requested metadata, traversal up the dependency hierarchy may continue.
[0147] At decision block 1116, in one or more of various embodiments, if traversal of the dependency hierarchy can continue, control may loop back to 1106. Otherwise, control may proceed to block 1120. In one or more of various embodiments, the lineage engine may be configured to continue traversing the dependency hierarchy upward until all ancestor nodes of the anchor field have been visited, unless the query terminates by encountering a qualifying metadata, computational node, or multi-input flow node. Note that the traversal termination conditions may vary depending on the metadata query or metadata query type.
[0148] Additionally, in some embodiments, the lineage engine may be configured to employ various safety / performance measures to limit the amount of time or resources consumed by the execution of a query. Thus, in some embodiments, the lineage engine may be configured to terminate a query if it takes too long. Thus, in some embodiments, the lineage engine may be configured to determine a timeout value from configuration information or the query itself. Also, in some embodiments, the lineage engine may be configured to limit a query by limiting the total number of nodes in the dependency hierarchy that may be visited during the query. For example, in some embodiments, the lineage engine may be configured to limit the traversal of the dependency hierarchy by limiting the number of nodes that may be visited to, for example, 20,000 nodes. Note that in some embodiments, the lineage engine may be configured to determine a timeout value or traversal / node count limit based on configuration information to take local requirements or circumstances into account. Furthermore, in some embodiments, different types of queries may have different timeouts or node visitation limits.
[0149] At block 1118, in one or more of various embodiments, the lineage engine can be configured to determine metadata information for the visited nodes. In one or more of various embodiments, the metadata requested by the query can be determined from a first visited field node that can be associated with the requested metadata information. In some embodiments, if an anchor field has the requested metadata information, this can be the anchor field. Otherwise, in some embodiments, the metadata information can be determined from a field node in the dependency hierarchy that can be an ancestor of the anchor field. Note that for the first query type described herein, the metadata value can be obtained from the first field node that has a value for the metadata requested by the query.
[0150] At block 1120, in one or more of various embodiments, the lineage engine may be configured to return query results. In one or more of various embodiments, the lineage engine may be configured to return result information including the value of the requested metadata, if it can be determined. Otherwise, in some embodiments, a result indicating that the requested metadata could not be determined may be provided. In some embodiments, the result information may indicate why the requested metadata value could not be determined. Thus, in some embodiments, the result information may indicate whether a computational node or a multi-input flow node was encountered. Also, in some embodiments, the result information may indicate that none of the field nodes in the dependency hierarchy contained the requested metadata.
[0151] Additionally, as mentioned above and described in more detail below, the lineage engine may be configured to modify metadata query results based on information security considerations such as field / data access restrictions, user role, client type, client source, etc.
[0152] In one or more of various embodiments, the lineage engine may be configured to provide the result information synchronously or asynchronously to the client that provided the query. In some embodiments, the result information may be returned in an API parameter or return value. In some embodiments, the result information may include one or more individual parameters in various data structures consistent with the API being used. For example, in some cases, the result information may be provided as one or more JSON objects, XML files, HTTP responses, etc. In some embodiments, the lineage engine may be configured to support two or more protocols or formats for returning query result information to the client that provided the query information.
[0153] Control may then be returned to the calling process in one or more of various embodiments.
[0154] 12 illustrates a flowchart of a process 1200 for metadata inheritance for a data asset, according to one or more of various embodiments. After a start block, in one or more of various embodiments, at start block 1202, a metadata query for chained metadata information may be provided to a lineage engine. In this example, the chained metadata information metadata query requests metadata information for some or all of the ancestors on which the anchor field may depend. Otherwise, providing the query information for the chained metadata query can be considered similar to that described for block 1102.
[0155] Also, in one or more of various embodiments, a query for chained metadata may be considered a query type, and thus, in some embodiments, determining eligible metadata or other results associated with a query may differ from other query types.
[0156] In one or more of the various embodiments, the lineage engine may be configured to determine an anchor field based on the query at block 1204. See the description of block 1104.
[0157] In one or more of various embodiments, the lineage engine may be configured to visit the next higher node in the dependency hierarchy at block 1206. See the description of block 1106.
[0158] At decision block 1208, in one or more of various embodiments, if the visited node may be a computational node, control may flow to decision block 1218. Otherwise, control may proceed to decision block 1210. For reasons similar to those described for the FIRST query type shown in FIG. 11, encountering a computational node may terminate the upward traversal of the dependency hierarchy. However, in some embodiments, a chained metadata query may traverse multiple branches in a dependency hierarchy. Thus, in some cases, in some embodiments, encountering a computational node while performing a chained metadata query may not terminate the query performance as described above for the first metadata query.
[0159] At decision block 1210, in one or more of various embodiments, if the visited node may be a flow node, control may proceed to decision block 1212. Otherwise, control may proceed to decision block 1214. Encountering a flow node while traversing upward through a dependency hierarchy may be considered similar to that described for the FIRST query type shown in FIG.
[0160] At decision block 1212, in one or more of various embodiments, if the visited flow node may have one input, control may proceed to decision block 1114. Otherwise, control may proceed to decision block 1218. For reasons similar to those described for the FIRST query type shown in FIG. 11 , the lineage engine may be configured to terminate traversing up the dependency hierarchy if a multi-input flow node may be encountered during traversal. However, in some embodiments, a chained metadata query may traverse multiple branches in a dependency hierarchy. Thus, in some cases, in some embodiments, encountering a multi-input flow node while performing a chained metadata query may not terminate query execution as described above for the first metadata query.
[0161] At decision block 1214, in one or more of various embodiments, if the visited node has the requested metadata information, control may proceed to block 1216. Otherwise, control may proceed to decision block 1218. Similarly, as described above for process 1100, the lineage engine may be configured to determine whether the metadata information associated with the visited field includes a metadata value that may be responsive to the query.
[0162] At block 1216, in one or more of various embodiments, the lineage engine may be configured to determine metadata information for the visited blocks. Similar to what is described above for process 1100, the lineage engine may be configured to collect metadata values that may be responsive to the query from the visited field nodes. However, in contrast to process 1100, the lineage engine may be configured to collect metadata values from each visited field node, rather than terminating the query after finding the first field node associated with metadata that satisfies the query. Thus, in some embodiments, the lineage engine may be configured to collect eligible metadata values from two or more field nodes encountered during traversal of the dependency hierarchy.
[0163] In one or more of various embodiments, at decision block 1218, if traversal of the dependency hierarchy should continue, control may loop back to block 1206. Otherwise, control may proceed to block 1220. In one or more of various embodiments, the lineage engine may be configured to continue traversing the dependency hierarchy along two or more branches or paths, depending on the configuration of the dependency hierarchy. Also, similar to process 1100, the lineage engine may be configured to implement one or more safety / performance protections, such as timeouts or node count limits. See the discussion above for process 1100.
[0164] At block 1220, in one or more of various embodiments, the lineage engine may be configured to return query results. In one or more of various embodiments, the lineage engine can be configured to provide chained metadata information that may have been encountered during traversal of the dependency hierarchy. In some embodiments, the results of the chained metadata query can provide client information for the anchor field as well as some or all of its ancestors in the dependency hierarchy. Otherwise, in some embodiments, providing the query result information can be considered similar to that described above for block 1120, including modifying the metadata query results based on information security considerations such as field / data access restrictions, user role, client type, client source, etc.
[0165] Control may then be returned to the calling process in one or more of various embodiments.
[0166] 13 illustrates a flowchart of a process 1300 for information security associated with metadata inheritance of data assets, according to one or more of various embodiments. After a start block, at start block 1302, as described above, in one or more of various embodiments, a lineage engine may be configured to generate metadata query results based on performing a metadata query.
[0167] At block 1304, in one or more of the various embodiments, the lineage engine can be configured to determine an associated security policy associated with the metadata query results.
[0168] In one or more of various embodiments, one or more fields in a dependency hierarchy or corresponding data objects may be associated with permissions or access rules that may be enforced by a lineage engine.
[0169] In one or more of various embodiments, the lineage engine may be configured to apply one or more security policies that may be associated with one or more users, clients, data sources, data objects, etc. to prevent unauthorized access to protected / sensitive information.
[0170] In some embodiments, the lineage engine may be configured to determine a security policy associated with a metadata query or a metadata query result. In some embodiments, the lineage engine can be configured to determine the security policy based on configuration information, which allows different organizations or data owners to establish global or local security policies tailored to their local requirements or circumstances. For example, in some embodiments, the security policy may be configured differently depending on various factors, including the query type, the data source, the data object, the client, the client type, the user, etc. Thus, in some embodiments, the lineage engine can be configured to determine the security policy based on rules, maps, instructions, etc., to select a security policy based on configuration information to take into account the local circumstances or requirements.
[0171] In one or more of various embodiments, the lineage engine may be configured to execute / evaluate one or more rules, conditions, instructions, etc. that may be defined in a security policy.
[0172] At block 1306, in one or more of various embodiments, the lineage engine can be configured to evaluate the metadata query results taking into account the determined security policy.
[0173] In one or more of various embodiments, at decision block 1308, if a security policy issue can be determined, control can flow to decision block 1310. Otherwise, control can proceed to block 1314.
[0174] In one or more of various embodiments, at decision block 1310, if the metadata query should be aborted, control may be returned to the calling process. Otherwise, control may proceed to block 1312.
[0175] In some cases, in some embodiments, a security policy may indicate that a metadata query should be immediately aborted or rejected so that no information or metadata query results are returned to the client that submitted the query. For example, if a security violation is determined to be associated with a malicious actor, the security policy may act as if the request was ignored rather than responding with a result.
[0176] In one or more of various embodiments, the lineage engine can be configured to modify the metadata query results according to the determined security policy at block 1312. In some embodiments, the lineage engine can be configured to hide / obscure one or more portions of the metadata query results, such as metadata values, field names, data sources, etc.
[0177] For example, in some embodiments, the lineage engine may be configured to remove results associated with restricted / protected fields and leave accessible fields in the results. In some cases, non-sensitive information, such as data type, data source, and hierarchy information, may be included in the metadata query results, and the restricted information is removed.
[0178] In one or more of various embodiments, the particular filtering or obfuscation of the initial metadata query results may depend on the rules, conditions, instructions, etc. declared in the determined security policy being implemented.
[0179] At block 1314, in one or more of various embodiments, the lineage engine may be configured to provide the modified metadata query results to the client / user that submitted the metadata query.
[0180] Control may then be returned to the calling process in one or more of various embodiments.
[0181] It will be understood that each block of each flowchart diagram, and combinations of blocks in each flowchart diagram, can be implemented by computer program instructions. These program instructions can be provided to a processor to manufacture a machine, such that the instructions, when implemented on the processor, create means for performing the operations specified in one or more of the flowchart blocks. The computer program instructions, when executed by the processor, can cause the processor to perform a series of operational steps to produce a computer-implemented process, such that the instructions executing on the processor provide steps for performing the operations specified in one or more of the flowchart blocks. The computer program instructions can also cause at least some of the operational steps shown in each flowchart block to be performed in parallel. Furthermore, some of the steps may be performed across two or more processors, as may occur in a multiprocessor computer system. Furthermore, one or more blocks or combinations of blocks in each flowchart diagram can also be performed simultaneously with other blocks or combinations of blocks, or in an order different from that illustrated, without departing from the scope or spirit of the present invention.
[0182] Therefore, each block of each flowchart diagram supports a combination of means for performing the specified operations, a combination of steps for performing the specified operations, and program instruction means for performing the specified operations. It will also be understood that each block of each flowchart diagram and combination of blocks in each flowchart diagram can be performed by a dedicated hardware-based system that performs the specified operations or steps, or a combination of dedicated hardware and computer instructions. The foregoing examples should not be construed as limiting or exhaustive, but rather as illustrative examples for illustrating implementation of at least one of various embodiments of the present invention.
[0183] Additionally, in one or more embodiments (not shown), the logic of the exemplary flowcharts may be executed using an embedded logic hardware device instead of a CPU, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable array logic (PAL), or the like, or a combination thereof. The embedded logic hardware device may directly execute its embedded logic to perform operations. In one or more embodiments, a microcontroller may be configured to directly execute its own embedded logic to perform operations and access its own internal memory and its own external input / output interface (e.g., hardware pins or a wireless transceiver) to perform operations such as a system on a chip (SOC).
[0184] Illustrated use cases 14 illustrates a logical representation of a query 1400 for metadata inheritance of a data asset, in accordance with one or more of various embodiments. In this example, the query 1400 includes a metadata query 1402 and a metadata query result 1404.
[0185] In this example, a query type "FIRST" is declared in the metadata query 1402. In this example, a query of this type can be considered to be configured to return the first inheritable metadata that matches the query. In this example, the metadata query requests a metadata attribute named "description" for the anchor field "Field C."
[0186] Thus, in this example, result 1404 indicates that the metadata description for field C can be inherited from field B, which indicates that the metadata "description" directly associated with "field C" is null (because it is not defined for the field). Thus, in this example, the metadata is inherited from "field B," which has a value of "BBBBB."
[0187] 15 illustrates a logical representation of a query 1500 for metadata inheritance of a data asset, according to one or more of various embodiments. In this example, the query 1500 includes a metadata query 1502 and a metadata query result 1504.
[0188] In this example, a query type "CHAIN" is declared in the metadata query 1402. In this example, a query of this type can be considered to be configured to return a portion of the dependency hierarchy (e.g., a subtree) that contains metadata for fields that may be ancestors of the anchor field "Field C." Thus, two metadata values are returned, one for each field in the dependency hierarchy that may be determined to be an ancestor of the anchor field (Field C).
[0189] 16 illustrates a logical representation of a metadata query result 1600 for metadata inheritance, according to one or more of various embodiments. In this example, the metadata query result 1600 represents the result of discovering inheritable metadata. Thus, in this example, the metadata attribute "description" is reported as being null (e.g., empty). In the example above (e.g., process 1100), a null result such as the metadata query result 1600 may be generated if a computational node or a multi-input flow node may be encountered within the dependency hierarchy.
Claims
1. 1. A method for managing data, implemented by a computer including one or more processors, comprising: generating a hierarchical model including one or more edges representing dependencies between one or more field nodes, one or more computational nodes, or one or more flow nodes; in response to a query to determine one or more values of metadata associated with an anchor field that is a field node within the hierarchical model, traversing the hierarchical model upward from the anchor field based on the query and the hierarchical model, and in response to visiting one or more field nodes in the hierarchical model, collecting the one or more values of the metadata corresponding to the one or more field nodes; if the query type is a first query type requesting a first value of the metadata corresponding to a first ancestor field node of the anchor field obtained by the traversal, terminating the traversal of the hierarchical model in response to visiting the first ancestor field node; if the query type is a second query type requesting the one or more values of the metadata corresponding to the one or more field nodes obtained by the traversal, terminating the traversal of the hierarchical model in response to visiting a computational node within the hierarchical model; if the query type is the second query type, terminating the traversal of the hierarchical model in response to visiting a flow node that depends on two or more other nodes in the hierarchical model; providing a response to the query that includes the one or more collected values of the metadata associated with the anchor field; and A method comprising:
2. The method described in claim 1, wherein if the query type is the first query type, the first value of the metadata collected from the first ancestor field node is provided as the value of the metadata associated with the anchor field.
3. The method described in claim 1, wherein when the query type is the second query type, the one or more values of the metadata are sorted based on one or more of the dependencies in the hierarchical model corresponding to the anchor field and the one or more field nodes.
4. traversing the hierarchical model determining one or more security policies associated with the hierarchical model based on the query and the client that submitted the query; comparing the one or more field nodes and the client to the one or more security policies; determining one or more restricted field nodes within the hierarchical model based on the comparison, wherein the one or more security policies exclude the client from accessing information associated with the one or more restricted field nodes; 10. The method of claim 1, further comprising: excluding the one or more values of the metadata corresponding to the one or more restricted field nodes from the response to the query, wherein one or more of an identifier or a data type associated with the one or more restricted field nodes is included in the response to the query.
5. determining one or more of a timeout value or a node visitation limit value based on a query type or one or more of the queries; responsive to the time for providing the response to the query exceeding the timeout value, terminating the traversal of the hierarchical model and providing a partial response to the query including the one or more collected values of the metadata; 10. The method of claim 1, further comprising: responsive to a number of visited nodes in the hierarchical model exceeding the node visitation limit, terminating the traversal of the hierarchical model and providing a partial response including the one or more collected values of the metadata.
6. A computer system comprising one or more processors and a memory storing a program, The program, when executed by the one or more processors, causes the computer system to: generating a hierarchical model including one or more edges representing dependencies between one or more field nodes, one or more computational nodes, or one or more flow nodes; in response to a query to determine one or more values of metadata associated with an anchor field that is a field node within the hierarchical model, traversing the hierarchical model upward from the anchor field based on the query and the hierarchical model, and in response to visiting one or more field nodes in the hierarchical model, collecting the one or more values of the metadata corresponding to the one or more field nodes; if the query type is a first query type requesting a first value of the metadata corresponding to a first ancestor field node of the anchor field obtained by the traversal, terminating the traversal of the hierarchical model in response to visiting the first ancestor field node; if the query type is a second query type requesting the one or more values of the metadata corresponding to the one or more field nodes obtained by the traversal, terminating the traversal of the hierarchical model in response to visiting a computational node within the hierarchical model; if the query type is the second query type, terminating the traversal of the hierarchical model in response to visiting a flow node that depends on two or more other nodes in the hierarchical model; providing a response to the query that includes the one or more collected values of the metadata associated with the anchor field; and A computer system that executes the above.
7. The computer system of claim 6, wherein when the query type is the first query type, the first value of the metadata collected from the first ancestor field node is provided as the value of the metadata associated with the anchor field.
8. The computer system of claim 6, wherein when the query type is the second query type, the one or more values of the metadata are sorted based on one or more of the dependencies in the hierarchical model corresponding to the anchor field and the one or more ancestor field nodes.
9. traversing the hierarchical model determining one or more security policies associated with the hierarchical model based on the query and the client that submitted the query; comparing the one or more field nodes and the client to the one or more security policies; determining one or more restricted field nodes within the hierarchical model based on the comparison, wherein the one or more security policies exclude the client from accessing information associated with the one or more restricted field nodes; 7. The computer system of claim 6, further comprising: excluding the one or more values of the metadata corresponding to the one or more restricted field nodes from the response to the query, wherein one or more of an identifier or a data type associated with the one or more restricted field nodes is included in the response to the query.
10. The program, when executed by the one or more processors, causes the computer system to: determining one or more of a timeout value or a node visitation limit value based on the type of query or one or more of the queries; responsive to the time for providing the response to the query exceeding the timeout value, terminating the traversal of the hierarchical model and providing a partial response to the query including the one or more collected values of the metadata; 7. The computer system of claim 6, further configured to: terminate the traversal of the hierarchical model in response to a number of visited nodes in the hierarchical model exceeding the node visitation limit; and provide a partial response including the one or more collected values of the metadata.
11. A computer-readable storage medium storing one or more programs, the one or more programs being configured to, when executed by a computer system, cause the computer system to: generating a hierarchical model including one or more edges representing dependencies between one or more field nodes, one or more computational nodes, or one or more flow nodes; in response to a query to determine one or more values of metadata associated with an anchor field that is a field node within the hierarchical model, traversing the hierarchical model upward from the anchor field based on the query and the hierarchical model, and in response to visiting one or more field nodes in the hierarchical model, collecting the one or more values of the metadata corresponding to the field nodes; if the query type is a first query type requesting a first value of the metadata corresponding to a first ancestor field node of the anchor field obtained by the traversal, terminating the traversal of the hierarchical model in response to visiting the first ancestor field node; if the query type is a second query type requesting the one or more values of the metadata corresponding to the one or more field nodes obtained by the traversal, terminating the traversal of the hierarchical model in response to visiting a computational node within the hierarchical model; if the query type is the second query type, terminating the traversal of the hierarchical model in response to visiting a flow node that depends on two or more other nodes in the hierarchical model; providing a response to the query that includes the one or more collected values of the metadata associated with the anchor field; and A computer-readable storage medium that causes the computer to execute the method.
12. A computer-readable storage medium as described in claim 11, wherein if the query type is the first query type, the first value of the metadata collected from the first ancestor field node is provided as the value of the metadata associated with the anchor field.
13. A computer-readable storage medium as described in claim 11, wherein when the query type is the second query type, the one or more values of the metadata are sorted based on one or more of the dependencies in the hierarchical model corresponding to the anchor field and the one or more ancestor field nodes.
14. traversing the hierarchical model determining one or more security policies associated with the hierarchical model based on the query and the client that submitted the query; comparing the one or more field nodes and the client to the one or more security policies; determining one or more restricted field nodes within the hierarchical model based on the comparison, wherein the one or more security policies exclude the client from accessing information associated with the one or more restricted field nodes; 12. The computer-readable storage medium of claim 11, further comprising: excluding the one or more values of the metadata corresponding to the one or more restricted field nodes from the response to the query, wherein one or more of an identifier or a data type associated with the one or more restricted field nodes is included in the response to the query.
15. The one or more programs, when executed by the computer system, cause the computer system to: determining one or more of a timeout value or a node visitation limit value based on the type of query or one or more of the queries; responsive to the time for providing the response to the query exceeding the timeout value, terminating the traversal of the hierarchical model and providing a partial response to the query including the one or more collected values of the metadata; 12. The computer-readable storage medium of claim 11, further causing the computer to perform: in response to a number of visited nodes in the hierarchical model exceeding the node visitation limit, terminating the traversal of the hierarchical model and providing a partial response including the one or more collected values of the metadata.
Citation Information
Patent Citations
Device and / or method for providing query responses based on ephemeral data
JP2018518737A
Elimination of common subexpressions in complex database queries
US10901990B1
Interactive lineage analyzer for data assets
US20200334277A1