Rapidly develop user intent and analytical specifications in complex data spaces
By receiving and structuring user stories as phrase entities within the template, applying natural language processing and constructing knowledge graphs, the problem of rapid development of user analytical intentions in complex data spaces is solved, and the acceleration and specificity improvement of analytical requirements are achieved.
Patent Information
- Application Number
- CN202210876414.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-26
- Filing Date
- 2022-07-25
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-07-25
AI Technical Summary
The difficulty in developing and understanding user analytical intentions in the prior art in complex data spaces leads to a lack of specificity in analytical requirements and increases the risk of product development time and communication errors.
By receiving and structuring user stories as phrase entities within the template, natural language processing is used to discover data relationships, build knowledge graphs and link entities, train models to match technical requirements, and automatically configure the Q&A system.
Accelerate the cataloging and understanding of user analytical intentions in complex data spaces, improve the specificity of analytical requirements, and reduce product development time and communication errors.
Smart Images

Figure CN115687631B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to machine learning, and more particularly to cataloging, understanding, and accelerating the establishment of user analytical intent and development requirements within complex data spaces. Background Art
[0002] The creation of conventional analytical wizards configured to classify data typically focuses on parsing natural language processing (NLP) searches to retrieve data from multiple data sources to be displayed in data visualization or tabular formats. Classification methods can also be used to create prototypes by combining data from multiple sources. Some conventional configurable analytical models and visualizations enable end users to flexibly turn on / off "switches" or set parameter values to meet different information needs. Many conventional systems recommend or automatically create visualizations based on theoretical foundations such as data models and visualization reference models. Data attribute-based systems rely on data characteristics to select visual representations. Summary of the Invention
[0003] According to some embodiments of the present invention, a method for creating a question-answering system includes: receiving multiple user stories, wherein each user story is structured as multiple first phrase entities within a template (MLSS); applying natural language processing (NLP) to discover first data relationships between the first phrase entities and first contextual relationships between the first phrase entities, constructing a knowledge graph (KG) that captures second data relationships and second contextual relationships of multiple second phrase entities extracted from a data corpus, enriching the KG by linking the first phrase entities to the second phrase entities to form multiple enriched phrase entities in the KG, receiving a selection of an enriched phrase entity from the enriched phrase entities for completing the story template, identifying a technical requirement based on the selection of the enriched phrase entity from the enriched phrase entities, and training a model that matches at least one user story in the user stories with the technical requirement, wherein the model is stored in an analytical task library.
[0004] In accordance with at least one embodiment, a computer-implemented method of operating a question-answering system includes receiving a plurality of user stories, wherein each user story is structured as a first plurality of phrase entities within a template (MLSS), discovering a first data relationship between the phrase entities, discovering a first contextual relationship between the phrase entities, accessing a knowledge graph (KG) that captures a second plurality of entities and a second contextual relationship, enriching the KG by linking the first phrase entity to the second entity to form a plurality of enriched phrase entities in the KG, providing a display of a selection of an enriched phrase entity among the enriched phrase entities, and receiving a selection of an enriched phrase entity among the displayed enriched phrase entities, wherein the selected enriched phrase entity completes the story template.
[0005] As used herein, "facilitating" an action includes performing the action, making the action easier, assisting in performing the action, or causing the action to be performed. Thus, by way of example and not limitation, instructions executing on one processor may facilitate an action performed by instructions executing on a remote processor by sending appropriate data or commands to cause or assist in performing the action. For the avoidance of doubt, where an actor facilitates an action by an action other than performing the action, the action is still performed by some entity or combination of entities.
[0006] One or more embodiments of the present invention, or elements thereof, can be implemented in the form of a computer program product comprising a computer-readable storage medium having computer-usable program code for performing the indicated method steps. In addition, one or more embodiments of the present invention, or elements thereof, can be implemented in the form of a system (or apparatus) comprising a memory and at least one processor coupled to the memory and operable to perform the exemplary method steps. Still further, in another aspect, one or more embodiments of the present invention, or elements thereof, can be implemented in the form of an apparatus for performing one or more of the method steps described herein; the apparatus can include (i) a hardware module, (ii) a software module stored in a computer-readable storage medium (or multiple such media) and implemented on a hardware processor, or (iii) a combination of (i) and (ii); any of (i)-(iii) implementing the specific techniques set forth herein.
[0007] The technology of the present invention can provide substantial and beneficial technical effects. For example, one or more embodiments can provide:
[0008] Catalog, understand, and accelerate the establishment of user analytical intent and development requirements within complex data spaces; and
[0009] Automatically configure question answering systems.
[0010] These and other features and advantages of the present invention will become apparent from the following detailed description of illustrative embodiments of the invention, which is to be read in connection with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings:
[0012] Figure 1 Describes a cloud computing environment according to an embodiment of the present invention;
[0013] Figure 2 Depicts the abstract model layers according to an embodiment of the present invention;
[0014] Figure 3 Describing a composable analytics architecture according to an embodiment of the present invention;
[0015] Figure 4 is an illustration of a method for operating a composable analytics architecture according to an embodiment of the present invention;
[0016] Figure 5 is a diagram of an interconnected data system according to an embodiment of the present invention;
[0017] Figure 6 is a diagram of a method for determining analytical intent according to an embodiment of the present invention;
[0018] Figure 7 is a diagram of a collaborative framework supporting question answering according to an embodiment of the present invention;
[0019] Figure 8 FIGURE 1 illustrates the mapping of Mad-lib user stories (MUS) to analytical tasks according to an embodiment of the present invention;
[0020] Figure 9 illustrates an example user interface according to an embodiment of the present invention;
[0021] Figure 10 is an example implementation of a user interface (UI) and method according to an embodiment of the present invention;
[0022] Figure 11 A method for creating a question-answering system according to an embodiment of the present invention is shown; and
[0023] Figure 12 A computer system useful for implementing one or more aspects and / or elements of the present invention is depicted. DETAILED DESCRIPTION
[0024] According to example embodiments, the systems and methods described herein enable rapid sharing of expert knowledge across multiple disciplines for shared mental models and composable analytical frameworks, which increases the time-to-market of goods and services (see Figure 3 ).
[0025] Working on narrow, linear use cases can be complex, time-consuming, and expensive. Furthermore, the complexity and cost of discovering insights within data-rich industries necessitates AI systems configured to answer narrow questions. Furthermore, data visualization methods have allowed end users to explore data to answer adjacent questions. However, these experiences are often handled by subject matter experts (SMEs) within a single specialized field.
[0026] According to some embodiments, repeatable analytical workflows enable rapid retrieval of insights to satisfy analytical intent. The workflows facilitate rapid data-to-analytic intent mapping and metadata for internal / external experience and data visualization mapping (see Figure 4 According to at least one embodiment, the development of analytical requirements can be accelerated (see Figure 5 ) and can guide users in the combination of analytical insights based on the intent of changing business needs.
[0027] The present application will now be described in more detail with reference to the following discussion and the accompanying drawings. It should be noted that the drawings of the present application are provided for illustrative purposes only and, therefore, are not drawn to scale. It should also be noted that identical and corresponding elements are designated by identical reference numerals.
[0028] In the following description, many specific details are set forth, such as specific structures, components, materials, dimensions, processing steps and techniques, in order to provide an understanding of the various embodiments of the present application. However, one of ordinary skill in the art will appreciate that the various embodiments of the present application can be practiced without these specific details. In other cases, well-known structures or processing steps are not described in detail to avoid obscuring the present application.
[0029] It is understood in advance that although the present disclosure includes detailed descriptions about cloud computing, the implementation of the teachings cited herein is not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.
[0030] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services). These configurable resources can be quickly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0031] Features are as follows:
[0032] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities, such as server time and network storage, as needed without manual interaction with the service provider.
[0033] Broad network access: Capabilities are available over the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0034] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. There is a sense of location independence, as consumers typically do not have control or knowledge of the exact location of the provided resources, but are able to specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0035] Rapid elasticity: The ability to quickly and elastically provision capacity, in some cases automatically scaling down and releasing capacity to scale up quickly. To the consumer, the capacity available for provisioning typically appears unlimited and can be purchased in any quantity at any time.
[0036] Metered Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the utilized services.
[0037] The service model is as follows:
[0038] Software as a Service (SaaS): The ability provided to consumers is to use the provider's applications running on a cloud infrastructure. Applications are accessible from various client devices through a thin client interface such as a web browser (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.
[0039] Platform as a Service (PaaS): The capability provided to consumers is to deploy applications created or acquired using programming languages and tools supported by the provider onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but do have control over the deployed applications and the configuration of the application hosting environment.
[0040] Infrastructure as a Service (IaaS): The capabilities provided to consumers are processing, storage, networking, and other basic computing resources on which consumers can deploy and run arbitrary software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but rather have control over the operating system, storage, deployed applications, and potentially limited control over selected networking components (e.g., host firewalls).
[0041] The deployment modes are as follows:
[0042] Private cloud: Cloud infrastructure is operated solely for an organization. It can be managed by the organization or a third party and can exist on-premises or off-premises.
[0043] Community cloud: Cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist on-premises or off-premises.
[0044] Public cloud: Cloud infrastructure is made available to the public or large industry groups and is owned by the organization that sells cloud services.
[0045] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0046] The cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure consisting of a network of interconnected nodes.
[0047] Now refer to Figure 1, depicts an illustrative cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers can communicate, such as, for example, personal digital assistants (PDAs) or cellular phones 54A, desktop computers 54B, laptop computers 54C, and / or automobile computer systems 54N. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or a combination thereof, as described above. This allows the cloud computing environment 50 to provide infrastructure, platforms, and / or software as services for which cloud consumers do not need to maintain resources on local computing devices. It should be understood that Figure 1 The types of computing devices 54A-54N shown in FIG are intended to be illustrative only, and computing node 10 and cloud computing environment 50 may communicate with any type of computerized device over any type of network and / or network-addressable connection (eg, using a web browser).
[0048] Now refer to Figure 2 , showing the cloud computing environment 50 ( Figure 1 ) provides a set of functional abstraction layers. It should be understood in advance that Figure 2 The components, layers, and functions shown in are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As described, the following layers and corresponding functions are provided:
[0049] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: mainframes 61; servers based on RISC (Reduced Instruction Set Computer) architecture 62; servers 63; blade servers 64; storage devices 65; and network and networking components 66. In some embodiments, software components include web application server software 67 and database software 68.
[0050] Virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71 ; virtual storage 72 ; virtual networks 73 , including virtual private networks; virtual applications and operating systems 74 ; and virtual clients 75 .
[0051] In one example, the management layer 80 may provide the functionality described below. Resource provisioning 81 provides dynamic procurement of computing and other resources for performing tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking when utilizing resources within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management so that required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-arrangement and procurement of cloud computing resources in anticipation of future demand according to the SLA.
[0052] The workload layer 90 provides examples of functionality that can utilize a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analytical processing 94; transaction processing 95; and automated configuration question answering systems 96.
[0053] In complex systems (e.g., question answering systems), the interactions between the system's components and the data they generate can make it difficult to extract a clear analytical intent for the system's users (see Figure 5 ). The user's analytical intent describes the intended functionality of the analytical (e.g., widget, application, system, etc.) to support the user's ability to inform multivariate decisions, identify patterns, and / or perform trend analysis on multivariate data across multiple dimensions. According to some embodiments, mining analytical intent 704 in complex interconnected data systems (see Figure 7 ) requires knowledge of the industry / discipline involved (where), the analytical task to be accomplished (what), and the data needed to inform the analytical intent (how) (see Figure 6 ).
[0054] As an example, data democratization (e.g., making data accessible to a wide range of users) has led to advances across different data-intensive applications. For example, in this context, the healthcare field has adopted visual data mining tools to access and analyze data (e.g., to determine the cost breakdown associated with conditions for different procedures). Conversational experiences designed to understand user intent have emerged in parallel with data democratization. These conversational experiences attempt to infer the user's analytical intent from conversational discourse using, for example, natural language processing (NLP). Further, according to some examples, conversational discourse can include data captured from chatbots, chat logs, emails, other electronic communication media, and the like.
[0055] Without an effective way to capture and express analytical intent in a repeatable manner, product development teams invest significant time iterating on use case requirements and their mapping to analytical requirements. The resulting analytical requirements often lack specificity. This lack of specificity can lead to communication errors and increase development time. Furthermore, without a common analytical intent capture methodology, each use case often results in a customized technical implementation, limiting the reusability of analytical components across product lines.
[0056] According to some embodiments of the present invention, a Mad-lib Sentence Structure (MLSS) is a template that includes concepts of content or entities (e.g., <concept>> indicates), and a Mad-libs User Story (MUS) is an MLSS that has been populated with the selected content or entity (e.g., by the < <entity>> indicates). Figure 8 Examples of MLSS and MUS are shown. For example, in Example 1, 801, MLSS includes four concepts, including the first concept <<Industry A> >, and MUS includes the populated first concept < <electronicchain>>.
[0057] According to some embodiments of the present invention and Figure 3 The system 300 includes a capture module 301, a mapping module 302, and a construction module 303. The system 300 also includes a knowledge base (KB) 304 storing a knowledge graph, an analytic task library (ATL) 305, and a model and visualization repository 306. According to some examples, the system 300 is used to output an analytic specification and select a visualization for the analytic specification.
[0058] According to some embodiments, the knowledge base 304 captures data relationships and contextual relationships of entities derived from the data corpus and from continued use of the system, which introduces new data. According to some embodiments, the knowledge base 304 supports the development of the MUS by the capture module 301 by providing word clouds (e.g., clusters of highly related entities) and inferences of user intent in interactive scenarios. For example, user intent is used to select entities that can be placed in the MLSS to create the MUS.
[0059] According to some embodiments, the knowledge base 304 may be developed by ingesting published ontologies of entities in the domain, analyzing data models of analysis datasets, analyzing metadata of analytical methods (e.g., machine language (ML) models, visualization templates), etc. In some embodiments, the knowledge base 304 is developed by ingesting / analyzing data from multiple different sources, which may be from the same or different domains.
[0060] According to some embodiments, the knowledge base 304 may be further enhanced by applying NLP to metadata, data / analytic descriptions, and / or mining statistical relationships of data elements in analytical datasets. It should be understood that as used herein, analytical datasets are distinguished from training datasets.
[0061] According to some embodiments, the knowledge base 304 includes a knowledge graph of entities. These entities can be grouped. For example, entities can be grouped by industry type, starting word(s), job to be completed, dataset name, etc. According to some embodiments, the entities in the knowledge graph are connected by links, which are determined by using the MUS in the capture module 301 as training data.
[0062] According to some embodiments, such training includes incorporating analytical intents that may be added by the SME into the knowledge graph. Then, for test data (e.g., data for a real-world application), the knowledge base 304 may be utilized to create connections with the previously trained analytical intents.
[0063] According to some embodiments, the knowledge graph and domain knowledge can grow as new analytical data sets are added. For example, a new data set added to the knowledge graph can include a comprehensive data dictionary that can include descriptions of the data fields and, in some cases, corresponding summary statistics. Thus, the knowledge graph can be updated to include the new information.
[0064] According to some embodiments, the analytical task library 305 captures actions (e.g., analytical action descriptions) including analytical tasks such as implementing data selection and management. Analytical tasks include actions required to implement data selection, data management, and method development processes. Example analytical tasks may include summarizing data by geography, calculating the average of a parameter, etc. The analytical task library 305 supports analytical specification development by allowing users to navigate and identify relevant actions for analytical tasks for each MUS.
[0065] Example analytical tasks in the analytical task library 305 may be categorized as, for example:
[0066] Data selection: Use MUS content to help make choices (e.g., data sources);
[0067] Data management includes:
[0068] Data filtering: using MUS input to help establish filtering criteria to create data subsets; and
[0069] Data grouping: using MUS input to assist in determining grouping criteria to create aggregated data;
[0070] Method development (including method selection): Use MUS input to help select the analytical method selection process. Example analytical methods include methods for performing visualization, analytical models, etc.
[0071] According to some embodiments, analytical task library 305 may be improved over time as use cases, analytical needs, and analytical techniques are added.
[0072] Now refer to Figure 4 and a method 400 for operating a composable analytic architecture according to an embodiment of the present invention, at 401 , the capture module 301 captures a MUS, at block 402 , the mapping module 302 maps the MUS to an analytic task, and at block 403 , the construction module 303 constructs an analytic specification.
[0073] According to some embodiments, at block 401, the capture module 301 receives a selection of one or more Mad-lib sentence structures (MLSS) 411, user input 412 (such as conversational input from a chatbot), and an indication of the user's intent 413. According to some embodiments, the user input 412 may be used (e.g., using a trained classifier) to determine the user's intent 413. Using these inputs, the capture module 301 uses the selected MLSS to determine the MUS content, where a set of sample MUSs (see Figure 8 , MUS 802 and 804) were developed.
[0074] According to at least one embodiment, one or more MLSS are provided by system 411. Examples of MLSS are Figure 8 As shown in (801 and 803), MLSS includes one or more concepts. These concepts can be identified by appropriate characters (e.g., "<<>>") that can be recognized by the system.
[0075] According to some embodiments, the capture module 301 samples the way to construct the initial MUS. For example, the initial MUS sentence can be expressed as:
[0076] A<<Business Model> ><<Industry Sector> >needs to<<Starter Words> ><<description of job to be done> >for its<<Location Types> >using<<AvailableData&Analytics> >.
[0077] The mapping module 302 develops the initial MUS. The development process is shown in the following three different developments of the selected MLSS:
[0078] 1) B2C<any industry> > Data (retail store information, transaction data, news searches, mobility, employment and unemployment, new cases) needs to be used to compare customer purchasing power at their retail locations pre-COVID and post-COVID.
[0079] 2) B2C<any industry> >Employment and unemployment data are needed to compare current unemployment with workforce performance for caregiver occupations by geographic region.
[0080] 3) Focus on employee health and availability, B2B<<any industry> >Overall workforce risk, infection rates, and availability across all locations and all organizations needs to be monitored or trended in order to assess the effectiveness of current policies and recommend changes based on the following data in their dashboard: RTWA Work Permit Status, Counts, etc., RTWA Workforce Availability (risk of transmission between employees), Number of people who have gone red and returned to work after quarantine, WCM, Caseworker Time to First Contact, Case Volume by Status).
[0081] In the example shown, static text such as "need to," "for its," and "using" remain unchanged, while the variables indicated by <<>> are filled in. It's clear from the example that the development of the initial sentence is flexible. This development is a human-guided task.
[0082] Mapping module 302 facilitates enterprise design thinking sessions (see Figure 9 ), for example, constructing an initial knowledge graph (KG) with SMEs or design and data scientists to store in the knowledge base 304.
[0083] According to some embodiments, and with reference to Figure 9 , the mapping module 302 causes an enterprise design thinking session UI 900 to be displayed, comprising one or more widgets (or UI elements), including user-provided instructions 901 , background information and concepts 902 , a set of predefined entity groups 903 - 906 , and a sandbox 907 for constructing sentences.
[0084] In box 414, the capture module 301 develops the MUS as is or as an extended MUS. According to some embodiments, the starter words in the knowledge graph are well suited for data visualization and machine learning. For example, analytical starter words are mapped to one or more data visualizations and analytical actions that include analytical tasks. For example, a timeline is a well-considered choice for a data visualization with the starter word "trend", while a pie chart is not. In another example, the starter word "classify" can be mapped to an analytical action of "dividing" an item in a dataset and a histogram data visualization. The system can learn these mappings using a knowledge corpus (including, for example, previous user selections).
[0085] According to some embodiments, at block 414, the capture module 301 groups the MUS based on an understanding of the user intent determined from the user input (i.e., based on the discovered / mined user intent 413), and the grouping is used as the basis for a specific user interface that may be supported by an intent-specific wizard (see Figure 10 ). More specifically, a list 1001 of different MUSs is selected, prioritized (e.g., based on confidence scores), and provided to a user interface (UI) wizard 1002 for user manipulation. In one example, MUSs are selected for a group based on similarities in one or more user inputs. One example selection logic includes grouping MUSs by industry (e.g., healthcare vs. media). Another example selection logic includes grouping MUSs by starting word (e.g., all MUSs involving the identifier "trend" can be grouped together, regardless of industry).
[0086] At block 401, the capture module 301 designs a wireframe to support analytical development at 303 / 402. The wireframe can be directly derived from the MLSS, allowing the user to navigate to a specific dashboard based on the analytical intent of the corresponding MLSS. In one example, navigation is facilitated by mapping the MUS (e.g., user input) to metadata or a lookup table for the dashboard. The MUS can be derived directly from the wireframe via a UI wizard 1002, where the user supplies data input or data selection for a variable (e.g., represented by "<<>>").
[0087] According to some embodiments, at block 401, the capture module 301 operates to further mature the knowledge graph as the number of MUSs grows. For example, when additional MUSs are made available through the UI wizard 1002, such as when a user creates a new MUS, these new MUSs can be directly translated / mapped to user intents. For example, the MUSs created by the user can be mapped to intents via an appropriate NLP or knowledge-based model. In an example interaction scenario, when a user makes a query, a topic model (e.g., Latent Dirichlet Allocation (LDA)) is invoked to map the query to entities of the knowledge graph. Once an entity is identified, the knowledge graph, along with its analytical intent link, is used as a module for identifying analytical intents and returning relevant data fields cut across multiple datasets (e.g., entities found in the knowledge graph and / or analytical content based on the entity, such as dashboards, data, etc.).
[0088] According to some embodiments, at 402, the mapping module 302 maps the MUS to analytical tasks. For example, for each MUS, the mapping module 302 analyzes the initial mad-lib concepts (i.e., the concepts replaced by the selected entities) and identifies matching analytical tasks in the analytical task library 305. For example, at box 415, given the concepts, the mapping module 302 determines the task description and, at box 416, annotates the MUS with the task description. According to one example, the matching of analytical tasks can be performed by utilizing NLP techniques, such as named entity recognition (NER) of concepts and standardization of analytical tasks. Figure 7 , the mapping is shown by arrow 701. The concept-to-task mapping can be one-to-one, one-to-many, or many-to-many. For example, at 702, < <industry>>Concept notification "DataFiltering" ("data filtering") task. In another example, at 703, <<starter word> > and <<job to bedone> > concepts together inform "Method Selection".
[0089] According to some embodiments, at 403, the construction module 303 constructs analytical specifications. According to one example, these analytical specifications include functions such as categorize, classify, recognize, compare and contrast, correlate (relationship), cluster or group (relationship), etc. For each framework, the construction module 303 promotes the corresponding MUS (or MUS group), using the knowledge graph to find similar concepts for all Mad-lib entities. According to one example, similarity can be determined by finding direct matches to concepts in the knowledge graph (e.g., the same word or a word synonymous with the user input). In another example, similarity is determined by looking at adjacent nodes in the knowledge graph (e.g., concepts related to the concept). For each analytical task, the construction module 303 constructs an analytical action specification (e.g., a function to be completed based on given data). According to some embodiments, these specifications are used as technical requirements for the analytical development team or aligned with an automated analytical pipeline to perform actions like the Mad-lib wizard.
[0090] According to some embodiments, construction of an analytical action specification includes selecting data, managing data, and developing data at 403. At blocks 417 and 418, MUS concepts for discovering extensions are selected, managed, and developed.
[0091] According to some embodiments, at 417, the selection method includes searching the data model metadata for a "Data Selection" concept as a way to determine which dataset(s) to further investigate for analytical development. For example, given a "Data Selection" action for the phrase "Mobility Data," the system searches the knowledge graph (or some other data source processed by the system) to find all data sources with "mobility" in their metadata. For example, the method may search for mobility data on the web (e.g., Google Mobility Data, Apple Mobility Data, etc.). According to one example, the method may look for structured / unstructured data sources that have already been processed for the system. According to at least one embodiment, the search is initially performed on data sources that have already been processed for the system, and then performed on unprocessed data (e.g., the internet).
[0092] According to some embodiments, at 417, the management method includes analyzing the "Data Filtering" concept to identify data fields and data values for filtering. For example, under the geographic coverage, a "Data Filtering" action is performed on the phase "West Coast." The knowledge base provides similar concepts to "West Coast" (which may include California associated with the state, Seattle associated with the city, etc.), and similar concepts form the basis of the data filtering criteria. Analysis of the "Data Grouping" concept identifies (one or more) data fields and (one or more) data values to aggregate the data.
[0093] According to some embodiments, at 418 , the development method includes searching a model repository by mapping a “Method Selection” concept with model metadata to find reusable / similar models, and searching a visualization template repository by mapping a “Method Selection” concept with visualization metadata to find reusable / similar visualizations.
[0094] Reference is made to searching the model repository by mapping the "Method Selection" concept to model metadata to find reusable / similar models. In one example, the "Method Selection" action for the "predict demand" stage and the "Method Selection" action for the "historical sales data" stage are mapped to a pre-built trained model that uses historical sales transaction data to predict future demand. If no suitable model is found, analytical question sentences can be developed to drive model development and tagged with MUS concepts and keywords derived from the knowledge graph.
[0095] Reference is made to searching a visualization template repository by mapping "Method Selection" concepts with visualization metadata to find reusable / similar visualizations. In one example, if no suitable model is found, UX development requirements are developed to drive visualization template development, where the visualization templates are tagged with MUS concepts and keywords derived from the knowledge graph.
[0096] According to one or more embodiments, visualizations are trained in parallel with training of the model. According to some embodiments, models and visualizations are linked in an analytical task library, e.g., if a model is determined to be relevant to a user's entity selection, there are one or more visualizations that are automatically suggested (output).
[0097] According to at least one embodiment and referring to Figure 5 Mining for analytical intent in complex, interconnected data systems requires knowledge of the industry / discipline involved (where) 501, knowledge of the job to be done (what) 502, and knowledge of the data needed to inform the analytical intent (how) 503. Answering narrow business questions can be difficult due to the complexity and cost of discovering insights within data-rich industries. According to some embodiments, data visualization enables end users to explore data to answer adjacent questions.
[0098] According to at least one embodiment and with reference to Figure 6 , analytical intent 601 describes analytical intent that supports the end user's ability to inform complex multivariate decisions across multiple dimensions, identify complex patterns, or perform trend analysis on multivariate data. Figure 6 , at 602-605, an analytical intent 601 is developed, wherein at 602, for each industry, an affinity graph of applicable characteristics of the industry (e.g., industry "A") is determined, at 603, for each industry, a starting word is identified (e.g., as a link to a set of commonly understood analytical operations / concepts related to the corresponding industry), and at 604 and 605, the work to be done and specific data are identified, respectively.
[0099] refer to Figure 6 , using summarization / sharing of their expertise about a given problem space 600 at each contributing discipline (e.g., SME 606, Design 607, and Data 608), working within this collaborative framework, the team’s understanding is extended across all three disciplines. For example, Figure 6 As shown in the brackets in [ ], subject matter experts (SMEs) know their industry and organization, but may have difficulty describing their goals or tasks in analytical terms. Data scientists need to create appropriate models from these analytical terms to support the SME's tasks. Choosing from starting words defined by design professionals (such as user researchers) helps bridge any communication gaps more quickly.
[0100] According to at least one embodiment and again with reference to Figure 7 , shown by arrow 701. The mapping may take different forms depending on the task being performed. For example, the mapping may include querying data (e.g., knowledge base 304 and analytical task library 305) and performing tasks such as aggregation, filtering, searching, etc. The concept-to-task mapping may be one-to-one, one-to-many, or many-to-many. For example, at 702, < <industry>>Concept notification "Data Filtering" task. In another example, at 703, <<starter word> > and <<job to be done> > concepts together inform "Method Selection".
[0101] According to some embodiments, a set of sample MUS is developed (see Figure 8 ). Example 801 shows a mapping where <<Industry A> > is mapped to <<electronic chain> >,<<Starter Word> > is mapped to < <predict>>、<<Job to be Done> > is mapped to<SKU demand> >and<<Specific Data> > is mapped to <<historic sales data> >, where <<electronic chain> >、< <predict>> etc. are mad-lib entities in the knowledge graph and are used to create MUS 802. Example 803 shows a mapping in which additional elements of the mad-lib sentence structure are mapped to entities in the knowledge graph and used to create Mad-lib User Story (MUS) 804.
[0102] According to some embodiments, the Enterprise Design Thinking Session U1 900 (see Figure 9 ) is facilitated. According to some embodiments, and with reference to Figure 10 , this method prioritizes mad-lib entities 1001 from design thinking sessions (see Figure 9 , 907 ), presents an initial constrained sentence case UI wizard 1002 and receives user selections for each concept, and links to a dashboard capable of answering mad-lib questions 1003 developed using the UI wizard 1002 .
[0103] Overview:
[0104] According to some embodiments of the present invention and with reference to Figure 11 A method for creating a question-answering system 1100 includes: receiving a plurality of user stories 1101, wherein each user story is structured as a plurality of first phrase entities within a template (MLSS); applying natural language processing (NLP) to discover first data relationships between the first phrase entities and first contextual relationships between the first phrase entities 1102; constructing a knowledge graph (KG) that captures second data relationships and second contextual relationships of a plurality of second phrase entities extracted from a data corpus 1103; enriching the KG by linking the first phrase entities to the second phrase entities to form a plurality of enriched phrase entities in the KG 1104; receiving a selection of an enriched phrase entity from the enriched phrase entities for completing a story template 1105; identifying a technical requirement based on the selection of the enriched phrase entity from the enriched phrase entities 1106; and training a model that matches at least one user story to the technical requirement, wherein the model is stored in an analytical task library 1107. According to some embodiments, a model may be selected after receiving another user story and used to answer or prepare a response to a corresponding technical requirement 1108.
[0105] Overview:
[0106] According to one or more embodiments of the present application, a computer-implemented method for creating a question-answering system includes receiving a plurality of user stories, wherein each user story is structured as a first plurality of phrase entities within a template (MLSS), applying natural language processing (NLP) to discover a first data relationship between the phrase entities and a first contextual relationship between the phrase entities, constructing a knowledge graph (KG) that captures a second data relationship and a second contextual relationship of a second plurality of entities extracted from a data corpus, enriching the KG by linking the first phrase entity to the second entity to form a plurality of enriched phrase entities in the KG, receiving a selection of an enriched phrase entity in the enriched phrase entities to complete the story template, identifying a technical requirement based on the selection of the enriched phrase entity in the enriched phrase entities, and training a model that matches at least one of the user stories to the technical requirement, wherein the model is stored in an analytical task library.
[0107] In accordance with at least one embodiment, a computer-implemented method of operating a question-answering system includes receiving a plurality of user stories, wherein each user story is structured as a first plurality of phrase entities within a template (MLSS), discovering a first data relationship between the phrase entities, discovering a first contextual relationship between the phrase entities, accessing a knowledge graph (KG) that captures a second plurality of entities and a second contextual relationship, enriching the KG by linking the first phrase entity to the second entity to form a plurality of enriched phrase entities in the KG, providing a display of a selection of an enriched phrase entity among the enriched phrase entities, and receiving a selection of an enriched phrase entity among the displayed enriched phrase entities, wherein the selected enriched phrase entity completes the story template.
[0108] The methods of the embodiments of the present disclosure may be particularly suitable for use in electronic devices or alternative systems. Accordingly, embodiments of the present invention may take the form of entirely hardware embodiments or embodiments combining software and hardware aspects, which may be collectively referred to herein as "processors," "circuits," "modules," or "systems."
[0109] In addition, it should be noted that any of the methods described herein may include the additional step of providing a computer system for organizing and servicing the resources of a computer system. Further, a computer program product may include a tangible computer-readable recordable storage medium having code adapted to be executed to perform one or more method steps described herein, including providing a system having different software modules.
[0110] One or more embodiments of the present invention, or elements thereof, can be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operable to perform the exemplary method steps. Figure 12 Depicting a computer system that can be used to implement one or more aspects and / or elements of the present invention, and also representing a cloud computing node according to an embodiment of the present invention. Figure 12 , cloud computing node 10 is only one example of a suitable cloud computing node and is not intended to suggest any limitation on the scope of use or functionality of the embodiments of the invention described herein. Regardless, cloud computing node 10 is capable of implementing and / or performing any of the functions set forth above.
[0111] In cloud computing node 10, there is a computer system / server 12, which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, fat clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.
[0112] Computer system / server 12 may be described in the general context of computer system-executable instructions (e.g., program modules) executed by a computer system. Generally speaking, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. Computer system / server 12 may be practiced in a distributed cloud computing environment, where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0113] like Figure 12 As shown, computer system / server 12 in cloud computing node 10 is shown in the form of a general-purpose computing device. Components of computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and bus 18 that couples various system components, including system memory 28, to processor 16.
[0114] Bus 18 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0115] Computer system / server 12 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by computer system / server 12, and includes both volatile and nonvolatile media, removable and non-removable media.
[0116] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer system / server 12 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (not shown and commonly referred to as a "hard drive"). Although not shown, a magnetic disk drive for reading from or writing to a removable non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In such cases, each may be connected to the bus 18 via one or more data media interfaces. As will be further depicted and described below, the memory 28 may include at least one program product having a set (e.g., at least one) program modules configured to perform the functions of embodiments of the present invention.
[0117] A program / utility 40 having a set (at least one) of program modules 42, as well as an operating system, one or more application programs, other program modules, and program data, may be stored in memory 28 by way of example and not limitation. Each or some combination of the operating system, one or more application programs, other program modules, and program data may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of embodiments of the present invention as described herein.
[0118] Computer system / server 12 may also communicate with one or more external devices 14 (e.g., a keyboard, pointing device, display 24, etc.); and / or any device that enables computer system / server 12 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication may occur via input / output (I / O) interface 22. In addition, computer system / server 12 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via network adapter 20. As shown, network adapter 20 communicates with other components of computer system / server 12 via bus 18. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, and external disk drive arrays, RAID systems, tape drives, and data archiving storage systems.
[0119] Thus, one or more embodiments may utilize software running on a general purpose computer or workstation. Figure 12 Such an implementation may, for example, employ a processor 16, memory 28, and an input / output interface 22 to a display 24 and external devices 14 (such as a keyboard, pointing device, etc.). As used herein, the term "processor" is intended to include any processing device, for example, a processing device including a CPU (central processing unit) and / or other forms of processing circuitry. Furthermore, the term "processor" may refer to more than one individual processor. The term "memory" is intended to include memory associated with a processor or CPU, for example, RAM (random access memory) 30, ROM (read-only memory), fixed storage devices (e.g., hard drive 34), removable storage devices (e.g., disks), flash memory, etc. Furthermore, as used herein, the phrase "input / output interface" is intended to encompass, for example, an interface to one or more mechanisms for inputting data to the processing unit (e.g., a mouse), as well as an interface to one or more mechanisms for providing results associated with the processing unit (e.g., a printer). The processor 16, memory 28, and input / output interface 22 may be interconnected, for example, via a bus 18 that is part of the data processing unit 12. Suitable interconnections (e.g., via bus 18) may also be provided to a network interface 20 (such as a network card) and a media interface (such as a floppy disk or CD-ROM drive). The network interface 20 may be provided for connecting to a computer network interface, and the media interface may be provided for connecting to a suitable media interface.
[0120] Thus, computer software including instructions or codes for executing the methods of the present invention described herein may be stored in one or more associated memory devices (e.g., ROM, fixed or removable memory), and when ready to be used, partially or completely loaded (e.g., loaded into RAM) and implemented by the CPU. Such software may include, but is not limited to, firmware, resident software, microcode, etc.
[0121] A data processing system suitable for storing and / or executing program code will include at least a processor 16 coupled directly or indirectly to memory elements 28 through a system bus 18. The memory elements may include local memory used during actual implementation of the program code, bulk storage, and cache memories 32 which provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from bulk storage during implementation.
[0122] Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system either directly or through intervening I / O controllers.
[0123] Network adapter 20 may also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks.Modems, cable modem and Ethernet cards are just a few of the currently available types of network adapters.
[0124] As used herein (including in the claims), a "server" includes a physical data processing system (e.g., a Figure 12 System 12 is shown.) It will be understood that such a physical server may or may not include a display and keyboard.
[0125] One or more embodiments may be implemented at least in part in the context of a cloud or virtual machine environment, but this is exemplary and non-limiting. Figure 1-Figure 2 For example, consider the database application in layer 66.
[0126] It should be noted that any of the methods described herein may include the additional step of providing a system comprising different software modules contained on a computer-readable storage medium; these modules may include, for example, any or all appropriate elements depicted in the block diagrams and / or described herein; by way of example and not limitation, any, some, or all of the modules / boxes and / or sub-modules / boxes described. The method steps may then be performed using the different software modules and / or sub-modules of the system described above executed on one or more hardware processors (such as 16). Further, a computer program product may include a computer-readable storage medium having code suitable for implementation to perform one or more method steps described herein, including providing a system having different software modules.
[0127] An example of a user interface that may be employed in some cases is Hypertext Markup Language (HTML) code provided to a browser of a user's computing device by a server, etc. The HTML is parsed by the browser on the user's computing device to create a graphical user interface (GUI).
[0128] Exemplary System and Article of Manufacture Details
[0129] The present invention may be a system, method, and / or computer program product.The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon for causing a processor to perform various aspects of the present invention.
[0130] Computer readable storage medium can be a tangible device that can retain and store the instructions used by the instruction execution device.Computer readable storage medium can be, for example but not limited to, electronic storage device, magnetic storage device, optical storage device, electromagnetic storage device, semiconductor storage device or any suitable combination of the above.The non-exhaustive list of more specific examples of computer readable storage medium includes the following: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device such as punch card or the protrusion structure in the groove with the instruction recorded thereon and any suitable combination of the above.Computer readable storage medium as used herein should not be interpreted as temporary signal itself, such as radio wave or other free propagation electromagnetic wave, electromagnetic wave propagated by waveguide or other transmission medium (for example, light pulse passing through fiber optic cable) or electric signal emitted by wire.
[0131] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or downloaded to an external computer or external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.
[0132] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, the configuration data of integrated circuit or source code or the object code written in any combination of one or more programming languages, these programming languages include object-oriented programming languages (such as Smalltalk, C++ etc.) and process programming languages (such as " C " programming languages or similar programming languages). The computer-readable program instructions can be performed completely on the user's computer, partly on the user's computer, performed as an independent software package, partly on the user's computer, partly on a remote computer or fully on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer by any type of network (including local area network (LAN) or wide area network (WAN)), or can be connected to an external computer (for example, using an internet service provider through the internet). In certain embodiments, the electronic circuit comprising for example programmable logic circuit, field programmable gate array (FPGA) or programmable logic array (PLA) can make the electronic circuit personalized and perform the computer-readable program instructions by utilizing the state information of the computer-readable program instructions, so as to perform various aspects of the present invention.
[0133] The present invention will be described below with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0134] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in the flowchart and / or block diagram or multiple blocks. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to operate in a specific manner. Thus, the computer-readable storage medium having the instructions stored therein includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in the flowchart and / or block diagram or multiple blocks.
[0135] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in or in multiple boxes in the flowchart and / or block diagram.
[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations of possible implementations of the systems, methods and computer program products according to different embodiments of the present invention. To this end, each box in the flowchart or block diagram may represent a module, segment or portion of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions annotated in the box may not occur in the order annotated in the figure. For example, depending on the functions involved, the two boxes shown in succession may actually be executed substantially simultaneously, or the boxes may sometimes be executed in the opposite order. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.
[0137] The description of various embodiments of the present invention has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.< / predict> < / predict> < / industry> < / industry> < / electronicchain> < / entity> < / concept>
Claims
1. A computer-implemented method for creating a question-answering system, the computer-implemented method comprising: receiving a plurality of user stories, wherein each of the user stories is structured as a plurality of first phrase entities within a template; Applying natural language processing (NLP) to discover a first data relationship between the first phrase entities and a first contextual relationship between the first phrase entities; Constructing a knowledge graph KG, wherein the knowledge graph KG captures second data relations and second context relations of a plurality of second phrase entities extracted from the data corpus; enriching the knowledge graph KG by linking the first phrase entity to the second phrase entity to form a plurality of enriched phrase entities in the knowledge graph KG, wherein, in the linking, a link between the first phrase entity and the second phrase entity is determined by using the user story as training data, and the training of determining the link using the user story as training data includes incorporating analytical intent into the knowledge graph KG; receiving a selection of an enriched phrase entity of the enriched phrase entities for completing the story template; identifying a technical requirement based on a selection of an enriched phrase entity among the enriched phrase entities; as well as A model is trained that matches at least one of the user stories to the technical requirements, wherein the model is stored in an analytical task library.
2. The method of claim 1 , further comprising using the model to process data related to a technical requirement of another user story. 3 . The method of claim 1 , wherein each of the enriched phrase entities describes one of data selection, transformation, model configuration, and report design specification. The method of claim 1 , further comprising training at least one visualization using the technical requirements.
5. The method according to claim 4, further comprising: The model and the at least one visualization are stored in a searchable repository based on textual elements of the phrase entity. The method of claim 5 , wherein the text elements are each categorized as at least one of an industry type, a starting word, a role of an actor, and a data type. The method of claim 1 , wherein the user stories are stored in a library of user stories. The method of claim 1 , wherein the enriched phrase entities are mapped to analytical tasks in the analytical task library.
9. The method of claim 1, wherein the analytical task is utilized to annotate technical requirements of the user story. 10 . The method according to claim 1 , further comprising iteratively updating the knowledge graph (KG) based on received user feedback.
11. A computer-implemented method of operating a question-answering system, the method comprising: receiving a plurality of user stories, wherein each of the user stories is structured as a plurality of first phrase entities within a template; discovering a first data relationship between the first phrase entities; discovering a first contextual relationship between the first phrase entities; Accessing a knowledge graph KG, wherein the knowledge graph KG captures a second data relationship and a second context relationship of a plurality of second phrase entities; enriching the knowledge graph KG by linking the first phrase entity to the second phrase entity to form a plurality of enriched phrase entities in the knowledge graph KG, wherein, in the linking, a link between the first phrase entity and the second phrase entity is determined by using the user story as training data, and the training of determining the link using the user story as training data includes incorporating analytical intent into the knowledge graph KG; providing a display of a selection of an enriched phrase entity among the enriched phrase entities; as well as A selection of an enriched phrase entity among the displayed enriched phrase entities is received, wherein the selected enriched phrase entity completes the story template. 12 . The method of claim 11 , wherein each of the enriched phrase entities describes one of data selection, transformation, model configuration, and report design specification.
13. The method according to claim 11, further comprising: identifying a technical requirement based on the selected enriched phrase entity; as well as A model is trained that matches at least one of the user stories to the technical requirements, wherein the model is stored in an analytical task library.
14. The method of claim 13, further comprising using the model to process data related to a technical requirement of another user story. The method of claim 13 , wherein the analytical task is utilized to annotate technical requirements of the user story.
16. The method according to claim 13, further comprising: Access data associated with the user story; as well as The data associated with the user story is displayed using at least one visualization selected according to the technical requirements.
17. The method according to claim 16, further comprising: The model and the at least one visualization are stored in a searchable repository based on textual elements of the phrase entity.
18. The method according to claim 11, wherein The enriched phrase entities are mapped to analytical tasks in an analytical task library.
19. A non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed by a computer, cause the computer to perform a method of operating a question-answering system, the method comprising: receiving a plurality of user stories, wherein each of the user stories is structured as a plurality of first phrase entities within a template; discovering a first data relationship between the first phrase entities; discovering a first contextual relationship between the first phrase entities; Accessing a knowledge graph KG, wherein the knowledge graph KG captures a second data relationship and a second context relationship of a plurality of second phrase entities; enriching the knowledge graph KG by linking the first phrase entity to the second phrase entity to form a plurality of enriched phrase entities in the knowledge graph KG, wherein, in the linking, a link between the first phrase entity and the second phrase entity is determined by using the user story as training data, and the training of determining the link using the user story as training data includes incorporating analytical intent into the knowledge graph KG; providing a display of a selection of an enriched phrase entity among the enriched phrase entities; as well as A selection of an enriched phrase entity among the displayed enriched phrase entities is received, wherein the selected enriched phrase entity completes the story template.
20. The computer-readable storage medium of claim 19, wherein the method further comprises: identifying technical requirements based on the selected enriched phrase entities; as well as A model is trained that matches at least one of the user stories to the technical requirements, wherein the model is stored in an analytical task library.
21. The computer-readable storage medium of claim 20, wherein the method further comprises using the model to process data related to a technical requirement of another user story.
22. The computer-readable storage medium of claim 20, wherein the method further comprises: Access data associated with the user story; as well as The data associated with the user story is displayed using at least one visualization selected according to the technical requirements.
23. The computer-readable storage medium of claim 19, wherein each of the enriched phrase entities describes one of data selection, transformation, model formulation, and report design specification.
24. A system comprising modules respectively configured to perform the steps of the method according to any one of claims 1 to 18.
25. A computer program product comprising a computer-readable storage medium, wherein the computer-readable storage medium embodies program instructions, wherein the program instructions are executable by a computing device to cause the computing device to perform the steps of the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Intelligent reading support
US20210109918A1