Feature selection system

CN116542253BActive Publication Date: 2026-09-01ACCENTURE GLOBAL SOLUTIONS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210848709.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-01-25
Filing Date
2022-07-19
Publication Date
2026-09-01
Estimated Expiration
2042-07-19

Smart Images

  • Figure CN116542253B_ABST
    Figure CN116542253B_ABST
Patent Text Reader

Abstract

A feature selection system. This document describes a computer-implemented method comprising: receiving, via a network, at least one of text, audio, image, or video data associated with an entity of interest; identifying a candidate feature set for a specific entity based on the received data; loading a feature library comprising multiple features, each of which is assigned to one or more feature spaces; and using a feature selection engine, selecting one or more features from each feature space based on the candidate feature set for the specific entity.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] In general, products and services can be defined by their characteristics. These characteristics typically include numerical metrics. For complex products and services, web-based platforms can be used to identify and continuously track the growing number of relevant characteristics and metrics. Summary of the Invention

[0002] This specification provides an overview of a system that uses natural language processing to process information from a data source to identify a set of features for a specific entity (e.g., a company). The feature selection engine uses machine learning techniques to select relevant features from a feature library based on the characteristics of the specific entity. The selected features can be displayed on a user interface, for example, for editing and updating features.

[0003] In general, an innovative aspect of the subject matter described in this specification can be implemented in methods comprising the actions of: receiving, via a network, at least one of text, audio, image, or video data associated with an entity of interest; identifying a candidate feature set for a particular entity based on the received data; loading a feature library comprising multiple features, each of which is assigned to one or more feature spaces; and using a feature selection engine to select one or more features from each feature space based on the candidate feature set for the particular entity. Other implementations of this aspect include corresponding systems, apparatus, and computer programs configured to perform actions of methods encoded on a computer storage device.

[0004] These and other implementations may each optionally include one or more of the following features.

[0005] In some implementations, the candidate feature set for identifying a particular entity includes: extracting text from text, audio, image, or video data associated with the entity; and using a natural language processing (NLP) model to create vectors of feature-related words based on the usage and context of words in the data associated with the entity.

[0006] In some implementations, these operations include receiving, via a network, at least one of text, audio, image, or video data associated with multiple entities; identifying an entity domain including the entity of interest and multiple additional entities based on the received data; and loading a candidate feature set for a specific domain from a feature library for that specific domain, wherein a feature selection engine selects one or more features from each feature space in the feature space based on the candidate feature set for the specific entity and the candidate feature set for the specific domain.

[0007] In some cases, these operations may include assigning weighted scores to each of the selected features using interpretable AI techniques; and filtering the selected features based on the weighted scores.

[0008] In some implementations, the feature selection engine selects a first feature set, which includes one or more features, from each feature space in the feature space at a first time, based on the candidate feature set of a specific entity; and at a second time after the first time, the feature selection engine selects a second feature set, which includes one or more features, from each feature space in the feature space based on the candidate feature set of the specific entity and the candidate feature set of a specific domain.

[0009] In some implementations, each candidate feature among the loaded candidate features of a particular domain also includes a baseline measurement, and these operations include identifying the baseline measurement based on the baseline measurement of the candidate features of the particular domain for each of the selected features.

[0010] In some implementations, these operations may include identifying one or more custom features that are included in a candidate feature set of a particular entity or a candidate feature set of a particular domain and are not included in a feature library; and storing one or more custom features in a feature library.

[0011] In some implementations, these operations may include generating a visualization for each feature space, which presents both the features selected by the feature selection engine and one or more unselected features; receiving user input to select or deselect one or more features; and updating the selected features based on the user input.

[0012] It should be understood that the methods according to this disclosure may include any combination of the aspects and features described herein. That is, for example, the apparatus and methods according to this disclosure are not limited to the combinations of aspects and features specifically described herein, but may also include any combination of the provided aspects and features.

[0013] Details of one or more implementations of this disclosure are set forth in the accompanying drawings and the description below. Other features and advantages of this disclosure will become apparent from the specification, drawings, and claims. Attached Figure Description

[0014] Figure 1 An example system for selecting features is described.

[0015] Figure 2 An example feature tree is depicted.

[0016] Figure 3 This is a flowchart of an example process for selecting features.

[0017] Figure 4 An example system is described that can execute an implementation of this disclosure.

[0018] The same reference numerals and names in various figures represent the same elements. Detailed Implementation

[0019] In most technology and business contexts, “success” has traditionally been defined by numerical metrics, often enshrined in customer-supplier contracts or service level agreements (SLAs). The parties to such agreements are generally able to assess the success of the collaboration. In recent years, however, perspectives have shifted towards recognizing that the parties to these agreements have responsibilities to other stakeholders outside the contractual parties (e.g., shareholders, employees, customers, and suppliers) and the public. While parties may continue to define success with numbers, the expectations of other stakeholders are beginning to shape customer-supplier relationships.

[0020] For example, tire manufacturers and car manufacturers may have an agreement whereby the tire manufacturer supplies the car manufacturer with a specific model of tire. Traditionally, the agreement might specify performance metrics, such as braking distance, and financial metrics, such as unit price. However, with the advent of electric vehicles, the number of metrics may increase to include, for example, road noise requirements that could influence tread pattern design and material selection. Furthermore, electric vehicles may imply certain sustainability expectations from external stakeholders, which could also affect material selection and tire manufacturing methods. When faced with a sudden increase in complexity and a shift in priorities, parties may overemphasize some areas of concern while neglecting others, resulting in an incomplete or unbalanced overview of the engineering project.

[0021] Similar challenges may arise in business settings, such as consulting. While consultants are traditionally hired to improve financial metrics, there is now an expectation of improving performance simultaneously through various means. At the outset of a contract, the client's management meets with a senior consultant to identify areas for improvement, actions and measures to be taken within those areas, and the metrics to measure those improvements. Generally, the client clearly articulates the areas for improvement, and the consultant recommends actions, measures, and metrics. When faced with a variety of potential issues, the client's upper management may become overly focused on a particular area, such as diversity or sustainability, for example, due to a series of negative public relations events. Conversely, the senior consultant may clearly favor tracking improvements through specific metrics familiar from past contracts with other clients.

[0022] Therefore, flexible and robust tools are needed to accurately identify and track a balanced set of expectations for numerous projects. In this disclosure, expectations can be defined by features, such as goals, actions, and metrics. More specifically, this disclosure describes a set of feature trees, each containing structured customer-related information, and how this set can be leveraged using artificial intelligence to identify features specific to an individual customer. As illustrated in the examples, feature trees are applied to a variety of contexts and situations. In addition to providing a balanced spectrum of actions and metrics, implementations of this disclosure are also suitable for automated tracking and can even interact with the customer's operating system (e.g., an enterprise resource planning system) to automatically implement the identified actions as part of the customer's actions.

[0023] Figure 1 An example system 100 for selecting features is depicted. Example system 100 includes a data source 110, a feature library 120, a feature selection engine 130, and a user interface 140. As described in more detail below, the data source 110 may include public and / or internal information associated with entities of interest. Information from the data source 110 is used to compile a list of candidate features 150 for a particular entity. The feature selection engine 130 can be used to select one or more features 160 loaded from the feature library 120. In some cases, the feature library 120 includes multiple feature trees, each belonging to a feature space or domain of interest. As described in more detail below, some implementations of system 100 are configured to exchange data with an enterprise resource planning system 180 or other information consumers. In some implementations, system 100 may also include a relevance engine 190.

[0024] Feature trees can be used to organize information according to a hierarchical structure. For example, a feature tree can group the features of an electric vehicle into groups such as chassis, powertrain, battery, electronics, and vehicle interior. Each of these groups can be decomposed into subgroups with progressively increasing levels of detail. For tires, a feature tree could include groups such as tread pattern, construction, and rubber composition. Software architecture can be similarly decomposed into a hierarchical structure of subsystems. In some cases, each node of a feature tree can be associated with one or more numerical metrics. For example, an electric vehicle's battery might need to provide a driving range measured in miles or kilometers, or a weight below a certain limit. Therefore, feature trees provide a standardized format for storing information. For example, the aforementioned car manufacturer could have a feature tree for every configuration of every vehicle. Collections of feature trees can be stored in a format that... Figure 1 The feature library is available on multiple network-based non-transient storage devices.

[0025] In some cases, the hierarchical structure of a feature tree can indicate the relationships between different parts of a system, rather than just simple structural relationships. Figure 2An example value tree 200 is depicted, which links strategic objectives to metrics through a hierarchical structure of sub-objectives and transformative actions. For example, the strategy domain 202 of sustainability can be broken down into several objectives 204, such as waste reduction and environmental concern. Each objective domain 204 includes one or more sub-objectives 206. For example, waste reduction could include both increasing recycling and reducing landfill waste. Environmental concern could include reducing carbon emissions. As indicated by the ellipsis, for clarity, a particular implementation of the value tree may include other nodes and branches omitted in the diagram.

[0026] Sub-objectives can be achieved through specific customer actions 208 measured using metric 210. In some cases, the top 204 through 206 of the value tree are generic (e.g., spanning an entire department), while the actions and metrics can be industry-specific or even company-specific. For example, in some cases, a manufacturing company can increase recycling by installing a waste heat recovery system that captures heat generated in a manufacturing process and uses that heat to drive a refrigeration cycle, thereby cooling another part of the manufacturing system. The workload associated with this type of recycling can be measured, for example, by the cooling capacity provided by the waste heat recovery system. However, this type of action may only be relevant to large-scale production, rather than small-scale production relying on 3D printing. Other actions that can increase recycling include, for example, using reusable packaging within the supply chain. If this sub-objective is considered for the IT security department rather than the manufacturing department, these targeted actions may not be appropriate.

[0027] Even within the manufacturing sector, Figure 2 The actions illustrated are merely examples. In some cases, analyzing and modifying existing manufacturing lines to reduce waste heat released into the environment may be appropriate, rather than simply increasing cooling capacity. Similarly, analyzing machine settings and tolerances to reduce the number of discarded parts may be appropriate, rather than simply focusing on the recycling of such parts. These additional actions involve, for example, another strategic area—the technical aspects of modernizing or improving manufacturing processes—which are not in themselves sustainability goals. Nevertheless, this example demonstrates that actions and metrics are specific and may involve multiple dimensions or strategic areas.

[0028] While the example value tree 200 is depicted as having five levels, other examples may include additional levels. For instance, certain actions may be broken down into a hierarchical structure of subtasks. In some cases, a particular action may be associated with multiple alternative metrics. These metrics may be associated with baseline and target values ​​based on industry information or past contracts.

[0029] For a given contract or project, the feature tree set (e.g., Figure 1The selected features (160) can represent relevant target domains, actions, and metrics in a structured manner. In some cases, identifying the feature tree set includes identifying suitable candidate feature trees stored in the feature tree set based on customer needs. In some implementations, identifying the feature tree set may include identifying multiple strategy domains and providing one or more feature trees for each strategy domain. Strategy domains may include financial, experience, sustainability, inclusion and diversity, and talent, to name just a few. In some cases, one or more feature trees, along with other value trees in the set, may be created and identified for custom strategy domains not included in the feature tree set.

[0030] Conversely, the feature tree stored in the collection can be updated (or a new feature tree can be created) as feedback from client contracts or projects, such as Figure 1 The arrows from user interface 140 to feature library 120 are shown in the diagram. In some cases, a feature tree can be complete in the sense that the number of features includes detailed metrics with updated baselines and target values. Other feature trees can be incomplete in the sense that only higher-level objectives are defined, while actions and metrics may be incomplete. Incomplete feature trees can capture new trends or evolving situations whose appropriate operations and metrics are not yet fully understood. These updated feature trees and newly created feature trees can be identified in subsequent client contracts or, in some cases, used to update ongoing contracts.

[0031] In some implementations, the feature selection engine 130 can identify a set of features 160 related to an entity or customer by identifying suitable candidate feature trees stored in a feature tree set. Candidate feature trees are identified based on, for example, customer needs. Customer needs are blocks of information expressing a customer's priorities or concerns related to a contract or project. Generally, a need will correspond to one or more nodes within a feature tree, although some interpretation is usually required to match the customer's expression with the exact content of the feature tree nodes.

[0032] The implementation of this disclosure can select requirements by assembling information associated with the client. This information can be obtained from information related to the client ( Figure 1 The data source 110 is associated with data in the form of text data, audio data, image data, or video data. The data is "associated" with the customer in the sense that it was created by the customer and / or describes the customer and their activities. For example, the data may include meeting minutes or records typically used to capture needs, but this process originates from a broader range of sources.

[0033] In some cases, this information may be publicly available. Publicly available data can include data from a customer's website, press releases, earnings calls, social media presence, and technical product documentation (e.g., user manuals), to name just a few. For example, publicly available data created by a customer could refer to a PDF manual uploaded to the customer's website. Articles describing new product launches on business publication websites could be examples of publicly available information from third parties. For the purposes of this disclosure, data created or provided by the customer themselves may reflect needs more accurately than third-party data. For example, a press release issued by the customer themselves may be more reliable than third-party social media posts tagged with the customer.

[0034] Implementations of this disclosure may also include techniques for identifying and correcting biases in the collected data. Such techniques can, for example, help prevent manipulation of the underlying model by publishing a large number of carefully crafted articles that could significantly alter the model and its performance. Anti-bias techniques may include, for example, data preprocessing before training, processing during training itself, or post-processing after training.

[0035] Using publicly available information can have the advantage of compiling requirements periodically at appropriate times. For example, requirements can be compiled while preparing for the initial client meeting, potentially shortening the client onboarding timeline. In fact, requirements can be compiled independently of client contracts.

[0036] In the case of client contracts, internal information may be used to replace or supplement publicly available data. This internal information is generally not available outside the entity of interest. Examples of internal information may include meeting notes, recordings of phone calls or meetings, internal company memos, pages on the company intranet, technical specifications, and lab notebooks. This information can be useful for contracts dealing with confidential applications that have not yet been made public (e.g., product development or product launch).

[0037] In some implementations, the data source may include a mixture of internal and publicly available information and / or a mixture of information created by the customer and information about the customer written by a third party.

[0038] Depending on the format of the original data, optical character recognition (OCR) and text-to-speech technologies can be used to convert the raw data into text. For example, an OCR engine such as Tesseract can be used to convert image data into text. In some cases, the image data can be resized or modified to remove noise and increase contrast to improve OCR accuracy. In some implementations, the extracted text is manually verified (e.g., for a digitized image of handwritten notes on a whiteboard).

[0039] Then, Natural Language Processing (NLP) techniques are used to process the text data. For example, text fragments can be fed into an NLP model using algorithms such as GloVe, word embeddings, etc. In some cases, a proprietary dictionary that matches keywords encoded from customer data with entries or nodes in a feature tree set can be used for NLP processing. NLP processing can create word vectors corresponding to potential customer needs. As mentioned earlier, needs will typically correspond to nodes within one or more feature trees.

[0040] The output of NLP processing is a list of 150 candidate features for a specific entity, each of which can represent a customer need. This list can be used to filter a feature tree set. An example of potential features extracted from a customer manual PDF might be "shifting to 40% renewable energy by 2026". For example, data from word embeddings can be aggregated to create a list of candidate features. This list can also be validated and stored for later use.

[0041] The implementation can identify suitable candidate feature trees in the feature tree set based on customer needs, such as a list of candidate features 150 for a specific entity. For example, an AI-based feature selection engine 130 can be used to filter the feature tree set included in the feature library 120 based on the candidate features 150.

[0042] In some cases, feature trees are grouped within a set by areas of interest (e.g., based on strategic objectives), and feature selection engine 130 is configured to provide at least one candidate feature tree for each area from a pre-selected set of areas of interest. For example, a user can specify financial, customer or employee experience, sustainability, inclusion and diversity, and talent as areas of interest, and feature selection engine 130 will provide at least one value tree for each strategic area. Another implementation may relate to the design of hybrid electric vehicles, and areas of interest may include, for example, chassis, electronics, batteries, powertrains, and vehicle interiors. In yet another example, each area of ​​interest may correspond to a step in a process, such as deposition, removal, patterning, and modification of electrical properties in semiconductor manufacturing.

[0043] In some cases, the list of candidate features 150 is evaluated by referencing other entities occupying the same technological or business space as the customer. For example, the customer might be a company in the energy industry. The feature selection engine 130 can be configured to use machine learning algorithms (e.g., k-nearest neighbors) to identify similar companies based on market size, industry, region, revenue segment, and area of ​​interest. The feature selection engine 130 can augment the list of candidate features compiled for the customer, or weight the individual features within the list to more closely reflect the market's overall focus.

[0044] In some implementations, the list of candidate features 150 can be used to search within the feature tree set and return matching feature trees. For example, the entry "shifting to 40% renewable energy by 2026" could return all feature trees related to renewable energy. Depending on the size of the feature tree set and the breadth of candidate features, the number of feature trees returned by the search may be greater than expected. Therefore, in some cases, the feature selection engine 130 can recommend specific features or feature trees for a specific contract.

[0045] Individual features within a feature tree of a set can be evaluated by creating a Single-Valued Factorization (SVD) matrix and applying weighted scores to each feature in the matrix. SVD is a matrix factorization technique that uses collaborative filtering mechanisms and latent factor models to generate recommendations. The weighted scores can be determined using interpretable AI techniques that learn from past contracts (i.e., feature selection and feature trees). The weighted scores can consider, for example, the selection of specific features by similar customers and the probability of achieving a target metric based on the current benchmark.

[0046] For example, a collection of feature trees can depict different designs for a particular vehicle and all its subsystems and constituent components. The goal of a particular contract might be to reduce the overall vehicle weight. Customer-related information can be processed according to the techniques disclosed herein to obtain a candidate list of features (i.e., components) for weight reduction. While the weight of identified features can be reduced simply, the above suggestions can be used to identify feature trees for vehicles similar to those of interest, and to identify features within these feature trees that contribute to weight reduction for those vehicles. In this example, the selected feature 160 can be input into a design platform to modify the weight of the associated component.

[0047] In some implementations, the feature selection engine 130 can be configured to recommend baseline and target values ​​for each of the recommended features. For example, the feature selection engine 130 can access a database storing baseline and target values ​​for commonly used features across different industries. The baseline and target values ​​can be updated during the ongoing contract period. For instance, a pilot client in a specific industry might initially set a target value of 40% for a given feature. Within a year of the contract's commencement, unforeseen circumstances might arise in the market, causing the client's competitors to set a 70% target for the same feature. In this case, the feature selection engine 130 can be configured to automatically generate a message containing the updated target for the feature whenever updated target information is stored, and send this message to the relevant users for the initial client account.

[0048] System 100 can be configured to generate a graphical user interface 140 that displays features and / or feature trees selected by feature selection engine 130. In some cases, the graphical user interface 140 can be used to create custom feature trees that are not included in the feature tree set, i.e., not included in feature library 120. In other cases, the user can use interface 140 to modify portions of the selected feature tree 160, for example, by adding individual features to or deleting features from an existing tree. Figure 1 As indicated by the corresponding arrow, the new or modified feature tree can be saved to feature library 120.

[0049] The user interface 140 can also be used to track the progress of the selected features 160 over a period of time. Users can manually update the values ​​corresponding to feature 160 via the interface 140. In some examples, machine learning techniques can be used to determine the correlations between the selected features 160 and help users update these values. In other examples, updated feature values ​​can be automatically entered based on update information from the data source 110.

[0050] For example, system 100 may include a correlation engine 190 configured to determine correlations between selected features 160. In some implementations, the correlation engine 190 receives the selected features 160 and automatically clusters them. For example, customer clusters can be created using k-means and centroid-based clustering. In some cases, users (e.g., data scientists) can evaluate whether a given correlation identified by the correlation engine is a valid correlation. In some cases, the evaluation of whether an identified correlation is valid depends on a minimum random sample of that correlation. Valid correlations can be stored in a correlation store. System 100 can use Euclidean distance to compare individual customers with correlations stored in the correlation store. An encoder can be used to transform the data into a single numeric vector for comparison. In some cases, system 100 requests user input via an interface regarding whether a customer exhibits a specific correlation identified by the correlation engine 190.

[0051] The Relevance Engine 190 can also use k-means to find relevance patterns for specific individual customers. These relevance patterns exist within features selected for the customer but are not yet displayed in the clusters (e.g., across industries or market segments). The Relevance Engine 190 can be configured to monitor these relevances until a sufficient set of samples within a cluster exhibits a relevance, and then save that relevance to a relevance repository.

[0052] In some implementations, feature selection engine 130 is configured to select features based on the correlation identified by correlation engine 190.

[0053] Figure 3This is a flowchart of an example process 300 for selecting a feature set for an entity of interest (e.g., a customer). This process can be, for example, by... Figure 1 The feature selection system 100 is used to perform this.

[0054] Feature selection system 100 receives at least one of text, audio, image, or video data associated with an entity of interest via a network (302). Feature selection system 100 identifies a candidate feature set for a specific entity based on the received data (304). In some cases, feature selection system 100 extracts text from the text, audio, image, or video data associated with the entity and uses a natural language processing (NLP) model to create vectors of feature-related words based on the usage and context of words in the data associated with the entity. Feature selection system 100 loads a feature library comprising multiple features, each of which is assigned to one or more feature spaces (306). For example, as described above, the feature library may include a collection of feature trees. Each feature tree assigns multiple features to a given feature space (e.g., a domain of interest). Feature selection system 100 uses feature selection engine 130 to select one or more features from each feature space in the feature space based on the candidate feature set for a specific entity (308).

[0055] In some implementations, feature selection engine 130 selects one or more features from each feature space in the feature space based on both a candidate feature set for a specific entity and a candidate feature set for a specific domain. In this case, feature selection system 100 may receive at least one of text, audio, image, or video data associated with multiple entities other than the initial entity of interest. Based on the received data, feature selection system 100 identifies entity domains common to the entity of interest and multiple additional entities. For example, an entity domain may contain shared technologies implemented by the entities (e.g., blockchain technology). In other cases, an entity domain may correspond to a specific type of device (plug-in hybrid electric vehicle). In other examples, an entity domain may include a specific industry (e.g., the energy industry or the semiconductor manufacturing industry).

[0056] Once the entity domain is identified, the feature selection system 100 can load a candidate feature set for the specific domain from a feature library for that domain. Then, the feature selection engine 130 can select one or more features from each feature space in the feature space based on both the candidate feature set for the specific entity and the candidate feature set for the specific domain. In some cases, each of the loaded candidate features for the specific domain also includes a baseline measurement, and the feature selection system 100 identifies the baseline measurement for each of the selected features based on the baseline measurement of the candidate features for the specific domain.

[0057] In some cases, feature selection engine 130 selects a first feature set, comprising one or more features, from each feature space in the feature space based on a candidate feature set for a specific entity. Then, at a second time after the first time, feature selection engine 130 selects a second feature set, comprising one or more features, from each feature space in the feature space based on both the candidate feature set for the specific entity and the candidate feature set for the specific domain. For example, the second time could be a month, a quarter, or a year after the first time. In some cases, the second feature set can be used to update or overwrite the first feature set. For example, feature selection engine 130 can continuously select new feature sets at regular intervals after the second time.

[0058] In some implementations, feature selection system 100 uses interpretable AI techniques to assign weighted scores to each of the selected features and filters the selected features based on the weighted scores.

[0059] In some cases, feature selection system 100 can identify one or more custom features that are included in the candidate feature set of a specific entity or the candidate feature set of a specific domain but not included in the feature library, and store the one or more custom features in the feature library.

[0060] In some implementations, a visualization can be generated that presents both the features selected by the feature selection engine 130 and one or more unselected features for each feature space. The visualization can be presented to the user on a user terminal device. User input can be received to select or deselect one or more features, and the selected features can be updated based on the user input.

[0061] Figure 4 An example system 400 capable of implementing the present disclosure is depicted. Example system 400 includes a computing device 402, a back-end system 408, and a network 406. In some examples, network 406 includes a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof, and connects websites, devices (e.g., computing device 402), and back-end systems (e.g., back-end system 408). In some examples, network 406 can be accessed via wired and / or wireless communication links.

[0062] In some examples, computing device 402 may include any suitable type of computing device, such as a desktop computer, laptop computer, handheld computer, tablet computer, personal digital assistant (PDA), cellular phone, network appliance, camera, smartphone, enhanced general packet radio (EGPRS) mobile phone, media player, navigation device, email device, game console, or a suitable combination of any two or more of these devices or other data processing devices.

[0063] In the depicted examples, backend system 408 includes at least one server system 412 and data storage 414 (e.g., a database and knowledge graph structure). In some examples, at least one server system 412 hosts one or more computer-implemented services that a user can interact with using a computing device. For example, according to an implementation of this disclosure, server system 412 may host one or more applications provided as part of a feature selection system. For example, user 420 (e.g., a supplier) may interact with the feature selection system using computing device 402. In some examples, user 420 may provide data associated with entities of interest to select one or more features using a feature selection engine, as described in more detail herein. In some examples, user 420 may create or update one or more feature trees, as described in more detail herein. In some cases, updated or new feature trees(s) ... Figure 1 Enterprise Resource Planning System 180.

[0064] Figure 1 An Enterprise Resource Planning (ERP) system 180 can be implemented as a software system for collecting, storing, managing, and interpreting data from various departments within an enterprise (e.g., manufacturing, purchasing, accounting, sales and marketing, human resources, etc.). For example, the ERP system 180 can be used to share data collected by various user platforms across different departments within an enterprise. Data collected by one department is often stored locally on a computer or server in a non-standard format defined by the hardware or software platform used by that department. This discrepancy can cause the ERP system to be unable to process fragmented or incomplete data in real time. Conversely, some data instances may be duplicated, overloading the ERP system's processor and consuming unnecessary bandwidth for data transmission.

[0065] For example, an ERP system might have data related to the battery of a hybrid electric vehicle and data related to the powertrain of the same vehicle. If this data cannot be integrated due to inconsistencies in format, the ERP system may be unable to plan the manufacturing steps for installing the battery in each of a series of vehicles. Similarly, if data is stored in separate locations and cannot be shared promptly or easily across ERP systems, difficulties in production planning may arise. In some cases, the work of integrating data from such non-standard or incomplete records may lead to production errors, such as the wrong battery being installed in the wrong vehicle. The implementation of this disclosure addresses these and other problems by collecting, transforming, and integrating information across ERP systems into a standardized format. Although Figure 1 System 100 and ERP system 180 are depicted as separate entities, but in some implementations, system 100 can be directly integrated into ERP system 180 itself.

[0066] The implementations and all functional operations described in this specification can be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. These implementations can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of substances affecting machine-readable propagation signals, or a combination of one or more of them. The term "computing system" includes all means, devices, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. In addition to hardware, the means may also include code that creates an execution environment for the computer program in question (e.g., code constituting processor firmware, protocol stacks, database management systems, operating systems, or any suitable combination thereof). Propagation signals are artificially generated signals (e.g., machine-generated electrical, optical, or electromagnetic signals) that are generated to encode information for transmission to a suitable receiver device.

[0067] A computer program (also known as a program, software, software application, script, or code) can be written in any suitable programming language, including assembly or interpreted languages, and can be deployed in any suitable form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored as a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), in a single file specific to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer, or on multiple computers located at a site or distributed across multiple sites and interconnected by a communication network.

[0068] The processes and logic flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. These processes and logic flows can also be executed by special-purpose logic circuitry (e.g., FPGA (Field Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit)), and the apparatus can also be implemented as special-purpose logic circuitry.

[0069] Processors suitable for executing computer programs include, for example, both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors of any suitable type of digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. Components of a computer may include a processor for executing instructions and one or more storage devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled thereto to receive data or transfer data to or from them, or both. However, a computer does not need to have such devices. Furthermore, a computer may be embedded in another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio player, a global positioning system (GPS) receiver). Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM discs. Processors and memory can be supplemented by dedicated logic circuits or incorporated into dedicated logic circuits.

[0070] To provide interaction with the user, these implementations can be carried out on a computer with display devices for showing information to the user (e.g., CRT (cathode ray tube), LCD (liquid crystal display) monitors) and keyboards and pointing devices (e.g., mouse, trackball, touchpad), through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any suitable form of sensory feedback (e.g., visual feedback, auditory feedback, tactile feedback); and input from the user can be received in any suitable form, including acoustic, speech, or tactile input.

[0071] These implementations can be implemented in computing systems that include backend components (e.g., as a data server), middleware components (e.g., an application server), and / or frontend components (e.g., a client computer with a graphical user interface or web browser through which a user can interact with the implementation), or any suitable combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any suitable form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), such as the Internet.

[0072] A computing system may include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is established through computer programs running on their respective computers that have a client-server relationship with each other.

[0073] While this specification contains numerous details, these details should not be construed as limiting the scope of this disclosure or the scope of any claims, but rather as descriptions of features specific to a particular implementation. Some features described in this specification in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately in multiple implementations or in any suitable sub-combination. Furthermore, although these features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may refer to a sub-combination or a variation of the sub-combination.

[0074] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or sequentially, or requiring all of the shown operations to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the above implementation should not be interpreted as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or encapsulated in multiple software products.

[0075] Many implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of this disclosure. For example, various forms of processes shown above can be used, in which steps are reordered, added, or removed. Therefore, other implementations are within the scope of the appended claims.

Claims

1. A computer-implemented method, comprising: Receive, via a network, at least one of text, audio, image, or video data associated with multiple entities, including entities of interest; Based on the received data, a set of candidate features is used to identify a specific entity; Based on the received data, identify the entity fields that include the plurality of entities; Load a feature library containing multiple features, each of which is assigned to one or more feature spaces; as well as Load a candidate feature set for a specific domain from a feature library for that specific domain of the entity domain; Using a feature selection engine, (i) at a first time, based on the candidate feature set of the specific entity, a first feature set including one or more features is selected from each feature space in the feature space; and (ii) at a second time after the first time, based on the candidate feature set of the specific entity and the candidate feature set of the specific domain, a second feature set including one or more features is selected from each feature space in the feature space. Using interpretable AI technology, a weighted score is assigned to each of the selected features in the first feature set and the second feature set; The selected features of the first feature set and the second feature set are filtered based on the weighted scores. By presenting the features of the filters and visualizing the selection of each filter's features, the system receives user input that selects features of one or more specific filters from the first feature set and the second feature set. as well as The interpretable AI technology is updated to learn from user input that selects features of one or more specific filters, thereby generating future weighted scores for features associated with other contracts of entities similar to the entity of interest.

2. The computer-implemented method according to claim 1, wherein the candidate feature set identifying the specific entity includes: Extract text from the text, audio, image, or video data associated with the entity; as well as Using a Natural Language Processing (NLP) model, vectors of feature-related words are created based on the usage and context of words in the data associated with the entity.

3. The computer-implemented method of claim 1, wherein, Each candidate feature among the loaded candidate features of the specific domain also includes a baseline measurement, and the method further includes: identifying the baseline measurement based on the baseline measurement of the candidate features of the specific domain for each of the selected features.

4. The computer-implemented method according to claim 1 further includes: Identify one or more custom features that are included in the candidate feature set of the specific entity or the candidate feature set of the specific domain, but are not included in the feature library; as well as The one or more customized features are stored in the feature library.

5. A system comprising: One or more processors; as well as A computer-readable storage device coupled to the one or more processors and storing instructions thereon, which, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: Receive, via a network, at least one of text, audio, image, or video data associated with multiple entities, including entities of interest; Based on the received data, a set of candidate features is used to identify a specific entity; Based on the received data, identify the entity fields that include the plurality of entities; Load a feature library comprising multiple features, each of which is assigned to one or more feature spaces; and Load a candidate feature set for a specific domain from a feature library for that specific domain of the entity domain; Using a feature selection engine, (i) at a first time, based on the candidate feature set of the specific entity, a first feature set including one or more features is selected from each feature space in the feature space; and (ii) at a second time after the first time, based on the candidate feature set of the specific entity and the candidate feature set of the specific domain, a second feature set including one or more features is selected from each feature space in the feature space. Using interpretable AI technology, a weighted score is assigned to each of the selected features in the first feature set and the second feature set; The selected features of the first feature set and the second feature set are filtered based on the weighted scores. By presenting the features of the filters and visualizing the selection of each filter's features, the system receives user input selecting features of one or more specific filters from the first feature set and the second feature set; and The interpretable AI technology is updated to learn from user input that selects features of one or more specific filters, thereby generating future weighted scores for features associated with other contracts of entities similar to the entity of interest.

6. The system of claim 5, wherein the candidate feature set identifying the specific entity comprises: Extract text from the text, audio, image, or video data associated with the entity; as well as Using a Natural Language Processing (NLP) model, vectors of feature-related words are created based on the usage and context of words in the data associated with the entity.

7. The system of claim 5, wherein, Each candidate feature in the loaded candidate features of the specific domain also includes a baseline measurement, and the operation further includes: For each of the selected features, a baseline measurement is identified based on the baseline measurement of the candidate features of the specific domain.

8. The system of claim 5, wherein the operation further comprises: Identify one or more custom features that are included in the candidate feature set of the specific entity or the candidate feature set of the specific domain, but are not included in the feature library; as well as The one or more customized features are stored in the feature library.

9. A computer-readable storage medium coupled to one or more processors and storing instructions thereon, the instructions, when executed by the one or more processors, causing the one or more processors to perform an operation, the operation comprising: Receive, via a network, at least one of text, audio, image, or video data associated with multiple entities, including entities of interest; Based on the received data, a set of candidate features is used to identify a specific entity; Based on the received data, identify the entity fields that include the plurality of entities; Load a feature library containing multiple features, each of which is assigned to one or more feature spaces; as well as Load a candidate feature set for a specific domain from a feature library for that specific domain of the entity domain; Using a feature selection engine, (i) at a first time, based on the candidate feature set of the specific entity, a first feature set including one or more features is selected from each feature space in the feature space; and (ii) at a second time after the first time, based on the candidate feature set of the specific entity and the candidate feature set of the specific domain, a second feature set including one or more features is selected from each feature space in the feature space. Using interpretable AI technology, a weighted score is assigned to each of the selected features in the first feature set and the second feature set; The selected features of the first feature set and the second feature set are filtered based on the weighted scores. By presenting the features of the filters and visualizing the selection of each filter's features, the system receives user input that selects features of one or more specific filters from the first feature set and the second feature set. as well as The interpretable AI technology is updated to learn from user input that selects features of one or more specific filters, thereby generating future weighted scores for features associated with other contracts of entities similar to the entity of interest.

10. The storage medium of claim 9, wherein the candidate feature set identifying the particular entity comprises: Extract text from the text, audio, image, or video data associated with the entity; as well as Using a Natural Language Processing (NLP) model, vectors of feature-related words are created based on the usage and context of words in the data associated with the entity.

Citation Information

Patent Citations

  • Omnichannel, intelligent, proactive virtual agent

    US20190042988A1