Methods for accessing data records in a master data management system
By integrating multiple search engines and adapting to user feedback, the system optimizes data access in master data management systems, addressing inefficiencies and enhancing user experience through improved search performance and relevance.
Patent Information
- Application Number
- JP2021557224
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-02
- Filing Date
- 2020-03-19
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2040-03-19
AI Technical Summary
Existing master data management systems face challenges in efficiently accessing and managing data records due to limitations in search engine capabilities and performance, leading to repeated searches and suboptimal user experiences.
The system enhances master data management systems with multiple search engines, selects the best combination based on attribute performance, and processes requests using these engines to provide optimized results, which are then refined through user interaction feedback for continuous improvement.
This approach improves data access efficiency, reduces redundant searches, and enhances user experience by leveraging diverse search engine capabilities and adapting to user behavior, resulting in higher-quality search outcomes.
Smart Images

Figure 0007740839000001 
Figure 0007740839000002 
Figure 0007740839000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of digital computer systems, and more particularly to a method for accessing data records in a master data management system. [Background technology]
[0002] Enterprise data matching deals with matching and linking customer data received from different sources to create a single version of the truth. Master data management (MDM)-based solutions work with enterprise data to index, match, and link the data. Master data management systems may enable access to these data. However, there is a continuous need to improve access to data in master data management systems. Summary of the Invention [Means for solving the problem]
[0003] Various embodiments provide a method, a computer system and a computer program product for accessing data records in a master data management system as described in the subject matter of the independent claims. Advantageous embodiments are described in the dependent claims. The embodiments of the invention may be freely combined with one another if they are not mutually exclusive.
[0004] In one aspect, the present invention relates to a method for accessing a data record in a master data management system, the data record including a plurality of attributes, the method comprising: augmenting the master data management system with one or more search engines to enable access to the data records; receiving a request for data in a master data management system; identifying a set of one or more attributes from a plurality of attributes referenced in the received request; selecting a combination of one or more search engines from the master data management system whose performance for searching at least some values of the set of attributes satisfies the current selection rules; processing the request using a combination of search engines; Providing at least a portion of the results of the processing.
[0005] In another aspect, the present invention relates to a computer system for providing access to a data record, the data record including a plurality of attributes, the computer system including a plurality of search engines for providing access to the data record, a user interface configured to receive a request for data, an entity identifier configured to identify a set of one or more attributes from the plurality of attributes referenced in the received request, an engine selector configured to select a combination of one or more search engines from the search engines such that their performance for searching values of at least some of the set of attributes satisfies current selection rules, the search engines configured to process the request, and a results provider configured to provide at least some of the results of the processing.
[0006] In another aspect, the present invention relates to a computer program product including a computer readable storage medium having computer readable program code embodied thereon, the computer readable program code configured to access data records in a master data management system, the data management system including a search engine for enabling access to the data records, the data records including a plurality of attributes, the computer readable program code being further configured to: receive a request for data at the master data management system; identify a set of one or more attributes from the plurality of attributes referenced in the received request; select a combination of one or more search engines from the master data management system whose performance for searching values of at least some of the set of attributes satisfies current selection rules; process the request using the combination of search engines; and provide at least a portion of the results of the processing.
[0007] In the following, embodiments of the invention will be explained in more detail, by way of example only, with reference to the drawings, in which: [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a flow diagram illustrating a method for accessing data records in a master data management system. [Figure 2] 1 is a flow diagram illustrating a method for providing search results for a set of search engines. [Figure 3] 1 is a flow diagram illustrating a method for providing search results from multiple search engines. [Figure 4A] FIG. 1 illustrates a table containing normalized and merged search results from different engines. [Figure 4B] FIG. 10 illustrates a table containing example engine weights. [Figure 4C] FIG. 10 illustrates a table containing example attribute weights based on the reliability with which an entity recognizes and identifies attribute types. [Figure 4D]FIG. 10 illustrates a table containing example integrity weights. [Figure 4E] FIG. 10 illustrates a table containing example freshness weights. [Figure 4F] FIG. 1 shows a table containing outcome records and associated weights and scores. [Figure 5] 1 is a flow diagram illustrating a method for updating weights used to weight matching scores of data records resulting from processing a search request by multiple search engines. [Figure 6A] FIG. 1 shows a table containing the number of user clicks as a function of the completeness of the data record. [Figure 6B] FIG. 1 shows a table containing the percentage of user clicks as a function of the completeness of the data record. [Figure 6C] FIG. 10 shows a graph of the distribution of click percentage as a function of the completeness of the data record. [Figure 7] 7 is a block diagram illustrating a computer system 700 according to an example of the present disclosure. [Figure 8] 1 is a flow diagram of a method illustrating an example of the operation of a master data management system. [Figure 9] FIG. 1 illustrates an example of request processing according to the present subject matter. DETAILED DESCRIPTION OF THE INVENTION
[0009] The descriptions of various embodiments of the present invention are provided for illustrative purposes and are not intended to be exhaustive or limiting to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications or technical improvements to technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0010] The subject matter may enable efficient access to data stored in a master data management system. The subject matter may improve the performance of a master data management system. The subject matter may reduce the number of repeated or retried search requests because the subject matter may use multiple search engines to provide the best possible results, eliminating the need for users to retry or reformulate search queries as may occur with other systems.
[0011] A master data management system may use a single type of search engine. The present subject matter may also allow a master data management system to use different types of search engines. The type of search engine may be determined by the technology used to perform the search, such as full-text search or structured probabilistic search. For example, an additional search engine added by the present method may be of a different type than the type of search engine originally included in the master data management system. The present subject matter may thus provide a collective search and matching engine that aims to take advantage of the best of all the different capabilities of multiple search and indexing engines based on the type of input data or the type of query made. Because different indexing or search engines have different capabilities, they work best for different types of input or different requirements. The present subject matter may enable better ways to search data by using multiple different indexing engines and search engines, which improves the user experience without affecting the performance of machine-based interactions.
[0012] For example, the steps of identifying, selecting, processing, and providing may be performed automatically upon receiving a request for data. In one example, the steps of identifying, selecting, processing, and providing may be repeated automatically upon receiving a further request for data, each iteration using updated selection rules resulting from a previous execution of the method.
[0013] The results may include data records. Provisioning the data records may include displaying data representing the data records in a graphical user interface. For example, a row may be displayed for each data record, which may be a hyperlink or link that a user can click to access more information about that data record.
[0014] A data record, or record, is a collection of related data items, such as a particular user's name, date of birth (DOB), and class. A record represents an entity; the entity represents the user, object, or concept about which information is stored in the record.
[0015] According to one embodiment, the method further includes updating the selection rules based on user actions on the provided results, where the updated selection rules become the current selection rules, and repeating the identifying, selecting, processing, and providing steps using the current selection rules when another request for data is received. In one example, the selection rules may be updated after a predetermined period of time, e.g., during which the method may be performed multiple times, based on a combination of user actions on the provided results during that period. This may enable a self-improving search system based on user input and experience. A search engine whose performance for searching at least some values of the set of attributes satisfies the current selection rules is a search engine that is part of a predetermined table of the data management system associated with at least some of the set of attributes. For example, the table includes multiple entries. Each entry i in the table includes a search engine SEi and one or more associated attributes Ti that are preferably searched by the search engine. In one example, each association between Ti and SEi may be assigned an updated score that may be changed or updated. The selected search engine is a search engine SEi of a table associated with one or more attributes of the set of attributes; for example, when the set of attributes includes T1 and T2, the table may be searched to identify entries having T1 and T2, and the selected search engine is the search engine of those identified entries. Updating the selection rules may include updating the table, for example, when the number of clicks on displayed results derived from search engine SEx and associated with a given searched attribute Tx is less than a threshold, thereby, for example, removing the association between Tx and SEx, or modifying the update score if Tx and SEx are associated with an update score, for example, by lowering the update score.For example, removal may occur when the same combination of Tx and SEx has been found to perform poorly at least once in the past, e.g., when the number of clicks on associated results is less than a threshold on multiple occasions, and thus the associated update score is less than a given threshold. In one example, the table may initially have many or all possible combinations of attributes and search engines, and entries that do not perform for a predetermined period of time may be removed.
[0016] According to one embodiment, the results include data records in the master data management system associated with respective matching scores obtained by the search engine's scoring engine, and the provided results include non-duplicate data records having matching scores above a predetermined score threshold. The matching score may indicate a level or degree of match between the data record and the requested data.
[0017] By providing only results that meet the matching score selection criteria, this embodiment may further improve the performance of the master data management system. For example, irrelevant results may not be provided to the user. This may save processing resources, such as display resources and data transmission resources, that would otherwise be used for irrelevant results. For example, score weighting may be performed as described in the following embodiments.
[0018] According to one embodiment, the results include data records in the master data management system associated with respective matching scores obtained by the search engine's scoring engine, and the method further includes weighting the matching scores according to the performance of components involved in generating the results, where the components include the method steps and elements used to generate the results as well as at least a portion of the results, and the provided results include non-duplicate data records having weighted matching scores higher than a predetermined score threshold. The weighting step may, for example, include assigning, for each data record in the results, a weight to each component of the components that provided or generated that data record, which may include the provided data record itself, combining the weights, and weighting the data record's matching score using the combined weights.
[0019] For example, generating search results for a received data request involves executing a search process (the method may include a search process). This search process may have multiple process steps, and each process step may be performed by a system element, such as a search engine or a scoring engine. The search process may have components, which may be process steps, system elements, or results provided by the search process, or a combination thereof. Each component may have a function that contributes to obtaining the search results. Each of these components of the search process may affect the quality of the results obtained. For example, when a component of the search process is not functioning properly, this may affect the search results. For example, if a component is a process step that identifies attributes for a received request, and this component does not efficiently identify attributes of a particular type, this process step may not correctly identify attributes of this type. Thus, when a request for data referencing attributes of this type is received, the results obtained may be affected because they may include irrelevant and unnecessary search results for the incorrectly identified attributes. The performance of components of the search process may have different contributions to the results obtained by the search process. This embodiment may take into account at least some of those contributions by weighting the matching scores accordingly. For example, each component of at least some of the components of the search process of this embodiment may be assigned a weight indicating its performance in performing its respective function. The weights may be defined by a user, for example, the weights may be initially defined by a user (e.g., for a first run of the method) and then automatically updated by a weight update method as described herein. The weights may be used to weight the matching scores. This embodiment may further increase the performance of the data management system.For example, further irrelevant results may not be provided to the user, which may save processing resources, such as display resources and data transmission resources.
[0020] Examples of components that are considered in weighting the search process may be described in the following embodiment, which may be advantageous for identifying and weighting components whose performance may have a greater impact on search results.
[0021] According to one embodiment, the component includes a search engine, an identification step, and results. The method further includes the steps of assigning an engine weight to each of the search engines and attribute weights to a set of attributes, where the attribute weight of an attribute indicates a level of confidence with which the attribute is identified; assigning, for each data record in the results, a completeness weight indicating the completeness of the data record and a freshness weight indicating the freshness of the data record; combining, for each data record in the results, the respective engine weight, attribute weight, completeness weight, and freshness weight; and weighting the score of the data record by the combined weight. Attribute weights may be generated at the attribute level and applied to the entire result set (and all attributes), which is returned in response to the received request. This may allow the result set to be considered less useful when the automatically determined search-entity-type itself is incorrect.
[0022] The following embodiments provide weight update methods for updating weights used by the present subject matter, which allow for efficient and systematic processing of the weighting procedure.
[0023] According to one embodiment, the method further includes providing user parameters quantifying user behavior on the provided results, and for each of at least some of the components, defining values of the user parameters and associated values of component parameters describing the component, and using the defined associations to update a weight assigned to the component. The component parameters may include, for example, at least one of completeness, freshness of data records, search engine ID, and trustworthiness of identifying attributes.
[0024] For example, user actions or interactions may be monitored by an activity monitor in the master data management system. In one example, a user action may be a user click on a provided result. Associated values of user parameters and component parameters may be provided in the form of a distribution that can be fitted or modeled to derive weights. For example, a distribution of click counts for various features of rows representing data records (e.g., features may indicate, for example, which search engine the data record came from, how reliable the entity type detection was, how complete the record was, how fresh the record was, etc.) may be provided and analyzed to find weights. This embodiment may be performed, for example, for every new click, and the distribution can be modified, for example, as every new click is fed back into the system, thereby facilitating weight reallocation. This embodiment may allow weights used in previous iterations of the method to be updated. This embodiment may allow the data management system to continue self-improvement based on its own experience with data searches. For example, all weights used in the above embodiment may be updated. In another example, only some of the weights used (e.g., the completeness weight) may be updated. Updating the weights may include determining new weights and replacing used weights with the respective new weights, which may be determined according to this embodiment by monitoring user activity with respect to results provided to users.
[0025] According to one embodiment, the method further comprises providing a look-up table relating values of the user parameters to values of the component parameters, and using the look-up table to update the weights assigned to the components.
[0026] According to one embodiment, the method further includes modeling variation in values of the user parameters as a function of values of the component parameters using a predetermined model, using the model to determine update weights for the components, and using the update weights to update the weights assigned to the components. For example, the predetermined model may be configured to receive the component parameter values as inputs and output the respective weights. This may enable accurate weighting techniques according to the present subject matter.
[0027] In one embodiment, a user action among the user actions includes a mouse click on a displayed result among the provided results, and the user parameters include at least one of the number of clicks, the frequency of clicks, and the duration of accessing a given result among the results. For example, the activity monitor may use a click count, check the time spent on an individual result (e.g., from when that result is clicked to when the back / restart button is used), check traversal of the result set and consider the last selected record on which the user spent more than some threshold time as the "user-preferred result," or some combination thereof.
[0028] According to one embodiment, for each attribute of the set of attributes, the selection rule includes the steps of: determining, for each of the search engines, a value of a performance parameter indicative of the search engine's performance for searching the value of the attribute; weighting the determined values by their respective current weights; and selecting search engines having a performance parameter value higher than a predetermined performance threshold.
[0029] For example, in a first or initial execution of the method of this embodiment, the current weight may be set to 1. In another example, when a set of attributes includes three attributes att1, att2, and att3, the performance of each search engine, such as search engine 1 (SE1), may be evaluated. This may result in three performance parameter values, Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1, for each search engine. The current weight for search engine SE1 may be determined from Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1, resulting in weights W1_SE1, W2_SE1, and W2_SE1. These weights may be used to weight the performance parameter values Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1. To determine whether to select search engine SE1, a weighted combination of Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1 may be determined, and SE1 may be selected if the combined value (e.g., average) is higher than a performance threshold. In another example, each of the weighted performance values Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1 may be compared to a performance threshold, and SE1 may be selected only if each of their values is higher than the performance threshold.
[0030] According to one embodiment, the performance parameters include at least one of the number of results and the level of matching of the results to what is expected or required.
[0031] In one embodiment, the selection rule uses a table associating attributes with corresponding search engines, and updating the selection rule includes determining a value for a user parameter that quantifies user behavior for each search engine in the combination of search engines; using the determined value associated with each search engine in the combination of search engines to identify values of the user parameter that are less than a predetermined threshold; for each identified value of the user parameter, determining an attribute from the set of attributes and a search engine associated with the identified value; and updating the table with the determined attribute and search engine. In one example, the table initially contains many or all possible combinations of attributes and search engines. For example, after a predetermined period of time, unperformed entries may be removed. For example, the user parameter may be the number of clicks for each result in the combination of search engines, i.e., there is a value for the user parameter for each displayed result. These values may be compared to a predetermined threshold (e.g., 10 clicks), and displayed results associated with values less than the threshold may be identified. Each of those identified results is the result of a search by a given search engine X for one or more attributes, such as attribute T1 of the set of attributes, such that X and T1 may be used to update the tables described herein.
[0032] According to one embodiment, processing of requests is performed in parallel by a combination of search engines, which may accelerate the subject search process.
[0033] According to one embodiment, the search engine combination is a ranked list of search engines, and requests are processed sequentially according to the ranked list until a minimum number of results is exceeded. This may conserve processing resources. If the engine selection rules suggest only Engine 1 (SE1), but the actual search does not produce enough results, SE2 (next in the ranked list) may be used.
[0034] According to one embodiment, the results provided include data records that are filtered depending on the sender of the request, for example, obtaining a list of matches for a given data input, providing role-based visibility, applying consent filters, and then applying data governance rules, providing better quality matches and search flexibility while respecting privacy.
[0035] In one embodiment, identifying the set of attributes includes inputting the received request into a predetermined machine learning model and receiving a classification of the request from the machine learning model, the classification indicating the set of attributes.
[0036] According to one embodiment, the selection rule includes inputting a set of attributes into a predetermined machine learning model and receiving one or more search engines from the machine learning model that can be used to search the set of attributes.
[0037] According to one embodiment, the method further includes receiving a training set indicating one or more different sets of attributes, each set of attributes labeled to indicate a search engine suitable for conducting a search on the set of attributes, and generating a machine learning model by training a predetermined machine learning algorithm using the training set.
[0038] Figure 1 is a flow diagram of a method for accessing a data record in a master data management system. The data record includes a plurality of attributes.
[0039] For example, the master data management system may process records received from client systems and store the data records in a central repository. For example, the client systems may communicate with the master data management system over a network connection, including, for example, a wireless local area network (WLAN) connection, a wide area network (WAN) connection, a local area network (LAN) connection, or a combination thereof.
[0040] The data records stored in the central repository may have a predetermined data structure, such as a data table with multiple columns and rows. The predetermined data structure may include multiple attributes (e.g., each attribute represents a column in the data table). In another example, the data records may be stored in a graph database as entities with relationships. The predetermined data structure may include a graph structure, where each record may be assigned to a node in the graph. Examples of attributes may be name, address, etc.
[0041] The master data management system may include a search engine (called the initial search engine), which searches the data records stored in the central repository based on the received search query using a single technique, such as a probabilistic structured search. Like any other search engine, the initial search engine may be highly suited to certain types of attributes and less suited to others. That is, the performance of the initial search engine may depend on the type of attribute value being searched. For example, the attribute "name" may be well searched by a probabilistic search engine because of its nickname and phonetic equivalents, while an address attribute such as city may be partial and therefore work well with a free-text search engine. Therefore, in step 101, the master data management system may be enhanced with one or more search engines to enable access to the data records in the central repository. This may result in multiple search engines, including the initial search engine and additional search engines. For example, each search engine in a master data management system may be associated with a respective API through which search queries may be received. This may enable a collective search and matching engine aimed at leveraging the best of all the different capabilities of multiple search and indexing engines based on the type of input data or the type of query made. Because different indexing or search engines have different capabilities, they work best for different types of input or with different requirements.
[0042] In step 103, the master data management system may receive a request for data. This request may be received, for example, in the form of a search query. For example, the search query may be used to search for an attribute value, a collection of attribute values, or any combination thereof. The search query may be, for example, an SQL query. The received request may indicate one or more attributes of data records in the central repository. This may be done, for example, by explicitly indicating the attribute in the request, indirectly indicating the attribute, or both. For example, the search query may be a structured search, in which comparison or range predicates are used to restrict the values of a particular attribute. A structured search may provide an explicit reference to an attribute. In another example, the search query may be an unstructured search, such as a keyword search that filters out records that do not contain some form of a specified keyword. An unstructured search may indirectly reference an attribute. In one example, the received request may include a name, an entity type, or a mathematical expression and time expression, or a combination thereof, in an unstructured format.
[0043] Upon receiving a request, in step 105, entity identifiers in the master data management system may be used to identify a set of one or more attributes referenced in the received request. Identifying the set of attributes may further include identifying the entity type of each attribute for at least some of the set of attributes. For example, the received request may be analyzed, e.g., parsed to search for attributes having the searched value. For example, the entity identifiers may identify entity names and types, mathematical expressions, and time expressions coming in as unstructured text and map them to attributes in the master data management system with specific probabilities, thereby enabling them to be used to perform structured searches.
[0044] An entity identifier may be, for example, a token recognizer that identifies strings, numbers, pattern names, locations, etc. For example, email identification may use the following email structure: abc@uvw.xyz. Phone number identification may be based on the fact that phone numbers are 10 digits long. Social Security number (SSN) identification may be based on the fact that SSNs have the following structure: AAA-BB-CCCC.
[0045] In one example, the entity identifier may use a machine learning (ML) model generated by a machine learning (ML) algorithm. The ML algorithm may be configured to read enterprise data and identify / learn portions of the data to identify attributes. The entity identifier may use the ML model to determine whether a given input text could be a name, address, phone number, SSN, etc., with a certain probability. The engine selector may also use the ML model generated by the ML algorithm to make its selection.
[0046] Using the set of identified attributes (e.g., and / or associated entity types), the engine selector of the master data management system may select a combination of one or more search engines from the search engines of the master data management system in step 107. For example, the performance of each search engine of the master data management system for searching values of each attribute of the attributes may be evaluated. The performance of the search engines may be determined by evaluating a performance parameter. For example, the performance parameter may be the average number of results obtained by the search engine for searching different values of the attribute and clicked or used by users. The performance parameter may alternatively or additionally include the average matching score of the results obtained by the search engine for searching different values of the attribute and clicked or used by users.
[0047] Using the current selection rules, a selection of one or more search engine combinations may be performed. For example, the selection rules may be applied to each given attribute of the set of attributes as follows: For each search engine of the search engines of the master data management system, a value of a performance parameter may be determined that indicates the performance of the search engine for searching values of the given attribute. As a result, for example, when the set of attributes includes two attributes, each search engine may have two performance values associated with the two attributes, and therefore multiple values may be obtained for each search engine of the search engine combination.
[0048] For example, when the set of attributes includes name and date of birth attributes, a structured probabilistic search engine may yield better results for this set of inputs and may therefore be selected. In addition, a free text search engine may be selected. Execution of the request may be performed using two engines as follows: If no results are found by the probabilistic search engine, a free text search may also be performed. In another example, both search engines may be used to execute the request, regardless of their respective results. In another example, the set of attributes may include year of birth and phone number. In this case, both engines may be selected because the probabilistic search engine can handle edit distance values and the year of birth can be well served by the free text engine as a partial text of the date of birth. When the received request specifically invokes AND or NOT logic, the full text search engine may be used.
[0049] After selecting the search engine combination, the request may be processed using the search engine combination in step 109. For example, the engine selector may determine, based on pre-built heuristics, to use the search engine combination in parallel or sequentially to process the data. Based on the engine selector's rules, the search engine combination is used to obtain a list of candidates.
[0050] In step 111, at least some results of processing the request by the combination of search engines may be provided, such as by a results provider in a master data management system. For example, rows of resulting data records may be displayed in a graphical user interface to allow the user to access one or more data records of the results. For example, the user may perform a user action on the provided results. The user action may include, for example, a mouse click, or a touch gesture, or another action that allows the user to access the provided results.
[0051] The results provided may include all results obtained after processing the request through the combination of search engines, or only a predetermined subset of all those results. For example, search results from the combination of search engines may be aggregated and duplicates removed, resulting in a candidate list of data records. The resulting candidate list of data records may then be scored. For example, multiple scoring engines in the master data management system may be used. For example, depending on the attribute, scoring functionality may or may not be available. Multiple scoring engines are used because a PME-based scorer may not be able to score all types of entities (e.g., data contract types). Of all the results obtained, one set of results may go to one scorer and the other set may go to some other scoring engine. To improve efficiency, these scoring engine invocations may occur in parallel.
[0052] In step 113, the selection rules may be updated based on user actions taken on the provided results. The updated selection rules become the current selection rules and may therefore be used for further received requests for master data management system data. For example, upon receiving a subsequent request for master data management system data of step 103, steps 105-113 may be repeated, and during this iteration, the updated selection rules may be used in selection step 107.
[0053] For example, the selection rules may initially be based primarily on the capabilities / applicability of search engines for a given set of attributes, but the selection rules may continue to improve based on, for example, user clicks, feedback, and results (quality and performance) of previous searches. Alternative search engines may also be dynamically selected when previous selections of search engines do not yield results.
[0054] 2 is a flow diagram of a method for providing search results from a set of one or more search engines. The method of FIG. 2 may be applied, for example, to the data management system of FIG. 1 (e.g., FIG. 2 may provide details of step 111 of FIG. 1) or to other search systems.
[0055] For example, a set of search engines may process a search request for data, and the search results may include, for example, data records. In step 201, each data record in the results may be associated with or assigned a matching score. The matching score may be obtained by one or more scoring engines. For example, the matching scores of the data records in the results may be obtained by one or more scoring engines. In the case of two or more scoring engines, the matching score may be a combination (e.g., an average) of the matching scores obtained by the two or more scoring engines. In one example, of all the results obtained, one set of results may be processed by one scoring engine, and the other set may be processed by some other scoring engine. At least some of the one or more scoring engines used to score the results of a given search engine may or may not be part of the given search engine.
[0056] For example, each search engine in the set of search engines may include a scoring engine configured to score the results of the respective search engine. In another example, one or more common scoring engines may be used to score the results obtained by the set of search engines. For example, each search engine in the set of search engines may be configured to connect to a scoring engine and receive scores of data records from the scoring engine.
[0057] In step 203, the matching scores may be weighted. The weighting of the matching scores may be performed according to the performance of the components involved in generating the results. For example, a search process is performed to generate search results. The search process may include process steps performed by system elements, such as a search engine, to obtain the search results. Thus, the search process may have process steps, system elements, and components that are search results. Each of these components of the search process may have its own performance to perform its respective function. The performance of a component indicates how well the component performs its function or task. The performance of each component may be quantified by evaluating its respective performance parameters. The performance may affect the search results. In other words, each component of the search process has a contribution or influence on the quality of the obtained search results. At least some of the contributions of at least some of the components of the search process may be taken into account by defining and assigning weights to them. The weight assigned to a component may be indicative of (e.g., proportional to) the performance of the component, e.g., if the efficiency of a method step for identifying attributes is 80%, the weight may be 0.8. In one example, a weight may be assigned to each component of a search process. In another example, some of the components of a search process may be selected or identified (e.g., by a user), and those identified components may be associated with respective weights. In one example, the weights may be user-defined weights. As a result of the weighting step, each data record in the search results may be associated with the weight of the component of the search process that resulted in that data record. The matching score of that data record may be weighted by a combination of its associated weights, e.g., the combination may be a multiplication of the weights.
[0058] In step 205, results may be provided by using the weighted matching scores to remove duplicate data records from the results and retaining non-duplicate data records having a weighted matching score higher than the resulting predetermined score threshold. For example, the results may be displayed in a user interface, where the user may see a list of rows, each row associated with a data record from the provided results.
[0059] The provided results may be acted upon or used by a user. For example, a user may perform user actions on the provided results. Those user actions may be monitored, for example, by an activity monitor. For example, after presenting a user with a list of results in a user interface, the activity monitor may track the user's clicks on the presented results. A click on a result row may be considered to be what the user thinks he or she was looking for.
[0060] In step 207, user actions may be optionally processed and analyzed, for example. For example, the distribution of click counts for various characteristics of the data record (e.g., which engine the data record came from, how reliable the entity type detection was, how complete the record was, how fresh the record was, etc.) may be analyzed. This data is captured to find correlations, and weights are then calculated based on a lookup table or derived from an equation predicted by an ML-based regression model. Thus, every new click can change the distribution as it is fed back into the system, thereby aiding in weight reassignment. The calculated weights may be used to update the weights used to obtain search results in step 209; for example, the calculated weights may replace the corresponding weights used to obtain search results. The updated weights may then be used when processing further search requests to provide further search results.
[0061] FIG. 3 is a flow diagram of a method for providing search results from multiple search engines. The method of FIG. 3 may be applied, for example, to the data management system of FIG. 1, and FIG. 3 may provide details of step 111 of FIG. 1. For clarity, FIG. 3 is described with reference to the example of FIGS. 4A-F, which show two search engines, Engine 1 and Engine 2, and a set of five attributes. One search engine performs a probabilistic search, while the other performs a free-text search. Further, assume that the received request or input token is given as Name + Date of Birth (DOB), and that the entity identifier identifies the first token as Name with 90% confidence and sends it to Search Engine 1, and identifies the second token as Date of Birth with 60% confidence and sends it to Search Engine 2.
[0062] In this example, components of a search process performed, for example, by the method of Figure 1 may include a search engine, an identification step 105, and results. Example result data records R1-R6 are provided in tables 401 and 402 of Figure 4A. The results R1-R6 from the two search engines are aggregated and their matching scores are normalized to result in the matching scores in table 403.
[0063] In step 301, each of the search engines may be assigned an engine weight. Examples of engine weights are shown in Figure 4B. For example, search engines Engine 1 and Engine 2 may be assigned an initial weight of 0.5.
[0064] In step 303, each of a set of five attributes, name, date of birth, address, identifier, and email, is assigned an attribute weight indicating the confidence level at which the attribute is identified. The attribute weights shown in FIG. 4C may be an initial set of weights, which may be updated after the search request is executed. For example, as shown in FIG. 4C, the attribute weight for the name attribute and confidence levels of 0% to 10% is 0.1. In one example, the attribute weight may be obtained using the confidence level value, e.g., if the confidence level is less than 10%, the attribute weight may be equal to 0.1. However, other weight determination methods may be used.
[0065] In step 305, each resulting data record may be assigned a completeness weight indicating the completeness of the data record and a freshness weight indicating the freshness of the data record. The table in FIG. 4D shows example completeness weight values for a given data record. The completeness weights shown in FIG. 4D may be an initial set of weights, which may be updated after a search request is performed. For example, as shown in FIG. 4D, the completeness weight for a given data record may be provided as a function of the completeness of the data record. For example, the completeness weight for a completeness between 10% and 20% is 0.2. In one example, the completeness weight may be derived using the completeness value; for example, if the completeness is less than 10%, the completeness weight may be equal to 0.1. However, other example weighting methods may be used.
[0066] The table in FIG. 4E shows example freshness weight values for a given data record. The freshness weights shown in FIG. 4E may be an initial set of weights, which may be updated after a search request is performed. For example, as shown in FIG. 4E, the freshness weight for a given data record may be provided as a function of the freshness of the data record. For example, the freshness weight for a data record having a freshness of 3-5 years is 0.8. However, other example weighting methods may be used.
[0067] In step 307, the respective engine weight, attribute weight, completeness weight, and freshness weight for each data record in the results may be combined, and the data record's score may be weighted by the combined weight. The combined weight may be, for example, a multiplication of the four weights. The resulting weighted score, or final score, is shown in the table of FIG. 4F. The final score may be used to filter the results and provide them to the user. For example, because the final scores of data records R1, R2, and R6 are higher than threshold 1, only those data records may be provided to the user. The table of FIG. 4F indicates that records R1, R2, and R3 originate from engine 1, so their engine weight Wa is 0.5. Records R4, R5, and R6 originate from engine 2, so their engine weight Wa is 0.5. The attribute weight Wb (associated with the name attribute) is 0.9 for R1, R2, and R3 because this is an entity recognizer result set that identifies the name attribute with 90% confidence. The attribute weight Wb (related to the DOB attribute) is 0.6 for R4, R5, and R6 because these are the result sets of the entity recognizer that identify DOB with 60% confidence. The completeness weight Wc is based on the completeness of each record. For example, R1 is 80% complete, so its completeness weight is 0.8. The freshness weight Wd is based on the freshness of each record. For example, R1 is fresh, i.e., the last modified date is less than one year, so its freshness weight is 1. The final score may be obtained as follows: Final Score = Initial Normalized Score * (A * Wa) * (B * Wb) * (C * Wc) * (D * Wd), where A, B, C, and D are weights assumed to be 1 for simplicity.
[0068] FIG. 5 is a flow diagram of a method for updating weights used to weight matching scores of data records resulting from processing a search request by multiple search engines. For purposes of simplicity, FIG. 5 illustrates updating completeness weights. However, this weight updating method may be used for other weights. FIG. 5 may be described with reference to the example of FIG. 4.
[0069] When providing the results to the user, in step 501, the activity monitor may monitor user actions taken on the provided results. For example, the activity monitor may count the number of clicks made on each data record displayed to the user. This may result in the table of FIG. 6A, which shows the number of clicks made by the user on data records of different completeness. For example, the user made one mouse click on a row representing a data record with 80% completeness.
[0070] In step 503, the results of the monitoring operation shown in FIG. 6A may be processed or analyzed to find updated completeness weights. To that end, a lookup table, shown in FIG. 6B, may be generated. The lookup table contains an association between the completeness ranges used for weighting (see FIG. 4D) and the percentage of clicks users made on data records having the listed range of completeness. In this example, the data indicates that users rarely click on records with less than 30% completeness, while approximately 40% of clicks occur on records with more than 80% completeness. Based on the weights in the lookup table, a new record with 60% completeness would be given a proportional weight of 12%. For example, as can be seen from the tables in FIGS. 6A-B, the percentage of clicks on data records with 50%-60% completeness is 12%. This percentage of clicks may then be used to determine the updated weights. For example, the completeness weight for the 50%-60% completeness range would be 0.12 instead of the initial weight of 0.6 (in Figure 4D).
[0071] In another example, analysis of user behavior may be performed by modeling the variation in completeness as a function of click rate, as illustrated in Figure 6C. An example model 601 is shown in Figure 6C. This model 601 may be used to determine update weights for a given value of completeness. The model 601 is described by an equation that can be predicted by an ML-based regression model.
[0072] The result of this method may be updated weights that may be used to replace the initial weights, for example as provided in Figure 4. The updated weights may be used to weight the matching scores of data records that result from making a new search request.
[0073] FIG. 7 illustrates a block diagram of a computer system 700 according to an example of the present disclosure. The computer system 700 may be configured to perform, for example, master data management. The computer system 700 includes a master data management system 701 and one or more client systems 703. The client systems 703 may have access to data sources 705. The master data management system 701 may control access (e.g., read and write access) to a central repository 710. The master data management system 701 may use index data 711 to process fuzzy searches.
[0074] The master data management system 701 may process data records received from client systems 703 and store the data records in a central repository 710. For example, the client systems 703 may obtain data records from different data sources 705. The client systems 703 may communicate with the master data management system 701 via a network connection including, for example, a wireless local area network (WLAN) connection, a wide area network (WAN) connection, a local area network (LAN) connection, or a combination thereof.
[0075] The master data management system 701 may further be configured to process data requests or queries to access data stored in the central repository 710. The queries may be received, for example, from the client systems 703. The master data management system 701 includes an entity recognizer 721 for identifying attributes or entities in received data requests. The entity recognizer 721 may identify, for example, names and types of entities, mathematical expressions and time expressions coming from user input as unstructured text, and map them to attributes of data records stored in the central repository 710 with a particular probability or confidence, thereby enabling them to be used to perform structured searches of attributes. For example, the entity recognizer 721 may be a token recognizer that identifies strings / numbers or pattern names, locations, e.g., that emails should follow abc@uvw.xyz, phone numbers should follow 10-digit numbers, SSNs should follow the AAA-BB-CCCC structure, etc. The entity recognizer 721 may be configured to use machine learning models to classify or identify attributes of data records stored in the central repository 710 in the input data. The master data management system 701 further includes an engine selector 722 for selecting one or more suitable engines for processing a received search request. The engine selector 722 may determine whether to use one or more engines in parallel or sequentially to process the data based on pre-built heuristics. For example, the rules used to initially select an engine are primarily based on the engine's ability / applicability to a given set of attributes and entity types. After initial processing of the first request, the engine selector continues to refine its rules based on user clicks, feedback, and the results (quality and performance) of previous searches. The engine selector 722 may also dynamically select an alternative engine if previous selections of search engines do not produce results.Based on the rules of the engine selector 722, multiple search engines may be selected and used to obtain a list of good candidates. The search results from all engines are aggregated and duplicates are removed. The resulting list of candidates is then scored. Multiple scoring engines are used. Depending on the attribute, scoring functionality may or may not be available. In addition to the PME-based scorer, other scoring engines are used to score the search results. For example, of all the results obtained, one set of results may go to one scorer and the other set may go to some other scoring engine. To improve efficiency, these engine calls may be made in parallel.
[0076] The master data management system 701 further includes a weight provider and results aggregator 723 for weighting and aggregating the results obtained by the search engines. When scoring is performed by all scorers, the aggregation of the results may be based on a weighted average of the scores.
[0077] Weights are derived and refined over time by finding correlations between patterns and result set features and match quality. The analyzer may use machine learning to recognize these correlations. The analyzed result set features may include (but are not limited to): the match engine used to obtain the score (e.g., certain scoring engines may have a wider score range or be less reliable than others); the certainty with which the input data type was detected by the entity recognizer; the completeness of the records, indicating, for example, how many fields were populated and the freshness of the data (last updated date). Weights are a set of numbers used to modify the score of the result set. Match quality is indicated by an analysis of user clicks. Clicks on a displayed result indicate a user perceives a better match. Match quality may also be based on explicit feedback about match quality, which can be explored in the UI. The analysis of correlations is fed back to improve the weight provider 723. Results obtained by the search engine are aggregated using weights and then brought to the next stage based on comparison to a threshold record.
[0078] The master data management system 701 further includes different APIs to enable storage and access to data in the central repository 710. For example, the master data management system 701 includes a Create, Read, Update, and Delete (CRUD) API 724 to enable access to data, such as storage of new data records in the central repository 710. The master data management system 701 also includes APIs related to the search engine it contains. For purposes of illustration, Figure 7 shows two types of APIs: a structured search API 725 and a fuzzy search API 726.
[0079] The master data management system 701 further includes components that enable filtering of results to be provided to users. For example, the master data management system 701 includes a component 727 for applying visibility rules and another component 728 for applying consent management. The master data management system 701 includes a component 729 for applying standardization rules to data to be stored in the central repository 710. Because data security and privacy are paramount in a master data management solution, filtering may be advantageous. While full-text searches may attempt to cast a wide net to find matches, such overreach may remain within the system and ensure that information is not inadvertently disclosed to unwanted users. To this end, multiple filters may check whether the querying user has access to the returned fields and whether the resulting records have the relevant consent from the data owners required to be used for the processing purposes provided by the user. Filtering occurs later in the search process to allow for proper matching by all possible attributes. The result of filtering may be a list of records in decreasing order of match score, including records for which the required consent was provided, and only those rows may be allowed or visible to the user who initiated the search.
[0080] The master data management system 701 further includes indexing, matching, scoring, and linking services 730. Each client system 703 may include a stewardship search user interface (UI) 741 for submitting search queries to query data in the central repository 710. Each client system may further include services such as a messaging service 742 and a batch load service 743.
[0081] With reference to FIG. 8, the operation of computer system 700 will now be described in detail.
[0082] FIG. 8 shows a flow diagram for a method illustrating an example of the operation of the master data management system 701. In block 801, a free text search may be entered into a browser, which may be, for example, an example of the Stewardship Search UI 741. The entity recognizer 721 may receive the free text search request (block 802) and process the received request to identify attributes or entities, as described herein, for example, in FIG. 1 . The engine selector 722 may then be used (block 803) to select a suitable search engine for the identified attributes. As illustrated in FIG. 8, two search engines are selected and used to execute the received search request (blocks 804 and 805). Results of the execution of the search request may be scored (block 806) using a matching and scoring service of the master data management system 701. Scoring may also use an add-on scoring mechanism (block 807). The results are then aggregated, and the scores are normalized (block 808). Before providing the results to the user, several filters may be applied (block 809). These filters may include, for example, at least one of a visibility rule filter, a consent-based data filter, and a custom filter. The filtered results are then displayed in a browser (e.g., the browser that received the free text search) (block 810). The displayed results may be monitored (block 811) and analyzed by a user click and quality feedback analyzer. For example, the analyzer may use a machine learning model to determine weights based on user activity on the results. The weights may be used to update the engine selector 722 and the weight provider 723, as indicated by arrows 812 and 813. The weights provided by the weight provider 723 may then be used for the scoring block 808 in the next iteration of the method.
[0083] FIG. 9 shows a diagram illustrating an example of request processing according to the present subject matter. A first column 901 in FIG. 9 shows example content of a received request or input token. For example, a received request may include “Robert,” “Bangalore,” and the number “123-45-6789.” A second column 902 shows the results of entity recognition when processing the received request. For example, “Robert” is identified as a name attribute, “Bangalore” is identified as an address attribute, and the number “123-45-6789” is identified as an SSN attribute. Columns 902 and 904 indicate that the engine selector selected search engine “Search Engine 1” to process the request for “Robert.” Columns 902 and 904 further indicate that the engine selector selected search engine “Search Engine 2” to process the request for “Bangalore.” Columns 902 and 904 further indicate that the engine selector selected both search engines, "Search Engine 1" and "Search Engine 2," to process the request for "123-45-6789." As shown in column 905, the results of processing the request are processed, for example, aggregated, before being served. For example, column 905 indicates that search engine "Search Engine 1" found records R1, R2, and R3 when searching for "Robert." Column 905 further indicates that search engine "Search Engine 2" found records R4 and R5 when searching for "Bangalore." Column 905 further indicates that search engine "Search Engine 1" found record R6 when searching for "123-45-6789," and that search engine "Search Engine 2" found record R7 when searching for "123-45-6789." Results R1-R7 may need to be filtered using a data governance filter before being provided to the user, as shown in column 906. After being filtered, the results may then be output to the user, as shown in column 907.As shown in column 907, the date of birth value has been filtered out from records R1-R7 because the user submitting the results does not have access to it.
[0084] It is understood that one or more of the above-described embodiments of the present invention may be combined, provided that the combined embodiments are not mutually exclusive.
[0085] In the following embodiments, various embodiments are identified.
[0086] 1. A method for accessing a data record in a master data management system, the data record including a plurality of attributes, the method comprising: augmenting the master data management system with one or more search engines to enable access to the data records; receiving a request for data in a master data management system; identifying a set of one or more attributes from a plurality of attributes referenced in the received request; selecting a combination of one or more search engines from the master data management system whose performance for searching at least some values of the set of attributes satisfies the current selection rules; processing the request using a combination of search engines; Providing at least a portion of the results of the processing. A method comprising:
[0087] 2. The method of item 1, further comprising: updating the selection rules based on user actions on the provided results, the updated selection rules becoming the current selection rules; and repeating the identifying, selecting, processing, and providing steps using the current selection rules when another request for data is received.
[0088] 3. The method of item 1, wherein the results include data records in a master data management system associated with respective matching scores obtained by the scoring engine of the search engine, the method further comprising a step of weighting the matching scores according to the performance of components involved in providing the results, the components including method steps, elements used to provide the results, and at least a portion of the results, and the provided results include non-duplicate data records having weighted matching scores higher than a predetermined score threshold.
[0089] 4. The component includes a search engine, an identification step, and a result, and the method further includes: assigning an engine weight to each of the search engines; assigning attribute weights to a set of attributes, the attribute weight of an attribute indicating a level of confidence with which said attribute is identified; assigning to each resulting data record a completeness weight indicating the completeness of the data record and a freshness weight indicating the freshness of the data record; combining the respective engine weights, attribute weights, completeness weights, and freshness weights for each resulting data record; and weighting the score of the data record by the combined weights. Item 3. The method according to item 3, comprising:
[0090] 5. Providing user parameters quantifying user behavior; for each of at least some of the components, determining values of user parameters and associated values of component parameters describing the component; and using the determined associations to update weights assigned to the component. Item 5. The method according to item 4, further comprising:
[0091] 6. The method of item 5, further comprising the steps of providing a lookup table relating values of user parameters to values of component parameters, and using the lookup table to update the weights assigned to the components.
[0092] 7. The method of item 5, further comprising the steps of modeling the variation of the values of the user parameters with the values of the component parameters using a predetermined model, using the model to determine update weights for the components, and using the update weights to update the weights assigned to the components.
[0093] 8. The method of item 5, wherein one of the user actions includes an indication of result selection, the indication including a mouse click on a displayed result from the provided results, and the user parameters include at least one of the number of clicks, the frequency of clicks, and the duration for accessing a given result from the results.
[0094] 9. The method according to item 1, wherein the results include data records in the master data management system associated with respective matching scores obtained by the search engine's scoring engine, and the provided results include non-duplicate data records having matching scores higher than a predetermined score threshold.
[0095] 10. For each attribute in the set of attributes, the selection rule is determining, for each of the search engines, a value of a performance parameter indicative of the search engine's performance for searching the value of the attribute; selecting a search engine having a performance parameter value higher than a predetermined performance threshold; The method according to item 1, comprising:
[0096] 11. The method of claim 10, wherein the performance parameters include at least one of the number of outcomes and the level of matching of the outcomes to expectations.
[0097] 12. Selection rules use a table that associates attributes with corresponding search engines, and selection rule updates are determining values for user parameters that quantify user behavior with respect to the results provided by each search engine of the search engine combination; using a predetermined value associated with each search engine of the combination of search engines to identify values of the user parameter that are less than a predetermined threshold; for each identified value of the user parameter, determining an attribute from the set of attributes and a search engine associated with the identified value; and updating the table using the determined attribute and search engine. Item 11. The method according to item 10, comprising:
[0098] 13. The method according to item 1, wherein the processing of requests is performed in parallel by a combination of search engines.
[0099] 14. The method of item 1, wherein the search engine combination is a ranked list of search engines, and the processing of requests is performed sequentially according to the ranked list until a minimum number of results is exceeded.
[0100] 15. The method of item 1, wherein identifying the set of attributes includes inputting the received request into a predetermined machine learning model, and receiving a classification of the request from the machine learning model, the classification indicating the set of attributes.
[0101] 16. The method described in item 1, further comprising the steps of inputting the set of attributes into a predetermined machine learning model and receiving one or more search engines from the machine learning model that can be used to search the set of attributes.
[0102] 17. The method of claim 16, further comprising receiving a training set indicating one or more different sets of training attributes, each set of training attributes being labeled to indicate a search engine suitable for conducting a search on the set of training attributes, and generating a machine learning model by training a predetermined machine learning algorithm using the training set.
[0103] 18. The method of item 1, wherein the provided results include data records that are filtered depending on the sender of the request.
[0104] 19. A method for providing search results of a search engine according to a predetermined search process, the method comprising: receiving results of the search request obtained by the search engine, each result being associated with a matching score; For each result, determining a set of one or more components of the search process that participated in providing the result, and assigning a predetermined weight to each component of the set of components; weighting the matching scores using weights; providing results having a weighted matching score higher than a predetermined score threshold. A method comprising:
[0105] 20. Analyzing user behavior relative to the provided results by evaluating user parameters that quantify the user behavior; determining, for each component of at least some of the set of components, one or more values of component parameters describing that component and associated values of the user parameters; determining update weights using the determined associations; replacing weights assigned to at least some of the components with the determined weights; using the updated weights to repeat the method for further received search results; 20. The method of claim 19, further comprising:
[0106] 21. The method of item 20, further comprising the steps of providing a table relating values of user parameters to values of component parameters, and using the table to update weights assigned to components.
[0107] 22. The method of item 20, further comprising the steps of modeling the associations between the values using a predetermined model, using the model to determine update weights for the components, and using the update weights to update the weights assigned to the components.
[0108] 23. The method of item 20, wherein one of the user actions includes a mouse click on a displayed result of the provided results, and the user parameters include at least one of the number of clicks, the frequency of clicks, and the duration for accessing a given result of the results.
[0109] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0110] The present invention may be a system, a method, or a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0111] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: Portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge structures in grooves with recorded instructions, and any suitable combination of the foregoing. Computer-readable storage media, as used herein, should not be construed to refer to transitory signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted through wires.
[0112] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium into each computing / processing device, or may be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to a computer-readable storage medium within the respective computing / processing device for storage.
[0113] The computer readable program instructions for carrying out the operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, or C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user computer system computer, partially on the user computer system computer as a stand-alone software package, partially on the user computer system computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user computer system's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection to the external computer may be made (e.g., over the Internet using an Internet service provider). In some embodiments, electronic circuitry, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions by using state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.
[0114] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0115] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, cause the implementation of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored includes instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0116] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps to create a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0117] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. In addition, it will be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be realized by a special-purpose hardware-based system that performs the specified functions or operations or that implements or executes a combination of special-purpose hardware and computer instructions.
Claims
1. 1. A method executed by a master data management system for accessing a data record in the master data management system, the data record including a plurality of attributes, the method comprising: augmenting said master data management system with a plurality of search engines for enabling access to said data records; receiving a request for data at the master data management system; identifying a set of one or more attributes of the plurality of attributes referenced in the received request, wherein the set of one or more attributes and one or more entity types associated with the one or more attributes are identified by an entity recognizer using a machine learning model; selecting a combination of one or more search engines from the plurality of search engines of the master data management system whose performance for searching at least some values of the set of attributes satisfies current selection rules, the combination of one or more search engines being selected based on the set of one or more attributes and the one or more entity types; configuring the master data management system to configure the combination of one or more search engines; processing the request using a combination of the one or more search engines, wherein the master data management system accesses the plurality of search engines using a plurality of interfaces, including a structured search application programming interface (API) and a fuzzy search application programming interface (API); and providing at least a portion of the results of the processing based on displaying the results of the processing in a browser; A method comprising:
2. inputting the set of attributes into a second predetermined machine learning model; receiving one or more search engines that can be used to search the set of attributes from the second machine learning model; The method of claim 1 further comprising:
3. receiving a training set indicating one or more different sets of training attributes, each set of training attributes labeled to indicate a preferred search engine for conducting the search on the set of training attributes; generating the second machine learning model by training a predetermined machine learning algorithm using the training set; The method of claim 2 further comprising:
4. 1. A method executed by a master data management system for accessing a data record in the master data management system, the data record including a plurality of attributes, the method comprising: augmenting said master data management system with one or more search engines for enabling access to said data records; receiving a request for data at the master data management system; identifying a set of one or more attributes of the plurality of attributes referenced in the received request; selecting a combination of one or more of the search engines of the master data management system whose performance for searching at least some values of the set of attributes satisfies current selection rules; processing the request using the combination of search engines; and providing at least a portion of the results of said processing. Including, the results include data records in the master data management system associated with respective matching scores obtained by a scoring engine of the search engine, the method further comprising weighting the matching scores according to the performance of components involved in providing the results, the components including method steps, elements used to provide the results, and at least a portion of the results, the provided results including non-duplicate data records having weighted matching scores higher than a predetermined score threshold; The components involved in providing the results include the search engine, the identifying step, and the results, and the method further comprises: assigning an engine weight to each of said search engines; assigning attribute weights to the set of attributes, the attribute weight of an attribute indicating a level of confidence with which the attribute is identified; assigning to each of the resulting data records an integrity weight indicating the integrity of the data record and a freshness weight indicating the freshness of the data record; for each of the resulting data records, combining an engine weight, an attribute weight, a completeness weight, and a freshness weight, and weighting the matching score of the data record by the combined weight. A method comprising:
5. providing user parameters quantifying user actions; for each component of at least a portion of the components, determining values of the user parameters and associated values of component parameters describing the component; and using the determined associations to update the weights assigned to the components. The method of claim 4 further comprising:
6. 6. The method of claim 5, further comprising providing a lookup table relating values of the user parameters to the values of the component parameters, and using the lookup table to update the weights assigned to the components.
7. 6. The method of claim 5, further comprising the steps of: modeling variations in values of the user parameters with the values of the component parameters using a predetermined model; using the model to determine update weights for the components; and using the update weights to update the weights assigned to the components.
8. 8. The method of claim 1, further comprising: updating the selection rules based on user actions on the provided results, the updated selection rules becoming the current selection rules; and repeating the identifying, selecting, processing, and providing steps using the current selection rules when another request for data is received.
9. 9. The method of claim 1, wherein a user action among the user actions on the provided results includes an indication of selection of a result, the indication including a mouse click on a displayed result among the provided results, and user parameters quantifying the user action include at least one of the number of the mouse clicks, the frequency of the mouse clicks, and the duration of accessing a given result among the results.
10. For each attribute in the set of attributes, the selection rule is: determining, for each of said search engines, a value of a performance parameter indicative of said performance of said search engine for searching values of said attribute; selecting the search engines having a performance parameter value higher than a predetermined performance threshold; The method according to any one of claims 1 to 9, comprising:
11. The method of claim 10 , wherein the performance parameters include at least one of the number of results and a level of matching of the results to an expectation.
12. updating the selection rules based on user actions on the provided results; The selection rules use a table that associates attributes with corresponding search engines, and the step of updating the selection rules includes: determining the value of a user parameter that quantifies the user's behavior on the provided results of each search engine of the combination of search engines; using the determined values associated with each search engine of the combination of search engines to identify the values of the user parameters that are less than a predetermined threshold; for each identified value of the user parameters, determining the attribute of the set of attributes and the search engine associated with the identified value; and updating the table with the determined attributes and search engines.
12. The method of claim 10 or 11, comprising:
13. The method of any one of claims 1 to 12, wherein the processing of the requests is performed in parallel by the combination of the search engines.
14. 14. The method of claim 1, wherein the combination of search engines is a ranked list of search engines, and the processing of the request is performed sequentially according to the ranked list until a minimum number of results is exceeded.
15. The method of claim 1, claim 2 or claim 3, or any one of claims 8 to 14 (only when citing claim 1), wherein identifying the set of attributes comprises inputting the received request into a predetermined machine learning model and receiving a classification of the request from the machine learning model, the classification being indicative of the set of attributes.
16. The method of any one of claims 1 to 15, wherein the provided results comprise data records that are filtered depending on the sender of the request.
17. A computer program product causing a processor to carry out the steps of the method according to any one of claims 1 to 16.
18. A computer system for enabling access to data records in a master data management system, the data records including a plurality of attributes, the computer system comprising: a user interface configured to receive a request for data; a plurality of search engines for enabling access to said data records, said plurality of search engines configured to process said requests; an entity identification means configured to identify a set of one or more attributes of the plurality of attributes referenced in the received request, the entity identification means identifying the set of one or more attributes and one or more entity types associated with the one or more attributes using a machine learning model; an engine selector configured to select a combination of one or more search engines from the plurality of search engines whose performance for searching at least some values of the set of attributes satisfies a current selection rule, the plurality of search engines being accessed using a plurality of interfaces including a structured search application programming interface (API) and a fuzzy search application programming interface (API), and the engine selector selecting the combination of one or more search engines based on the set of one or more attributes and the one or more entity types; means for configuring the master data management system to search using a combination of the one or more search engines; and a results provider configured to provide at least a portion of the results of the process, the results provider providing the results of the process based on displaying the results of the process in a browser; 1. A computer system comprising:
19. 20. The computer system of claim 18, wherein the engine selector inputs a set of attributes into a predetermined second machine learning model and receives from the second machine learning model one or more search engines that can be used to search the set of attributes.
20. 20. The computer system of claim 19, wherein the computer system receives a training set indicating different sets of one or more training attributes, each set of training attributes labeled to indicate a search engine suitable for conducting the search on the set of training attributes, and wherein the second machine learning model is generated by training a predetermined machine learning algorithm using the training set.
Citation Information
Patent Citations
Multimedia information retrieving device, multimedia information retrieving method and computer readable recording medium in which program to make computer execute its method is recorded
JP2002007418A
Local item extraction
JP2008527502A
Apparatus and method for processing information, program, and storage medium
JP2010267075A
Method and system for search engine selection and optimization
JP2019501466A
Method and system for validating information
US20160063052A1