Method for accessing data records of a master data management system
By introducing multiple search engines and corresponding components into the main data management system, and automatically selecting and combining appropriate search engines to process data requests, the existing system's shortcomings in data access efficiency and performance are solved, and an efficient and user-friendly data access experience is achieved.
Patent Information
- Application Number
- CN202080026536.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-02
- Filing Date
- 2020-03-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-03-19
AI Technical Summary
The existing master data management system has shortcomings in data access efficiency and performance, especially when processing multi-attribute data records, it is difficult to effectively utilize the capabilities of multiple search engines, resulting in users needing to retry or re-develop search queries.
By enhancing the master data management system, multiple search engines are introduced, and through components such as entity identifiers, engine selectors and result providers, appropriate search engines are automatically selected and combined to process data requests and provide processing results.
Improve efficient access to data stored in the main data management system, improve system performance, reduce duplicate or retry search requests, and improve user experience.
Smart Images

Figure CN113661488B_ABST
Abstract
Description
Background Art
[0001] The present invention relates to the field of digital computer systems, and more particularly to a method for accessing data records of a master data management system.
[0002] Enterprise data matching involves matching and linking customer data received from different sources and creating a single version of the real data. Solutions based on master data management (MDM) work with enterprise data and perform indexing, matching and linking of data. Master data management systems can allow access to these data. However, there is a continuous need to improve access to data in master data management systems. Summary of the invention
[0003] Various embodiments provide a method, a computer system and a computer program product for accessing data records of a master data management system as described by the subject matter of the independent claims. Advantageous embodiments are described in the dependent claims. If the embodiments of the invention are not mutually exclusive, they can be freely combined with each other.
[0004] In one aspect, the present invention relates to a method for accessing a data record of a master data management system, the data record comprising a plurality of attributes. The method comprises:
[0005] enhancing the master data management system with one or more search engines for enabling access to the data records;
[0006] receiving a request for data at a master data management system;
[0007] identifying a set of one or more attributes of the plurality of attributes referenced in the received request;
[0008] selecting a combination of one or more search engines among the search engines of the master data management system, the performance of the one or more search engines for searching for values of at least a portion of the attribute set satisfying a current selection rule;
[0009] Use a combination of search engines to handle the request;
[0010] At least a portion of the results of the processing are provided.
[0011] On the other hand, the present invention relates to a computer system for enabling access to a data record, the data record comprising multiple attributes, the computer system comprising multiple search engines for enabling access to the data record; a user interface configured to receive a data request; an entity identifier configured to identify a set of one or more attributes among the multiple attributes that are referenced in the received request; an engine selector configured to select a combination of one or more search engines among the search engines, the performance of the one or more search engines for searching for values of at least a portion of the attribute set satisfies a current selection rule; wherein the search engine is configured to process the request; and a result provider configured to provide at least a portion of the processed results.
[0012] On the other hand, the present invention relates to a computer program product having a computer readable program code implemented therewith, the computer readable program code being configured to access data records of a master data management system, the data management system including a search engine for enabling access to the data records, the data records including a plurality of attributes, the computer readable program code being further configured to: receive a request for data at the master data management system; identify a set of one or more attributes among the plurality of attributes that are referenced in the received request; select a combination of one or more search engines among the search engines of the master data management system, the performance of the one or more search engines for searching for values of at least a portion of the attribute set satisfying a current selection rule; process the request using the combination of search engines; and provide at least a portion of the processed results. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Embodiments of the present invention are explained in more detail below, by way of example only, with reference to the accompanying drawings, in which:
[0014] Figure 1 is a flow chart of a method for accessing data records of a master data management system,
[0015] Figure 2 is a flow chart of a method for providing search results of a set of search engines,
[0016] Figure 3 is a flow chart of a method for providing search results of multiple search engines,
[0017] Figure 4A depicts a table including search results from different engines, the search results being normalized and merged,
[0018] Figure 4B A table including examples of engine weights is depicted,
[0019] Figure 4Cdepicts a table including examples of attribute weights for identifying attribute types based on confidence of entity recognition,
[0020] Figure 4D A table depicting examples including completion weights,
[0021] Figure 4E A table depicting an example including freshness weighting,
[0022] Figure 4F depicts a table containing result records and associated weights and scores,
[0023] Figure 5 is a flow chart of a method for updating weights for weighting matching scores of data records of results of processing a search request by a plurality of search engines,
[0024] Fig. 6A depicts a table including the number of user clicks as a function of data record completion,
[0025] Figure 6B depicts a table including user click scores as a function of data record completion,
[0026] Figure 6C is a graph of the distribution of click scores as a function of data record completion,
[0027] Figure 7 A block diagram representation of a computer system 700 according to an example of the present disclosure is shown.
[0028] Figure 8 depicts a flowchart describing an example method of operation of a master data management system,
[0029] Fig. 9 A schematic diagram depicting an example of processing a request in accordance with the present subject matter. DETAILED DESCRIPTION
[0030] The description of various embodiments of the present invention will be presented for the purpose of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements existing in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0031] The present subject matter can enable efficient access to data stored in a master data management system. The present subject matter can improve the performance of a master data management system. The present subject matter can reduce the number of repeated or retried search requests because it can use multiple search engines to provide the best possible results, and thus users do not have to retry or reformulate search queries as is the case with other systems.
[0032] A master data management system may use a single type of search engine. Using the present subject matter, a master data management system may use different types of search engines. The type of search engine may be defined by the technology it uses to perform searches such as full-text searches or structured probabilistic searches. For example, the additional search engines added by the present method may be of a different type than the type of search engine originally included in the master data management system. Therefore, the present subject matter may provide an integrated search and matching engine that aims to utilize the best of all the different capabilities of multiple search and index engines based on the type of input data or the type of query being performed. Different indexes or search engines do have different capabilities, so they work at best on different types of inputs or different requirements. The present subject matter may enhance the user experience without affecting the performance of machine-based interactions by implementing a better way to search for data using multiple different indexes and search engines.
[0033] For example, the identification, selection, processing and providing steps can be automatically performed when a data request is received. In one example, the identification, selection, processing and providing steps can be automatically repeated when further data requests are received, wherein in each repetition, the updated selection rules generated by the immediately previous execution of the method are used.
[0034] The results may include data records. Providing the data records may include displaying data indicating the data records on a graphical user interface. For example, for each data record, a row may be displayed, where the row may be a hyperlink or link that a user can click to access detailed information of the data record.
[0035] A data record is a collection of related data items such as a specific user's name, date of birth (DOB), and category. A record represents an entity, where an entity refers to a user, object, or concept, about which information is stored in a record.
[0036] According to one embodiment, the method further comprises updating the selection rule based on the user operation on the provided result, the updated selection rule becomes the current selection rule, and when another data request is received, the identification, selection, processing and providing steps are repeated using the current selection rule. In one example, the updating of the selection rule can be performed after a predefined time period, for example, during the time period, the method may have been performed multiple times, and the updating is performed based on the combination of user operations on the provided results during the time period. This can realize a self-improving search system based on user input and experience. The search engine is a search engine that is part of a predefined table of a data management system associated with at least a part of the attribute group, and the performance of the search engine searching for the value of at least a part of the attribute group satisfies the current selection rule. For example, the table includes multiple entries. Each entry i of the table includes a search engine SEi and one or more related attributes Ti appropriately searched by the search engine. In one example, each association of Ti and SEI can be assigned an update score that can be changed or updated. The selected search engine is a search engine SEI of a table associated with one or more attributes in the attribute group, for example, if the attribute group includes T1 and T2, the table can be searched to identify entries with T1 and T2, and the selected search engine is the search engine of those identified entries. The updating of the selection rules may include updating a table, for example, if the number of clicks on a displayed result from a search engine SEx and associated with a given attribute Tx of the search is less than a threshold, the table may be updated accordingly, for example, by deleting the association between Tx and SEx, or if Tx and SEx are associated with an update score, by changing the update score, for example, by lowering the update score. For example, if the same combination Tx and SEx was previously found at least once and the performance was not good, for example, the number of clicks on the associated result was less than a threshold multiple times, and therefore the associated update score was below a given threshold, then the deletion may be performed. In one example, the table initially has many or all possibilities for combinations between attributes and search engines, and within a predefined period, non-performing entries may be removed.
[0037] According to one embodiment, the results include data records of the master data management system associated with corresponding match scores obtained by a scoring engine of the search engine, wherein the provided results include non-duplicate data records having match scores above a predefined score threshold. The match score may indicate a level or degree of match between the data record and the requested data.
[0038] By providing only results that meet the selection criteria of the matching score, this embodiment can further improve the performance of the master data management system. For example, irrelevant results may not be provided to the user. This can save processing resources, such as display resources and data transmission resources that would be used for irrelevant results. For example, the weighting of the scores can be performed as described in the following embodiments.
[0039] According to one embodiment, the result includes data records of a master data management system associated with corresponding match scores obtained by a scoring engine of a search engine, the method further comprising weighting the match scores according to performance of components involved in producing the result, the components comprising method steps, elements for producing the result, and at least a portion of the result, wherein the provided result includes non-repeating data records having weighted match scores above a predefined score threshold. The weighting may, for example, include: for each data record of the result, assigning a weight to each of the components that provided or produced the data record, wherein the components may include the provided data record itself, combining the weights and weighting the match score of the data record using the combined weights.
[0040] For example, the generation of search results of the received data request involves the execution of a search process (the method may include a search process). The search process has a plurality of process steps, each of which may be performed by a system element such as a search engine or a scoring engine. The search process may have components as process steps and / or system elements and / or the results provided by them. Each component may have a function that it performs to contribute to the acquisition of search results. Those components of the search process may each have an impact on the quality of the results obtained. For example, if the components of the search process do not operate correctly, this may affect the search results. For example, if the component is a processing step that identifies the attributes in the received request, and the component may not be effective in identifying the attributes of a particular type, it may occur that the processing step does not correctly identify the attributes of that type. Therefore, when a request for data with attributes of this type referenced therein is received, the results obtained may be affected because they may include irrelevant, unwanted search results of the attributes of the wrongly identified. The performance of the components of the search process may have different contributions to the results obtained by the search process. This embodiment may take into account at least a portion of these contributions by weighting the matching score accordingly. For example, each component in at least a portion of the components of the search process of this embodiment can be assigned a weight indicating the performance of its corresponding function. Weight, for example, can be user-defined, such as weight can be initially defined by the user (for example, for the first execution of this method), and can be automatically updated later using a weight update method as described herein. These weights can be used to weight the matching score. This embodiment can further improve the performance of the data management system. For example, further irrelevant results may not be provided to the user. This can save processing resources, such as display resources and data transmission resources.
[0041] An example of components that are considered in the weighting of the search process may be described in the following embodiment. This embodiment may be advantageous because it identifies and weights components whose performance may have a greater impact on the search results.
[0042] According to one embodiment, the component includes a search engine, an identification step, and a result. The method also includes: assigning an engine weight to each of the search engines; assigning an attribute weight to the attribute set, wherein the attribute weight of an attribute indicates a confidence level that the attribute is identified; assigning a completeness weight indicating the data record and a freshness weight indicating the data record to each data record of the result; for each data record of the result, combining the corresponding engine weight, attribute weight, completeness weight, and freshness weight, and weighting the score of the data record by the combined weight. Attribute weights can be generated at the attribute level and applied to the complete result set (and all attributes) returned for the received request. This can make the result set less useful if the automatically determined search entity type itself is incorrect.
[0043] The following embodiments provide a weight updating method for updating weights used according to the present subject matter. They enable efficient and systematic processing of the weighting process.
[0044] According to one embodiment, the method further comprises: providing a user parameter that quantifies a user action on the provided result; for each component in at least a portion of the components, determining a value of the user parameter and an associated value of a component parameter describing the component; and updating a weight assigned to the component using the determined association. For example, the component parameter may include at least one of a degree of completion, a freshness of a data record, an ID of a search engine, and a confidence level that may identify an attribute.
[0045] For example, user actions or interactions can be monitored by an activity monitor of the master data management system. In one example, user actions can be the results provided by user clicks. The associated values of user parameters and component parameters can be provided in the form of distributions, which can be fitted or modeled to derive weights. For example, the distribution of click counts relative to various characteristics of rows representing data records (e.g., characteristics can, for example, indicate which search engine the data record comes from, what the confidence of entity type detection is, how complete the record is, how fresh the record is, etc.) can be provided and analyzed to find weights. For example, this embodiment can be performed for each new click, for example, when each new click is fed back to the system, the distribution can be changed, and thus help to redistribute weights. This embodiment can make it possible to update the weights used in the previous iteration of this method. This embodiment can enable the data management system to maintain self-improvement based on its own experience of data search. For example, all weights used in the above-mentioned embodiments can be updated. In another example, only a part of the weights used (e.g., completion weights) can be updated. Updating weights can include determining new weights and replacing the used weights with corresponding new weights. According to this embodiment, new weights can be determined by monitoring user activities related to the results provided to the user.
[0046] According to one embodiment, the method further comprises providing a lookup table associating values of the user parameters with values of the component parameters, and using the lookup table to update the weights assigned to the components.
[0047] According to one embodiment, the method further includes modeling the change of the user parameter value as a function of the value of the component parameter using a predefined model, and using the model to determine the updated weight of the component and using the updated weight to update the weight assigned to the component. For example, the predefined model can be configured to receive the component parameter value as input and output the corresponding weight. This can implement accurate weighting techniques according to the present subject matter.
[0048] According to one embodiment, the user operation in the user operation includes a mouse click on a displayed result in the provided results, wherein the user parameter includes at least one of the number of clicks, the frequency of clicks, and the duration of accessing a given result in the results. For example, the activity monitor can use the click count and / or can check the time spent on each result (e.g., after it is clicked until the back / restart button is used) and / or it can check the back and forth operation on the result set, and the last selected record in which the user spends more than a certain threshold can be considered as a "result liked by the user."
[0049] According to one embodiment, for each attribute in the attribute set, the selection rule includes: for each search engine in the search engines, determining a value indicating a performance parameter of the search engine for searching for the value of the attribute; weighting the determined value with a corresponding current weight; and selecting a search engine whose performance parameter value is higher than a predetermined performance threshold.
[0050] For example, in the first or initial execution of the method of this embodiment, the current weight may be set to 1. In another example, if the attribute set includes three attributes ATT1, ATT2, and ATT3, the performance of each search engine (e.g., search engine 1 (SE1)) may be evaluated. For each search engine, this may result in three performance parameter values Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1. The current weight of search engine SE1 may be determined from Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1 to obtain weights W1_SE1, W2_SE1, and W2_SE1. These weights may be used to weight the performance parameter values Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1. In order to decide whether to select search engine SE1, a combination of weighted Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1 may be determined, and if the combined value (e.g., average value) is above a performance threshold, SE1 may be selected. In another example, each weighted performance value Perf_att1_SE1, Perf_att2_SE1, and Perf_att3_SE1 is compared with a performance threshold, and SE1 may be selected only if each of them is above the performance threshold.
[0051] According to one embodiment, the performance parameter includes at least one of: the number of results and the degree to which the results match the desired or requested content.
[0052] According to one embodiment, the selection rule uses a table that associates attributes to corresponding search engines, and the updating of the selection rule includes: determining the value of a user parameter that quantifies the user operation of the provided results of each engine in the combination of the search engines; and using the determined value associated with each search engine in the combination of search engines to identify the value of the user parameter that is less than a predefined threshold, and for each identified value of the user parameter, determining the attribute in the attribute set and the search engine associated with the identified value, and using the determined attribute and search engine to update the table. In one example, the table initially has many or all possibilities of combinations between attributes and search engines. For example, after a predetermined period of time, non-executing entries can be removed. For example, the user parameter can be the number of clicks on each result in the provided results, that is, for each displayed result, there is a value of the user parameter. These values can be compared with a predetermined threshold (e.g., 10 clicks), and the displayed results associated with the value less than the threshold can be identified. Each of these identified results is obtained by a given search engine X as a result of searching for one or more attributes, such as attribute T1 of the attribute group. Therefore, X and T1 can be used to update the table as described herein.
[0053] According to one embodiment, the processing of requests is performed in parallel by a combination of search engines. This can speed up the search process of the present subject matter.
[0054] According to one embodiment, the combination of search engines is a ranked list of search engines, where processing of requests is performed serially following the ranked list until a minimum number of results is exceeded. This can save processing resources. If the engine selection rule only suggests engine 1 (SE1), but the actual search does not produce enough results, SE2 (next in the ranked list) can be used.
[0055] According to one embodiment, the results provided include data records that are filtered based on the sender of the request. For example, data control rules are applied after obtaining a list of matches for a given data input and providing role-based visibility and applying consent-related filters; thereby respecting privacy while providing better matching quality and search flexibility.
[0056] According to one embodiment, identifying the set of attributes includes inputting the received request into a predefined machine learning model; receiving a classification of the request from the machine learning model, the classification indicating the set of attributes.
[0057] According to one embodiment, selecting a rule includes: inputting the set of attributes into a predefined machine learning model, and receiving from the machine learning model one or more search engines that can be used to search the set of attributes.
[0058] According to one embodiment, the method further includes: receiving a training set indicating different sets of one or more attributes, wherein each attribute set is labeled to indicate a search engine suitable for executing the attribute set; and using the training set to train a predefined machine learning algorithm to generate the machine learning model.
[0059] Figure 1 The invention is a flow chart of a method for accessing a data record of a master data management system. The data record includes a plurality of attributes.
[0060] For example, the master data management system can process records received from the client system and store the data records in a central repository. The client system can communicate with the master data management system, for example, via a network connection, including, for example, a wireless local area network (WLAN) connection, a WAN (wide area network) connection, a LAN (local area network) connection, or a combination thereof.
[0061] The data records stored in the central repository may have a predefined data structure, such as a data table with multiple columns and rows. The predefined data structure may include multiple attributes (e.g., each attribute represents a column of the data table). In another example, the data records may be stored in a graph database as entities with relationships. The predefined data structure may include a graph structure, in which each record may be assigned to a node of the graph. Examples of attributes may be names, addresses, etc.
[0062] The master data management system may include a search engine (referred to as an initial search engine) that performs a search for data records stored in a central repository based on a received search query using a single technique such as a probabilistic structural search. The initial search engine, like any other search engine, may work well for a particular type of attribute, but not for other attributes. That is, the performance of the initial search engine may depend on the type of attribute value being searched. For example, due to nicknames and phonetics, the attribute "name" may be well searched by a probabilistic search engine, while an attribute address such as a city may work well with a free text search engine because it is partial. To this end, in step 101, the master data management system may be enhanced with one or more search engines that are capable of accessing the data records of the central repository. This may result in multiple search engines, including the initial search engine and the added search engine. For example, each search engine of the master data management system may be associated with a respective API through which a search query may be received. This may enable the aggregate search and matching engine to utilize the best of all the different capabilities of multiple search and index engines based on the type of input data or the type of query being performed. Different indexes or search engines do have different capabilities, so they work on different types of input or different requirements at best.
[0063] The master data management system can receive a data request in step 103. For example, the request can be received in the form of a search query. For example, the search query can be used to retrieve an attribute value, a set of attribute values, or any combination thereof. The search query can be, for example, an SQL query. The received request can involve one or more attributes of a data record of a central repository. This can be performed, for example, by explicitly referencing the attributes in the request and / or indirectly referencing the attributes. For example, the search query can be a structured search in which comparison or range predicates are used to limit the values of certain attributes. A structured search can provide an explicit reference to an attribute. In another example, the search query can be an unstructured search, for example, a keyword search that filters out records that do not contain a certain form of a specified keyword. An unstructured search can indirectly reference an attribute. In one example, the received request can include a name, an entity type, and / or numeric and time expressions in an unstructured format.
[0064] Upon receiving the request, in step 105, an entity identifier of the master data management system may be used to identify a set of one or more attributes referenced in the received request. The identification of the attribute group may further include identifying the entity type of each attribute of at least a portion of the attribute group. For example, the received request may be analyzed, such as parsing the received request to search for attributes whose values are being searched. For example, the entity identifier may identify the name and type of the entity, numeric and time expressions in user input entered as unstructured text, and map them to attributes of the master data management system with a certain probability, which allows them to be used to perform structured searches.
[0065] The entity identifier may be, for example, a token identifier that identifies a string, a numeric value, a pattern name, a location, etc. For example, the identification of an email may use the following email structure ABC@UVW.XYZ. The identification of a phone number may be based on the fact that a phone number is 10 digits. The identification of a social security account number (SSN) may be based on the fact that an SSN has the following structure AAA-BB-CCCC.
[0066] In one example, the entity identifier may use a machine learning (ML) model generated by a machine learning (ML) algorithm. The ML algorithm may be configured to read enterprise data, identify / learn portions of the data, and identify attributes. Using the ML model, the entity identifier may determine with a certain probability whether the input text may be a name or an address or a phone number or an SSN, etc. The engine selector may also use the ML model generated by the ML algorithm to perform the selection.
[0067] At step 107, using the identified set of attributes (e.g., and / or associated entity types), an engine selector of the master data management system may select a combination of one or more search engines of the search engines of the master data management system. For example, the performance of each search engine of the master data management system may be evaluated to search for values of each attribute of the attribute. The performance of the search engine may be determined by evaluating a performance parameter. The performance parameter may, for example, be an average number of results obtained by the search engine for searching for different values of the attribute and clicked on or used by a user. The performance parameter may alternatively or additionally include an average match score of results obtained by the search engine for searching for different values of the attribute and clicked on or used by a user.
[0068] The selection of the combination of one or more search engines may be performed using the current selection rule. For example, the selection rule may be applied for each given attribute in the attribute set as follows: For each search engine in the search engines of the master data management system, a value of a performance parameter may be determined that indicates the performance of the search engine for searching for a value of the given attribute. This may result in multiple values for each search engine in the combination of search engines, for example, if the attribute set includes two attributes, each search engine may have two performance values associated with the two attributes.
[0069] For example, if the attribute set includes name and birthday attributes, a structured probabilistic search engine may obtain better results for this input set and may therefore be selected. In addition, a free text search engine may be selected. Furthermore, both engines may be used to execute the request as follows: a free text search may also be performed when the probabilistic search engine finds no results. In another example, both search engines may be used to execute the request regardless of their respective results. In another example, the set of attributes may include birthday year and phone number. In this case, both engines may be selected because the probabilistic search engine can handle edit distance values and the birth year may be well satisfied by the free text engine as partial text of the date of birth. If the received request specifically calls for AND or NOT logic, a full text search engine may be used.
[0070] After selecting the combination of search engines, the request may be processed using the combination of search engines in step 109. For example, the engine selector may decide to use the combination of search engines to process data in parallel or sequentially based on a pre-established heuristic. The combination of search engines is used to obtain a candidate list based on the rules of the engine selector.
[0071] In step 111, at least a portion of the results of processing the request by the combination of search engines may be provided, for example, by a result provider of the master data management system. For example, rows of data records of the results may be displayed on a graphical user interface to enable a user to access one or more data records of the results. For example, a user may perform a user action on the provided results. The user action may, for example, include a mouse click or a touch gesture or another action that enables a user to access the provided results.
[0072] The results provided may include all results obtained after the request is processed by the combination of search engines, or may include only a predefined portion of all those results. For example, search results from the combination of search engines are aggregated and duplicates are removed, thereby generating a candidate list of data records. The resulting candidate list of data records may be scored. For example, multiple scoring engines of a master data management system are used. For example, depending on the attributes, a scoring function may be available or unavailable. Since a PME-based scorer may not be able to score all types of entities (e.g., contract-type data), multiple scoring engines are used. Of all the results obtained, one set of results may go to one scorer, while another set may go to some other scoring engines. The calls to these scoring engines may be performed in parallel to improve efficiency.
[0073] Based on the user actions performed on the provided results, the selection rules may be updated in step 113. The updated selection rules become the current selection rules and may therefore be used for further received data requests of the master data management system. For example, upon receiving a subsequent request to the received request of step 103 for data of the master data management system, steps 105-113 may be repeated and during this repetition, the updated selection rules may be used in the selection step 107.
[0074] For example, the selection rules are initially based primarily on the capabilities / suitability of the search engines corresponding to a given set of attributes, but the selection rules keep refining the rules based on, for example, user clicks, feedback, and the results (quality and performance) of searches conducted so far. If the previous selection of a search engine did not deliver results, an alternative search engine may also be dynamically selected.
[0075] Figure 2 is a flow chart of a method for providing search results of a collection of one or more search engines. Figure 2 The method can be applied, for example, Figure 1 data management systems (e.g. Figure 2 Can provide Figure 1 The details of step 111) may be applied to other search systems.
[0076] For example, the group of search engines may process a search request for data, and the search results may, for example, include data records. In step 201, each data record of the result may be associated with a match score or assigned a match score. The match score may be obtained by one or more scoring engines. For example, the match score of the data record of the result may be obtained by one or more scoring engines. In the case of more than one scoring engine, the match score may be a combination (e.g., an average) of the match scores obtained by more than one scoring engine. In one example, of all the results obtained, one group of results may be processed by one scoring engine, while another group may be processed by some other scoring engines. At least a portion of the one or more scoring engines used to score the results of a given search engine may or may not be part of a given search engine.
[0077] For example, each search engine in the search engine group may include a scoring engine configured to score the results of the corresponding search engine. In another example, one or more common scoring engines may be used to score the results obtained by the search engine set. For example, each search engine in the search engine group may be configured to connect to the scoring engine and receive the score of the data record from the scoring engine.
[0078] In step 203, the match score may be weighted. The weighting of the match score may be performed according to the performance of the components involved in producing the result. For example, in order to produce the search results, a search process is performed. The search process may include process steps performed by a system element such as a search engine to obtain the search results. The search process may therefore have components as process steps, system elements, and search results. Each of these components of the search process may have its own performance for performing the corresponding function. The performance of the component indicates how good the component is in performing its function or task. The performance of each component may be quantified by evaluating a corresponding performance parameter. The performance may affect the search results. In other words, each component of the search process has a contribution or influence on the quality of the search results obtained. At least a portion of these contributions may be considered by determining and assigning weights to at least a portion of the components of the search process. The weight assigned to the component may indicate (e.g., be proportional to) the performance of the component, for example, if the efficiency of the method step for identifying the attribute is 80%, the weight may be 0.8. In one example, a weight may be assigned to each of the components of the search process. In another example, a portion of the components of the search process may be selected or identified (e.g., by a user), and those identified components may be associated with corresponding weights. In one example, the weights may be user defined weights. The weighting step may result in each data record of the search result being associated with a weight of a component of the search process that resulted in the data record. The match score of the data record may be weighted by a combination of its associated weights, for example, the combination may be the product of the weights.
[0079] Using the weighted matching score, the results may be provided by removing duplicate data records of the results and retaining non-duplicate data records of the results having a weighted matching score above a predefined score threshold in step 205. For example, the results may be displayed on a user interface, such as a user may see a list of rows, each row being associated with a data record of the provided results.
[0080] The provided results can be operated or used by the user. For example, the user can perform user actions on the provided results. These user actions can be monitored, for example, by an activity monitor. For example, after a list of results is shown to the user on a user interface, the activity monitor can track the user's clicks on the shown results. A click on a row of results can be considered as the row that the user thinks she / he is looking for.
[0081] User actions, for example, may optionally be processed and analyzed in step 207. For example, the distribution of click counts relative to various characteristics of the data record (e.g., which engine it comes from, what is the confidence of the entity type detection, how complete the record is, how fresh the record is, etc.) may be analyzed. This data is captured to find correlations, and weights are calculated accordingly based on a lookup table or derived from an equation predicted by an ML-based regression model. Thus, as each new click is fed back to the system, the distribution may be changed, and thus help to redistribute the weights. The calculated weights may be used to update the weights used to obtain search results in step 209, for example, the calculated weights may replace the corresponding weights used to obtain search results. The updated weights may then be used when providing further search results to process further search requests.
[0082] Figure 3 is a flow chart of a method for providing search results of multiple search engines. Figure 3 The method can be applied, for example, Figure 1 Data management systems such as Figure 3 Can provide Figure 1 For details of step 111, for clarity, refer to Figures 4A-4F The example in , refers to two search engines 1 and engine 2 and a set of five properties to describe Figure 3 . One search engine implements probabilistic search and the other implements free text search. Further assume that the received request or input token is given as name+date of birth (Name+DOB), and the entity identifier identifies the first token as name with 90% confidence and sends it to search engine 1, and identifies the second token as DOB with 60% confidence and sends it to search engine 2.
[0083] In this example, for example, Figure 1 The components of the search process performed by the method may include a search engine, an identification step 105 and a result. Examples of the data records R1 to R6 of the result are shown in Figure 4A The results R1 to R6 of the two search engines are aggregated and their matching scores are normalized, thereby generating the matching scores of Table 403.
[0084] In step 301, an engine weight may be assigned to each of the search engines. Examples of engine weights are described in Figure 4B For example, an initial weight of 0.5 may be assigned to search engine 1 and search engine 2.
[0085] In step 303, each of the set of four attributes: name, DOB, address, identifier, and email is assigned an attribute weight indicating the confidence in identifying the attribute. Figure 4CThe attribute weights shown in may be an initial set of weights that may be updated after a search request is executed. Figure 4C As shown, for attribute names and confidence levels between 0% and 10%, the attribute weight is 0.1. In one example, the value of the confidence level can be used to obtain the attribute weight, for example, if the confidence level is less than 10%, the attribute weight can be equal to 0.1. However, other weight determination methods can be used.
[0086] In step 305 , a weight indicating the degree of completion of the data record and a weight indicating the degree of freshness of the data record may be allocated to each data record of the result. Figure 4D The table shows example values of completion weights for given data records. Figure 4D The completion weights shown in may be an initial set of weights that may be updated after a search request is executed. Figure 4D As shown, the completion weight of a given data record can be provided as a function of the completion of the data record. For example, for a completion between 10% and 20%, the completion weight is 0.2. In one example, the completion weight can be obtained using the value of the completion, for example, if the completion is less than 10%, the completion weight can be equal to 0.1. However, other example weighting methods can be used.
[0087] Figure 4E The table shows the freshness-weighted instance values for a given data record. Figure 4E The freshness weights shown in may be an initial set of weights that may be updated after a search request is executed. Figure 4E As shown, a freshness weight for a given data record may be provided based on the freshness of the data record. For example, for data records with a freshness between 3 and 5 years, the freshness weight is 0.8. However, other example weighting methods may be used.
[0088] For each data record of the result, the corresponding engine weight, attribute weight, completion weight and freshness weight may be combined in step 307, and the score of the data record may be weighted by the combined weight. The combined weight may be, for example, the product of the four weights. The final result score as the weighted score is as follows: Figure 4F Using the final score, the results can be filtered and provided to the user. For example, only data records R1, R2 and R6 can be provided to the user when their final scores are higher than a threshold value 1. Figure 4FThe table shows that for records R1, R2, and R3, the engine weight Wa is 0.5 because they come from engine 1, and for records R4, R5, and R6, the engine weight Wa is 0.5 because they come from engine 2. For R1, R2, and R3, the attribute weight (associated with the name attribute) Wb is 0.9 because they are the result set of the entity identifier that identifies the name attribute with 90% confidence. The attribute weight (associated with the DOB attribute) Wb of R4, R5, and R6 is 0.6 because they are the result set of the entity identifier that identifies the DOB with 60% confidence. The completion weight wc is based on the completion of each record. For example, R1 is 80% complete, so 0.8 is the completion weight. The freshness weight Wd is based on the freshness of each record. For example, R1 is fresh, that is, the last modified date is less than 1 year, so 1 is the freshness weight. The final score can be obtained as follows: Final score = initial normalized score * (A*Wa) * (B*Wb) * (C*Wc) * (D*Wd), where A, B, C and D are weights assumed to be 1 for simplicity.
[0089] Figure 5 is a flow chart of a method for updating weights for weighting matching scores of data records of results of processing a search request by multiple search engines. For simplification purposes, Figure 5 The update of the completion weight is described. However, the weight update method can be used for other weights. Figure 5 .
[0090] When providing results to the user, the activity monitor may monitor user actions performed on the provided results in step 501. For example, the activity monitor may count the number of clicks that have been performed for each data record displayed to the user. This may yield Fig. 6A of table. Fig. 6A The table shows the number of clicks performed by the user for different degrees of completion of the data record. For example, the user performs one mouse click on the row representing the data record with 80% completion.
[0091] In step 503, the following may be processed or analyzed: Fig. 6A The results of the monitored operations shown in order to find the updated completion weight. To this end, the following can be generated: Figure 6B The lookup table includes the range of completions used for weighting (see Figure 4D) and the percentage of clicks performed by users on data records with a degree of completion in the listed range. In this example, the data shows that users almost never click on records that are less than 30% complete, while ~40% of clicks occur on records that are greater than 80% complete. According to the weights in the lookup table, a new record with a degree of completion of 60% would be given a weight proportional to 12%. For example, for data records with a degree of completion between 50% and 60%, the score for the click is calculated from Figure 6A-6B For example, the completion weight for a 50% to 60% completion range would become 0.12 instead of ( Figure 4D The initial weight is 0.6.
[0092] In another example, if Figure 6C As illustrated in , analysis of user operations can be performed by modeling changes in completion as a function of click scores. Figure 6C An example model 601 is shown in . The model 601 can be used to determine updated weights for a given value of completion. The model 601 is described by an equation that can be predicted by an ML-based regression model.
[0093] The result of the method may be updated weights, which may be used to replace the initial weights provided in FIG. 4 , for example. The updated weights may be used to weight the matching scores of data records generated by executing new search requests.
[0094] Figure 7 A block diagram representation of a computer system 700 according to an example of the present disclosure is depicted. The computer system 700 can be configured to perform master data management, for example. The computer system 700 includes a master data management system 701 and one or more client systems 703. The client systems 703 can access a data source 705. The master data management system 701 can control access (read and write access, etc.) to a central repository 710. The master data management system 701 can utilize index data 711 to process fuzzy searches.
[0095] The master data management system 701 can process the data records received from the client systems 703 and store the data records in the central repository 710. The client systems 703 can obtain the data records, for example, from different data sources 705. The client systems 703 can communicate with the master data management system 701 via a network connection, including, for example, a wireless local area network (WLAN) connection, a WAN (wide area network) connection, a LAN (local area network) connection, or a combination thereof.
[0096] The master data management system 701 may also be configured to process data requests or queries for accessing data stored in the central repository 710. For example, a query may be received from a client system 703. The master data management system 701 includes an entity identifier 721 for identifying attributes or entities in the received data request. The entity identifier 721 may, for example, identify the names and types of entities, numbers, and time expressions in user input entered as unstructured text, and map them to the attributes of data records stored in the central repository 710 with a specific probability or confidence, which allows them to be used to perform structured search attributes. For example, the entity identifier 721 may be a token identifier that identifies a string / value or a pattern name, location, such as an email should follow ABC@UVW. xyz or a telephone number after 10 digits or an SSN after the AAA-BB-CCCC structure. The entity identifier 721 may be configured to use a machine learning model to classify or identify input data attributes of data records stored in the central repository 710. The master data management system 701 also includes an engine selector 722 for selecting one or more engines suitable for executing the received search request. The engine selector 722 can decide to use one or more engines to process data in parallel or sequentially based on pre-established heuristics. For example, the rules initially used to select an engine are mainly based on the capabilities / suitability of the engine corresponding to a given set of attributes and entity type. After the initial processing of the first request, the engine selector keeps improving its rules based on the results (quality and performance) of the user's clicks, feedback, and searches done so far. If the previous selection of the search engine does not deliver results, the engine selector 722 can also dynamically select a replacement engine. Based on the rules of the engine selector 722, multiple search engines can be selected and used to obtain a good candidate list. Search results from all engines are aggregated and duplicates are removed. Then the obtained candidate list is scored. Multiple scoring engines are used. Depending on the attributes, the scoring function may be available or unavailable. In addition to the PME-based scorer, other scoring engines are also used to score the search results. For example, among all the results obtained, one group of results may go to one scorer, while another group may go to some other scoring engine. The calls of these engines can be performed in parallel to improve efficiency.
[0097] The master data management system 701 also includes a weight provider and result aggregator 723 for weighting and aggregating the results obtained by the search engine. Once all raters have completed the rating, the aggregation of the results can be based on a weighted average of the ratings.
[0098] By looking for correlations between the characteristics of the pattern and result set and the quality of the match, the weights are derived and refined over a period of time. The analyzer can use machine learning to identify these correlations. The characteristics of the result set under analysis may include (but are not limited to) at least one of the following: the matching engine used to obtain the score, such as a specific scoring engine may have a wider scoring range or be less reliable than other scoring engines; the certainty of the input data type detected by the entity recognizer; the completeness of the record, such as indicating how many fields are filled and the freshness of the data (last updated date). The weight is a set of numbers used to modify the score of the result set. The quality of the match is indicated by the analysis of the user's clicks. Clicks on the results shown indicate that the user understands a better match. The quality of the match can also be based on explicit feedback about the quality of the match that can be found on the UI. The analysis of the correlation is fed back to improve the weight provider 723. The results obtained by the search engine are aggregated using weights, and then continued to the next stage based on comparisons with threshold records.
[0099] The master data management system 701 also includes various APIs for allowing storage and access to data in the central repository 710. For example, the master data management system 701 includes a create, read, update, and delete (CRUD) API 724 for enabling access to data such as storing new data records in the central repository 710. The master data management system 701 also includes an API associated with a search engine that it includes. Figure 7 Two APIs, a structured search API 725 and a fuzzy search API 726, are shown for example purposes.
[0100] The master data management system 701 also includes components capable of filtering the results to be provided to the user. For example, the master data management system 701 includes a component 727 for applying visibility rules and another component 728 for applying consent management. The master data management system 701 includes a component 729 for applying standardized rules to the data to be stored in the central repository 710. Filtering may be advantageous because data security and privacy are extremely important in master data management solutions. Although full-text searches attempt to cast a wide net to find matches, it can be ensured that such an excessive range remains within the system and that information is not inadvertently disclosed to unsolicited users. To this end, multiple filters will check whether the returned fields can be accessed by the queried user, and whether the resulting records have the necessary associated permissions from the data owner for the processing purpose provided by the user. Filtering is performed at a later stage in the search process to allow appropriate matching with all possible attributes. The result of the filtering can be a list of records in descending order of matching scores, containing only those records for which the required consent is provided, where those columns are allowed or visible to the user who initiated the search.
[0101] The master data management system 701 also includes indexing, matching, scoring, and linking services 730. Each client system 703 may include an administrative search user interface (UI) 741 for submitting search queries for querying data in the central repository 710. Each client system may also include services such as a messaging service 742 and a bulk loading service 743.
[0102] Reference Figure 8 The operation of computer system 700 is described in detail.
[0103] Figure 8 A flow chart describing an example method of operation of the master data management system 701 is depicted. In block 801, a free text search may be entered in a browser, for example, which may be an example of the management search UI 741. The entity recognizer 721 may receive (block 802) the free text search request and may, as herein, for example, Figure 1 The received request is processed as described in to identify the attribute or entity. The engine selector 722 can then be used (block 803) to select a search engine suitable for the identified attribute. Figure 8 As illustrated in , two search engines (boxes 804 and 805) are selected and used to execute the received search request. The execution results of the search request can be scored (box 806) using the matching and scoring services of the master data management system 701. The scoring can also use an additional scoring mechanism (box 807). Then, the results are aggregated and the scores are normalized (box 808). Before providing the results to the user, some filters can be applied (box 809). These filters can, for example, include at least one of visibility filters and rules based on agreed data filters and custom filters. Then the filtered results are displayed in a browser (e.g., a browser receiving a free text search) (box 810). The displayed results can be monitored (box 811) and analyzed by user clicks and quality feedback analyzers. For example, the analyzer can use a machine learning model to determine weights based on the user's actions on the results. As shown in arrows 812 and 813, weights can be used to update engine selector 722 and weight provider 723. Then, the weight provided by weight provider 723 can be used for scoring box 808 in the next iteration of the method.
[0104] Fig. 9 A diagram illustrating an example of processing a request in accordance with the present subject matter is depicted. Fig. 9The first column 901 shows example contents of a received request or input token. For example, the received request may include "Robert", "Bangalore", and the number "123-45-6789". The second column 902 shows the result of entity identification when processing the received request. For example, "Robert" is identified as a name attribute, "Bangalore" is identified as an address attribute, and the number "123-45-6789" is identified as an SSN attribute. Columns 902 and 904 indicate that the engine selector has selected the search engine "Search Engine 1" for processing the request "Robert". Columns 902 and 904 also indicate that the engine selector has selected the search engine "Search Engine 2" for processing the request "Bangalore". Columns 902 and 904 also indicate that the engine selector has selected both search engines "Search Engine 1" and "Search Engine 2" for processing the request "123-45-6789". The results of processing the request are processed, e.g., aggregated, before being provided as shown in column 905. For example, column 905 shows that search engine "Search Engine 1" has found records R1, R2, and R3 when searching for "Robert". Column 905 also shows that search engine "Search Engine 2" has found records R4 and R5 when searching for "Bangalore". Column 905 also shows that search engine "Search Engine 1" has found record R6 when searching for "123-45-6789", and search engine "Search Engine 2" has found record R7 when searching for "123-45-6789". Before being provided to the user, results R1 to R7 may need to be filtered using a data control filter as shown in column 906. After being filtered, as shown in column 907, the results can then be output to the user. As shown in column 907, the date of birth values are filtered out from records R1 to R7 because the user who submitted the results is not allowed to access them.
[0105] It should be understood that one or more of the above-described embodiments of the present invention may be combined as long as the combined embodiments are not mutually exclusive.
[0106] Various embodiments are specified in the following examples.
[0107] 1. A method for accessing a data record of a master data management system, the data record comprising a plurality of attributes, the method comprising:
[0108] enhancing the master data management system with one or more search engines to enable access to the data records;
[0109] receiving a request for data at a master data management system;
[0110] identifying a set of one or more attributes of the plurality of attributes referenced in the received request;
[0111] selecting a combination of one or more search engines among the search engines of the master data management system, the performance of the one or more search engines for searching for values of at least a portion of the attribute set satisfying a current selection rule;
[0112] Use a combination of search engines to handle the request;
[0113] At least a portion of the results of the processing is provided.
[0114] 2. The method according to clause 1 also includes updating the selection rules based on user operations on the provided results, the updated selection rules becoming the current selection rules, and when another data request is received, repeating the identification, selection, processing and providing steps using the current selection rules.
[0115] 3. A method according to clause 1, wherein the results include data records of the master data management system associated with corresponding match scores obtained by a scoring engine of the search engine, and the method also includes weighting the match scores based on the performance of components involved in providing the results, the components including method steps, elements for providing the results, and at least a portion of the results, wherein the provided results include non-duplicate data records having weighted match scores above a predefined score threshold.
[0116] 4. The method of clause 3, wherein the components include a search engine, an identification step, and a result, and the method further comprises:
[0117] assigning an engine weight to each of the search engines;
[0118] assigning attribute weights to the set of attributes, wherein the attribute weight of an attribute indicates a confidence level that the attribute is identified;
[0119] assigning, to each data record of the result, a completion weight indicating the data record and a freshness weight indicating the data record;
[0120] For each data record of the result, the corresponding engine weight, attribute weight, completion weight, and freshness weight are combined, and the score of the data record is weighted by the combined weight.
[0121] 5. The method according to clause 4, further comprising:
[0122] Provide user parameters to quantify user actions;
[0123] For each component of at least a portion of the components, determining a value of the user parameter and an associated value of a component parameter describing the component; and updating a weight assigned to the component using the determined association.
[0124] 6. The method of clause 5, further comprising providing a lookup table that associates values of user parameters with values of component parameters, and using the lookup table to update weights assigned to components.
[0125] 7. The method according to clause 5 further includes modeling the change in the value of the user parameter with the value of the component parameter using a predefined model, and using the model to determine the updated weight of the component, and using the updated weight to update the weight assigned to the component.
[0126] 8. A method according to clause 5, wherein the user action in the user operation includes an indication of a result selection, which includes a mouse click on a displayed result in the provided results, and wherein the user parameters include at least one of the number of clicks, the frequency of clicks, and the duration of accessing a given result in the results.
[0127] 9. A method according to clause 1, wherein the results include data records of the master data management system associated with corresponding matching scores obtained by a scoring engine of the search engine, wherein the provided results include non-duplicate data records having matching scores above a predefined score threshold.
[0128] 10. The method of clause 1, wherein, for each attribute in the attribute set, the selection rule comprises:
[0129] for each of the search engines, determining a value of a performance parameter indicative of performance of the search engine for use in searching for a value of the property;
[0130] A search engine having a performance parameter value higher than a predetermined performance threshold is selected.
[0131] 11. The method of clause 10, wherein the performance parameter comprises at least one of the following: the number of results and the degree to which the results match expectations.
[0132] 12. The method of clause 10, wherein the selection rule uses a table associating attributes to corresponding search engines, and the updating of the selection rule comprises:
[0133] determining a value of a user parameter, the value of the user parameter quantifying a user action on the provided results of each search engine in the combination of search engines; and
[0134] Using the determined values associated with each search engine in the combination of search engines to identify values of the user parameter that are less than a predefined threshold, and for each identified value of the user parameter, determining the attributes in the attribute set and the search engine associated with the identified value, and updating the table using the determined attributes and search engines.
[0135] 13. The method of clause 1, wherein processing of the requests is performed in parallel by a combination of the search engines.
[0136] 14. The method of clause 1, wherein the combination of search engines is a ranked list of search engines, wherein processing of the requests is performed serially following the ranked list until a minimum number of results is exceeded.
[0137] 15. A method according to clause 1, wherein identifying the set of attributes includes inputting the received request into a predefined machine learning model; receiving a classification of the request from the machine learning model, the classification indicating the set of attributes.
[0138] 16. The method of clause 1, wherein the attribute set is input into a predefined machine learning model, and one or more search engines that can be used to search the attribute set are received from the machine learning model.
[0139] 17. The method according to clause 16 also includes: receiving a training set indicating different sets of one or more training attributes, wherein each training attribute set is labeled to indicate a search engine suitable for executing the training attribute set; using the training set to train a predefined machine learning algorithm to generate the machine learning model.
[0140] 18. The method of clause 1, wherein the provided results include data records filtered based on a sender of the request.
[0141] 19. A method for providing search results of a search engine according to a predefined search process, the method comprising
[0142] receiving results of a search request obtained by the search engine, each of the results being associated with a matching score;
[0143] for each of the results, determining a set of one or more components of the search process involved in providing the result, and assigning a predefined weight to each component in the set of components;
[0144] weighting the matching score using the weight;
[0145] Results with a weighted match score above a predefined score threshold are provided.
[0146] 20. The method according to clause 19, further comprising:
[0147] Analyze user actions on the provided results by evaluating user parameters that quantify user actions;
[0148] For each component in at least a portion of the set of components, determining one or more values of a component parameter describing the component and an associated value of the user parameter; determining an updated weight using the determined association; and
[0149] replacing weights assigned to the at least a portion of the components with the determined weights;
[0150] The method is repeated for further received search results using the updated weights.
[0151] 21. The method of clause 20, further comprising providing a table that associates values of user parameters with values of component parameters, and using the table to update weights assigned to components.
[0152] 22. The method of clause 20, further comprising modeling the associations between the values using a predefined model, and using the model to determine updated weights for the components, and using the updated weights to update the weights assigned to the components.
[0153] 23. A method according to clause 20, wherein the user operation in the user operation includes a mouse click on a displayed result in the provided results, and wherein the user parameters include at least one of the number of clicks, the frequency of clicks, and the duration of accessing a given result in the results.
[0154] Various aspects of the present invention are described herein with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to embodiments of the present invention. It will be understood that each frame of the flow chart and / or block diagram and the combination of frames in the flow chart and / or block diagram can be implemented by computer-readable program instructions.
[0155] The present invention may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon, the computer-readable program instructions being used to cause a processor to perform various aspects of the present invention.
[0156] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or a raised structure in a groove on which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a temporary signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.
[0157] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.
[0158] The computer-readable program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (e.g., Smalltalk, C++, etc.) and conventional procedural programming languages (e.g., "C" programming languages or similar programming languages). The computer-readable program instructions may be executed completely on the computer of the user's computer system, partially on the computer of the user's computer system, executed as an independent software package, partially on the computer of the user's computer system and partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer may be connected to the computer of the user's computer system via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider via the Internet). In some embodiments, in order to perform various aspects of the present invention, an electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute a computer-readable program instruction to personalize the electronic circuit by utilizing the state information of the computer-readable program instructions.
[0159] Various aspects of the present invention are described herein with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to embodiments of the present invention. It will be understood that each frame of the flow chart and / or block diagram and the combination of frames in the flow chart and / or block diagram can be implemented by computer-readable program instructions.
[0160] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can guide the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0161] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0162] Flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention.In this regard, each frame in flow chart or block diagram can represent the module, segment or part of instruction, and it comprises one or more executable instructions for realizing the logical function of appointment.In some alternative embodiments, the function mentioned in the frame may not occur in the order mentioned in the figure.For example, two frames shown continuously can actually be performed substantially simultaneously, or these frames can sometimes be performed in reverse order, depending on the function involved.It will also be noted that the combination of the frame in each frame of block diagram and / or flow chart illustration and block diagram and / or flow chart illustration can be realized by the dedicated hardware-based system of performing specified function or action or performing the combination of special hardware and computer instruction.
Claims
1. A method for accessing a data record of a master data management system, the data record comprising a plurality of attributes, the method comprising: enhancing the master data management system with one or more search engines for enabling access to the data records; receiving a request for data at a master data management system; an attribute set identifying one or more attributes of the plurality of attributes referenced in the received request; selecting a combination of one or more search engines among the search engines of the master data management system, the performance of the one or more search engines for searching for values of at least a portion of the attribute set satisfying a current selection rule; Processing the request using a combination of search engines; providing at least a portion of the results of said processing, The result includes a data record of the master data management system associated with a corresponding match score obtained by a scoring engine in the search engine, and the components involved in providing the result include the search engine, the identification step and the result. The method further includes: assigning an engine weight to each of the search engines; assigning attribute weights to the set of attributes, wherein the attribute weights of attributes indicate a confidence level that the attributes are identified; assigning, to each data record of the result, a completion weight indicating the data record and a freshness weight indicating the data record; and For each data record of the result, the corresponding engine weight, attribute weight, completion weight, and freshness weight are combined, and the score of the data record is weighted by the combined weight.
2. The method according to claim 1, characterized in that: It also includes updating the selection rule based on a user operation on the provided result, the updated selection rule becoming the current selection rule, and when another data request is received, repeating the identification, selection, processing and providing steps using the current selection rule.
3. The method according to claim 1, wherein: The method also includes weighting the match score based on performance of components involved in providing the result, the components including method steps, elements for providing the result, and at least a portion of the result, wherein the provided result includes non-duplicate data records having a weighted match score above a predefined score threshold.
4. The method according to claim 2, further comprising: Provide user parameters to quantify user actions; for each component of at least a portion of the components, determining a value of the user parameter and an associated value of a component parameter describing the component; And using the determined association to update the weight assigned to the component.
5. The method of claim 4, further comprising providing a lookup table that relates values of the user parameters to the values of the component parameters, and using the lookup table to update the weights assigned to the components.
6. The method according to claim 4 also includes using a predefined model to model the change in the value of the user parameter and the value of the component parameter, and using the model to determine the updated weight of the component and using the updated weight to update the weight assigned to the component.
7. The method according to claim 4, wherein: The user operation in the user operation includes an indication of selection of a result, the indication including a mouse click on a displayed result in the provided results, wherein the user parameter includes at least one of the number of clicks, the frequency of clicks, and the duration of accessing a given result in the results.
8. The method according to claim 1, wherein: The results include data records of the master data management system associated with corresponding matching scores obtained by a scoring engine of the search engine, wherein the provided results include non-duplicate data records having matching scores above a predefined score threshold.
9. The method according to claim 1, wherein: For each attribute in the attribute set, the selection rule comprises: for each of the search engines, determining a value of a performance parameter indicative of performance of the search engine for use in searching for a value of the property; A search engine having a performance parameter value higher than a predetermined performance threshold is selected.
10. The method of claim 9, wherein the performance parameter comprises at least one of the following: the number of results and the degree to which the results match expectations.
11. The method according to claim 9, wherein the selection rule uses a table that associates attributes to corresponding search engines, and the updating of the selection rule comprises: determining a value of a user parameter that quantifies a user action on the results provided by each search engine in the combination of search engines; as well as Using the determined values associated with each search engine in the combination of search engines to identify values of the user parameter that are less than a predefined threshold, and for each identified value of the user parameter, determining the attributes in the attribute set and the search engine associated with the identified value, and updating the table using the determined attributes and search engines.
12. The method according to claim 1, wherein: The processing of the requests is performed in parallel by the combination of search engines.
13. The method according to any one of claims 1 to 12, wherein: The combination of search engines is a ranked list of search engines, wherein processing of the request is performed successively following the ranked list until a minimum number of results is exceeded.
14. A method according to any one of claims 1-12, wherein identifying the set of attributes includes inputting the received request into a predefined machine learning model; receiving a classification of the request from the machine learning model, the classification indicating the set of attributes.
15. According to the method described in any one of claims 1-12, the attribute set is input into a predefined machine learning model, and one or more search engines that can be used to search the attribute set are received from the machine learning model.
16. The method according to claim 15, further comprising: receiving a training set indicating different sets of one or more training attributes, wherein each training attribute set is labeled to indicate a search engine suitable for performing a search for the training attribute set; The training set is used to train a predefined machine learning algorithm to generate the machine learning model.
17. The method according to any one of claims 1-12, wherein the provided results include data records filtered according to a sender of the request.
18. A computer program product comprising a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code being configured to access a data record of a master data management system, the data management system comprising a search engine for enabling access to the data record, the data record comprising a plurality of attributes, the computer readable program code being further configured to: receiving a request for data at the master data management system; an attribute set identifying one or more attributes of the plurality of attributes referenced in the received request; selecting a combination of one or more search engines among the search engines of the master data management system, the performance of the one or more search engines for searching for values of at least a portion of the attribute set satisfying a current selection rule; processing the request using a combination of search engines; providing at least a portion of the processing results, Wherein, the result comprises a data record of the master data management system associated with a corresponding match score obtained by a scoring engine in the search engine, the components involved in providing the result comprise the search engine, the identifying step and the result, and the computer readable program code is further configured to: assigning an engine weight to each of the search engines; assigning attribute weights to the set of attributes, wherein the attribute weights of attributes indicate a confidence level that the attributes are identified; assigning, to each data record of the result, a completion weight indicating the data record and a freshness weight indicating the data record; and For each data record of the result, the corresponding engine weight, attribute weight, completion weight, and freshness weight are combined, and the score of the data record is weighted by the combined weight.
19. A computer system for enabling access to a data record, the data record comprising a plurality of attributes, the computer system comprising a plurality of search engines for enabling access to the data record; A user interface configured to receive a request for data; an entity identifier configured to identify an attribute set of one or more attributes of the plurality of attributes referenced in the received request; an engine selector configured to select a combination of one or more search engines among the search engines, wherein the performance of the one or more search engines for searching for values of at least a portion of the attribute set satisfies a current selection rule; wherein the search engine is configured to process the request; and a result provider configured to provide at least a portion of the result of the processing, wherein the result comprises a data record of the master data management system associated with a corresponding match score obtained by a scoring engine in the search engine, and the components involved in providing the result comprise the search engine, the identifying step and the result, The scoring engine is configured to: assigning an engine weight to each of the search engines; assigning attribute weights to the set of attributes, wherein the attribute weights of attributes indicate a confidence level that the attributes are identified; assigning, to each data record of the result, a completion weight indicating the data record and a freshness weight indicating the data record; and For each data record of the result, the corresponding engine weight, attribute weight, completion weight, and freshness weight are combined, and the score of the data record is weighted by the combined weight.
20. The computer system of claim 19, the computer system being a master data management system.
21. The computer system of claim 19, wherein: The computer system also includes a weight provider configured to weight the match score based on performance of components involved in providing the result, the components including method steps, elements for providing the result, and at least a portion of the result, wherein the provided result includes non-duplicate data records having a weighted match score above a predefined score threshold.
Citation Information
Patent Citations
Search method and device based on query word
CN102043833A