Methods, devices, equipment and storage media for searching data assets

By determining the position and priority of keywords in the data asset management platform and sorting them based on various value factors of data assets, the shortcomings of existing intelligent search and recommendation technologies are addressed, enabling personalized data asset search and recommendation and improving the accuracy and efficiency of data search.

CN116186097BActive Publication Date: 2026-04-03AVATR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing data asset management platforms struggle to achieve intelligent search and recommendation, failing to accurately meet users' personalized needs.

Method used

By determining the position and priority of keywords in data assets, and combining the usage frequency, click popularity, dependence on other data assets, and topic relevance of data assets, multi-level sorting and filtering are performed to generate a personalized data asset set.

Benefits of technology

It enables more accurate fulfillment of user needs, helps users quickly find the data assets they need, and improves the intelligence and efficiency of data search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116186097B_ABST
    Figure CN116186097B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for searching data assets. The search method includes: determining and sorting data assets in a data asset set based on the position of keywords in each data asset and the priority of the position of keywords in the data assets, to obtain a first data asset set containing keywords and having a first ranking; and sorting data assets in the first data asset set with the same position of keywords in the data assets based on the value of each first data asset in the first data asset set, to obtain a second data asset set with a second ranking, wherein the value of the first data asset is used to characterize at least one of the following: the frequency of use of the data asset, the click popularity, the dependence of other data assets on the data asset, and the topic relevance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of computer technology, and in particular to a method, apparatus, device, and storage medium for searching data assets. Background Technology

[0002] A data asset management platform, based on big data processing and storage, centered on data asset management, and with data insights as its value proposition, transforms enterprise data from being difficult to understand and controllable to being manageable and operational. Building a data asset platform enables unified data processing, storage, authorization, and search, as well as full lifecycle management. However, how to achieve intelligent search and recommendation of data assets to more accurately meet people's needs remains a research topic. Summary of the Invention

[0003] In view of this, embodiments of this application provide at least one method, apparatus, device, and storage medium for searching data assets.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] In a first aspect, embodiments of this application provide a method for searching data assets. The method includes: determining the position of a keyword carried in a search request within each data asset of a data asset set; determining the priority of the keyword's position within the data asset; filtering and sorting the data assets in the data asset set based on the keyword's position within each data asset and the keyword's priority, to obtain a first data asset set containing the keyword and having a first ranking; sorting the data assets in the first data asset set with the same keyword position based on the value of each first data asset in the first data asset set, to obtain a second data asset set with a second ranking, and outputting the second data asset set with the second ranking, wherein the value of the first data asset is used to characterize at least one of the following: the data asset's usage frequency, click popularity, the dependence of other data assets on the data asset, and topic relevance.

[0006] Secondly, embodiments of this application provide a data asset search device, the device comprising: a first determining module, configured to determine the position of a keyword carried in a search request in each data asset of a data asset set; a second determining module, configured to determine the priority of the position of the keyword in the data asset; a filtering and sorting module, configured to filter and sort the data assets in the data asset set based on the position of the keyword in each data asset and the priority of the position of the keyword in the data asset, to obtain a first data asset set containing the keyword and having a first ranking; and a first ranking module, configured to sort the data assets in the first data asset set with the same position of the keyword in the data asset set based on the value of each first data asset in the first data asset set, to obtain a second data asset set with a second ranking, and output the second data asset set with the second ranking, wherein the value of the first data asset is used to characterize at least one of the following: the usage frequency of the data asset, the click popularity, the dependence of other data assets on the data asset, and the topic relevance.

[0007] Thirdly, embodiments of this application provide a computer device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.

[0009] In this embodiment, keywords are first sorted according to their priority within the data assets, and then sorted according to their value, resulting in a second set of data assets with a second ranking. The value of a data asset considers at least one of the following: usage frequency, click volume, dependence of other data assets on that data asset, and topic relevance. Therefore, the final second set of data assets with the second ranking can combine factors such as data asset popularity, user habits, importance, and topic relevance, more accurately meeting the needs of different users, enabling personalized search, and helping users find the data assets they need more quickly.

[0010] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0012] Figure 1 A schematic diagram illustrating the implementation process of a data asset search method provided in this application embodiment;

[0013] Figure 2 A schematic diagram illustrating the implementation process of another data asset search method provided in this application embodiment;

[0014] Figure 3 A schematic diagram illustrating an implementation process for determining the value of a first data asset, provided as an embodiment of this application;

[0015] Figure 4 A schematic diagram illustrating the implementation process of a data asset recommendation method provided in this application embodiment;

[0016] Figure 5A A schematic diagram of a data asset catalog provided in an embodiment of this application;

[0017] Figure 5B A schematic diagram of a framework structure for implementing intelligent recommendation of data assets, provided for an embodiment of this application;

[0018] Figure 5C A schematic diagram of a knowledge graph provided in an embodiment of this application;

[0019] Figure 6 A schematic diagram illustrating the composition of a data asset search device provided in an embodiment of this application;

[0020] Figure 7 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0023] The terms “first / second / third” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first / second / third” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this application.

[0025] This application provides a method for searching data assets, which can be executed by a processor of a computer device. The computer device refers to a device with data processing capabilities, such as a server, laptop, tablet, desktop computer, smart TV, set-top box, or mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device). Figure 1 This application provides a schematic diagram illustrating the implementation process of a data asset search method, as shown in the embodiments below. Figure 1 As shown, the method includes the following steps S101 to S104:

[0026] Step S101: Determine the location of the keywords carried in the search request in each data asset of the data asset set;

[0027] Here, a search request refers to a request entered by an object (e.g., a user) for searching. Data assets can include table assets, report assets, metric assets, and Asset Priority Index (API) assets. Locations within a data asset can include: the data asset's name, content, and comments. The data asset's name refers to its title, its content to the main text, and its comments to explain the content. For example, in the case of a table asset, locations within the data asset include the table name, fields, and comments.

[0028] The location of a keyword in each data asset can include the following: the keyword does not exist, the keyword is in the name, the keyword is in the content, or the keyword is in the comments.

[0029] In some embodiments, the data asset search method provided in this application is applied to a data asset management platform. Then, the implementation of step S101 may include: after the object clicks the "Data Search" button in the data asset management platform, the above step S101 is executed.

[0030] Step S102: Determine the priority of the keyword's position in the data asset;

[0031] Here, priority refers to the order in which the searched keywords are located within the data assets. For example, the search prioritizes data assets where the keyword is in the name; then it searches data assets where the keyword is in the content; and finally, it searches data assets where the keyword is in the annotations. During implementation, this can be pre-configured according to requirements.

[0032] Step S103: Based on the position of the keyword in each data asset and the priority of the position of the keyword in the data asset, filter and sort the data assets in the data asset set to obtain a first data asset set containing the keyword with a first sort.

[0033] Here, the implementation of step S103 may include: first, filtering out data assets containing keywords from the data asset set, and then sorting the filtered data assets according to the priority of the keyword's position in the data assets to obtain a first data asset set containing keywords with a first sort.

[0034] Correspondingly, in some embodiments, the location of keywords in the data asset includes: name, content, and annotation. The implementation of step S103 may include the following steps S1031 and S1032:

[0035] Step S1031: Based on the position of the keyword in each data asset, filter out the target data asset set containing the keyword from the data asset set;

[0036] Here, the target data asset set is the set of data assets that have been selected and contain keywords.

[0037] Step S1032: Based on the priority of the keyword in the data assets from high to low, namely name, content, and annotation, sort the data assets in the target data asset set to obtain a first data asset set containing the keyword with a first sort.

[0038] Here, the implementation of step S1032 may include: sorting the data assets in the target data asset set in descending order of priority as name, content, and annotation, prioritizing the display of data assets whose keywords are in the name, then displaying data assets whose keywords are in the content, and finally displaying data assets whose keywords are in the annotation, thus obtaining the first data asset set.

[0039] Step S104: Based on the value of each first data asset in the first data asset set, sort the data assets in the first data asset set whose keywords are in the same position in the data assets to obtain a second data asset set with a second sort, and output the second data asset set with the second sort, wherein the value of the first data asset is used to characterize at least one of the following: the usage frequency of the data asset, the click popularity, the dependence of other data assets on the data asset, and the topic relevance.

[0040] Here, the value of the first data asset is used to characterize at least one of the following: the frequency of use of the data asset, click popularity, the dependence of other data assets on the data asset, and topic relevance. That is, the value of the first data asset is used to characterize any one, two, three, or four of the following: the frequency of use of the data asset, click popularity, the dependence of other data assets on the data asset, and topic relevance.

[0041] The usage frequency of data assets is used to characterize how frequently a data asset is used. In implementation, data assets can be divided into frequently used data assets and infrequently used data assets based on their usage frequency. For example, frequently used and infrequently used data assets can be obtained through object collection, where collected data assets are frequently used and uncollected data assets are infrequently used.

[0042] Click popularity is used to characterize whether a data asset has been clicked by multiple people. In implementation, data assets can be sorted according to the number of clicks on an object to obtain click popularity data.

[0043] The dependency of other data assets on a given data asset is used to characterize the lineage relationship between data assets. For example, if data asset A is a subordinate of data asset B, and data asset B is a subordinate of data asset C, then there is a lineage relationship between data assets A, B, and C. In practice, this can be represented by the number of other data assets that depend on that data asset. For instance, if data asset A depends on data assets B and C, totaling two, then the dependency of other data assets on that data asset is 2.

[0044] Topic relevance refers to whether something is related to the topic being searched. It can be categorized as either relevant or irrelevant. In practice, relevant topics can be set as favorites, while unfavorited topics are considered irrelevant. Alternatively, relevant and irrelevant topics can be preset.

[0045] In some embodiments, the implementation of step S104 may include: setting different levels for the usage frequency, click popularity, dependence of other data assets on the data asset, and topic relevance of the data asset, and setting different scores for each level; then obtaining the value of the data asset by adding the scores of different items; and finally sorting the data assets from high to low according to their value to obtain a second set of data assets with a second ranking.

[0046] In this embodiment, firstly, the position of the keywords carried in the search request in each data asset of the data asset set is determined; then, the priority of the keyword position in the data asset is determined; subsequently, based on the position of the keyword in each data asset and the priority of the keyword position in the data asset, the data assets in the data asset set are filtered and sorted to obtain a first data asset set containing the keywords and having a first ranking; finally, based on the value of each first data asset in the first data asset set, the data assets in the first data asset set with the same keyword position in the data asset are sorted to obtain a second data asset set with a second ranking, and the second data asset set with the second ranking is output, wherein the value of the first data asset is used to characterize at least one of the following: the usage frequency of the data asset, the click popularity, the dependence of other data assets on the data asset, and the topic relevance.

[0047] As can be seen, in this embodiment, keywords are first sorted according to their priority within the data assets, and then sorted according to their value, resulting in a second set of data assets with a second ranking. The value of a data asset considers at least one of the following: usage frequency, click volume, dependence of other data assets on that data asset, and topic relevance. Therefore, the final second set of data assets with the second ranking can combine factors such as data asset popularity, user habits, importance, and topic relevance, more accurately meeting the needs of different users, enabling personalized search, and helping users find the data assets they need more quickly.

[0048] In some embodiments, such as Figure 2 As shown, before step S103 "based on the value of each first data asset in the first data asset set, sort the data assets in the first data asset set whose keywords are in the same position in the data assets", the following steps S201 and S202 are also included:

[0049] Step S201: Based on the priority of the database to which each first data asset belongs in the first data asset set, perform a first sorting on the data assets in the first data asset set whose keywords are in the same position in the data assets, and obtain a third data asset set with a third sorting.

[0050] Here, since different data assets come from different databases, such as the HIVE database and business databases, and different databases have different focuses and concerns; and since the object may have preferences for data in a certain database, after obtaining the first set of data assets, the data assets in the first set with the same keyword position in the data assets can be sorted again to obtain a third set of data assets with a third sort, so that databases with higher priority are displayed first, and databases with lower priority are displayed later. In implementation, for example, if the keyword in the first 100 data assets is in the name of the data asset, then step S201 is implemented by sorting the first 100 data assets, so that databases with higher priority are displayed first, and databases with lower priority are displayed later.

[0051] Step S202: Use the third data asset set as the first data asset set;

[0052] Here, in step S202, the third data asset set is used as the first data asset set to perform the above step S104.

[0053] Correspondingly, the implementation of step S104 may include: based on the value of each third data asset in the third data asset set, performing a second sorting on adjacent data assets belonging to the same database in the third data asset set to obtain the second data asset set with the second sorting.

[0054] Here, since step S201 has already sorted the data assets according to database priority, data assets with the same keyword position in the data assets belong to the same database and are grouped together. Therefore, step S104 involves a second sorting of adjacent data assets belonging to the same database in the third data asset set, resulting in a second data asset set with a second sort. For example, if the keywords in the first 100 data assets are all in the asset names, then step S201 involves sorting the first 100 data assets first, making the first 30 data assets belong to the higher-priority database, HIVE. Step S104 then involves sorting the first 30 data assets again according to their value, resulting in a second data asset set with a second sort.

[0055] In this embodiment of the application, before sorting according to the value of the data assets, the data assets are first sorted by database priority, so that databases frequently used by the object or database content that better meets the object's needs are displayed first, thereby making the dataset obtained by the search more in line with the object's needs and helping the object find the required data assets more quickly.

[0056] In some embodiments, such as Figure 3 As shown, the value of the first data asset is used to characterize the data asset's usage frequency, click popularity, and the dependence of other data assets on the data asset. Correspondingly, the method further includes the following steps S301a to S303a:

[0057] Step S301a: Obtain the usage frequency level, query volume ranking, and number of directly dependent data assets for each first data asset in the first data asset set;

[0058] Here, the usage frequency level can be obtained by classifying the usage frequency of the aforementioned data assets; it can also be set according to the user's usage preferences. In some embodiments, the usage frequency level can include frequently used data assets and infrequently used data assets. In practice, frequently used data assets can be set through methods such as favorites, while data assets not favorited are considered infrequently used data assets. In some embodiments, the usage frequency level can also be set by the user in advance, such as frequently used, infrequently used, or occasionally used.

[0059] The query volume ranking, also known as the click popularity, can be determined by statistically analyzing the number of queries to data assets over a period of time and ranking them accordingly.

[0060] The number of directly dependent data assets refers to the degree of dependence of the aforementioned other data assets on this data asset. In some embodiments, the number of directly dependent data assets is the number of directly dependent data assets that are related to this data asset.

[0061] Step S302a: Based on the usage frequency level, query volume ranking, and number of directly dependent data assets for each first data asset, determine the usage frequency score, popularity score, and dependency score of the corresponding data asset respectively;

[0062] Here, the implementation of step S302a may include: defining different scores for different usage frequency levels, query volume rankings, and the number of directly dependent data assets, thereby obtaining the usage frequency score, popularity score, and dependency score of the data assets. For example, a frequently used data asset score of 50 points and an infrequently used data asset score of 0 points; query volume ranking: top 10: 30 points, top 11-top 30: 20 points, top 31-top 100: 10 points; number of directly dependent data assets: 1. Quantity > 20, 20 points; 2. 10 < quantity ≤ 20, 10 points; 3. 5 < quantity ≤ 10, 5 points. This application embodiment does not limit the method of setting different level scores.

[0063] Step S303a: Determine the value of the corresponding first data asset based on the usage frequency score, popularity score, and dependence score of each first data asset.

[0064] Here, the implementation of step S303a may include: adding the usage frequency score, popularity score and dependence score of each first data asset to obtain the value of the corresponding first data asset.

[0065] In this embodiment of the application, the value of the corresponding first data asset is obtained by acquiring the usage frequency level, query volume ranking, and number of directly dependent data assets of each first data asset in the first data asset set, as well as their corresponding scores. Then, the usage frequency score, popularity score, and dependency score of each first data asset are added together to obtain the value of the corresponding first data asset, thereby realizing the quantification of the value of data assets.

[0066] In some embodiments, the value of the first data asset is used to characterize the relevance of the topic, and the method further includes the following steps S301b to S303b:

[0067] Step S301b: Obtain the topic relevance of each first data asset in the first data asset set;

[0068] Here, the relevance of a topic can include whether it is related to or unrelated to the topics preferred by the audience. In implementation, the audience's preferred topics (i.e., related to the topic) and unpreferred topics (i.e., unrelated to the topic) can be set in advance; alternatively, topics can be gradually added to the list of favorites, with unfavorited topics representing unpreferred topics.

[0069] Step S302b: Based on the topic relevance of each first data asset, determine the topic relevance score of the corresponding data asset;

[0070] Here, the implementation of step S302b may also include: setting different scores for the relevance of the topic, thereby obtaining a topic relevance score for the data asset. For example, 30 points for topic relevance and 0 points for topic irrelevance.

[0071] Step S303b: Determine the value of the corresponding first data asset based on the topic relevance score of each first data asset.

[0072] Here, step S303b is implemented such that the value of the first data asset is equal to the topic relevance score of the first data asset. In some embodiments, where the value of the first data asset is also used to characterize at least one of the data asset's usage frequency, click popularity, and dependence on the data asset by other data assets, the value of the first data asset is equal to the topic relevance score of the first data asset plus the scores of the other items.

[0073] In some embodiments, the method is applied to a data asset management platform, which includes a data asset catalog covering subject domains, subtopics, tables, and fields, for enabling navigation search and searching of data assets by subject catalog, facilitating users to understand the overall picture of data assets and accurately find the assets they need to search for.

[0074] In some embodiments, such as Figure 4 As shown, the method further includes the following steps S401 to S404:

[0075] Step S401: In response to the data recommendation request, based on the knowledge graph between object groups, objects, business metadata assets, and business data assets, determine a subset of similar objects in the object set that meet the preset conditions for similarity with the target object that issued the search request, wherein the object set includes objects from at least two of the object groups;

[0076] Here, an object can be a user using the data asset management platform, and an object group is a group consisting of multiple objects, such as a sales group or a research and development group. An object set includes objects from at least two object groups; that is, an object set can include objects from some object groups or objects from all object groups.

[0077] Business metadata assets include metric metadata assets and report metadata assets. In some embodiments, metric metadata assets are also referred to as metric assets, such as assets related to the metric, sales amount, sales quantity, responsible person, etc. Report metadata assets are also referred to as report assets, such as assets related to report attributes, report name, report description, report identification number, etc.

[0078] Business data assets refer to data assets stored in a database that are related to business operations. In some embodiments, business data assets are also referred to as table assets, such as the actual data in a table.

[0079] The preset conditions can be set according to the actual usage. For example, members of the same object group can be called a similar object subset that meets the preset conditions; or objects with the same preferences can be called a similar object subset that meets the preset conditions.

[0080] In some embodiments, the implementation of step S401 may include: performing the above step S401 after the target object clicks the "Data Recommendation" button in the data asset management platform.

[0081] Step S402: Obtain the data assets viewed by each similar object in the subset of similar objects and the data assets viewed by the target object;

[0082] Here, the implementation of step S402 can obtain the data assets viewed by each similar object in the subset of similar objects within a preset time period and the data assets viewed by the target object, in order to reduce the amount of computation.

[0083] Step S403: Based on the data assets viewed by each similar object and the data assets viewed by the target object, determine the data assets that the target object has not viewed among the data assets viewed by each similar object;

[0084] Here, the data assets that the target object has not viewed in the data assets viewed by each similar object are the data assets that the target object has not viewed, which is the data assets viewed by each similar object minus the data assets viewed by the target object. For example, if similar object A has viewed data assets b and c, and target object C has viewed data assets b and e, then the data asset that the target object has not viewed in the data assets viewed by the similar object is c.

[0085] Step S404: Recommend the data assets that the target object has not viewed from the data assets viewed by each similar object to the target object.

[0086] Here, we will take data asset c, which is not viewed by the target object, as an example to illustrate this. In this case, data asset c will be recommended to the target object.

[0087] In this embodiment, firstly, a subset of similar objects similar to the target object that issued the search request is determined through a knowledge graph among object groups, objects, business metadata assets, and business data assets; then, data assets that the target object has not viewed are identified among the data assets viewed by each similar object in the subset of similar objects; finally, the data assets that the target object has not viewed among the data assets viewed by each similar object are recommended to the target object, so that the target object can view the data assets viewed by other similar objects, thereby learning more about data assets related to itself.

[0088] In some embodiments, the data assets viewed by the target object include business metadata assets or business data assets, wherein the business metadata assets are associated with the business data assets. Correspondingly, after step S502 "obtain the data assets viewed by the target object", the following steps S601 and S602 are further included:

[0089] Step S601: Based on the knowledge graph between the object group, object, business metadata asset, and business data asset, and the data asset viewed by the target object, determine the first associated data asset, wherein the first associated data asset is associated with but different from the data asset viewed by the target object;

[0090] Here, since the data assets viewed by the target object include business metadata assets or business data assets, meaning the target object only viewed a portion of the business metadata assets and business data assets, the implementation of step S601 can derive a first associated data asset that is related to but different from the data assets viewed by the target object based on the knowledge graph between the object group, the object, the business metadata assets, and the business data assets. For example, if the data asset viewed by the target object is a business data asset, then the first associated data asset that is related to but different from the data asset viewed by the target object is the business metadata asset. As another example, if the data asset viewed by the target object is a business metadata asset, then the first associated data asset that is related to but different from the data asset viewed by the target object is the business data asset. When the business metadata assets include indicator metadata assets and report metadata assets, if the data asset viewed by the target object is an indicator metadata asset, then the first associated data asset that is related to but different from the data asset viewed by the target object can be either the business data asset or the report metadata asset.

[0091] Step S602: Recommend the first associated data asset to the target object.

[0092] In this embodiment of the application, a first associated data asset is determined that is related to but different from the data asset viewed by the target object through a knowledge graph among object groups, objects, business metadata assets, and business data assets. The first associated data asset is then recommended to the target object, so that the target object can see more related data assets and learn more about data assets in the same direction.

[0093] In some embodiments, after step S503 "determine the data assets that the target object has not viewed in the data assets viewed by each similar object", the following steps S5031 and S5032 are further included:

[0094] Step S5031: Based on the knowledge graph between the object group, the object, the business metadata asset, and the business data asset, and the data asset that the target object has not viewed in the data assets viewed by each similar object, determine the second associated data asset, wherein the second associated data asset is associated with but different from the data asset that the target object has not viewed in the data assets viewed by each similar object;

[0095] Here, similarly, if a data asset viewed by a similar object but not viewed by the target object is a business data asset, then a second associated data asset that is different from the target data asset viewed by the similar object is a business metadata asset. For example, if a data asset viewed by a similar object but not viewed by the target object is a business metadata asset, then a second associated data asset that is different from the target data asset viewed by the similar object is a business data asset. When business metadata assets include indicator metadata assets and report metadata assets, if a data asset viewed by a similar object but not viewed by the target object is an indicator metadata asset, then a second associated data asset that is different from the target data asset viewed by the similar object can be either a business data asset or a report metadata asset.

[0096] Step S5032: Recommend the second associated data asset to the target object.

[0097] In this embodiment of the application, a second related data asset is identified that is associated with and different from the target data asset that the target object has not viewed among the data assets viewed by similar objects. The second related data asset is then recommended to the target object, so that the target object can see more related data assets and learn more about data assets in the same direction.

[0098] In some embodiments, the implementation of step S501, "based on the knowledge graph between object groups, objects, business metadata assets, and business data assets, determining a subset of similar objects in the object set whose similarity to the target object that issued the search request meets preset conditions," may include the following steps S5011a to S5013a:

[0099] Step S5011a: Based on the knowledge graph between the object group, the object, the business metadata asset, and the business data asset, determine the object group to which the target object belongs;

[0100] Here, because the knowledge graph contains object groups and relationships between objects, the object group to which the target object belongs can be determined through the knowledge graph after the target object is obtained.

[0101] Step S5012a: Determine the other objects in the object group besides the target object;

[0102] Here, if the object group containing the target object includes three objects A, B, and C, and the target object is object A, then the other objects in the object group besides the target object include B and C.

[0103] Step S5013a: Determine the other objects as a subset of similar objects whose similarity to the target object meets a preset condition.

[0104] That is, objects B and C mentioned above are subsets of similar objects that meet the preset conditions for similarity with the target object.

[0105] In this embodiment of the application, objects belonging to the same object group are identified by a knowledge graph containing object groups and relationships between objects. Since objects in the same object group usually need to view the same or similar data, recommending the data assets viewed by other objects in the same object group other than the target object to the target object can achieve the goal of recommending more and more relevant data assets to the target object.

[0106] In some embodiments, the implementation of step S501, "based on the knowledge graph between object groups, objects, business metadata assets, and business data assets, determining a subset of similar objects in the object set whose similarity to the target object that issued the search request meets preset conditions," may include the following steps S5011b to S5014b:

[0107] Step S5011b: Obtain the data assets viewed by each object in the object set within a preset time period and their first quantity;

[0108] Here, the preset time period can be one month or three months, and can be set according to needs during implementation. The object set refers to all or some of the objects using the data asset management platform. The data assets viewed by each object within the preset time period and their initial quantity refer to all data assets viewed by each object and their quantity. For example, if object A views data assets a, b, c, d, and e, then the initial quantity of data assets viewed by object A is 5.

[0109] Step S5012b: For each object in the object set other than the target object, determine a second number of identical data assets viewed by the target object and each of the other objects within a preset time period;

[0110] Here, identical data assets refer to assets that are viewed repeatedly by the target object and other objects. For example, if target object C views data assets a, b, c, d, and e, and object B views data assets a, b, and c, then the identical data assets viewed by target object C and object B are a, b, and c, and the second quantity is 3.

[0111] Step S5013b: Determine the maximum value of the third number of data assets viewed by the target object and the first number of each of the other objects;

[0112] Here, if the data assets viewed by target object C are a, b, c, d, and e, then the third quantity is 5. If the data assets viewed by object B are a, b, and c, then the first quantity of object B is 3. Correspondingly, the maximum value of the third quantity and the first quantity is 5.

[0113] Step S5014b: For each object in the object set other than the target object, based on the second quantity and the maximum value, determine a subset of similar objects in the object set that meet the preset conditions for similarity with the target object.

[0114] Here, the preset condition can be the number of times a certain other object and the target object view data assets repeatedly (i.e., the second quantity) divided by the maximum number of times the other object and the target object view data assets. If the ratio is greater than a threshold, then the other object is considered a similar object whose similarity meets the preset condition. For example, if the second quantity of target object C and object B is 3, and the maximum value of target object C and object B is 5, then the ratio of the second quantity to the maximum value is 3 / 5. If the threshold is 50%, then this ratio is greater than the threshold, and object B is considered a similar object whose similarity meets the preset condition. For all other objects in the object set, the ratio of the second quantity to the maximum value is determined to determine whether other objects are similar objects whose similarity meets the preset condition, thus obtaining a subset of similar objects.

[0115] In this embodiment, by determining the second number of data assets that are the same in the data assets viewed by each other object and the target object, and the maximum value of the data assets viewed by each other object and the target object, and then determining a subset of similar objects in the object set that meet the preset conditions of similarity with the target object based on the ratio of the second number to the maximum value of each other object, a virtual similar object group is obtained based on the repetition rate of the data assets viewed by two objects. Then, the data assets viewed by other objects in the virtual similar object group other than the target object are recommended to the target object, thereby recommending more and more relevant data assets to the target object.

[0116] In some embodiments, since there is a lineage relationship between data assets, if a target object or a similar object views data asset A, upstream and downstream data assets in the data asset set that are related to data asset A can also be recommended to the target object.

[0117] A data asset management platform, based on big data processing and storage, centered on data asset management, and with data insights as its value proposition, transforms enterprise data from being difficult to understand and controllable to being manageable and operational. Building a data asset platform enables unified data processing, storage, authorization, and search, and provides full lifecycle management. It allows for precise data searches based on usage frequency, data asset value, and data relevance, making data asset information retrieval more intelligent. By matching users (the aforementioned objects), user attributes, and data information, it enables data to automatically find users who need the information, maximizing data value. This achieves data recommendation and data-driven user matching, helping users find, understand, and use the data they need more quickly, shifting from passive data retrieval to proactive data push, increasing user engagement.

[0118] This application embodiment performs full-text search on data assets, treating data as an asset and automatically maintaining relationships, generating models, and performing intelligent retrieval. It achieves high retrieval and hit rates for asset catalogs generated from big data and for fields such as lineage, tags, and metrics within those catalogs. Furthermore, it effectively prioritizes the retrieval of frequently searched Top N data.

[0119] This application's embodiments firstly address data asset management by achieving unified management of assets such as table assets, indicator assets, and report assets (i.e., the aforementioned data assets). It implements standardized data collection and access capabilities, as well as the visualization of data maps. It supports entity relationships based on data assets such as teams, individuals, reports, indicator names, and tables (i.e., the aforementioned knowledge graph), supports full-text search of asset data, and implements data recommendations to help better understand and use this asset data. Based on the usage popularity and data asset value of data dictionaries, indicator assets, table assets, and API assets in platforms such as data querying, offline scheduling, and Business Intelligence (BI) data analysis, it makes data searching more intelligent.

[0120] Part 1: Data-Driven Intelligent Search

[0121] Unified data asset management, enabling users to quickly retrieve the data assets they care about and clarify the relationships between data assets.

[0122] Data-driven intelligent search rules (using data assets as a table asset as an example):

[0123] Step 1: Search for keywords (keyword matching rules: table name (i.e., the name above) > field (i.e., the content above) > comment). Enter the search keywords, and the system will first match tables whose table names contain the keywords, then tables whose field names match the keywords, and finally tables whose table comments contain the search keywords.

[0124] Step 2: Sort by the specified database name (i.e., the database mentioned above), that is, sort the search results from Step 1 by database name.

[0125] Step 3, Value (i.e., the value of the first data asset mentioned above) ranking rules: Based on the frequently used tables (i.e., the frequency of use of the data assets mentioned above), the number of queries in the data query platform (i.e., the click popularity mentioned above), and the number of direct subordinate dependent tables of the table (i.e., the dependence of other data assets on the data asset mentioned above), the ranking table scores are calculated according to importance. Tables with higher scores have higher value and are searched first.

[0126] In implementation, different scores can be assigned to different situations based on their importance. For example, the scoring method is as follows:

[0127] (1) Commonly used tables score: 50 points

[0128] (2) Ranking score of the table most frequently queried in the data query platform: maximum 30 points (top 10: 30 points, top 11-top 30: 20 points, top 31-top 100: 10 points)

[0129] (3) The more direct subordinate tables a table has, the higher its ranking priority, up to a maximum of 20 points (1. Number of tables > 20, 20 points; 2. 10 < number of tables ≤ 20, 10 points; 3. 5 < number of tables ≤ 10, 5 points).

[0130] The score (i.e., value) of each table is obtained by adding up the scores of each item in each table.

[0131] Part Two: Data Recommendation and Asset Navigation from a Business Perspective

[0132] From a business perspective, relevant topic tables are recommended to enable data asset navigation and topic-based directory searching. Related tables are aggregated by business themes, providing users with data navigation through topic and sub-topic directories. Data assets are managed, identified, and shared from multiple perspectives, including classification, theme, and application. Data classification plays a significant role in data asset management, improving management efficiency and enhancing user experience by being business-value-driven.

[0133] Steps for building a data asset catalog:

[0134] (1) Data resource inventory, sort out the correspondence between tables and functions, and comprehensively sort out the access business systems and table data dictionaries;

[0135] (2) Data link analysis: Based on the data dictionary, sort out the data flow path and identify the data source and link relationship;

[0136] (3) Construct a data asset catalog and release it to various business departments. Based on the classification of the company's data resources, formulate catalog structure specifications; construct a data asset catalog covering subject areas, sub-subjects, tables, and fields, further establish the correspondence between business and data, and release the data asset catalog.

[0137] like Figure 5A As shown, the subject areas include: Digital Marketing, Sales Operations, Service Experience, Finance, and Human Resources. In the case of the Digital Marketing subject area, the sub-sub ...

[0138] Part Three: Data-Driven Intelligent Recommendations

[0139] like Figure 5B As shown, the implementation of this part mainly includes data collection (data from HIVE library, business library, etc.), building asset models (including table assets, report assets, etc.), building a knowledge graph (used to represent the relationships between users, between users and reports, etc.), and data recommendation (recommending table assets, reports, etc.). Through establishing business and data mapping relationships, tracing the source of business views, and tracing the source of data views, a data asset map is constructed, including business views, data views, and the mapping relationships between business and data views. This identifies the overall picture of data assets and their relationships, and the knowledge graph is built using the data asset map. The knowledge graph can be used to represent the relationships between entities, such as the relationships between users, between users and data assets, and between data assets. Based on the relationships between entities such as teams, people, reports, indicator names, and tables, a knowledge graph is constructed (where teams are the aforementioned object groups, people are the aforementioned objects, reports and indicator names are the aforementioned business metadata assets, and tables are the aforementioned business data assets). By utilizing users' preferences for the same data assets, the similarity between two users is calculated, and then the required data is recommended to the user, realizing "data finding people."

[0140] For example, users in the same group may share common data viewing preferences (i.e., other members of the same team share the same data viewing preferences as the user). In user group 1, users A, B, and C, knowing user A's data viewing preferences, can generate recommendations based on this, suggesting data assets that user A prefers to users B and C. Another example is constructing virtual user groups and making data recommendations based on these virtual user groups. By utilizing whether users have viewed the same data assets, the similarity between two users is calculated, and virtual user groups are constructed (i.e., creating virtual teams for users with high repetition rates, and pushing data viewed by other members of the virtual team to the user).

[0141] The specific steps are as follows:

[0142] Step 1: Construct a knowledge graph based on entity relationships such as team, people, reports, indicator names, and tables (e.g., ...). Figure 5C As shown, this knowledge graph displays the relationships between teams, people, reports, indicator names, and tables, forming relationships between users, between users and reports, and between reports and table assets.

[0143] Step 2: Calculate the similarity between users by leveraging their preferences for the same data assets;

[0144] Step 3: Based on this, generate recommendations, for example: recommend data assets preferred by user A to users B and C who have the same attributes.

[0145] Compared with related technologies, it has the following advantages:

[0146] 1. Based on the relationships between entities such as teams, people, reports, indicators, and tables, it enables data asset recommendation and data-driven people matching.

[0147] 2. Utilize data usage popularity and data asset value to support smarter search of data assets.

[0148] 3. Apply data lineage and business logic data relationships, and recommend upstream and downstream data in detail tables, summary tables, and tables.

[0149] 4. Recommend relevant topic tables from a business perspective to enable data asset navigation and searching by topic directory.

[0150] 5. It connects source data (business systems), data warehouse, and data applications, recording the entire process of data from generation to consumption, helping users gain insights into data and unlock its value.

[0151] 6. Unified data management makes it more convenient and centralized for users to view and use data.

[0152] Key Inventions:

[0153] 1. Entity relationships based on user groups composed of business departments are a common type of data recommendation based on entity relationships. Calculating user similarity by leveraging user preferences for the same data assets is an area that requires continuous system optimization; accumulating historical data can improve recommendation effectiveness.

[0154] 2. Based on data usage popularity and data asset value, make data search more intelligent and help users find the data they need faster.

[0155] Based on the foregoing embodiments, this application provides a data asset search device, which includes various units and modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0156] Figure 6 This is a schematic diagram of the composition structure of a data asset search device provided in an embodiment of this application, as shown below. Figure 6 As shown, the data asset search device 600 includes: a first determining module 610, a second determining module 620, a filtering and sorting module 630, and a first sorting module 640, wherein:

[0157] The first determining module 610 is used to determine the position of the keywords carried in the search request in each data asset of the data asset set;

[0158] The second determining module 620 is used to determine the priority of the keyword's position in the data asset;

[0159] The filtering and sorting module 630 is used to filter and sort the data assets in the data asset set based on the position of the keyword in each data asset and the priority of the position of the keyword in the data asset, so as to obtain a first data asset set containing the keyword with a first sort.

[0160] The first sorting module 640 is used to sort data assets with the same keyword position in the first data asset set based on the value of each first data asset in the first data asset set, to obtain a second data asset set with a second sort, and output the second data asset set with the second sort, wherein the value of the first data asset is used to represent at least one of the following: the usage frequency of the data asset, the click popularity, the dependence of other data assets on the data asset, and the topic relevance.

[0161] In some embodiments, the apparatus further includes: a second sorting module, configured to, before sorting data assets in the first data asset set whose keywords are in the same position in the data assets based on the value of each first data asset in the first data asset set, perform a first sorting based on the priority of the database to which each first data asset belongs in the first data asset set, to obtain a third data asset set with a third sorting; a third determining module, configured to use the third data asset set as the first data asset set; correspondingly, the first sorting module 640 is further configured to, based on the value of each third data asset in the third data asset set, perform a second sorting based on the value of each third data asset in the third data asset set, to obtain the second data asset set with the second sorting.

[0162] In some embodiments, the position of the keyword in the data assets includes: name, content, and annotation. The filtering and sorting module includes: a filtering submodule, used to filter out a target data asset set containing the keyword in the data asset set based on the position of the keyword in each data asset; and a sorting submodule, used to sort the data assets in the target data asset set according to the priority of the keyword's position in the data assets from high to low, namely name, content, and annotation, to obtain a first data asset set containing the keyword with a first sort.

[0163] In some embodiments, the value of the first data asset is used to characterize the usage frequency, click popularity, and dependence of other data assets on the data asset. Correspondingly, the device further includes: a first acquisition module, used to acquire the usage frequency level, query volume ranking, and number of directly dependent data assets for each first data asset in the first data asset set; a fourth determination module, used to determine the usage frequency score, popularity score, and dependence score of the corresponding data asset based on the usage frequency level, query volume ranking, and number of directly dependent data assets of each first data asset; and a fifth determination module, used to determine the value of the corresponding first data asset based on the usage frequency score, popularity score, and dependence score of each first data asset.

[0164] In some embodiments, the value of the first data asset is used to characterize the relevance of a topic, and the apparatus further includes: a second acquisition module, configured to acquire the topic relevance of each first data asset in the first data asset set; a sixth determination module, configured to determine the topic relevance score of a corresponding data asset based on the topic relevance of each first data asset; and a seventh determination module, configured to determine the value of a corresponding first data asset based on the topic relevance score of each first data asset.

[0165] In some embodiments, the apparatus further includes: an eighth determining module, configured to, in response to a data recommendation request, determine a subset of similar objects in the object set whose similarity to the target object that issued the search request meets a preset condition, based on a knowledge graph among object groups, objects, business metadata assets, and business data assets, wherein the object set includes objects from at least two of the object groups; a third obtaining module, configured to obtain data assets viewed by each similar object in the subset of similar objects and data assets viewed by the target object; a ninth determining module, configured to, based on the data assets viewed by each similar object and data assets viewed by the target object, determine data assets in the data assets viewed by each similar object that the target object has not viewed; and a first recommending module, configured to recommend the data assets in the data assets viewed by each similar object that the target object has not viewed to the target object.

[0166] In some embodiments, the data assets viewed by the target object include business metadata assets or business data assets, the business metadata assets being associated with the business data assets. The apparatus further includes: a tenth determining module, configured to, after obtaining the data assets viewed by the target object, determine a first associated data asset based on the knowledge graph between the object group, the object, the business metadata assets, and the business data assets, and the data assets viewed by the target object, wherein the first associated data asset is associated with but different from the data assets viewed by the target object; and a second recommending module, configured to recommend the first associated data asset to the target object.

[0167] In some embodiments, the apparatus further includes: an eleventh determining module, configured to determine a second associated data asset based on the knowledge graph between the object group, the object, the business metadata asset, and the business data asset, and the data asset not viewed by the target object among the data assets viewed by each similar object, wherein the second associated data asset is associated with but different from the data asset not viewed by the target object among the data assets viewed by each similar object; and a third recommending module, configured to recommend the second associated data asset to the target object.

[0168] In some embodiments, the eighth determining module includes: a first determining submodule, used to determine the object group to which the target object belongs based on the knowledge graph between the object group, the object, the business metadata asset, and the business data asset; a second determining submodule, used to determine other objects in the object group besides the target object; and a third determining submodule, used to determine the other objects as a subset of similar objects whose similarity to the target object meets a preset condition.

[0169] In some embodiments, the eighth determining module includes: an acquisition submodule, configured to acquire data assets viewed by each object in the object set within a preset time period and their first quantity; a fourth determining submodule, configured to determine, for each object in the object set other than the target object, a second quantity of the same data assets viewed by the target object and each of the other objects within the preset time period; a fifth determining submodule, configured to determine the maximum value between a third quantity of data assets viewed by the target object and the first quantity of each of the other objects; and a sixth determining submodule, configured to, for each object in the object set other than the target object, based on the second quantity and the maximum value, determine a subset of similar objects in the object set whose similarity to the target object meets a preset condition.

[0170] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this disclosure can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0171] It should be noted that, in the embodiments of this application, if the above-mentioned data asset search method and recommendation method are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0172] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0173] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.

[0174] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0175] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0176] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0177] It should be noted that, Figure 7 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this application, such as... Figure 7 As shown, the hardware entity of the computer device 700 includes: a processor 701, a communication interface 702, and a memory 703, wherein:

[0178] Processor 701 typically controls the overall operation of computer device 700.

[0179] Communication interface 702 enables computer devices to communicate with other terminals or servers over a network.

[0180] The memory 703 is configured to store instructions and applications executable by the processor 701, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 701 and various modules in the computer device 700. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 701, the communication interface 702, and the memory 703 can be performed via bus 704.

[0181] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0182] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0183] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0184] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0185] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0186] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0187] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0188] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for searching data assets, characterized in that, include: Determine the location of the keywords carried in the search request within each data asset in the data asset collection; wherein, the location within the data asset includes: the name, content, and annotations of the data asset; Determine the priority of the keyword's position in the data asset; Based on the position of the keyword in each data asset and the priority of the position of the keyword in the data asset, the data assets in the data asset set are filtered and sorted to obtain a first data asset set containing the keyword and having a first sort. Based on the value of each first data asset in the first data asset set, the data assets with the same keyword position in the data assets in the first data asset set are sorted to obtain a second data asset set with a second sort. The second data asset set with the second sort is output, wherein the value of the first data asset is used to represent at least one of the following: the usage frequency of the data asset, the click popularity, the dependence of other data assets on the data asset, and the topic relevance.

2. The method according to claim 1, characterized in that, Before sorting data assets with the same keyword position in the first data asset set based on the value of each first data asset in the first data asset set, the method further includes: Based on the priority of the database to which each first data asset belongs in the first data asset set, the data assets with the same keyword position in the data assets in the first data asset set are sorted for the first time to obtain the third data asset set with the third sort. The third data asset set is used as the first data asset set; Correspondingly, the step of sorting data assets with the same keyword position in the first data asset set based on the value of each first data asset in the first data asset set to obtain a second data asset set with a second sort includes: Based on the value of each third data asset in the third data asset set, adjacent data assets belonging to the same database in the third data asset set are sorted a second time to obtain the second data asset set with the second sort.

3. The method according to claim 1 or 2, characterized in that, The location of the keyword in the data assets includes: name, content, and annotation. Based on the location of the keyword in each data asset and the priority of that location, the data assets in the data asset set are filtered and sorted to obtain a first data asset set containing the keyword and having a first ranking, including: Based on the position of the keyword in each data asset, a target data asset set containing the keyword is selected from the data asset set. Based on the priority of the keyword in the data assets, from high to low, the data assets in the target data asset set are sorted to obtain a first data asset set containing the keyword and having a first sort.

4. The method according to claim 1 or 2, characterized in that, The value of the first data asset is used to characterize the frequency of use, click popularity, and dependence of other data assets on the data asset. Correspondingly, the method further includes: Obtain the usage frequency level, query volume ranking, and number of directly dependent data assets for each first data asset in the first data asset set; Based on the usage frequency level, query volume ranking, and number of directly dependent data assets of each first data asset, the usage frequency score, popularity score, and dependency score of the corresponding data asset are determined respectively. The value of each first data asset is determined based on its usage frequency score, popularity score, and dependence score.

5. The method according to claim 3, characterized in that, The value of the first data asset is used to characterize the relevance of the topic, and the method further includes: Obtain the topic relevance of each first data asset in the first data asset set; Based on the topic relevance of each first data asset, a topic relevance score for the corresponding data asset is determined. The value of each first data asset is determined based on its topic relevance score.

6. The method according to claim 1 or 2, characterized in that, Also includes: In response to a data recommendation request, based on the knowledge graph between object groups, objects, business metadata assets, and business data assets, a subset of similar objects in the object set that meet preset conditions in similarity to the target object that issued the search request are determined, wherein the object set includes objects from at least two of the object groups; Obtain the data assets viewed by each similar object in the subset of similar objects and the data assets viewed by the target object; Based on the data assets viewed by each similar object and the data assets viewed by the target object, determine the data assets that the target object has not viewed among the data assets viewed by each similar object; The target object is recommended data assets that it has not viewed among the data assets viewed by each similar object.

7. The method according to claim 6, wherein the data asset viewed by the target object includes business metadata assets or business data assets, the business metadata assets being associated with the business data assets, and correspondingly, after obtaining the data asset viewed by the target object, the method further includes: Based on the knowledge graph between the object group, object, business metadata asset, and business data asset, and the data asset viewed by the target object, a first associated data asset is determined, wherein the first associated data asset is associated with but different from the data asset viewed by the target object; The first associated data asset is recommended to the target object.

8. The method according to claim 6, characterized in that, After determining the data assets that the target object has not viewed among the data assets viewed by each similar object, the method further includes: Based on the knowledge graph between the object group, object, business metadata assets, and business data assets, and the data assets that the target object has not viewed in the data assets viewed by each similar object, a second associated data asset is determined, wherein the second associated data asset is associated with but different from the data assets that the target object has not viewed in the data assets viewed by each similar object; The second associated data asset is recommended to the target object.

9. The method according to claim 6, characterized in that, The knowledge graph based on object groups, objects, business metadata assets, and business data assets determines a subset of similar objects in the object set whose similarity to the target object that issued the search request meets preset conditions, including: Based on the knowledge graph between the object groups, objects, business metadata assets, and business data assets, the object group to which the target object belongs is determined; Identify the other objects in the object group besides the target object; The other objects are identified as a subset of similar objects whose similarity to the target object meets a preset condition.

10. The method according to claim 6, characterized in that, The knowledge graph based on object groups, objects, business metadata assets, and business data assets determines a subset of similar objects in the object set whose similarity to the target object that issued the search request meets preset conditions, including: Obtain the data assets viewed by each object in the object set within a preset time period and their first quantity; For each object in the object set other than the target object, determine a second number of identical data assets that the target object and each of the other objects viewed within a preset time period. Determine the maximum value between the third number of data assets viewed by the target object and the first number of each of the other objects; For each object in the object set other than the target object, a subset of similar objects in the object set that meet the preset conditions for similarity with the target object is determined based on the second quantity and the maximum value.

Citation Information

Patent Citations

  • Digital asset search user interface

    US20190340252A1