Knowledge graph construction method and device, data query method and device, electronic equipment and storage medium

By establishing a knowledge graph in the data warehouse, calculating the similarity between tables and user behavior similarity, it solves the problem that existing search engines find it difficult to discover indirect related data tables, and achieves more accurate and relevant search results.

CN119940503APending Publication Date: 2025-05-06中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510025867.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing search engines have difficulty discovering data tables hidden in data warehouses that are indirectly related to search keywords, causing analysts to miss important non-directly relevant data resources when looking for data.

Method used

By establishing a knowledge graph, calculate the similarity between tables of all data tables in the library table and the matching similarity between user behavior and data tables, build a multi-dimensional search index to support users to search in multiple dimensions.

Benefits of technology

Improves the accuracy and relevance of search results, helping users discover data resources that are indirectly related to search keywords.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940503A_ABST
    Figure CN119940503A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method and device, a data query method and device, electronic equipment and a storage medium, and the knowledge graph construction method comprises the steps: calculating the inter-table similarity of all data tables in a library table according to a first similarity calculation rule; calculating the matching similarity between the user behavior and the data table according to a second similarity calculation rule; according to the inter-table similarity of all the data tables and the matching similarity of the user behaviors and the data tables, a knowledge graph is established, the first similarity calculation rule calculates the similarity based on table attributes, and the second similarity calculation rule calculates the similarity based on actual use of the user. On one hand, the knowledge graph is established based on the similarity between the tables, and on the other hand, the relationship between the entities in the knowledge graph is optimized and updated according to each query result. The invention further provides a data query method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of knowledge graphs and query engines, and in particular to a knowledge graph construction method, a data query method, a device and an electronic device, and a storage medium. Background Art

[0002] A massive amount of data tables are stored in the data warehouse, and there may be complex associations and dependencies between these tables.

[0003] Search engines mainly rely on keyword matching, and analysts find it difficult to find tables hidden in the data warehouse that are indirectly related to the search keywords. This causes analysts to miss some important but not directly related data resources when searching for data. Summary of the invention

[0004] The embodiments of the present application provide a knowledge graph construction method, a data query method, an apparatus and an electronic device, and a storage medium to improve the accuracy and relevance of search results by establishing a knowledge graph.

[0005] The present application embodiment adopts the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a knowledge graph construction method, wherein the construction method comprises:

[0007] According to the first similarity calculation rule, calculate the inter-table similarity of all data tables in the library table;

[0008] According to the second similarity calculation rule, the matching similarity between the user behavior and the data table is calculated;

[0009] A knowledge graph is established based on the similarity between all the data tables and the matching similarity between the user behavior and the data tables.

[0010] The first similarity calculation rule calculates the similarity based on table attributes, and the second similarity calculation rule calculates the similarity based on actual use by users.

[0011] In some embodiments, the establishing of the knowledge graph according to the inter-table similarity of all the data tables and the matching similarity between the user behavior and the data tables includes:

[0012] If the similarity between any two data tables is within a preset threshold range according to the inter-table similarity of all the data tables, it is determined whether there is a relationship with the knowledge graph triple table;

[0013] If yes, update the weight to the similarity between any two data tables;

[0014] If the similarity between the user's actual use and the data table is within a preset range according to the matching similarity between the user behavior and the data table, it is determined whether there is a relationship with the knowledge graph triple table;

[0015] If yes, the update weight is the similarity between the user's actual usage and the data table;

[0016] The similarity is used as the weight of the relationship between the two data tables and is used to evaluate the degree of association between any two data tables.

[0017] In some embodiments, the method further comprises:

[0018] Determine the relationship between the data tables according to the inter-table similarities of all the data tables;

[0019] Based on the matching similarity between the user behavior and the data table, the business scenarios and business fields to which the data table is applicable, or the actual application in a specific business scenario or business field, are determined.

[0020] In some embodiments, calculating the inter-table similarity of all data tables in the library table according to the first similarity calculation rule includes:

[0021] Calculate similarity based on table attributes, using sim tf (A,B) represents the field similarity value between table A and table B;

[0022] Calculate the word vector A for each field in table A i And the word vector B of each word in table B j The cosine similarity of

[0023]

[0024] Among them A i ·B j represents their dot product, |A i ||B j | represents vector A i and vector B j mold;

[0025] For each field vector A in table A i Find the field word vector B with the largest cosine similarity with table B j ,

[0026] sim f (A i )=max(sim f (A i ,B j ))

[0027] Calculate all simf (A i ) The average value is used as the similarity between Table A and Table B

[0028]

[0029] Among them, Table A has a fields and Table B has b fields.

[0030] In some embodiments, the method further includes:

[0031] Assume that the number of blood relationship links between Table A and Table B is q, C represents the degree of kinship between Table A and Table B, and the value range of C is 0 < C ≤ 1. The larger C is, the farther the blood relationship is. When C = 1, it means there is no blood relationship. Then the calculation formula for the degree of blood relationship is as follows:

[0032] C = 1 - e -k·q

[0033] where k is a constant, k > 0, which is used to adjust the curve shape and determines the growth rate of the curve;

[0034] Consider the matching degree of multiple factors such as the business field, business scenario, and labels between Table A and Table B, and adjust the contribution of each dimension to the final similarity metric value according to the importance of each dimension. Assume that the weighted average S1, S 2, S3...S q represents the matching degree scores of q different factors, w1, w 2, w3...w q represents the corresponding weights, and the score of each factor is represented by S j means,

[0035]

[0036] where 0 ≤ S j ≤ 1, 0 means complete match, and 1 means complete mismatch;

[0037] 0 ≤ D ≤ 1, the closer D is to 0, the higher the matching degree. When D is equal to 1, it means that Table A and Table B have no correlation in these features;

[0038] Finally, the similarity between Table A and Table B is obtained:

[0039] sim t (A,B) = sim tf (A,B) × C × D.

[0040] In some embodiments, calculating the matching similarity between user behavior and the data table according to the second similarity calculation rule includes:

[0041] According to the second similarity calculation rule, the usage frequency of the data table is converted into a similarity index with the business scenario or business field;

[0042] Assuming W represents the business scenario or business field, the degree of match between Table A and W in the graph is

[0043]

[0044] Where m means that table A is used in m analysis tasks for W, and α is a positive number used to control the effect of m on sim u (A,W) sensitivity of impact;

[0045] The sim u The value range of (A,W) is 0≤sim u (A,W)≤1;

[0046] When table A is pre-marked with the applicable business scenario or business domain W, sim u (A,W)=0.

[0047] In a second aspect, an embodiment of the present application further provides a data query method, which is applied to the knowledge graph construction method as described in the first aspect, and the query method includes:

[0048] In response to the user's query request, a search is performed in the knowledge graph based on a preset shortest path calculation algorithm to find a data table with the shortest distance in the link that meets the requirements and return it as a query result.

[0049] In a third aspect, an embodiment of the present application further provides a knowledge graph construction device, wherein the device comprises:

[0050] A first similarity calculation module is used to calculate the inter-table similarity of all data tables in the library table according to a first similarity calculation rule;

[0051] A second similarity calculation module is used to calculate the matching similarity between the user behavior and the data table according to a second similarity calculation rule;

[0052] The knowledge graph building module is used to build a knowledge graph based on the similarity between all the data tables and the matching similarity between the user behavior and the data tables.

[0053] The first similarity calculation rule calculates the similarity based on table attributes, and the second similarity calculation rule calculates the similarity based on actual use by users.

[0054] In a fourth aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer executable instructions, wherein the executable instructions, when executed, cause the processor to perform the above method.

[0055] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes the above method.

[0056] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: according to the first similarity calculation rule, the inter-table similarity of all data tables in the library table is calculated; according to the second similarity calculation rule, the matching similarity between the user behavior and the data table is calculated. Thus, a knowledge graph is established based on the inter-table similarity of all data tables and the matching similarity between the user behavior and the data table. Since the first similarity calculation rule calculates the similarity based on the table attributes, basic inter-table similarity information can be obtained. Since the second similarity calculation rule calculates the similarity based on the actual use of the user, similarity information related to the user behavior can be obtained. The knowledge graph obtained by the above method incorporates multiple features such as business fields, business scenarios, field attributes, tags, etc., supports users to conduct multi-dimensional searches, thereby further improving the accuracy and relevance of search results. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0058] Figure 1 This is a schematic diagram of the process of constructing a knowledge graph in an embodiment of the present application;

[0059] Figure 2 This is a schematic diagram of the structure of the knowledge graph construction device in the embodiment of the present application;

[0060] Figure 3 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0062] The technical terms involved in the embodiments of this application are as follows:

[0063] Knowledge graph: Knowledge graph is a technology that displays a large amount of knowledge and its relationships in a graphical way. It organizes knowledge in the form of "entity-relationship-entity" triples and is widely used in intelligent search, text analysis and other fields.

[0064] Entity: In the knowledge graph, an entity refers to a concrete or abstract thing that is distinguishable and exists independently.

[0065] Entity recognition: Entity recognition is a key step in building a knowledge graph. Its purpose is to identify entities with specific meanings from unstructured text data, such as names of people, places, institutions, dates and times, proper nouns, etc.

[0066] Attributes: Attributes in a knowledge graph are information used to describe the characteristics and properties of an entity, usually in the form of a key-value pair, where the key is the name of the attribute and the value is the specific value of the attribute. Attributes can enrich the description of an entity, help distinguish different entities, and support complex query and analysis operations.

[0067] Relationship: In the knowledge graph, a relationship describes the connection or interaction between entities, such as closeness, subordination, spatial relationship, and temporal relationship. Relationships define the connection method and direction between entities, which helps to understand the association between entities.

[0068] Relation extraction: Relation extraction is the process of automatically extracting semantic relationships between entities from text, such as the closeness of tables, and expressing them in a structured form.

[0069] Triple: In the knowledge graph, a triple is the basic unit for describing the relationship between entities, consisting of two entities and a relationship between them.

[0070] Graph database: Graph database is a new type of database based on graph theory. It stores and queries data in a graph structure consisting of nodes and edges, and can efficiently process complex related queries. It is suitable for scenarios such as social networks, recommendation systems, and bioinformatics, and supports distributed storage and computing of large-scale data.

[0071] Word vector: Word vector is a way to convert words in a text into numerical vectors. It captures the semantic information of words and makes similar words similar in the vector space. It is often used in text analysis in natural language processing and machine learning.

[0072] Word2Vec: Word2Vec is a word vector training tool launched by Google in 2013. It converts words in a text into vectors of fixed length through a neural network model, which can represent the semantic relationship between words. Word2Vec is often used in natural language processing tasks, such as information retrieval, sentiment analysis, and recommendation systems.

[0073] After analysis, it is known that due to the complexity of data warehouses and the hidden relationships between data tables, existing search engines often find it difficult to discover hidden correlations, resulting in incomplete and inaccurate search results. Existing search methods mainly rely on keyword matching, ignoring the importance of user behavior in the search process, lack of analysis of user behavior, and unable to provide personalized search results.

[0074] Since existing recommendation algorithms mainly focus on mining users' potential interests and matching items, they have been widely used in e-commerce, social networks and other fields, but their application in database table recommendation is still in the exploratory stage. Database tables have specific data structures and contents, which are significantly different from conventional items. Current recommendation algorithms are not well adapted. Therefore, it is necessary to develop specialized recommendation algorithms based on the characteristics of database tables. At the same time, in data warehouses, due to the huge amount of data and complex table structure, traditional recommendation technologies have high overhead for single recommendation and cannot meet users' needs for rapid response.

[0075] That is to say, although recommendation algorithms perform well in mining users' potential interests and matching items' potential associations, they still fail to adapt well to the field of database table recommendation. The data structure and content of database tables are fundamentally different from other items, requiring more accurate and professional recommendation strategies. At the same time, traditional recommendation technologies usually face the cold start problem, and it is difficult to make effective recommendations for new users or new data that lack historical behavior records. In existing recommendation systems, in order to find a database table that matches user needs, a large amount of feature calculation and similarity comparison is required, which is time-consuming and does not meet the temporary characteristics of the analysis task.

[0076] In view of the above shortcomings, a knowledge graph construction method is provided in the embodiment of the present application, which is based on word vector and knowledge graph technology to quickly locate the associated database table. By applying the model and establishing the formula, the similarity between different database tables is calculated, the association relationship between the user and the database table is established, and the search process is optimized.

[0077] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0078] The present application embodiment provides a method for constructing a knowledge graph, such as Figure 1 As shown, a schematic diagram of the knowledge graph construction method in an embodiment of the present application is provided, and the method at least includes the following steps S110 to S130:

[0079] Step S110, calculating the inter-table similarities of all data tables in the library table according to the first similarity calculation rule.

[0080] In the process of constructing the knowledge graph, similarity is used as the basis for judging whether there is a relationship between entities, and the similarity value is the interval distance between entities.

[0081] The process of building a knowledge graph includes: table feature data collection. Specifically, based on the data warehouse data dictionary, extract the Chinese name and English name of the library table, as well as all the Chinese field names, English field names, business domain attribute labels, business scenario attribute labels, other annotated labels, and each table on the current table's lineage network link to prepare for subsequent processing.

[0082] It should be noted that the "business domain" and "business scenario" involved in the embodiments of this application all follow a standardized unified classification system, which is based on a core premise that all database tables of the relevant system and the analysis tasks launched by analysts are labeled with their business scenarios and business domains. Table labels are manually or otherwise annotated to indicate attributes such as the category or nature of the table.

[0083] The process of building a knowledge graph also includes: data processing. Specifically, when processing database table field data, you will usually encounter some common fields, such as ID, date, and status. These fields exist in almost all tables, but do not reflect specific business attributes. In order to focus more on business-related data analysis, common fields need to be removed.

[0084] Therefore, when processing data, it is preferred to automatically identify and remove common fields such as ID, date and status through preset filtering rules to ensure that only fields that truly reflect business characteristics are focused on during the analysis process, eliminating interference from irrelevant fields, thereby improving the efficiency and accuracy of data processing.

[0085] The process of building a knowledge graph also includes: entity and relationship definition, defining the entities of the knowledge graph and the relationships between them based on the semantics of the data warehouse and the business and efficiency needs of the search scenario.

[0086] Among them, entities include but are not limited to tables, table fields, table business scenarios, table business areas, and table tags.

[0087] The definition of relationship covers several aspects: the relationship between tables is "table-[similar]-table", which is determined by the first similarity calculation rule. The corresponding relationship between tables and fields is "table-[includes]-field", which describes which fields the table contains; "table-[used for]-business scenario".

[0088] It can be understood that similarity is used as the weight of the table-table relationship to evaluate the degree of association between two tables.

[0089] The process of building a knowledge graph also includes: building knowledge graph triples, processing the prepared data into a series of entity-relationship-entity triples according to the entity and relationship rules defined in the entity and relationship definitions, forming a structured triple list, and initializing the relationship weight to 0.

[0090] Step S120: Calculate the matching similarity between the user behavior and the data table according to the second similarity calculation rule.

[0091] The definition of the relationship also includes: "table-[used for]-business domain", indicating the business scenarios and business domains to which the table is applicable, or the actual application in a specific business scenario or business domain, which is determined by pre-labeling or analyzing user behavior and using the second similarity calculation rule. The relationship between a table and a tag is "table-[labeled as]-tag", which is used to describe the attributes and characteristics of the table.

[0092] In order to expand the user's vision, optimize the experience, and integrate user behavior in search, under the premise of unifying the classification of business scenarios and business fields, statistics are collected on which tables are used more frequently in each business scenario and business field. In order to unify the caliber, the frequency of use is converted into the similarity between the table and the business scenario and business field. To ensure uniformity, the frequency of use of the table is converted into a similarity index with the business scenario and business field as the second similarity calculation rule.

[0093] Step S130, establishing a knowledge graph based on the inter-table similarity of all the data tables and the matching similarity between the user behavior and the data tables, wherein the first similarity calculation rule calculates the similarity based on table attributes, and the second similarity calculation rule calculates the similarity based on actual user use.

[0094] Before calculating the similarity between tables, you need to calculate the table field vector value. By starting and loading the pre-trained Word2Vec model into memory, traverse each field name of each table and use it as the input of the model. For each field name, the model outputs a vector value. In this way, the field name in the table is converted into a mathematical vector to represent the semantic information of the field, providing support for subsequent similarity calculations.

[0095] Furthermore, according to the inter-table similarities of all the data tables and the matching similarities between the user behavior and the data tables, after the triples are constructed, they are imported into the graph database to generate a knowledge graph.

[0096] Through the above method, a table similarity calculation method based on a neural network word vector model is adopted, which comprehensively considers the similarity of field attributes between tables, as well as multi-dimensional attributes such as blood relationship, business field, applicable business scenarios and labels, to construct a more accurate table-table similarity model.

[0097] Through the above method, by integrating user behavior and item similarity (similarity first calculation rule), a similarity evaluation system is constructed, which not only helps users expand their horizons and enrich search content, but also improves the accuracy and quality of search results.

[0098] Through the above method, by analyzing user behavior (based on the second similarity calculation rule), the recognition of implicit relationships is further strengthened. Based on the user's table usage behavior in analysis tasks, a model of the relationship between business domains, business scenarios and tables is constructed to improve the link structure of the knowledge graph.

[0099] Different from the related technologies, it mainly relies on keyword matching, ignores the importance of user behavior in the search process, lacks analysis of user behavior, and cannot provide personalized search results.

[0100] In one embodiment of the present application, the knowledge graph is established based on the inter-table similarity of all the data tables and the matching similarity of the user behavior and the data tables, including: if the similarity between any two data tables obtained based on the inter-table similarity of all the data tables is within a preset threshold range, determine whether there is a relationship with the knowledge graph triple table; if so, update the weight to the similarity between any two data tables; if the similarity between the user's actual use and the data table obtained based on the matching similarity of the user behavior and the data table is within a preset range of values, determine whether there is a relationship with the knowledge graph triple table; if so, update the weight to the similarity between the user's actual use and the data table; wherein, the similarity is used as the weight of the relationship between the two data tables to evaluate the degree of association between any two data tables.

[0101] The knowledge graph is constructed based on similarity, the weight and similarity between tables. In subsequent searches, the shortest path method is used. The input indicates that according to the knowledge graph, the table with the shortest distance in the link is found and returned to the user as the recommended search table.

[0102] According to the above similarity calculation rules, the similarity between all tables is calculated. When the similarity between two tables is 0 <sim t When (A,B)<1, determine whether there is a relationship in the constructed triple table. If so, update the weight to sim t (A, B), if not, add a set of relationships and record the weights. Similarly, update the weights of the relationships between business scenarios or business areas and tables to sim u (A, W), which is convenient for subsequent shortest path calculation.

[0103] sim t (A,B)=sim tf(A,B)×C×D represents the formula for calculating the inter-table similarity of all data tables in the library table according to the first similarity calculation rule. It not only includes the similarity calculation formula between tables but also takes into account c blood relationship and D matching degree.

[0104] In one embodiment of the present application, the method further includes: determining the relationship between data tables based on the inter-table similarity of all data tables; determining the business scenarios and business fields to which the data tables are applicable, or their practical applications in specific business scenarios or business fields, based on the matching similarity between the user behavior and the data tables.

[0105] It means that the matching similarity between user behavior and data table is calculated according to the second similarity calculation rule, mainly considering the user's operation behavior on the table, or historical operation behavior, or actual application in specific business scenarios or business fields.

[0106] In one embodiment of the present application, the similarity between all data tables in the library table is calculated according to the first similarity calculation rule, including: calculating similarity based on table attributes, using sim tf (A,B) represents the field similarity value between table A and table B; calculate the word vector A for each field in table A i And the word vector B of each word in table B j The cosine similarity of

[0107]

[0108] Among them A i ·B j represents their dot product, |A i ||B j | represents vector A i and vector B j mold;

[0109] For each field vector A in table A i Find the field word vector B with the largest cosine similarity with table B j ,

[0110] sim f (A i )=max(sim f (A i ,B j ))

[0111] Calculate all sim f (A i ) as the similarity between Table A and Table B

[0112]

[0113] Among them, table A has a fields and table B has b fields.

[0114] In an embodiment of the present application, the method further includes: assuming that the number of blood relationship links between table A and table B is q, and C represents the degree of kinship between table A and table B. The value range of C is 0 < C ≤ 1. The larger C is, the farther the blood relationship is. When C = 1, it means there is no blood relationship. Then the calculation formula for the degree of kinship is as follows:

[0115] C = 1 - e -k·q

[0116] where k is a constant, k > 0, which is used to adjust the curve shape and determines the growth rate of the curve;

[0117] Consider the matching degree of table A and table B in multiple factors such as business domain, business scenario, and label, and adjust their contributions to the final similarity metric value according to the importance of each dimension. Assume that the weighted average S1, S 2, S3... S q represents the matching degree scores of q different factors, and w1, w 2, w3... w q represents the corresponding weights. Each factor score is represented by S j represent,

[0118]

[0119] where 0 ≤ S j ≤ 1, 0 means complete match, 1 means complete non - match; 0 ≤ D ≤ 1. The closer D is to 0, the higher the matching degree. When D equals 1, it means that table A and table B have no relevant relationship in these features; Finally, the similarity between table A and table B is obtained:

[0120] sim t (A, B) = sim tf (A, B) × C × D.

[0121] Specifically, in the embodiment of the present application, combining two recommendation methods, i.e., item - based recommendation and user - behavior - based recommendation, and aiming at the particularity of the table recommendation scenario, a similarity calculation model suitable for table recommendation is designed. At the same time, considering table attributes and actual user applications to construct the similarity between tables, two similarity calculation formulas based on table attributes and based on actual user usage are designed respectively.

[0122] First, for calculating the similarity based on table attributes, use sim tf (A, B) to represent the field similarity value between table A and table B. Assume that table A has a fields and table B has b fields. The calculation method is as follows:

[0123] For the word vector A of each field in Table A i , calculate its cosine similarity sim with the word vector B of each word in Table B j . The formula is as follows: f (A i , B j )

[0124]

[0125] where A i ·B j represents their dot product, and |A i ||B j | represents the norms of vector A i and vector B j .

[0126] For each field vector A in Table A i , find the field word vector B with the maximum cosine similarity to Table B j , that is:

[0127] sim f (A i ) = max(sim f (A i , B j ))

[0128] Calculate the average value of all sim f (A i ) as the similarity between Table A and Table B, that is:

[0129]

[0130] It should be noted that in order to more objectively reflect the similarity between the two tables, in the embodiments of this application, on the basis of analyzing the similarity of table fields, the matching degree of features such as the blood relationship, business domain, business scenario, and labels between the tables is further analyzed to make the similarity more accurate.

[0131] Assume that the number of blood relationship links between Table A and Table B is q, and C represents the degree of closeness of the blood relationship between Table A and Table B. The value range of C is 0 < C ≤ 1. The larger C is, the farther the blood relationship is. When C = 1, it means there is no blood relationship. The formula for the degree of closeness of the blood relationship is as follows:

[0132] C = 1 - e -k·q

[0133] where k is a constant, k > 0, used to adjust the curve shape and determine the growth rate of the curve. The larger k is, the faster C approaches 1, that is, the more sensitive it is to the link length.

[0134] It should be noted that k is used to adjust the impact of blood relationship on the recommended content: the smaller k is, the more distant blood relationship tables will be considered when making recommendations; the larger k is, the less dependent on blood relationship, and more consideration will be given to other recommendation factors.

[0135] In order to make the recommendation more objective, the embodiment of the present application flexibly integrates information from multiple dimensions, comprehensively considers the matching degree of Table A and Table B in terms of business field, business scenario, label and other factors, and adjusts the contribution of each dimension to the final similarity measurement value according to its importance. Assume that D represents the weighted average of the matching degree scores S1,S 2, S3...S q represents the matching scores of q different factors, w1,w 2, w3...w q Indicates the corresponding weight, and each factor score is represented by S j The calculation formula is as follows:

[0136]

[0137] Where 0≤S j ≤1, 0 means a perfect match, 1 means a complete mismatch. 0≤D≤1, the closer D is to 0, the higher the degree of match; when D is equal to 1, it means that Table A and Table B have no correlation in these features.

[0138] In summary, the final similarity calculation formula between Table A and Table B is as follows: t (A,B)=sim tf (A,B)×C×D.

[0139] In one embodiment of the present application, the calculation of the matching similarity between the user behavior and the data table according to the second similarity calculation rule includes: according to the second similarity calculation rule, converting the usage frequency of the data table into a similarity index with the business scenario or business field; assuming that W represents the business scenario or business field, then the matching degree between table A and W in the graph is

[0140]

[0141] Where m means that table A is used in m analysis tasks for W, and α is a positive number used to control the effect of m on sim u (A, W) sensitivity of the impact; the sim u The value range of (A,W) is 0≤sim u (A,W)≤1; when table A has been pre-marked with the applicable business scenario or business field W, sim u (A,W)=0.

[0142] In order to expand the user's vision and optimize the experience, the embodiments of this application integrate user behavior in the search. Under the premise of unifying the classification of business scenarios and business fields, statistics are made on which tables are used more frequently in each business scenario and business field. In order to unify the caliber, the frequency of use is converted into the similarity between the table and the business scenario and business field. To ensure uniformity, the frequency of use of the table is converted into a similarity index with the business scenario and business field.

[0143] Integrate user usage data to analyze the business scenarios or business fields in which users use a table. Assuming that W represents the business scenario or business field, the matching degree formula between table A and W in the graph is as follows:

[0144]

[0145] Where m means that for W, table A is used in m analysis tasks. u (A,W) means, 0≤sim u (A,W)≤1, when table A has been pre-marked as applicable to business scenarios or business areas W, sim u (A, W) = 0, α is a positive number used to control the m-sim u (A, W) sensitivity. The larger α is, the greater the increase in m is. u The greater the impact of (A,W).

[0146] In an embodiment of the present application, a data query method is also provided, wherein the method for constructing a knowledge graph described in item 1 is applied, and the query method includes:

[0147] In response to the user's query request, a search is performed in the knowledge graph based on a preset shortest path calculation algorithm to find a data table with the shortest distance in the link that meets the requirements and return it as a query result.

[0148] The data query method in the embodiment of the present application utilizes the query and analysis tools of the knowledge graph to perform data exploration, relationship analysis, visualization and other operations. Users can enter table names, field names, business areas, business scenarios, and tags for querying.

[0149] Preferably, the Dijkstra algorithm is used during the query to calculate the first n nodes with the shortest path (edge ​​length) from a node of the knowledge graph. The specific implementation of the algorithm is as follows:

[0150] First, locate the node corresponding to the search content, start the node with this node, mark it as visited, and set its distance to itself to 0; create a priority queue and put the distance value of the starting node into the queue;

[0151] Secondly, loop through the following steps until the first n nodes are found or the queue is empty: take the node with the smallest distance from the queue; visit all unvisited neighbor nodes of the node; update the distance values ​​of these neighbor nodes to the distance of the current node plus the edge weight to the neighbor node; add the updated neighbor nodes to the priority queue and sort them by distance value. Repeat updating the distance values ​​of these neighbor nodes to the distance of the current node plus the edge weight to the neighbor node until the first n nodes are found. Output the first n nodes found and their distances.

[0152] Finally, the node with the shortest distance is selected through the priority queue, and the distance of its neighbor nodes is updated. This ensures that the path to the first n nodes found is the shortest.

[0153] It can be understood that when adding, deleting or changing a table, follow the above steps to update the table relationship, attributes and relationship weights in the knowledge graph. When a new user accesses the data, find the corresponding table, update the relationship between the table and the application scenario, and implement the knowledge graph update.

[0154] The present application embodiment also provides a knowledge graph construction device 200, such as Figure 2 As shown, a schematic diagram of the structure of the knowledge graph construction device in an embodiment of the present application is provided, wherein the knowledge graph construction device 200 at least includes: a first similarity calculation module 210, a second similarity calculation module 220, and a knowledge graph establishment module 230, wherein:

[0155] In one embodiment of the present application, the first similarity calculation module 210 is specifically used to calculate the inter-table similarities of all data tables in the library table according to the first similarity calculation rule.

[0156] In the process of constructing the knowledge graph, similarity is used as the basis for judging whether there is a relationship between entities, and the similarity value is the interval distance between entities.

[0157] The process of building a knowledge graph includes: table feature data collection. Specifically, based on the data warehouse data dictionary, extract the Chinese name and English name of the library table, as well as all the Chinese field names, English field names, business domain attribute labels, business scenario attribute labels, other annotated labels, and each table on the current table's lineage network link to prepare for subsequent processing.

[0158] It should be noted that the "business domain" and "business scenario" involved in the embodiments of this application all follow a standardized unified classification system, which is based on a core premise that all database tables of the relevant system and the analysis tasks launched by analysts are labeled with their business scenarios and business domains. Table labels are manually or otherwise annotated to indicate attributes such as the category or nature of the table.

[0159] The process of building a knowledge graph also includes: data processing. Specifically, when processing database table field data, you will usually encounter some common fields, such as ID, date, and status. These fields exist in almost all tables, but do not reflect specific business attributes. In order to focus more on business-related data analysis, common fields need to be removed.

[0160] Therefore, when processing data, it is preferred to automatically identify and remove common fields such as ID, date and status through preset filtering rules to ensure that only fields that truly reflect business characteristics are focused on during the analysis process, eliminating interference from irrelevant fields, thereby improving the efficiency and accuracy of data processing.

[0161] The process of building a knowledge graph also includes: entity and relationship definition, defining the entities of the knowledge graph and the relationships between them based on the semantics of the data warehouse and the business and efficiency needs of the search scenario.

[0162] Among them, entities include but are not limited to tables, table fields, table business scenarios, table business areas, and table tags.

[0163] The definition of relationship covers several aspects: the relationship between tables is "table-[similar]-table", which is determined by the first similarity calculation rule. The corresponding relationship between tables and fields is "table-[includes]-field", which describes which fields the table contains; "table-[used for]-business scenario".

[0164] It can be understood that similarity is used as the weight of the table-table relationship to evaluate the degree of association between two tables.

[0165] The process of building a knowledge graph also includes: building knowledge graph triples, processing the prepared data into a series of entity-relationship-entity triples according to the entity and relationship rules defined in the entity and relationship definitions, forming a structured triple list, and initializing the relationship weight to 0.

[0166] In one embodiment of the present application, the second similarity calculation module 220 is specifically used to calculate the matching similarity between the user behavior and the data table according to the second similarity calculation rule.

[0167] The definition of the relationship also includes: "table-[used for]-business domain", indicating the business scenarios and business domains to which the table is applicable, or the actual application in a specific business scenario or business domain, which is determined by pre-labeling or analyzing user behavior and using the second similarity calculation rule. The relationship between a table and a tag is "table-[labeled as]-tag", which is used to describe the attributes and characteristics of the table.

[0168] In order to expand the user's vision, optimize the experience, and integrate user behavior in search, under the premise of unifying the classification of business scenarios and business fields, statistics are collected on which tables are used more frequently in each business scenario and business field. In order to unify the caliber, the frequency of use is converted into the similarity between the table and the business scenario and business field. To ensure uniformity, the frequency of use of the table is converted into a similarity index with the business scenario and business field as the second similarity calculation rule.

[0169] In one embodiment of the present application, the knowledge graph establishment module 230 is specifically used to establish a knowledge graph based on the inter-table similarity of all the data tables and the matching similarity between the user behavior and the data tables, wherein the first similarity calculation rule calculates the similarity based on table attributes, and the second similarity calculation rule calculates the similarity based on actual user use.

[0170] Before calculating the similarity between tables, you need to calculate the table field vector value. By starting and loading the pre-trained Word2Vec model into memory, traverse each field name of each table and use it as the input of the model. For each field name, the model outputs a vector value. In this way, the field name in the table is converted into a mathematical vector to represent the semantic information of the field, providing support for subsequent similarity calculations.

[0171] Furthermore, according to the inter-table similarities of all the data tables and the matching similarities between the user behavior and the data tables, after the triples are constructed, they are imported into the graph database to generate a knowledge graph.

[0172] In one embodiment of the present application, the knowledge graph building module 230 is also used to

[0173] If the similarity between any two data tables is within a preset threshold range according to the inter-table similarity of all the data tables, it is determined whether there is a relationship with the knowledge graph triple table;

[0174] If yes, update the weight to the similarity between any two data tables;

[0175] If the similarity between the user's actual use and the data table is within a preset range according to the matching similarity between the user behavior and the data table, it is determined whether there is a relationship with the knowledge graph triple table;

[0176] If yes, the update weight is the similarity between the user's actual usage and the data table;

[0177] The similarity is used as the weight of the relationship between the two data tables and is used to evaluate the degree of association between any two data tables.

[0178] In one embodiment of the present application, it also includes: an update module for

[0179] Determine the relationship between the data tables according to the inter-table similarities of all the data tables;

[0180] Based on the matching similarity between the user behavior and the data table, the business scenarios and business fields to which the data table is applicable, or the actual application in a specific business scenario or business field, are determined.

[0181] In one embodiment of the present application, the first similarity calculation module 210 is also used to

[0182] Calculate similarity based on table attributes, using sim tf (A,B) represents the field similarity value between table A and table B;

[0183] Calculate the word vector A for each field in table A i And the word vector B of each word in table B j The cosine similarity of

[0184]

[0185] Among them A i ·B j represents their dot product, |A i ||B j | represents vector A i and vector B j mold;

[0186] For each field vector A in table A i Find the field word vector B with the largest cosine similarity with table B j ,

[0187] sim f (A i )=max(sim f (A i ,B j ))

[0188] Calculate all sim f (A i ) as the similarity between Table A and Table B

[0189]

[0190] Among them, table A has a fields and table B has b fields.

[0191] In one embodiment of the present application, the first similarity calculation module 210 is also used to

[0192] Assume that the number of lineage links between table A and table B is q, and C represents the degree of kinship between table A and table B. The value range of C is 0 < C ≤ 1. The larger C is, the farther the blood relationship is. When C = 1, it means there is no blood relationship. Then the calculation formula for the degree of kinship is as follows:

[0193] C = 1 - e -k·q

[0194] where k is a constant, k > 0, which is used to adjust the curve shape and determines the growth rate of the curve;

[0195] Consider the matching degree of multiple factors such as the business domain, business scenario, and tags of table A and table B, and adjust their contributions to the final similarity metric value according to the importance of each dimension. Assume that D represents the weighted average of the matching degree scores S1, S 2, S3... S q represents the matching degree scores of q different factors, and w1, w 2, w3... w q represents the corresponding weights, and each factor score is represented by S j indicates.

[0196]

[0197] where 0 ≤ S j ≤ 1, 0 indicates a perfect match, and 1 indicates a complete mismatch;

[0198] 0 ≤ D ≤ 1. The closer D is to 0, the higher the matching degree. When D equals 1, it means that table A and table B have no correlation in these features;

[0199] Finally, the similarity between table A and table B is obtained:

[0200] sim t (A, B) = sim tf (A, B) × C × D.

[0201] In an embodiment of the present application, the second similarity calculation module 220 is further configured to

[0202] According to the second similarity calculation rule, convert the usage frequency of the data table into a similarity metric with the business scenario or business domain;

[0203] Assume that W represents the business scenario or business domain. Then the matching degree between table A and W in the graph

[0204]

[0205] where m represents that table A is used in m analysis tasks for W, and α is a positive number used to control the influence of m on sim u(A,W) sensitivity of impact;

[0206] The sim u The value range of (A,W) is 0≤sim u (A,W)≤1;

[0207] When table A is pre-marked with the applicable business scenario or business domain W, sim u (A,W)=0.

[0208] It can be understood that the above-mentioned knowledge graph construction device can implement the various steps of the knowledge graph construction method provided in the aforementioned embodiments, and the relevant explanations about the knowledge graph construction method are applicable to the knowledge graph construction device and will not be repeated here.

[0209] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 3 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. Of course, the electronic device may also include hardware required for other services.

[0210] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0211] The memory is used to store the program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0212] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a knowledge graph construction device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0213] According to the first similarity calculation rule, calculate the similarity between all data tables in the library table;

[0214] According to the second similarity calculation rule, the matching similarity between the user behavior and the data table is calculated;

[0215] A knowledge graph is established based on the similarity between all the data tables and the matching similarity between the user behavior and the data tables.

[0216] The first similarity calculation rule calculates the similarity based on table attributes, and the second similarity calculation rule calculates the similarity based on actual use by users.

[0217] The above application Figure 1 The method performed by the knowledge graph construction device disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in software form. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in a decoding processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0218] The electronic device may also perform Figure 1 The method executed by the knowledge graph construction device in Figure 1 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0219] The present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, enable the electronic device to execute Figure 1 The method performed by the knowledge graph construction device in the embodiment shown is specifically used to perform:

[0220] According to the first similarity calculation rule, calculate the similarity between all data tables in the library table;

[0221] According to the second similarity calculation rule, the matching similarity between the user behavior and the data table is calculated;

[0222] A knowledge graph is established based on the similarity between all the data tables and the matching similarity between the user behavior and the data tables.

[0223] The first similarity calculation rule calculates the similarity based on table attributes, and the second similarity calculation rule calculates the similarity based on actual use by users.

[0224] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0225] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0226] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0227] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0228] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0229] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0230] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0231] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0232] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0233] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A knowledge graph construction method, wherein: The construction method includes: Calculating the inter-table similarity of all data tables in the library table according to the first similarity calculation rule; Calculating the matching similarity between user behavior and data tables according to the second similarity calculation rule; Establishing a knowledge graph based on the inter-table similarity of all the data tables and the matching similarity between the user behavior and the data tables, wherein, the first similarity calculation rule calculates similarity based on table attributes, and the second similarity calculation rule calculates similarity based on actual user usage.

2. The method of claim 1, wherein: The establishing of the knowledge graph based on the inter-table similarity of all the data tables and the matching similarity between the user behavior and the data tables includes: If, according to the inter-table similarity of all the data tables, the similarity between any two data tables is within a preset threshold range, determining whether there is a relationship in the knowledge graph triple table; If so, updating the weight to the similarity between any two data tables; If, according to the matching similarity between the user behavior and the data tables, the similarity between the actual user usage and the data tables is within a preset range value, determining whether there is a relationship in the knowledge graph triple table; If so, updating the weight to the similarity between the actual user usage and the data tables; wherein, the similarity is used as the weight of the relationship between two data tables to evaluate the association degree between any two data tables.

3. The method of claim 2, wherein: The method further includes: Determining the relationship between data tables according to the inter-table similarity of all the data tables; Determining the business scenarios and business fields applicable to the data tables, or the actual applications in specific business scenarios or business fields, according to the matching similarity between the user behavior and the data tables.

4. The method of claim 1, wherein: The calculating of the inter-table similarity of all data tables in the library table according to the first similarity calculation rule includes: Calculate similarity based on table attributes, using sim tf (A,B) represents the field similarity value between table A and table B; Calculate the word vector A for each field in table A i And the word vector B of each word in table B j The cosine similarity of Among them A i ·B j represents their dot product, |A i ||B j | represents vector A i and vector B j mold; For each field vector A in table A i Find the field word vector B with the largest cosine similarity with table B j , Yes f (THE i )=max(yes f (THE i ,B j )) Calculate all sim f (A i ) as the similarity between Table A and Table B Among them, table A has a fields and table B has b fields.

5. The method according to claim 4, the method further includes: Assuming that the number of blood relationship links between table A and table B is q, c represents the degree of kinship between table A and table B, the value range of c is 0 < C ≤ 1, the larger c is, the farther the blood relationship is, and when C = 1, it means there is no blood relationship. Then the calculation formula of the degree of blood relationship is as follows: C=1-e -k·q where k is a constant, k > 0, which is used to adjust the curve shape and determines the growth rate of the curve; Consider the matching degree between Table A and Table B in terms of business field, business scenario, and label, and adjust the contribution of each dimension to the final similarity measurement value according to its importance. Suppose D is used to represent the weighted average of the matching scores S1, S2, S3...S q Represents the matching scores of q different factors, w1, w2, w3...w q Indicates the corresponding weight, and each factor score is represented by S j express, Where 0≤S j ≤1, 0 means a perfect match, 1 means a complete mismatch; 0 ≤ D ≤ 1, the closer D is to 0, the higher the matching degree. When D is equal to 1, it means that table A and table B have no relevant relationship in these features; Finally, the similarity between table A and table B is obtained: Yes t (A,B)=yes tf (A,B)×C×D。 6. The method of claim 1, wherein: The calculating of the matching similarity between user behavior and data tables according to the second similarity calculation rule includes: Converting the usage frequency of the data tables into a similarity index related to business scenarios or business fields according to the second similarity calculation rule; Assuming that W represents the business scenario or business field, then the matching degree between table A and W in the graph Where m means that table A is used in m analysis tasks for W, and α is a positive number used to control the effect of m on sim u (A,W) sensitivity of impact; The sim u The value range of (A,W) is 0≤sim u (A,W)≤1; When table A is pre-marked with the applicable business scenario or business domain W, sim u (A,W)=0.

7. A data query method, wherein: Applied to the knowledge graph construction method according to any one of claims 1 to 6, the query method includes: In response to a user's query request, searching in the knowledge graph based on a preset shortest path calculation algorithm, and finding the data table with the shortest distance that meets the requirements in the link as the query result to return.

8. A knowledge graph construction device, wherein: The device includes: A first similarity calculation module is used to calculate the inter-table similarity of all data tables in the library table according to a first similarity calculation rule; A second similarity calculation module is used to calculate the matching similarity between the user behavior and the data table according to a second similarity calculation rule; The knowledge graph building module is used to build a knowledge graph based on the similarity between all the data tables and the matching similarity between the user behavior and the data tables. The first similarity calculation rule calculates the similarity based on table attributes, and the second similarity calculation rule calculates the similarity based on actual use by users.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.