Knowledge graph presentation method, electronic equipment and storage medium
By obtaining key nodes in the knowledge graph, dividing the grid units, calculating the area allocation scores, and optimizing the node positions, the confusion problem caused by traditional layout algorithms is solved, and clear graph presentation and reasonable space utilization are achieved.
Patent Information
- Application Number
- CN202511344456.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Traditional force-directed layout algorithms can easily lead to chaotic layout, loss of focus, and overly crowded or scattered edge nodes in large-scale and complex knowledge graphs, making it difficult to achieve clear visualization.
By obtaining key nodes, dividing the grid units based on the number of key nodes, calculating the area allocation score and adjusting the grid unit area, and combining the force-directed layout algorithm to optimize the node position, it ensures that the key nodes are reasonably arranged in the central area and the non-key nodes are reasonably presented in the local area.
A reasonable layout of the knowledge graph is achieved, with key nodes and strongly correlated non-key nodes presented in a reasonable area, improving the clarity and space utilization of the graph.
Smart Images

Figure CN120821879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a knowledge graph presentation method, electronic device, and storage medium. Background Art
[0002] With the development of information technology, knowledge graphs, as an efficient knowledge representation and reasoning technology, have been widely used in a variety of fields, including search engines, intelligent question answering, and big data analysis. Knowledge graphs typically contain a large number of nodes and edges, and their visualization is crucial for users to intuitively understand the graph structure and discover potential connections.
[0003] Currently, common knowledge graph visualization methods primarily use force-directed layout algorithms. These algorithms automatically calculate node positions by simulating physical mechanics, such as the repulsive forces between nodes and the attractive forces between edges, thereby presenting a clear community structure. However, when graphs are large and complex, traditional force-directed layout algorithms can lead to problems such as disorganized layouts, loss of focus, and overcrowded or scattered edge nodes.
[0004] Therefore, there is an urgent need to provide a knowledge graph presentation method with a reasonable layout that can make the presented knowledge graph clearer. Summary of the Invention
[0005] In response to the above technical problems, the present invention provides a knowledge graph presentation method, electronic device and storage medium, which can optimize the node positions in the knowledge graph, achieve a reasonable layout of the presentation page, and make the presented knowledge graph clearer.
[0006] According to a first aspect of the present invention, a knowledge graph presentation method is provided, comprising the following steps: S100, obtaining several key nodes from a given knowledge graph; the key nodes refer to nodes in the given knowledge graph whose number of corresponding neighbor nodes is not less than n.
[0007] S200 , based on the number of key nodes, dividing the preset presentation page into an initial node grid including a plurality of grid units and obtaining a grid unit corresponding to each key node in the initial node grid.
[0008] S300 , calculating the area allocation score corresponding to each key node according to the number of neighboring nodes of different preset entity types connected to each key node and the preset presentation area corresponding to each node of the preset entity type.
[0009] S400 , based on the area allocation score corresponding to each key node, adjusting the area of the grid unit corresponding to each key node in the initial node grid to obtain a target node grid.
[0010] S500, based on the target node grid, adjust the presentation position of each key node and non-key node to obtain the final target knowledge graph; wherein, the non-key node refers to any node in the given knowledge graph except the key node.
[0011] According to the second aspect of the present invention, a non-transitory computer-readable storage medium is provided, in which at least one instruction or at least one program is stored. The at least one instruction or the at least one program is loaded and executed by a processor to implement the above-mentioned knowledge graph presentation method.
[0012] According to a third aspect of the present invention, there is provided an electronic device comprising a processor and the above-mentioned non-transitory computer-readable storage medium.
[0013] The present invention has at least the following beneficial effects: The present invention provides a knowledge graph presentation method, which first obtains several key nodes from a given knowledge graph, divides a preset presentation page into an initial node grid containing several grid units based on the number of key nodes, and obtains the grid unit corresponding to each key node in the initial node grid, thereby realizing a decentralized layout of the positions of the key nodes, and preferentially allocating the key nodes with many connected nodes to the center of the initial node grid, so that the presentation position of each key node is more reasonable; then calculates the area allocation score corresponding to each key node, and adjusts the area of the grid unit corresponding to each key node in the initial node grid based on the area allocation score corresponding to each key node to obtain a target node grid, and finally adjusts the presentation position of each key node and non-key node based on the target node grid, ensuring that the key nodes do not deviate from their corresponding cells, and that strongly associated non-key nodes are reasonably presented in a preset local area near the key nodes, thereby realizing a reasonable layout of the presentation page and making the presented knowledge graph clearer. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0015] Figure 1 A flowchart of the knowledge graph presentation method provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0017] The present invention provides a knowledge graph presentation method, such as Figure 1 As shown, the method includes the following steps: S100, obtain several key nodes from a given knowledge graph; the key nodes refer to nodes in the given knowledge graph whose corresponding number of neighbor nodes is not less than n; it can be understood as: taking nodes whose number of connected neighbor nodes exceeds a certain threshold as the main position adjustment nodes.
[0018] Specifically, a given knowledge graph includes nodes of several preset entity types. For example, when constructing a knowledge graph of corporate financing relationships or bidding relationships, there are two preset entity types: enterprise and legal person.
[0019] S200, based on the number of key nodes, divides the preset presentation page into an initial node grid containing several grid units and obtains the grid unit corresponding to each key node in the initial node grid; it can be understood that: the preset presentation page refers to a pre-built page for presenting a given knowledge graph.
[0020] Specifically, dividing the preset presentation page into an initial node grid including a plurality of grid units includes the following steps: S201: Calculate the preset aspect ratio R of the presentation page.
[0021] S202, set the number of rows and columns of the initial node grid to a and b respectively, traverse all values of a from 1 to L, and calculate the value of b respectively; where b=roundup(L / a), roundup() is a round-up function, and L is the number of key nodes.
[0022] S203, obtaining the a value and b value with the smallest absolute difference between the grid aspect ratio b / a and R to obtain the initial node grid after division; this can be understood as: the a value and b value with the smallest absolute difference between the grid aspect ratio b / a and R are used as the number of rows and columns of the initial node grid after division, respectively.
[0023] As described above, when dividing the preset presentation page, the number of grid units is guaranteed to ensure that all key nodes can be allocated to corresponding grid units, and the calculation of the aspect ratio is introduced to make the aspect ratio of the divided grid units closest to the aspect ratio of the preset presentation page, thereby making the divided grid units more reasonable and improving the space utilization of the preset presentation page.
[0024] Furthermore, obtaining the grid unit corresponding to each key node in the initial node grid includes the following steps: S210: Calculate the correlation strength between every two key nodes and construct a correlation matrix. A person skilled in the art knows how to calculate the correlation strength between any two nodes, which will not be described in detail here.
[0025] S220: Based on the association matrix, a force-directed layout algorithm is used to calculate the two-dimensional coordinates corresponding to each key node. Those skilled in the art are familiar with the specific implementation steps of calculating the two-dimensional coordinates of nodes using the force-directed layout algorithm, which will not be described in detail here.
[0026] S230, normalize the two-dimensional coordinates corresponding to each key node to the coordinate space of the initial node grid, and obtain the grid unit corresponding to each key node in the initial node grid; this can be understood as: assigning each key node to the grid unit where its normalized coordinates are located.
[0027] In the above, the grid unit adapted to each key node is found by calculating the association strength between every two key nodes, and the presentation positions of the key nodes are dispersed on the basis of maximizing the use of the preset presentation page. Since the key nodes are the nodes with more screened neighbor nodes, the presentation position of each key node is made more reasonable by dispersing the positions of the key nodes, and the key nodes with more connected nodes are preferentially allocated to the center of the initial node grid, thereby optimizing the presentation layout of the knowledge graph.
[0028] S300 , calculating the area allocation score corresponding to each key node according to the number of neighboring nodes of different preset entity types connected to each key node and the preset presentation area corresponding to each node of the preset entity type.
[0029] In one embodiment, the preset presentation area corresponding to the node of the preset entity type is proportional to the importance of the preset entity type; it can be understood that: importance is set for each preset entity type, the higher the importance, the larger the corresponding preset presentation area, and the preset presentation area corresponding to the node is generally a solid circle.
[0030] In another embodiment, the preset display area corresponding to the node of the preset entity type is proportional to the average number of characters in the name of the node corresponding to the preset entity type. For example, since the name of a company is generally longer, and the name of a legal person is generally shorter, a larger preset display area is set for the node corresponding to the company.
[0031] Preferably, the area allocation score corresponding to the key node meets the following conditions: , where Q represents the area allocation score corresponding to any key node, m is the number of preset entity types, S i is the preset presentation area corresponding to the node of the i-th preset entity type, N i The number of nodes of the i-th preset entity type connected to the key node.
[0032] As mentioned above, when calculating the area allocation score of the key nodes, since the preset presentation areas corresponding to different preset entity types are inconsistent, the proportion of the preset presentation area corresponding to each preset entity type and the number of nodes of different preset entity types connected are introduced to make the obtained area allocation score more reasonable. That is, the more nodes with larger preset presentation areas, the larger the corresponding area allocation score, which is conducive to optimizing the layout of the overall page.
[0033] S400 , based on the area allocation score corresponding to each key node, adjusting the area of the grid unit corresponding to each key node in the initial node grid to obtain a target node grid.
[0034] In a specific embodiment, step S400 includes the following steps: S401 , for any key node, calculating the ratio of the area allocation score corresponding to the key node to the total area allocation scores of all key nodes.
[0035] S402: The product of the total displayed area of the initial node grid and the ratio is used as the area of the grid cell corresponding to the key node, thereby adjusting the area of the grid cell corresponding to the key node in the initial node grid to obtain the target node grid. In one embodiment, during the adjustment, the key node closest to the center of the initial node grid is used as the adjustment reference. While maintaining the center of the grid cell corresponding to the key node unchanged, the area of the grid cell is expanded, and the area adjustment of each grid cell is sequentially extended outward. The center point of the outer grids may be offset outward accordingly, and the offset distance is successively reduced according to the preset offset distance. It should be noted that grid cells may overlap during the adjustment process.
[0036] As mentioned above, since the larger the area allocation score corresponding to the key node, the more neighboring nodes the key node has, a larger presentation cell should be allocated to it, which is conducive to the subsequent setting of the preset local area, so that when the positions of the key nodes and non-key nodes are subsequently optimized, the key nodes are ensured not to leave their corresponding cells, and the strongly associated non-key nodes can be reasonably presented in the preset local area near the key nodes.
[0037] S500, based on the target node grid, adjust the presentation position of each key node and non-key node to obtain the final target knowledge graph; wherein, the non-key node refers to any node in the given knowledge graph except the key node.
[0038] In a specific embodiment, step S500 includes the following steps: S501: For any key node, based on the average coordinates of several neighboring nodes connected to the key node, the key node is moved within its own grid cell in the direction of the average coordinates using a preset offset step size. It can be understood that the average coordinates of the neighboring nodes connected to the key node refer to the center of gravity of the several neighboring nodes connected to the key node. Those skilled in the art can set the preset offset step size based on actual needs, and this will not be further described here.
[0039] S502: Each non-critical node is initially assigned to a predetermined local region containing a critical node with the highest correlation strength with the non-critical node itself. A force-directed layout algorithm is then used to calculate the local position of each non-critical node, centered around the corresponding critical node. It should be noted that during this calculation, a global repulsive force is applied to prevent overlap between non-critical nodes in different local regions, and an edge constraint is applied to ensure that all nodes are within the predetermined presentation page.
[0040] S503, iteratively execute the above steps S501 and S502 until the average position change of all nodes is less than the preset change threshold or the maximum number of iterations is reached, and the final presentation positions of key nodes and non-key nodes are obtained; it can be understood that the average position change can be the average value of several position change distances.
[0041] As mentioned above, when the position of the key node has been preliminarily determined, the connection strength between each non-key node and the key node and the distribution center of the neighboring nodes are considered, and the positions of the key nodes and non-key nodes are fine-tuned to make the final knowledge graph more reasonable and clear.
[0042] Furthermore, after updating the target knowledge graph, since there is a lot of data in the target knowledge graph and the data is collected at different times or from different sources, in order to avoid the same entity from appearing repeatedly in the knowledge graph with the same name or different names, the method further includes the following steps: P100, based on the obtained first entity to be analyzed, the second entity to be analyzed, and the preset attribute parameter set corresponding to each entity to be analyzed, calculates the data similarity between the preset attribute parameter sets corresponding to the first entity to be analyzed and the second entity to be analyzed respectively; it can be understood as: according to the obtained preset attribute feature text corresponding to the first entity to be analyzed and the second entity to be analyzed respectively, extracting parameter values corresponding to several preset attribute indicators from each preset feature text. For example, when the entity to be analyzed is a legal person or a platform user, several preset attribute indicators include but are not limited to personal basic information such as height, gender, education level, graduation school, date of birth, and constellation. Data similarity can be understood as semantic data similarity.
[0043] Specifically, the first entity to be analyzed and the second entity to be analyzed are entities corresponding to any two nodes of a target entity type among several preset entity types.
[0044] Preferably, the first text to be analyzed and the second text to be analyzed are obtained by the following steps: P101, based on the preset attribute parameter set corresponding to each preset entity in the target knowledge graph, convert each preset attribute parameter set into a corresponding attribute parameter vector. This can be understood as follows: the preset entity is any entity of the target type in the target knowledge graph in the preset database. For example, when the target knowledge graph includes information such as a company name or a legal person name, the preset entity is any legal person name.
[0045] P102: Cluster the attribute parameter vectors using a preset clustering model to obtain a plurality of attribute parameter vector clusters. For example, a k-means clustering model or an adaptive iterative clustering model may be used. Those skilled in the art are familiar with the specific implementations of these two clustering models and will not be described in detail here.
[0046] P103: Use the preset entities corresponding to any two attribute parameter vectors in any attribute parameter vector cluster as the first entity to be analyzed and the second entity to be analyzed respectively.
[0047] As mentioned above, since the target knowledge graph may contain the same entity but with inconsistent names due to different data acquisition sources, the same entity needs to be identified. Through the above clustering process, entities with similar information can be clustered into one cluster. Only the entities in the same cluster need to be analyzed, and there is no need to judge whether one entity is the same entity as all other entities, which greatly reduces the amount of data calculation.
[0048] P200: If the name data corresponding to the first entity to be analyzed is the same as the name data corresponding to the second entity to be analyzed, and the data similarity is greater than a first preset similarity threshold, the first entity to be analyzed and the second entity to be analyzed are determined to be the same entity. Persons skilled in the art may set the first preset similarity threshold based on actual needs, such as 90%.
[0049] As mentioned above, since a node is usually used to represent an entity or object in the knowledge graph, when the name data corresponding to two entities to be analyzed are the same, it indicates that there is a certain probability that they are the same entity. However, in order to avoid coincidences, further judgment is required. On this basis, the comparison of data similarity and threshold of the preset attribute parameter set is introduced. When the data similarity is high, it is considered that the two entities to be analyzed are the same entity, which improves the accuracy of the judgment results.
[0050] P300: When the data similarity is not greater than the first preset similarity threshold and greater than the second preset similarity threshold, the historical activity areas and corresponding historical activity time periods corresponding to the first entity to be analyzed and the second entity to be analyzed, respectively, within a preset time period before the current moment are crawled from the given platform. For example, the historical activity areas and historical activity time periods are searched from a webpage or a given information platform. The historical activity areas may be specific location information such as the address of a company's office building, a hotel, or a factory. Those skilled in the art will set the second preset similarity threshold based on actual needs.
[0051] As mentioned above, when the preset attribute parameter information obtained is incomplete or the data changes due to different data uploaded at different times, there is a situation where the data similarity is not high enough. In this case, further judgment is required to improve the reliability of the judgment result.
[0052] Furthermore, the method further comprises: When the data similarity is not greater than a second preset similarity threshold, it is determined that the first entity to be analyzed and the second entity to be analyzed are not the same entity.
[0053] P400 , obtaining a first target device set of a first entity to be analyzed in a corresponding historical activity area and a corresponding historical activity period and a second target device set of a second entity to be analyzed in a corresponding historical activity area and a corresponding historical activity period.
[0054] When obtaining target devices in the historical activity area and historical activity period, all devices in the historical activity area during the historical activity period are obtained based on the anonymous location information uploaded by the devices.
[0055] As described above, since the location information of the entity to be analyzed cannot be directly obtained, when the historical activity area and historical activity time period are known, the first target device set obtained contains a device corresponding to the first entity to be analyzed, and the second target device set contains a device corresponding to the second entity to be analyzed. By judging whether the same device exists, it can provide another possibility for judging the same entity while ensuring privacy.
[0056] P500: If it is determined that the first target devices in the first target device set and the second target devices in the second target device set are identical, determine that the first entity to be analyzed and the second entity to be analyzed are the same entity.
[0057] In a preferred embodiment, the process of determining whether a plurality of first target devices in the first target device set and a plurality of second target devices in the second target device set are the same device is as follows: P501, based on the obtained unique identification codes corresponding to the first target device and the second target device, if the unique identification code corresponding to the first target device and the unique identification code corresponding to the second target device are the same, output a determination result that the first target devices and the second target devices are the same device. For example, the unique identification code is any one of an IMEI identification code, an OAID identifier, and an IDFV identifier, and the unique identification codes corresponding to the first target device and the second target device are of the same type.
[0058] P502: When the determination result that the plurality of first target devices and the plurality of second target devices are identical is not obtained, the trajectory data corresponding to each first target device and each second target device during the preset historical period is obtained. For example, the unique identification codes of all devices cannot be obtained due to insufficient permissions or the unique identification codes are presented in different formats due to different systems.
[0059] P503 , performing similarity calculation on the trajectory data corresponding to each first target device in the preset historical period and the trajectory data corresponding to each second target device in the preset historical period to generate a trajectory similarity matrix.
[0060] When calculating the similarity of trajectory data, the following steps are included: P5031, calculate the spatial distance between the position points of two trajectory data corresponding to the time points.
[0061] P5032: Calculate the trajectory similarity between the two trajectory data based on the time series of the spatial distance, for example, by using a dynamic time warping algorithm or a Fréchet distance algorithm.
[0062] P504: Based on a preset first trajectory similarity threshold, target device pairs having trajectory similarities exceeding the first trajectory similarity threshold are screened from the trajectory similarity matrix. The first trajectory similarity threshold is set based on the duration of a preset historical period and the corresponding average device density within the preset historical period.
[0063] When setting the first trajectory similarity threshold, consider that the shorter the preset historical period (i.e., the observation time), the more likely the similarity between two trajectories is to be accidental, and therefore the likelihood of similarity is greater. Therefore, the first trajectory similarity threshold should be set inversely proportional to the preset historical period. In densely populated areas such as city centers or train stations, the probability of two devices sharing a common trajectory is higher. To reduce the possibility of misjudgment, the first trajectory similarity threshold should be increased. Therefore, the first trajectory similarity threshold should be set proportionally to the average device density. Therefore, normalization or other methods can be used to convert the duration of the preset historical period and the average device density to the same data level. The first trajectory similarity threshold can then be calculated by taking a weighted sum of the weights corresponding to the duration of the preset historical period and the average device density, respectively.
[0064] P505: The preset historical period is divided into preset time windows. If the trajectory similarity between the first target device and the second target device in the target device pair within each time window exceeds a preset second trajectory similarity threshold, a determination result is output indicating that the plurality of first target devices and the plurality of second target devices are identical. Persons skilled in the art will determine the second trajectory similarity threshold based on actual needs, and this will not be further described here.
[0065] As mentioned above, when determining whether the same device exists in different areas at different time periods, the first consideration is to accurately identify the unique identification code of the device. If not all device unique identification codes can be identified, another method is used, that is, the trajectory data reported by the device, to first screen out device pairs with greater trajectory similarity. The trajectory is then segmented and the similarity of each small segment is compared separately. This improves the accuracy of the trajectory similarity calculation and thus the reliability of the judgment result. Since the probability of a device trajectory being continuously similar to an entity's activity location and activity time is very low, the identified identical devices are considered to be the devices corresponding to the two entities to be analyzed, which increases the credibility that the two entities to be analyzed are the same entity.
[0066] Furthermore, if the confidence level of the judgment result outputted in step P505 is lower than that of the judgment result outputted in step P501, the staff may decide whether to conduct further verification based on the accuracy requirements. By identifying two entities to be analyzed whose judgment results are the same entity, a data basis is also provided for further verification.
[0067] In another embodiment, the method further comprises the steps of: P10, if the name data corresponding to the first entity to be analyzed is different from the name data corresponding to the second entity to be analyzed, when the semantics of the preset attribute parameter set corresponding to the first entity to be analyzed and the preset attribute parameter set corresponding to the second entity to be analyzed are completely consistent, it is determined that the first entity to be analyzed and the second entity to be analyzed are the same entity.
[0068] P20, when the semantics of two preset attribute parameter sets are not completely consistent and the data similarity between the two preset attribute parameter sets is greater than the first preset similarity threshold, the historical activity areas and the corresponding historical activity time periods of the first entity to be analyzed and the second entity to be analyzed within the preset time period before the current moment are crawled from a given platform.
[0069] Furthermore, the method further comprises: When the data similarity between the preset attribute parameter set corresponding to the first entity to be analyzed and the preset attribute parameter set corresponding to the second entity to be analyzed is not greater than a first preset similarity threshold, it is determined that the first entity to be analyzed and the second entity to be analyzed are not the same entity.
[0070] P30 , obtaining a first target device set of the first entity to be analyzed in a corresponding historical activity area and a corresponding historical activity period and a second target device set of the second entity to be analyzed in a corresponding historical activity area and a corresponding historical activity period.
[0071] P40: If it is determined that the first target devices in the first target device set and the second target devices in the second target device set are identical devices, determine that the first entity to be analyzed and the second entity to be analyzed are the same entity.
[0072] In the above, we analyzed the case where the name data corresponding to two entities are different. When they are different, the probability that they belong to the same entity is relatively low. Therefore, compared with the case where the name data are the same, the strictness of the judgment condition is improved, ensuring the reliability of the judgment result.
[0073] Furthermore, the method further comprises the following steps: P1: If the first entity to be analyzed and the second entity to be analyzed are the same entity, obtain from a target knowledge graph stored in a preset database several objects that have connection relationships with the first entity to be analyzed and the second entity to be analyzed. These objects may include multiple types of objects, and the first entity to be analyzed and the second entity to be analyzed are objects of the same type.
[0074] P2: Associating and integrating a number of objects that are connected to the first entity to be analyzed and a number of objects that are connected to the second entity to be analyzed, and storing them in a preset database.
[0075] By finding and storing related data of the same entity, it has a positive effect on updating the target knowledge graph and statistical management of data.
[0076] On this basis, the method further comprises the following steps: P3, setting the association integration result in a preset hidden box in the preset presentation page; the association integration result includes several objects connected to the first entity to be analyzed, several objects connected to the second entity to be analyzed, and an association identifier.
[0077] P4, when receiving the user's instruction to click the preset hidden box, the associated integration result is presented according to the preset presentation area.
[0078] As mentioned above, when entities with the same name data or different name data in the target knowledge graph are identified as the same entity, while ensuring an intuitive display effect, the data complexity of the knowledge graph displayed on the page is limited and is restricted by spatial arrangement. Therefore, there is no need to convert the two entities into one entity on the page and connect the objects connected by the two entities with the converted entity. It is only necessary to store the association relationship and preset it in a preset hidden box so that users can query according to their needs, thereby obtaining the associated data and ensuring the page presentation effect and user convenience.
[0079] In another embodiment, before obtaining a given knowledge graph, how to accurately and comprehensively find the potential objects associated with each node is of great significance to the accurate construction of the given knowledge graph and subsequent data analysis. Based on this, the method further includes the following steps: F100, based on the information of the object to be processed uploaded by the server to be processed, obtain the parameter value corresponding to each of the preset parameters to be processed. It can be understood that the server to be processed is the server corresponding to the object to be processed. For example, the object to be processed can be an enterprise or a platform.
[0080] Specifically, the plurality of parameters to be processed include a plurality of technical capability parameters and a plurality of preset basic information parameters.
[0081] In specific implementations, when the subject to be processed is an enterprise, the technical capability information in the subject information includes the number of technical documents generated by the enterprise each year, and the preset basic information in the subject information includes basic information such as average education level, average salary, office area, years of establishment, and number of social security employees. When the subject to be processed is a platform, the technical capability information includes the number of technical documents generated by the platform each year, and the preset basic information includes information such as the total number of target users corresponding to the platform identifier, the number of target users each year, the average and maximum years of use by target users, etc. The number of technical documents can be the number of authorized patent documents.
[0082] Furthermore, the parameter values corresponding to several technical capability parameters include the total number of technical documents generated by the subject to be processed, the technical document quality score, the technical development trend score within a preset time period, and the current influence in the technical community. It can be understood that the current influence in the technical community refers to the current influence in the same technical field or industry. Those skilled in the art can set the preset time period based on actual needs, for example, 5-10 years.
[0083] In one specific embodiment, the technical document quality score is evaluated based on the total number of patent applications and the percentage of granted patents. For example, the total number of patent applications is normalized to a value between 0 and 1, and the technical document quality score is calculated by taking the weighted sum of the pre-set weights corresponding to the total number of patent applications and the percentage of granted patents.
[0084] In a specific embodiment, the technology development trend score within a preset time period is obtained through the following steps: F101, obtaining the number of technical documents each year within a preset time period, and determining the fluctuation of the number of technical documents; the fluctuation can be any one of rising, falling, stable, and wavy.
[0085] F102: Based on the fluctuation in the number of technical documents, several preset fitting curve models corresponding to the fluctuation are determined, and the model with the lowest degree of dispersion among the preset fitting curve models is used as the target fitting curve model. This can be understood as follows: a preset database stores several preset fitting curve models corresponding to each fluctuation. Each preset fitting curve model is used to fit the number of technical documents to obtain the preset fitting curve model with the lowest degree of dispersion.
[0086] F103: Find the preset importance level corresponding to the target fitting curve model; the preset importance levels decrease in descending order as the target fitting curve model represents the technological development trend. For example, when the target fitting curve model is a power function, it indicates a high technological development trend, and the corresponding preset importance level is also high.
[0087] F104: Perform data preprocessing on the preset importance corresponding to the target fitting curve model and the total number of technical documents within the preset time period, and calculate the weighted sum of the preset weights corresponding to the preset importance and the total number of technical documents to obtain the technology development trend score; data preprocessing refers to normalizing the preset importance and the total number of technical documents to values of the same magnitude, for example, both are between 0 and 1.
[0088] As mentioned above, by calculating the technology development trend score corresponding to the object to be processed, the development capability of the object to be processed can be represented. When calculating the technology development trend score, the most accurate target fitting curve model is selected by analyzing the fluctuation of the number of technical texts, and the technology development trend score is obtained comprehensively based on the two dimensions of the importance of the target fitting curve model and the total number of technical texts, so that the obtained technology development trend score is more reasonable and accurate, which is conducive to accurately evaluating the development capability of the object to be processed.
[0089] In a specific embodiment, the current influence of the object to be processed in the technical community is obtained through the following steps: F110 , extracting technical keywords from each technical text within a preset time period and forming a first keyword vector.
[0090] During implementation, in order to avoid a large number of keywords, core keywords that meet the preset conditions are extracted based on semantics, or keywords that are consistent with the business operations are filtered out from the extracted keywords.
[0091] F120 , obtaining a second keyword vector corresponding to each preset object in a preset database; it can be understood that the second keyword vector is obtained in the same manner as the first keyword vector, and will not be further elaborated here.
[0092] Specifically, the preset object refers to any object whose corresponding attribute label has a similarity with the attribute label of the object to be processed that is greater than a preset similarity threshold. For example, when the object to be processed is an enterprise, the attribute label is any industry field label to which the object to be processed belongs.
[0093] F130 , clustering the first keyword vector and the plurality of second keyword vectors is performed, and a first target cluster corresponding to the object to be processed is obtained based on the clustering result. This can be understood as follows: the first target cluster is the cluster where the object to be processed resides. In a specific implementation, a k-means clustering algorithm may be employed.
[0094] F140 , based on the distance between the object to be processed and the centroid of the first target cluster and the maximum distance corresponding to the first target cluster, calculate the current influence S of the object to be processed in the technical community; the current influence S meets the following conditions: S=1-(d / r), where d is the distance between the object to be processed and the centroid of the first target cluster, and r is the maximum distance corresponding to the first target cluster.
[0095] As described above, by calculating the current influence of the object to be processed in the technical community, the importance of the object to be processed in this technical field can be represented. When calculating the current influence, the object to be processed and several preset objects with the same attribute label are clustered. Objects clustered into a cluster indicate that they are objects with similar development potential or scale. Since the total distance between the cluster centroid and each vector is optimal, and the closer the distance to the cluster centroid, the higher the importance of the object, the current influence of the object to be processed in the technical community obtained by the above method is more reliable.
[0096] Furthermore, the plurality of preset basic information parameters include a plurality of parameters for characterizing the scale and resource allocation of the object to be processed. In one embodiment, the preset basic information parameters may be values obtained by performing data standardization processing on any parameter in the preset basic information.
[0097] F200: merging the parameter values corresponding to each parameter to be processed into a feature vector corresponding to the object to be processed, clustering the feature vectors corresponding to each preset object obtained from the preset database, and determining the reference object corresponding to the object to be processed based on the clustering result.
[0098] Preferably, the reference object corresponding to the object to be processed is determined by the following steps: F201 normalizes the feature vectors corresponding to the target object and the feature vectors corresponding to the preset object so that each feature in the feature vectors is at the same scale. For example, a Z-score normalization algorithm or a Min-Max normalization method may be used. Those skilled in the art are familiar with the specific implementations of these two normalization methods and will not be described in detail here.
[0099] F202 : Clustering the normalized feature vectors based on a preset number of clusters K, and determining a second target cluster corresponding to the object to be processed from the obtained K clusters.
[0100] In one embodiment, those skilled in the art set the K value according to actual needs, for example, setting K to 3 based on the importance of high, medium, and low, or obtaining the K value using the silhouette coefficient method. Those skilled in the art are aware of the implementation of the silhouette coefficient method and will not go into details here.
[0101] F203 : The first Z preset objects whose corresponding feature vectors in the second target cluster are closest to the feature vectors corresponding to the object to be processed are all used as reference objects.
[0102] Preferably, Z is proportional to the number of samples in the second target cluster.
[0103] As mentioned above, by obtaining the parameter values of the object to be processed corresponding to multiple indicators and merging them into feature vectors, the overall situation of the object to be processed can be reflected. Then, by clustering the feature vectors of preset objects with the same type of labels, several preset objects similar to the overall situation of the object to be processed can be accurately obtained.
[0104] F300 obtains several key objects associated with the reference object from a preset database, and obtains the feature tags corresponding to the reference object and each key object. This can be understood as follows: the key object and the reference object refer to an association relationship that meets preset requirements. For example, when it is necessary to find an investment object for the object to be processed, the key object can be an object that has been obtained after investing in the reference object.
[0105] Specifically, the step of obtaining several key objects associated with the reference object from a preset database includes the following steps: F301, based on the target knowledge graph stored in the preset database, several objects obtained from the target knowledge graph that have a connection relationship with the reference object are used as initial objects; it can be understood that: the target knowledge graph is a pre-constructed knowledge graph that includes a large number of objects and the relationship between objects.
[0106] F302: Obtain the type of each initial object and use the initial object with the same type as the reference object as the key object. For example, if the reference object is an enterprise, the initial object of the enterprise type is used as the key object. If the reference object also links to legal person information, the legal person information is not used as the key object.
[0107] As mentioned above, since the reference object is a screened object that is similar to the overall situation of the object to be processed, the key object associated with the reference object can also be considered to have a strong association with the object to be processed, which is conducive to the expansion of the associated tags of the object to be processed, and is conducive to finding more potential objects or potential users that are associated with the object to be processed, and improves the reliability of the potential objects or potential users found.
[0108] In F400, the feature labels corresponding to the reference object and each key object, along with the previously acquired feature labels of the object to be processed, are input into the pre-trained large model to obtain the expanded feature labels corresponding to the object to be processed. Based on the expanded feature labels, a list of target objects corresponding to the object to be processed is obtained. This can be understood as follows: when the pre-trained large model is used to expand the feature labels, semantic expansion of the labels is also included. The feature labels include technical field labels and business labels, among others.
[0109] For ease of understanding, let's take an example: the original feature tags for the object being processed include: lithium battery, energy storage system; the feature tags for the reference object and key object include: lithium battery, power battery, automotive industry, battery materials and recycling. The expanded output tags include: lithium battery, energy storage system, power battery, battery production equipment, automotive industry, and battery materials and recycling.
[0110] Specifically, the step of obtaining a target object list corresponding to the object to be processed according to the expanded feature tags includes the following steps: F401, obtaining a number of feature labels corresponding to each preset object.
[0111] F402 : For any preset object, if the feature tags corresponding to the preset object are consistent with the feature tags after expansion, the preset object is used as the target object corresponding to the object to be processed.
[0112] F403 , integrating the acquired target objects to obtain a target object list corresponding to the object to be processed.
[0113] As mentioned above, since the reference object and the key object are both potential objects with a high correlation with the object to be processed, by integrating and expanding the feature labels of the three, more and more reliable potential objects of the object to be processed can be found from the database based on the expanded labels, providing reliable data reference for subsequent operations of the object to be processed.
[0114] Furthermore, the method further comprises the following steps: F10, obtain several original objects associated with the object to be processed from the target knowledge graph stored in the preset database.
[0115] F20 , performing consistency comparison between a plurality of target objects in the target object list corresponding to the object to be processed and a plurality of original objects associated with the object to be processed, and identifying a target object that is different from each of the original objects from the plurality of target objects.
[0116] S30: Associating the identified target object with the object to be processed in the target knowledge graph to update the target knowledge graph.
[0117] As mentioned above, when the associated objects corresponding to the object to be processed are obtained, the associated objects of the object to be processed in the target knowledge graph are updated to ensure the timeliness and availability of the data, so that a more accurate and comprehensive data basis is available in the next calculation, ensuring the reliability of the calculation results.
[0118] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which can be set in an electronic device to store at least one instruction or at least one program related to implementing a method in a method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiment.
[0119] An embodiment of the present invention further provides an electronic device including a processor and the aforementioned non-transitory computer-readable storage medium.
[0120] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A knowledge graph presentation method, characterized in that: The method comprises the following steps: S100, obtaining a number of key nodes from a given knowledge graph; the key nodes are nodes in the given knowledge graph whose number of corresponding neighbor nodes is not less than n; S200, based on the number of key nodes, dividing the preset presentation page into an initial node grid including a plurality of grid units and obtaining a grid unit corresponding to each key node in the initial node grid; S300, calculating the area allocation score corresponding to each key node according to the number of neighboring nodes of different preset entity types connected to each key node and the preset presentation area corresponding to each node of the preset entity type; S400, based on the area allocation score corresponding to each key node, adjusting the area of the grid unit corresponding to each key node in the initial node grid to obtain a target node grid; S500, based on the target node grid, adjust the presentation position of each key node and non-key node to obtain the final target knowledge graph; wherein, the non-key node refers to any node in the given knowledge graph except the key node.
2. The knowledge graph presentation method according to claim 1, characterized in that: In step S200, dividing the preset presentation page into an initial node grid including a plurality of grid units includes the following steps: S201, calculating the aspect ratio R of the preset presentation page; S202, setting the number of rows and columns of the initial node grid to a and b respectively, traversing all values of a from 1 to L, and calculating the value of b respectively; where b = roundup(L / a), roundup() is a round-up function, and L is the number of key nodes; S203 , obtaining the a value and b value with the smallest absolute difference between the grid aspect ratio b / a and R, to obtain the initial node grid after division.
3. The knowledge graph presentation method according to claim 1, characterized in that: In step S200, obtaining the grid unit corresponding to each key node in the initial node grid includes the following steps: S210, calculating the correlation strength between every two key nodes and constructing a correlation matrix; S220, calculating the two-dimensional coordinates corresponding to each key node using a force-directed layout algorithm based on the association matrix; S230 , normalizing the two-dimensional coordinates corresponding to each key node to the coordinate space of the initial node grid, and obtaining the grid unit corresponding to each key node in the initial node grid.
4. The knowledge graph presentation method according to claim 1, characterized in that: The area allocation scores corresponding to the key nodes meet the following conditions: , where Q represents the area allocation score corresponding to any key node, m is the number of preset entity types, S i is the preset presentation area corresponding to the node of the i-th preset entity type, N i The number of nodes of the i-th preset entity type connected to the key node.
5. The knowledge graph presentation method according to claim 1, characterized in that: The preset presentation area corresponding to the node of the preset entity type is proportional to the importance of the preset entity type; or The preset presentation area corresponding to the nodes of the preset entity type is proportional to the average number of characters in the names of the nodes corresponding to the preset entity type.
6. The knowledge graph presentation method according to claim 1, characterized in that: Step S400 includes the following steps: S401, for any key node, calculating the ratio of the area allocation score corresponding to the key node to the total area allocation scores of all key nodes; S402 : Taking the product of the total presentation area of the initial node grid and the ratio as the area of the grid unit corresponding to the key node, so as to adjust the area of the grid unit corresponding to the key node in the initial node grid and obtain the target node grid.
7. The knowledge graph presentation method according to claim 1, characterized in that: Step S500 includes the following steps: S501, for any key node, based on the average coordinates of several neighboring nodes connected to the key node, the key node is moved within its own grid cell toward the direction of the average coordinates with a preset offset step size; S502, each non-critical node is initially assigned to a preset local area where a critical node with the highest correlation strength with the non-critical node is located, and a force-directed layout algorithm is used to calculate the local position of each non-critical node with the corresponding critical node as the center; S503, iteratively executing the above steps S501 and S502 until the average position change of all nodes is less than a preset change threshold or the maximum number of iterations is reached, and the final presentation positions of the key nodes and non-key nodes are obtained.
8. The knowledge graph presentation method according to claim 1, characterized in that: The method further comprises the steps of: P100, based on the obtained first entity to be analyzed, the second entity to be analyzed, and the preset attribute parameter sets corresponding to each entity to be analyzed, calculating the data similarity between the preset attribute parameter sets corresponding to the first entity to be analyzed and the second entity to be analyzed respectively; The first entity to be analyzed and the second entity to be analyzed are entities corresponding to any two nodes of a target entity type among several preset entity types; P200, if the name data corresponding to the first entity to be analyzed is the same as the name data corresponding to the second entity to be analyzed, and when the data similarity is greater than a first preset similarity threshold, determine that the first entity to be analyzed and the second entity to be analyzed are the same entity.
9. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the knowledge graph presentation method as described in any one of claims 1-8.
10. An electronic device, characterized in that: The device comprises a processor and the non-transitory computer-readable storage medium as claimed in claim 9.
Citation Information
Patent Citations
Knowledge graph visualization method and system
CN112597317A
Knowledge graph visualization method based on force guidance
CN113254670A
Knowledge graph visualization method
CN116090555A
Visualization method and device for subject knowledge graph
CN116108922A
Knowledge graph reasoning method and system based on space-time relationship
CN116362335A