Knowledge Graph-Based Network Generation Method, Reasoning Method, System, and Terminal

By dividing and indexing the knowledge graph database, a network audit model is generated, which solves the problems of high computing resource consumption and long response time in super-large-scale knowledge graphs, and efficient data inference and rapid data acquisition are achieved.

CN119990340BActive Publication Date: 2025-07-08HANGZHOU BYTE ARK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510480168.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-08
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

When faced with super-large-scale knowledge graphs, existing inference technologies consume high computing resources and have long response time, making it difficult to achieve efficient data inference.

Method used

By dividing the graph database based on preset entity categories, obtaining the data set to be marked, indexing and marking using the correlation frequency, generating a network audit model, combining low-dimensional embedding vectors and lightweight embedding representation strategies, data processing and inference processes are optimized.

Benefits of technology

When facing a large number of physical scenarios, it can quickly obtain target data, shorten data inference response time, improve data search accuracy and efficiency, and reduce computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990340B_ABST
    Figure CN119990340B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of knowledge graph construction, and in particular, to a network generation method, an inference method, a system, and an intelligent terminal based on a knowledge graph. The network generation method includes the following steps: dividing a preset graph database based on a preset entity category to obtain several sets of data to be labeled; obtaining a corresponding set of entities to be labeled according to the set of data to be labeled; obtaining associated entities based on attribute features, and obtaining an association frequency according to the entities to be labeled and the associated entities; obtaining feature-labeled entities according to the association frequency; generating a set of labeled entities corresponding to the set of entities to be labeled according to the feature-labeled entities, and obtaining a network audit model corresponding to the set of labeled entities based on index marking. This application performs different markings on the set of entities to be labeled, and can quickly locate according to the index marking when obtaining target data in the later stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of knowledge graph construction, and in particular to a network generation method, an inference method, a system, and a terminal based on a knowledge graph. Background Art

[0002] A knowledge graph is a modern theory that uses a visual graph to vividly display the core structure, development history, frontier fields, and overall knowledge architecture of a discipline to achieve the purpose of multi-disciplinary integration. It displays complex knowledge fields through data mining, information processing, knowledge metrology, and graphic drawing, reveals the dynamic development laws of knowledge fields, and provides practical and valuable references for disciplinary research.

[0003] With the development of deep learning and big data technologies, reasoning based on knowledge graphs has become one of the hotspots in AI research, especially showing great potential in improving search quality, personalized recommendations, etc. Most of the existing reasoning technologies based on knowledge graphs focus on theoretical modeling, using high-performance computing units such as GPUs for parallel computing, and reasoning through pre-trained neural network models.

[0004] However, when facing ultra-large-scale knowledge graphs, a high-performance GPU is required for reasoning. Therefore, there are problems such as high computational resource consumption and long response time. Summary of the Invention

[0005] To reduce the consumption of computing resources and improve the inference response time, this application provides a network generation method, an inference method, a system, and a terminal based on a knowledge graph.

[0006] In a first aspect, this application provides a network generation method based on a knowledge graph, adopting the following technical solution:

[0007] A network generation method based on a knowledge graph includes the following steps:

[0008] Partition a preset graph database based on a preset entity category to obtain several sets of data to be labeled;

[0009] Obtain a corresponding set of entities to be labeled according to the set of data to be labeled, where the set of entities to be labeled includes several entities to be labeled;

[0010] Obtain associated entities corresponding to the entities to be labeled based on the attribute features corresponding to the entities to be labeled, and obtain the association frequency according to the entities to be labeled and the associated entities;

[0011] Index and label the entities to be labeled and the associated entities according to the association frequency to obtain feature-labeled entities;

[0012] Generate a set of labeled entities corresponding to the set of entities to be labeled according to the feature-labeled entities, and obtain a network audit model corresponding to the set of labeled entities based on the index label.

[0013] By adopting the above technical solution, the preset graph database is divided based on the preset entity category to obtain different sets of data to be labeled, the corresponding set of entities to be labeled is obtained according to the set of data to be labeled of the same type, the corresponding associated entities are obtained based on the attribute features corresponding to the entities to be labeled, and thus the corresponding association frequency is obtained based on the two. The two are labeled based on the association frequency, thereby forming a network audit model with index labels. When facing more entity scenarios, the corresponding target data can be quickly obtained according to different index labels, shortening the overall data inference response time, and thus realizing efficient inference in real-time scenarios. Different labels are applied to the set of entities to be labeled, which can quickly locate according to the index label when obtaining target data in the later stage.

[0014] In some of the embodiments, the index label includes a category label and an association label. The indexing and labeling of the entity to be labeled and the associated entity based on the association frequency include the following steps:

[0015] Determine whether the association frequency is a first preset frequency;

[0016] If the association frequency is the first preset frequency, perform the association label on the entity to be labeled and the associated entity;

[0017] If the association frequency is not the first preset frequency, determine whether the association frequency is within a first preset frequency set;

[0018] If the association frequency is within the first preset frequency set, perform the category label and the association label on the entity to be labeled and the associated entity;

[0019] If the association frequency is not within the first preset frequency set, determine whether the association frequency is within a second preset frequency set;

[0020] If the association frequency is within the second preset frequency set, perform the category label on the entity to be labeled and the associated entity.

[0021] By adopting the above technical solution, the correlation degree between the entity to be labeled and the associated entity is determined based on the association frequency, so as to determine the index marking types of the entity to be labeled and the associated entity, generate different index marks for different entities to be labeled and associated entities, ensure the accuracy of the generated data for the network audit model, facilitate subsequent data search, improve the accuracy of the target data and quickly obtain the target data. For the entity to be labeled and the associated entity, they generally exist simultaneously, that is, when the correlation is relatively high, the two can be marked according to their association marks. For the entity to be labeled and the associated entity with a certain degree of correlation but not very high, they can be respectively marked with category marks and association marks. Subsequently, when searching for data to be found, the corresponding target data can be obtained more accurately, the search duration can be reduced, and the efficiency of obtaining the target data can be improved.

[0022] In some of the embodiments, partitioning the preset graph database based on the preset entity categories to obtain a plurality of sets of data to be labeled includes the following steps:

[0023] Obtain the data distribution characteristics of the preset graph database based on statistical analysis, and identify the high-density regions according to the data distribution characteristics;

[0024] Determine whether there are multiple entity categories in the high-density regions;

[0025] If there are multiple entity categories in the high-density regions, partition the preset graph database according to the entity categories to obtain a plurality of sets of the data to be labeled;

[0026] If there are not multiple entity categories in the high-density regions, obtain the entity features and the corresponding attribute features based on the high-density regions, and generate low-dimensional embedding vectors for the entity features and the corresponding attribute features based on a preset network model.

[0027] By adopting the above technical solution, the data distribution characteristics of the preset graph database are determined based on statistical analysis, the high-density regions are identified according to the data distribution characteristics, it is determined whether there are multiple entity categories in the high-density regions. If there are not multiple entity categories in the high-density regions, the entity features and the corresponding attribute features are obtained based on the high-density regions, and low-dimensional embedding vectors are generated for the entity features and the corresponding attribute features based on a preset network model. By reasonably obtaining the sets of data to be labeled with different data volumes, even in the case of extremely unbalanced data distribution, a good partitioning effect can be maintained, ensuring that the load of each set of data to be labeled is roughly equal, and improving the retrieval efficiency of the target data.

[0028] In some of the embodiments, after identifying the high-density regions according to the data distribution characteristics, the following steps are further included:

[0029] Obtain the corresponding number of data entities based on the high-density region, and compare the number of data entities with a preset compression number;

[0030] If the number of data entities is greater than the preset compression number, generate a low-dimensional embedding vector based on a preset network model for the high-density region and the corresponding attribute features.

[0031] By adopting the above technical solution, determine the data processing method for the high-density region according to the number of actual application scenarios. If the number of data entities is greater than the preset compression number, generate a low-dimensional embedding vector based on a preset network model for the high-density region and the corresponding attribute features, and compress the vector representations of entities and relationships through a lightweight embedding representation strategy, thereby reducing the GPU memory burden.

[0032] In some of the embodiments, obtaining the association frequency based on the entity to be labeled and the associated entity includes the following steps:

[0033] Obtain the regional positions corresponding to all the entities to be labeled based on the preset graph data, and divide the data to be scanned within the regional positions according to a preset range;

[0034] Determine whether the data to be scanned includes the associated entity;

[0035] If the data to be scanned includes the associated entity, perform a frequency mark on the regional position; obtain the association frequency based on the frequency mark and the total number of entities to be labeled.

[0036] In some of the embodiments, before obtaining the association frequency based on the frequency mark and the total number of entities to be labeled, the following steps are further included:

[0037] Obtain the information features between the entity to be labeled and the associated entity based on the data to be scanned;

[0038] Analyze the information features to obtain an analysis result, and determine whether the analysis result is a valid result;

[0039] If the analysis result is a valid result, remove the frequency mark of the entity to be labeled.

[0040] In a second aspect, the present application provides a network inference method based on a knowledge graph, adopting the following technical solution:

[0041] A network inference method based on a knowledge graph, based on the network audit model generated by the above-mentioned network generation method based on a knowledge graph, includes the following steps:

[0042] Obtain the information to be recognized, and based on the information to be recognized, obtain the main entity features and corresponding main attribute features;

[0043] Obtain the corresponding entity category according to the main entity features, and obtain the main network information corresponding to the main entity features in the network audit model according to the entity category;

[0044] Obtain the estimated association frequency according to the main entity features and the main attribute features, and obtain the estimated network data with corresponding index marks in the main network information according to the estimated association frequency;

[0045] Filter according to the main attribute features in the estimated network data to obtain the tail entity information corresponding to the information to be recognized.

[0046] By adopting the above technical solution, based on obtaining the main entity features and corresponding main attribute features in the information to be recognized, obtaining the corresponding entity category according to the main entity features, and filtering the data of the network audit model according to the entity category, so as to obtain the main network information with the same entity category mark, then obtaining the estimated association frequency according to the main entity features and the main attribute features, and obtaining the estimated network data with corresponding index marks in the main network information according to the estimated association frequency, and gradually obtaining the corresponding data through different index marks, reducing the search steps and improving the overall search efficiency; finally, filtering according to the main attribute features in the estimated network data to obtain the tail entity information corresponding to the information to be recognized, improving the accuracy of obtaining the tail entity information.

[0047] In some embodiments, after obtaining the main network information corresponding to the main entity features in the network audit model according to the entity category, the following steps are further included: obtaining the tail entity type according to the main attribute features, and judging whether the tail entity type belongs to the common attribute;

[0048] If the tail entity type belongs to the common attribute, obtain the common mark data in the main network information based on the main attribute features;

[0049] Filter according to the main attribute features in the common mark data to obtain the tail entity information corresponding to the information to be recognized.

[0050] By adopting the above technical solution, obtain the tail entity type based on the main attribute features, and then judge the tail entity type to obtain the type to which the corresponding tail entity information belongs, and narrow down the filtered data according to the type, thereby reducing the search duration of the tail entity information and improving the search efficiency.

[0051] In the third aspect, the present application provides a network generation system based on a knowledge graph, adopting the following technical solution:

[0052] A network generation system based on a knowledge graph, comprising:

[0053] A data partitioning module, which is used to partition a preset graph database based on preset entity categories to obtain several sets of to-be-labeled data sets;

[0054] An entity acquisition module, which is used to obtain a corresponding set of to-be-labeled entities according to the to-be-labeled data sets, and the set of to-be-labeled entities includes several to-be-labeled entities;

[0055] A frequency acquisition module, which is used to obtain associated entities corresponding to the to-be-labeled entities based on the attribute features corresponding to the to-be-labeled entities, and obtain an association frequency according to the to-be-labeled entities and the associated entities;

[0056] A label generation module, which is used to perform index labeling on the to-be-labeled entities and the associated entities according to the association frequency to obtain feature-labeled entities;

[0057] A model generation module, which is used to generate a set of labeled entities corresponding to the set of to-be-labeled entities according to the feature-labeled entities, and obtain a network audit model corresponding to the set of labeled entities based on the index labeling.

[0058] In a fourth aspect, the present application provides an intelligent terminal, adopting the following technical solution: An intelligent terminal, the intelligent terminal includes a processor and a memory that are coupled to each other, and a computer program capable of running on the processor is stored on the memory;

[0059] When the computer program is executed by the processor, it implements the network generation method based on the knowledge graph as described in the first aspect.

[0060] In summary, the present application includes at least one of the following beneficial technical effects:

[0061] 1. Partition the preset graph database based on preset entity categories, obtain different sets of to-be-labeled data, obtain corresponding sets of to-be-labeled entities according to the same type of to-be-labeled data sets, obtain corresponding associated entities based on the attribute features corresponding to the to-be-labeled entities, thereby obtaining corresponding association frequencies according to the two, and perform labeling on the two based on the association frequencies, so as to form a network audit model with index labeling. When facing a large number of entity scenarios, the corresponding target data can be quickly obtained according to different index labels, shortening the overall data reasoning response time, thereby achieving efficient reasoning in real-time scenarios. Performing different labels on the set of to-be-labeled entities can quickly locate according to the index labels when obtaining target data later;

[0062] 2. Determine the correlation degree between the entity to be labeled and the associated entity based on the association frequency, so as to determine the index label types of the entity to be labeled and the associated entity, generate different index labels for different entities to be labeled and associated entities, ensure the accuracy of the generated data in the network audit model, facilitate subsequent data search, improve the accuracy of target data and quickly obtain target data. For the entity to be labeled and the associated entity, they generally exist simultaneously, that is, when the correlation is high, the two can be labeled according to their association labels. For the entity to be labeled and the associated entity with a certain but not very high correlation, they can be respectively labeled with category labels and association labels. Subsequently, when searching for data to be searched, the corresponding target data can be obtained more accurately, reducing the search time and improving the efficiency of obtaining target data. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is a block diagram of a network generation method based on a knowledge graph provided by an embodiment of the present application;

[0064] Figure 2 is a block diagram of a method for obtaining a set of data to be labeled provided by an embodiment of the present application;

[0065] Figure 3 is a block diagram of a method for obtaining the association frequency provided by an embodiment of the present application;

[0066] Figure 4 is a schematic structural diagram of a network generation system provided by an embodiment of the present application;

[0067] Figure 5 is a block diagram of the structure of an intelligent terminal provided by this embodiment;

[0068] Figure 6 is a block diagram of a network inference method based on a knowledge graph provided by the present application.

[0069] Description of the reference numerals: 10, data division module; 20, entity acquisition module; 30, frequency acquisition module; 40, label generation module; 50, model generation module; 61, processor; 62, memory; 63, computer program. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] To more clearly understand the purpose, technical solution, and advantages of this application, the following describes and explains this application in conjunction with the accompanying drawings and embodiments. However, those of ordinary skill in the art should understand that this application can be implemented without these details. In some cases, to avoid unnecessary descriptions from obscuring aspects of this application, well-known methods, processes, systems, components, and / or circuits that have been described at a higher level will not be elaborated further. For those of ordinary skill in the art, it is obvious that various changes can be made to the disclosed embodiments of this application, and without departing from the principles and scope of this application, the general principles defined in this application can be applied to other embodiments and application scenarios. Therefore, this application is not limited to the illustrated embodiments, but conforms to the broadest scope consistent with the scope claimed in this application.

[0071] An embodiment of this application discloses a network generation method based on a knowledge graph.

[0072] Referring to Figure 1 , the network generation method based on a knowledge graph includes the following steps:

[0073] S100, partitioning a preset graph database based on a preset entity category to obtain several sets of data to be labeled.

[0074] Among them, the preset entity category represents different types of entities, the preset graph database includes data that needs to be sorted out in the knowledge graph, and the data to be labeled represents different data sets obtained by partitioning the preset graph database in chunks.

[0075] It should be noted here that an entity is the basic unit in a knowledge graph, representing an object in the real world, such as a person, a place, an organization, etc. The preset entity category is specifically based on the attributes, relationships, and context of the entity.

[0076] Exemplarily, in reality, entities can be classified into people, organizations, places, events, etc. according to specific properties. In natural language processing, entity recognition is a common application of entity classification, which involves identifying entities with specific categories from text.

[0077] Exemplarily, based on BiLSTM-CRF for entity recognition, the technology combines a bidirectional long short-term memory network (BiLSTM) and a conditional random field (CRF) model, which performs excellently in the NER task. BiLSTM can capture long-distance dependencies in text, while the CRF layer can utilize the constraint relationships between adjacent labels to improve the accuracy of annotation.

[0078] To facilitate the processing of graph data, a preset graph database stores a relational network graph, which is a graph generated by extracting entity features from the preset graph data and is specifically generated based on a preset model. The preset model is a pre-trained neural network model, and the neural network model uses an existing graph neural network, which will not be elaborated here.

[0079] Here, entity recognition is performed on the preset graph database, mainly using traditional entity recognition methods, specifically relying on a large number of rules and dictionaries. However, this method often has poor performance when dealing with complex texts. With the development of deep learning, neural network-based methods have become mainstream. For example, a model combining bidirectional long short-term memory network (BiLSTM) and conditional random field (CRF) performs well in the NER task.

[0080] It should be noted here that the data stored in the preset graph database also includes the relationships between entities and the attributes corresponding to each entity, that is, used to describe and supplement the entities. The amount of data stored in the obtained set of data to be labeled is roughly the same, mainly to facilitate the processing of data by intelligent terminals.

[0081] The intelligent terminal mentioned here processes and stores data for the entire preset graph database, facilitating subsequent clients to quickly obtain corresponding information based on this knowledge graph-based network inference method.

[0082] When the scale of the preset graph database is huge, containing hundreds of millions of nodes and relationships, the traditional way to obtain data is that when searching for information based on the knowledge graph generated by the preset graph database, it is necessary to traverse and retrieve the entire graph, resulting in a long retrieval time. At this time, we divide the entire knowledge graph into thousands of small graph blocks, and each graph has a corresponding multi-level index table to support fast positioning.

[0083] Combined Figure 2 , based on the preset entity categories, the preset graph database is divided to obtain several sets of data to be labeled, including the following steps:

[0084] S110, obtain the data distribution characteristics of the preset graph database based on statistical analysis, and identify the high-density areas according to the data distribution characteristics.

[0085] S120, determine whether there are multiple entity categories in the high-density areas.

[0086] S130, if there are multiple entity categories in the high-density areas, then divide the preset graph database according to the entity categories to obtain several sets of data to be labeled.

[0087] S140. If there are no multiple entity categories in the high-density region, obtain the entity features and corresponding attribute features based on the high-density region, and generate low-dimensional embedding vectors for the entity features and corresponding attribute features based on a preset network model.

[0088] Among them, the data distribution feature represents the data distribution shown in the relationship network diagram formed in the preset graph database. The high-density region represents the data region with multiple entity features. The entity category represents that the entity belongs to different types.

[0089] Here, it should be noted that first, the data distribution feature of the knowledge graph is determined through statistical analysis, and then a two-stage partitioning algorithm is designed. In the first stage, a clustering algorithm is used to identify the high-density region and the sparse region to form a preliminary partitioning suggestion. In the second stage, combined with the results of the preliminary partitioning, the genetic algorithm is used for iterative optimization to find the best partitioning scheme. In this way, even in the case of extremely unbalanced data distribution, good partitioning results can be maintained, ensuring that the load of each tile is roughly equal and maximizing the retrieval efficiency.

[0090] S200. Obtain the corresponding set of entities to be labeled according to the set of data to be labeled.

[0091] Among them, the set of entities to be labeled includes several entities to be labeled. The entity to be labeled represents the entity that needs to be subjected to subsequent labeling processing. The specific method for obtaining the corresponding set of entities to be labeled according to the set of data to be labeled is through a pre-trained language model.

[0092] The preset pre-trained language model can be BERT (Bidirectional Encoder Representations from Transformers). By understanding the context semantics, it can more accurately identify entities. Models such as BERT learn rich language features through pre-training on a large amount of unlabeled text, and thus can be effectively applied to the entity recognition task.

[0093] Exemplarily, specific characters and locations are identified in fairy tales. For example, by analyzing a fairy tale "Little Red Riding Hood", the entity recognition system can identify entities such as the protagonist, the story name, and important events mentioned in the text. This is of great significance to fields such as information aggregation, information retrieval, and analysis.

[0094] S300. Obtain the associated entities corresponding to the entities to be labeled based on the attribute features, and obtain the association frequency based on the entities to be labeled and the associated entities.

[0095] Among them, the attribute features characterize the relationship between entities, so as to confirm whether there is an association between entities. Specifically, traditional methods rely on rules and pattern matching, while machine learning-based attribute extraction methods can automatically identify attributes by learning patterns in data. Deep learning, especially methods based on pre-trained models such as RNN (Recurrent Neural Network) and BERT, has shown excellent performance in attribute extraction. These models can capture context information, thus more accurately identifying and classifying attributes. With the development of deep learning technology, especially the emergence of pre-trained language models (such as BERT), the accuracy and efficiency of attribute extraction have been significantly improved. These models can understand complex context information, thus more accurately extracting relevant attributes.

[0096] Refer to Figure 3 , the association frequency characterizes the degree of association between the entity to be labeled and the associated entity. The association frequency is obtained based on the entity to be labeled and the associated entity, including the following steps:

[0097] S310, obtain the regional positions corresponding to all entities to be labeled based on the preset graph data, and divide the data to be scanned within the preset range at the regional positions.

[0098] S320, determine whether the data to be scanned includes associated entities.

[0099] S330, if the data to be scanned includes associated entities, then perform frequency marking on the regional positions.

[0100] S340, obtain the association frequency based on the frequency marking and the total number of entities to be labeled.

[0101] Among them, the regional position characterizes the position obtained in the preset graph data, and this position is the area where the entity to be labeled is located. The data to be scanned characterizes the data obtained according to the preset range, and the preset range is a range set in advance, and this preset range can be set according to a sentence or a paragraph in the preset graph data.

[0102] Determine whether the data to be scanned includes associated entities. If the data to be scanned includes associated entities, it means that the associated entity and the entity to be labeled appear synchronously. Therefore, it is necessary to perform frequency marking on this regional position. After all entities to be labeled in the preset graph data have been executed the above steps S310 - step S330, then compare all frequency markings with the total number of all entities to be labeled according to step S340, so as to obtain the corresponding association frequency.

[0103] S400, perform index marking on the entity to be labeled and the associated entity according to the association frequency to obtain the feature-labeled entity.

[0104] Among them, the feature marking entity represents the entity with a retrieval index. The index marking includes a category marking and an association marking. Index marking the entity to be marked and the associated entity according to the association frequency includes the following steps:

[0105] S410, determine whether the association frequency is the first preset frequency.

[0106] S420, if the association frequency is the first preset frequency, perform association marking on the entity to be marked and the associated entity.

[0107] S430, if the association frequency is not the first preset frequency, determine whether the association frequency is within the first preset frequency set.

[0108] S440, if the association frequency is within the first preset frequency set, perform category marking and association marking on the entity to be marked and the associated entity.

[0109] S450, if the association frequency is not within the first preset frequency set, determine whether the association frequency is within the second preset frequency set.

[0110] S460, if the association frequency is within the second preset frequency set, perform category marking on the entity to be marked and the associated entity. It should be noted here that the index marking can specifically be a hash index.

[0111] Among them, the first preset frequency represents that the entity to be marked and the associated entity are completely related, that is, the associated entity will appear wherever the entity to be marked appears. Therefore, the first preset frequency can be set to 1. The first preset frequency set represents that there is a correlation between the entity to be marked and the associated entity. The first preset frequency set can be set to 0.4 - 1, and the first preset frequency set is a set that is closed at the front and open at the back. The second preset frequency set represents that there is an association between the entity to be marked and the associated entity, but the relevance is not high. Therefore, the association between the entity to be marked and the associated entity can be ignored. The second preset frequency set can be set to 0 - 0.4, and the second preset set is a set that is closed at the front and open at the back.

[0112] It should be noted here that the association marking corresponding to the feature marking entity is at the first - level marking, and the category marking is at the second - level marking. When performing a search, the search will be carried out according to the markings of different levels.

[0113] Specifically, in steps S410 - S460, it is described to determine whether the association frequency is the first preset frequency. If the association frequency is the first preset frequency, then perform an association mark on the entity to be marked and the associated entity. If the association frequency is not the first preset frequency, then determine whether the association frequency is within the first preset frequency set. If the association frequency is within the first preset frequency set, then perform a category mark and an association mark on the entity to be marked and the associated entity. If the association frequency is not within the first preset frequency set, then determine whether the association frequency is within the second preset frequency set. If the association frequency is within the second preset frequency set, then perform a category mark on the entity to be marked and the associated entity.

[0114] S500, generate a set of marked entities corresponding to the set of entities to be marked based on the feature - marked entities, and obtain a network audit model corresponding to the set of marked entities based on the index mark.

[0115] Among them, the set of marked entities indicates that there are multiple marked entities in this set. Therefore, based on the existing database, a network audit model is generated from the set of marked entities. This network audit model specifically includes a knowledge graph generated according to preset graph data, which is convenient for subsequent client users to obtain corresponding information based on this network audit model. Moreover, the index established according to this technical solution can reduce the time cost of customer retrieval.

[0116] In one of the embodiments, after identifying the high - density area according to the data distribution characteristics, the following steps are further included:

[0117] S600, obtain the corresponding number of data entities according to the high - density area, and compare the number of data entities with the preset compression number.

[0118] S610, if the number of data entities is greater than the preset compression number, then generate a low - dimensional embedding vector based on the preset network model for the high - density area and the corresponding attribute features.

[0119] Among them, the number of data entities represents the number of data entities existing in the high - density area, and the preset compression number represents the lowest data standard for determining that the data in the high - density area needs to be compressed. If the number of data entities is not greater than the preset compression number, then there is no need to compress the high - density area, and the intelligent network generation end can make corresponding processing. When the number of data entities is greater than the preset compression number, then generate a low - dimensional embedding vector based on the preset network model for the high - density area and the corresponding attribute features. Specifically, a lightweight embedding representation strategy is adopted to compress the vector representations of entities and relationships, reduce the GPU memory burden. For example, compress the vector dimensions of entities and relationships to below 64 or 32 dimensions, and generate a low - dimensional embedding vector through a neural network model.

[0120] In one of the embodiments, before obtaining the association frequency based on the frequency tag and the total number of entities to be tagged, the following steps are further included:

[0121] S700. Obtain the information features between the entity to be tagged and the associated entity according to the data to be scanned.

[0122] S710. Analyze the information features to obtain an analysis result, and determine whether the analysis result is a valid result.

[0123] S720. If the analysis result is a valid result, then remove the frequency tag of the entity to be tagged.

[0124] Wherein, the information features represent the data information existing between the entity to be tagged and the associated entity. If the analysis result is a valid result, it means that there are other entities in the information features, and the corresponding relationship between the entity to be tagged and the associated entity does not exist. Therefore, the frequency tag needs to be removed.

[0125] The embodiment of the present application also discloses a network generation system based on a knowledge graph.

[0126] Refer to Figure 4 , a network inference system based on a knowledge graph, including a data partitioning module 10, an entity acquisition module 20 network-connected to the data partitioning module 10, a frequency acquisition module 30 network-connected to the entity acquisition module 20, a tag generation module 40 network-connected to the frequency acquisition module 30, and a model generation module 50 network-connected to the tag generation module 40.

[0127] The data partitioning module 10 is used to partition a preset graph database based on a preset entity category to obtain several groups of data sets to be tagged. The entity acquisition module 20 is used to obtain a corresponding set of entities to be tagged according to the data set to be tagged, and the set of entities to be tagged includes several entities to be tagged. The frequency acquisition module 30 is used to obtain an associated entity corresponding to the entity to be tagged based on the attribute features corresponding to the entity to be tagged, and obtain an association frequency according to the entity to be tagged and the associated entity. The tag generation module 40 is used to index and tag the entity to be tagged and the associated entity according to the association frequency to obtain a feature-tagged entity. The model generation module 50 is used to generate a set of tagged entities corresponding to the set of entities to be tagged according to the feature-tagged entity, and obtain a network audit model corresponding to the set of tagged entities based on the index tag.

[0128] The model generation module 50 is used to generate a set of tagged entities corresponding to the set of entities to be tagged according to the feature-tagged entity, and obtain a network audit model corresponding to the set of tagged entities based on the index tag.

[0129] Other functions performed by the above data partitioning module 10, entity acquisition module 20, frequency acquisition module 30, tag generation module 40, and model generation module 50, as well as the technical details of each function, are the same as or similar to the corresponding features in the previously described knowledge graph-based network generation method, and thus will not be elaborated here.

[0130] An embodiment of the present application also discloses an intelligent terminal.

[0131] Refer to Figure 5 , the intelligent terminal includes a processor 61 and a memory 62 that are coupled to each other. A computer program 63 that can run on the processor 61 is stored on the memory 62. When the computer program 63 is executed by the processor 61, a knowledge graph-based generation method is implemented.

[0132] Among them, the memory 62 can be a ROM or other types of static storage devices that can store static information and instructions, a random access memory, or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory, a compact disc read-only memory, or other optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 62 can be an internal storage unit in some embodiments.

[0133] In addition, the processor 61 can be a central processing unit, a general-purpose processor, a data signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It is used to run the program code stored in the memory 62 or process data.

[0134] Specifically, the processor 61 and the memory 62 are connected by a bus. The bus can include a path for transmitting information between the above components. The bus can be a peripheral component interconnect standard bus or an extended industry standard architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only a thick line is used to represent it in, but it does not mean that there is only one bus or one type of bus.

[0135] Figure 5 Only the intelligent terminal with the memory 62, the processor 61, and the bus is shown. Those skilled in the art can understand that, Figure 5The structures shown do not constitute a limitation on the intelligent terminal. It can be a bus structure or a star structure. The intelligent terminal can also include more or fewer components than those shown, or combine certain components, or have different component deployments. Other existing or future electronic devices that are applicable should also be included within the scope of protection and are hereby incorporated by reference.

[0136] The implementation principle is as follows: First, based on the data distribution characteristics of the preset graph database obtained by the data processing end through statistical analysis, the intelligent terminal identifies high-density regions. By determining whether there are multiple entity categories in the high-density regions, if there are multiple entity categories in the high-density regions, the preset graph database is divided according to the entity categories to obtain several sets of data to be labeled. If there are not multiple entity categories in the high-density regions, the entity features and corresponding attribute features are obtained based on the high-density regions, and low-dimensional embedding vectors are generated based on the entity features and corresponding attribute features using the preset network model.

[0137] Next, the intelligent terminal obtains the corresponding set of entities to be labeled based on the set of data to be labeled, and obtains associated entities corresponding to the entities to be labeled based on the attribute features, and obtains the association frequency based on the entities to be labeled and the associated entities.

[0138] Finally, the intelligent terminal performs index marking on the entities to be labeled and the associated entities according to the association frequency to obtain feature-marked entities, generates a set of marked entities corresponding to the set of entities to be labeled based on the feature-marked entities, and obtains a network audit model corresponding to the set of marked entities based on the index marking.

[0139] Referring to Figure 6 , the embodiment of the present application also discloses a network inference method based on a knowledge graph. Based on the network audit model generated by the network generation method based on the knowledge graph provided in the above embodiment, it includes the following steps:

[0140] S800, obtain the information to be recognized, and obtain the main entity features and corresponding main attribute features based on the information to be recognized.

[0141] S810, obtain the corresponding entity category based on the main entity features, and obtain the main network information corresponding to the main entity features in the network audit model according to the entity category.

[0142] S820, obtain the estimated association frequency based on the main entity features and the main attribute features, and obtain the estimated network data with corresponding index markings in the main network information according to the estimated association frequency.

[0143] S830, screen according to the main attribute features in the estimated network data to obtain the tail entity information corresponding to the information to be recognized.

[0144] Among them, the information to be recognized represents the information input at the client usage end, specifically including the question of the target information that the client needs to obtain. The information to be recognized can be the information directly obtained by the intelligent terminal from the client input. The main entity feature represents the main body feature that the feature of the information to be recognized needs to obtain. Here, the feature extraction of the information to be recognized is the same as or similar to the description in the above embodiments, and will not be elaborated here.

[0145] It should be noted that since the information to be recognized includes several entity features, it is necessary to analyze the several entity features to obtain the main entity feature. The main attribute feature represents the attribute corresponding to the main entity feature. Here, there can be multiple main attribute features. Therefore, the intelligent terminal can specifically perform corresponding processing based on the information to be recognized.

[0146] The main network information represents the entity category obtained based on the main entity feature, and the information obtained in the same entity category of the network audit model based on the entity category. The estimated association frequency represents the estimated frequency obtained by the intelligent terminal based on the main entity feature and the main attribute feature. This estimated association frequency can be specifically obtained through the historical search data. Match the main entity feature and the main attribute feature in the search historical data. If a match is found, directly locate the position of the existing estimated network data, and perform a more specific search based on the main entity feature.

[0147] In one of the embodiments, after obtaining the main network information corresponding to the main entity feature in the network audit model according to the entity category, the following steps are further included:

[0148] S910, obtain the tail entity type according to the main attribute feature, and determine whether the tail entity type belongs to the common attribute.

[0149] S920, if the tail entity type belongs to the common attribute, obtain the common marked data in the main network information based on the main attribute feature.

[0150] S930, screen in the common marked data according to the main attribute feature to obtain the tail entity information corresponding to the information to be recognized.

[0151] Among them, the common attributes mainly include the attributes existing in the public area, such as country, location, etc. By asking where a person is from through the main attribute feature, through analyzing the main attribute feature, it can be obtained that the tail entity type is country, and country is a common attribute. Therefore, the common marked data with country marks can be directly obtained, and screening is performed in this country to screen out the country with the same mark as the main entity feature as the tail entity information corresponding to the main entity feature.

[0152] It should be noted here that the network inference method based on the knowledge graph is specifically applied to the customer usage end corresponding to the intelligent terminal. Furthermore, based on the network inference method based on the knowledge graph, the network inference based on the knowledge graph can be realized, which is convenient for obtaining the information of the tail entity.

[0153] Combined with Figure 5 , it should be noted here that the intelligent terminal includes a processor 61 and a memory 62 that are coupled to each other. The processor 61 is also used to execute the above steps S800 - S830 and execute the steps S900 - S830. A computer program 63 that can run on the processor 61 is stored on the memory 62. When the computer program 63 is executed by the processor 61, it is the network inference method based on the knowledge graph.

[0154] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and they can be executed in other orders.

[0155] The above are all the preferred embodiments of this application. It does not limit the protection scope of this application accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of this application should be covered within the protection scope of this application.

Claims

1. A network generation method based on a knowledge graph, characterized in that, Including the following steps: Dividing a preset graph database based on a preset entity category to obtain several sets of data to be labeled; Obtaining a corresponding set of entities to be labeled according to the set of data to be labeled, the set of entities to be labeled including several entities to be labeled; Obtaining associated entities corresponding to the entities to be labeled based on the attribute features corresponding to the entities to be labeled, and obtaining an association frequency according to the entities to be labeled and the associated entities; Indexing and labeling the entities to be labeled and the associated entities according to the association frequency to obtain feature-labeled entities; Generating a set of labeled entities corresponding to the set of entities to be labeled according to the feature-labeled entities, and obtaining a network audit model corresponding to the set of labeled entities based on the index labeling; Wherein, the index labeling includes a category label and an association label, and the indexing and labeling the entities to be labeled and the associated entities according to the association frequency includes the following steps: Judging whether the association frequency is a first preset frequency; If the association frequency is the first preset frequency, performing the association label on the entities to be labeled and the associated entities; If the association frequency is not the first preset frequency, judging whether the association frequency is within a first preset frequency set; If the association frequency is within the first preset frequency set, performing the category label and the association label on the entities to be labeled and the associated entities; If the association frequency is not within the first preset frequency set, judging whether the association frequency is within a second preset frequency set; If the association frequency is within the second preset frequency set, performing the category label on the entities to be labeled and the associated entities.

2. The network generation method based on a knowledge graph according to claim 1, wherein The dividing the preset graph database based on a preset entity category to obtain several sets of data to be labeled includes the following steps: Obtaining the data distribution characteristics of the preset graph database based on statistical analysis, and identifying high-density regions according to the data distribution characteristics; Judging whether there are multiple entity categories in the high-density region; If there are multiple entity categories in the high-density region, dividing the preset graph database according to the entity categories to obtain several sets of the data to be labeled; If there are not multiple entity categories in the high-density region, obtaining entity features and corresponding attribute features based on the high-density region, and generating low-dimensional embedding vectors based on the entity features and the corresponding attribute features by a preset network model.

3. The network generation method based on a knowledge graph according to claim 2, wherein After identifying the high-density region according to the data distribution characteristics, the following steps are further included: Obtaining the corresponding number of data entities according to the high-density region, and comparing the number of data entities with a preset compression number; If the number of data entities is greater than the preset compression number, generating low-dimensional embedding vectors based on the high-density region and the corresponding attribute features by a preset network model.

4. The network generation method based on a knowledge graph according to claim 1, characterized in that The obtaining the association frequency according to the entities to be labeled and the associated entities includes the following steps: Obtain the regional positions corresponding to all the to-be-labeled entities based on the preset atlas data, and divide the to-be-scanned data within the preset range at the regional positions; Determine whether the to-be-scanned data includes the associated entity; If the to-be-scanned data includes the associated entity, perform frequency labeling on the regional position; Obtain the association frequency based on the frequency labeling and the total number of the to-be-labeled entities.

5. The network generation method based on a knowledge graph according to claim 4, wherein Before obtaining the association frequency based on the frequency labeling and the total number of the to-be-labeled entities, the following steps are further included: Obtain the information features between the to-be-labeled entity and the associated entity based on the to-be-scanned data; Analyze the information features to obtain an analysis result, and determine whether the analysis result is a valid result; If the analysis result is a valid result, remove the frequency label of the to-be-labeled entity.

6. A network inference method based on a knowledge graph, characterized in that For the network audit model generated by the network generation method based on the knowledge graph according to any one of the above claims 1-5, the following steps are included: Obtain the to-be-identified information, and obtain the main entity feature and the corresponding main attribute feature based on the to-be-identified information; Obtain the corresponding entity category according to the main entity feature, and obtain the main network information corresponding to the main entity feature in the network audit model according to the entity category; Obtain the estimated association frequency according to the main entity feature and the main attribute feature, and obtain the estimated network data with corresponding index labels in the main network information according to the estimated association frequency; Filter according to the main attribute feature in the estimated network data to obtain the tail entity information corresponding to the to-be-identified information.

7. The network inference method based on a knowledge graph according to claim 6, wherein After obtaining the main network information corresponding to the main entity feature in the network audit model according to the entity category, the following steps are further included: Obtain the tail entity type according to the main attribute feature, and determine whether the tail entity type belongs to the common attribute; If the tail entity type belongs to the common attribute, obtain the common label data in the main network information based on the main attribute feature; Filter according to the main attribute feature in the common label data to obtain the tail entity information corresponding to the to-be-identified information.

8. A network generation system based on a knowledge graph, characterized in that, Include: A data division module (10), which is used to divide the preset atlas database based on the preset entity category to obtain several groups of to-be-labeled data sets; An entity acquisition module (20), which is used to obtain the corresponding to-be-labeled entity set according to the to-be-labeled data set, and the to-be-labeled entity set includes several to-be-labeled entities; A frequency acquisition module (30), which is used to obtain the associated entity corresponding to the to-be-labeled entity based on the attribute feature corresponding to the to-be-labeled entity, and obtain the association frequency according to the to-be-labeled entity and the associated entity; A label generation module (40), which is used to perform index labeling on the to-be-labeled entity and the associated entity according to the association frequency to obtain the feature-labeled entity; A model generation module (50), which is configured to generate a set of labeled entities corresponding to the set of entities to be labeled according to the feature-labeled entities, and obtain a network audit model corresponding to the set of labeled entities based on the index label; Wherein, the index label includes a category label and an association label, and the step of indexing and labeling the entity to be labeled and the associated entity according to the association frequency includes the following steps: Determine whether the association frequency is a first preset frequency; If the association frequency is the first preset frequency, perform the association label on the entity to be labeled and the associated entity; If the association frequency is not the first preset frequency, determine whether the association frequency is within a first preset frequency set; If the association frequency is within the first preset frequency set, perform the category label and the association label on the entity to be labeled and the associated entity; If the association frequency is not within the first preset frequency set, determine whether the association frequency is within a second preset frequency set; If the association frequency is within the second preset frequency set, perform the category label on the entity to be labeled and the associated entity.

9. An intelligent terminal, characterized in that, The intelligent terminal includes a processor (61) and a memory (62) that are coupled to each other, and a computer program (63) capable of running on the processor (61) is stored on the memory (62); When the computer program (63) is executed by the processor (61), it implements the knowledge graph-based network generation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • User question extension method and system and electronic equipment

    CN117171313A

  • Knowledge graph construction method in knowledge field

    CN117390201A