Network generation method and system based on knowledge graph, inference method and system and intelligent terminal
Through the network generation method of entity category division and correlation frequency marking of knowledge graphs, the problems of high computing resource consumption and long response time in super-large-scale knowledge graph inference are solved, and efficient and accurate data inference and retrieval are achieved.
Patent Information
- Application Number
- CN202510480168.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
Smart Images

Figure CN119990340A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of knowledge graph construction, and in particular to a network generation method, reasoning method, system and intelligent terminal based on knowledge graph. Background Art
[0002] Knowledge graphs use visual graphs to vividly display the core structure, development history, frontier fields and overall knowledge architecture of a discipline to achieve the modern theory of multidisciplinary integration. It displays complex knowledge fields through data mining, information processing, knowledge measurement and graphics drawing, reveals the dynamic development laws of knowledge fields, and provides practical and valuable references for disciplinary research.
[0003] With the development of deep learning and big data technology, reasoning based on knowledge graphs has become one of the hot topics in AI research, especially showing great potential in improving search quality and personalized recommendations. Most of the existing reasoning technologies based on knowledge graphs focus on theoretical modeling, using high-performance computing units such as GPUs for parallel computing, and reasoning through pre-trained neural network models.
[0004] Then, when faced with ultra-large-scale knowledge graphs, higher-performance GPUs are required for reasoning, which leads to problems such as high computing resource consumption and long response time. Summary of the invention
[0005] In order to reduce the consumption of computing resources and improve the reasoning response time, the present application provides a network generation method, reasoning method, system and intelligent terminal based on knowledge graph.
[0006] In the first aspect, the present application provides a network generation method based on a knowledge graph, which adopts the following technical solution: A network generation method based on knowledge graph includes the following steps: Divide the preset graph database based on preset entity categories to obtain several groups of data sets to be labeled; Acquire a corresponding entity set to be labeled according to the data set to be labeled, wherein the entity set to be labeled includes a plurality of entities to be labeled; Acquire an associated entity corresponding to the entity to be marked based on the attribute features corresponding to the entity to be marked, and acquire an associated frequency according to the entity to be marked and the associated entity; Indexing and marking the entity to be marked and the associated entity according to the association frequency to obtain a characteristic marked entity; A marked entity set corresponding to the to-be-marked entity set is generated according to the characteristic marked entity, and a network audit model corresponding to the marked entity set is obtained based on the index mark.
[0007] By adopting the above technical solution, the preset graph database is divided based on the preset entity categories, different sets of data to be marked are obtained, the corresponding set of entities to be marked is obtained based on the same type of set of data to be marked, and the corresponding associated entities are obtained based on the corresponding attribute characteristics of the entities to be marked, so as to obtain the corresponding association frequency based on the two, and mark the two based on the association frequency, thereby forming a network audit model with index marks. When facing more entity scenarios, the corresponding target data can be quickly obtained based on different index marks, shortening the overall data reasoning response time, thereby achieving efficient reasoning in real-time scenarios. Different markings of the entity sets to be marked can quickly locate the target data according to the index marks when obtaining the target data in the later stage.
[0008] In some embodiments, the index mark includes a category mark and an association mark, and the index mark of the to-be-marked entity and the associated entity according to the association frequency includes the following steps: Determining whether the associated frequency is a first preset frequency; If the association frequency is the first preset frequency, performing the association marking on the entity to be marked and the associated entity; If the associated frequency is not the first preset frequency, determining whether the associated frequency is within the first preset frequency set; If the associated frequency is within the first preset frequency set, performing the category marking and the association marking on the entity to be marked and the associated entity; If the associated frequency is not in the first preset frequency set, determining whether the associated frequency is in a second preset frequency set; If the associated frequency is within the second preset frequency set, the category marking is performed on the entity to be marked and the associated entity.
[0009] By adopting the above technical solution, the degree of correlation between the entity to be marked and the associated entity is determined based on the association frequency, thereby determining the index tag type of the entity to be marked and the associated entity, generating different index tags for different entities to be marked and associated entities, ensuring the generation of an accurate network audit model for data, facilitating subsequent data search, improving the accuracy of target data, and quickly acquiring target data. For entities to be marked and associated entities that generally exist at the same time, that is, have a high correlation, the two can be marked according to their associated tags. For entities to be marked and associated entities that are related but not very high, they can be categorized and associated, respectively. When searching for data to be found later, the corresponding target data can be more accurately acquired, the search time can be reduced, and the efficiency of acquiring target data can be improved.
[0010] In some embodiments, the method of dividing the preset graph database based on the preset entity categories to obtain a plurality of sets of data to be labeled includes the following steps: Acquire data distribution characteristics of the preset atlas database based on statistical analysis, and identify high-density areas according to the data distribution characteristics; Determining whether there are multiple entity categories in the high-density area; If there are multiple entity categories in the high-density area, the preset atlas database is divided according to the entity categories to obtain several groups of data sets to be marked; If there are no multiple entity categories in the high-density area, entity features and corresponding attribute features are obtained based on the high-density area, and low-dimensional embedding vectors are generated from the entity features and the corresponding attribute features based on a preset network model.
[0011] By adopting the above technical scheme, the data distribution characteristics of the preset graph database are determined based on statistical analysis, and the high-density area is identified based on the data distribution characteristics. It is determined whether there are multiple entity categories in the high-density area. If there are no multiple entity categories in the high-density area, the entity features and the corresponding attribute features are obtained based on the high-density area, and the entity features and the corresponding attribute features are generated into a low-dimensional embedding vector based on the preset network model. By reasonably acquiring data sets to be labeled with different data amounts, even in the case of extremely unbalanced data distribution, a good partitioning effect can be maintained, ensuring that the load of each data set to be labeled is roughly equal, thereby improving the retrieval efficiency of the target data.
[0012] In some of the embodiments, after identifying the high-density area according to the data distribution characteristics, the following steps are also included: Acquire the corresponding data entity quantity according to the high-density area, and compare the data entity quantity with a preset compression quantity; If the number of data entities is greater than the preset compression number, a low-dimensional embedding vector is generated for the high-density area and the corresponding attribute features based on a preset network model.
[0013] By adopting the above technical solution, the data processing method of the high-density area is determined according to the number of actual application scenarios. If the number of data entities is greater than the preset compression number, the high-density area and the corresponding attribute features are generated into a low-dimensional embedding vector based on the preset network model. Through a lightweight embedding representation strategy, the vector representation of entities and relationships is compressed, thereby reducing the GPU memory burden.
[0014] In some embodiments, the obtaining the association frequency according to the entity to be marked and the associated entity includes the following steps: Based on the preset atlas data, the regional positions corresponding to all the entities to be marked are acquired, and the data to be scanned are divided in the regional positions according to a preset range; Determining whether the data to be scanned includes the associated entity; If the data to be scanned includes the associated entity, frequency marking the area position; The associated frequency is obtained according to the frequency tag and the total number of the entities to be marked.
[0015] In some of the embodiments, before obtaining the associated frequency according to the frequency mark and the total number of the entities to be marked, the following steps are also included: Acquire information features between the entity to be marked and the associated entity according to the data to be scanned; Analyze the information characteristics to obtain analysis results, and determine whether the analysis results are valid results; If the analysis result is a valid result, the frequency mark of the entity to be marked is removed.
[0016] In the second aspect, the present application provides a network reasoning method based on knowledge graph, which adopts the following technical solution: A network reasoning method based on a knowledge graph, based on the network audit model generated by the network generation method based on the knowledge graph, comprises the following steps: Acquire information to be identified, and acquire a main entity feature and a corresponding main attribute feature based on the information to be identified; Acquire a corresponding entity category according to the main entity feature, and acquire main network information corresponding to the main entity feature in the network audit model according to the entity category; Acquire an estimated association frequency according to the main entity feature and the main attribute feature, and acquire estimated network data having a corresponding index mark in the main network information according to the estimated association frequency; The estimated network data is screened according to the main attribute feature to obtain tail entity information corresponding to the information to be identified.
[0017] By adopting the above technical scheme, based on obtaining the main entity characteristics and the corresponding main attribute characteristics in the information to be identified, the corresponding entity category is obtained according to the main entity characteristics, and the network audit model is screened according to the entity category to obtain the main network information with the same entity category mark, and then the estimated association frequency is obtained according to the main entity characteristics and the main attribute characteristics, and the estimated network data with the corresponding index mark is obtained in the main network information according to the estimated association frequency, and the corresponding data is gradually obtained through different index marks, so as to reduce the search steps and improve the overall search efficiency; finally, the estimated network data is screened according to the main attribute characteristics to obtain the tail entity information corresponding to the information to be identified, so as to improve the accuracy of obtaining the tail entity information.
[0018] In some of the embodiments, after obtaining the main network information corresponding to the main entity feature in the network audit model according to the entity category, the following steps are also included: Acquire a tail entity type according to the main attribute feature, and determine whether the tail entity type belongs to a common attribute; If the tail entity type belongs to a common attribute, obtaining common marking data in the main network information based on the main attribute feature; The shared tag data is screened according to the main attribute feature to obtain tail entity information corresponding to the information to be identified.
[0019] By adopting the above technical solution, the tail entity type is obtained based on the main attribute characteristics, and then the tail entity type is judged to obtain the type of the corresponding tail entity information, and the filtered data is narrowed down according to the type, thereby reducing the search time of the tail entity information and improving the search efficiency.
[0020] In the third aspect, the present application provides a network generation system based on knowledge graph, which adopts the following technical solutions: A network generation system based on knowledge graph, comprising: A data partitioning module, which is used to partition a preset graph database based on preset entity categories to obtain a plurality of sets of data to be labeled; An entity acquisition module, the entity acquisition module is used to acquire a corresponding entity set to be labeled according to the data set to be labeled, the entity set to be labeled includes a plurality of entities to be labeled; A frequency acquisition module, the frequency acquisition module is used to acquire the associated entity corresponding to the entity to be marked based on the attribute characteristics corresponding to the entity to be marked, and acquire the associated frequency according to the entity to be marked and the associated entity; A tag generation module, the tag generation module is used to index and tag the to-be-tagged entity and the associated entity according to the association frequency to obtain a characteristic tag entity; A model generation module is used to generate a set of marked entities corresponding to the set of entities to be marked based on the feature marked entities, and to obtain a network audit model corresponding to the set of marked entities based on the index mark.
[0021] In a fourth aspect, the present application provides an intelligent terminal, which adopts the following technical solution: An intelligent terminal, comprising a processor and a memory coupled to each other, wherein the memory stores a computer program that can be run on the processor; When the computer program is executed by the processor, the network generation method based on the knowledge graph as described in the first aspect is implemented.
[0022] In summary, the present application includes at least one of the following beneficial technical effects: 1. Divide the preset graph database based on preset entity categories to obtain different sets of data to be labeled, obtain the corresponding set of entities to be labeled based on the same type of data to be labeled, obtain the corresponding associated entities based on the corresponding attribute characteristics of the entities to be labeled, and then obtain the corresponding association frequency based on the two, and label the two based on the association frequency, thereby forming a network audit model with index labels. When facing a large number of entity scenarios, it can quickly obtain the corresponding target data based on different index labels, shorten the overall data reasoning response time, and thus achieve efficient reasoning in real-time scenarios. Different labels for the entity sets to be labeled can quickly locate the target data based on the index labels when obtaining the target data later; 2. Determine the degree of correlation between the entity to be marked and the associated entity based on the association frequency, thereby determining the index tag type of the entity to be marked and the associated entity, and generate different index tags for different entities to be marked and associated entities to ensure that the network audit model with accurate data is generated, which is convenient for subsequent data search, improves the accuracy of target data, and quickly obtains target data. For entities to be marked and associated entities that generally exist at the same time, that is, they have a high correlation, they can be marked according to their associated tags. For entities to be marked and associated entities that are related but not very high, they can be categorized and associated separately. When searching for data to be found later, the corresponding target data can be obtained more accurately, reducing the search time and improving the efficiency of obtaining target data. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a block diagram of a network generation method based on a knowledge graph provided in an embodiment of the present application; Figure 2 It is a block diagram of a method for obtaining a set of data to be labeled provided in an embodiment of the present application; Figure 3 is a block diagram of a method for obtaining associated frequencies provided in an embodiment of the present application; Figure 4 is a schematic diagram of the structure of a network generation system provided in an embodiment of the present application; Figure 5 is a structural block diagram of the intelligent terminal provided in this embodiment; Figure 6 This is a block diagram of the network reasoning method based on the knowledge graph provided in this application.
[0024] Explanation of the accompanying drawings: 10. Data partition module; 20. Entity acquisition module; 30. Frequency acquisition module; 40. Label generation module; 50. Model generation module; 61. Processor; 62. Memory; 63. Computer program. DETAILED DESCRIPTION
[0025] To more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. However, it should be understood by those of ordinary skill in the art that the present application can be implemented without these details. In some cases, in order to avoid unnecessary descriptions that make various aspects of the present application obscure, well-known methods, processes, systems, components and / or circuits that have been described at a higher level will not be described in detail. For those of ordinary skill in the art, it is obvious that various changes can be made to the embodiments disclosed in the present application, and without departing from the principles and scope of the present application, the general principles defined in the present application can be applied to other embodiments and application scenarios. Therefore, the present application is not limited to the embodiments shown, but conforms to the broadest scope consistent with the scope claimed for protection of the present application.
[0026] An embodiment of the present application discloses a network generation method based on a knowledge graph.
[0027] Reference Figure 1 ,The network generation method based on knowledge graph includes the following steps: S100, dividing the preset graph database based on preset entity categories to obtain several groups of data sets to be labeled.
[0028] Among them, the preset entity categories represent different types of entities, the preset graph database includes data that need to be sorted out through the knowledge graph, and the data sets to be labeled represent different data sets obtained by dividing the preset graph database into blocks.
[0029] It should be noted here that entities are the basic units in the knowledge graph, representing objects in the real world, such as people, places, organizations, etc. The preset entity categories are based on the attributes, relationships, and contexts of the entities.
[0030] For example, entities in reality can be classified into people, organizations, places, events, etc. according to specific properties. In natural language processing, entity recognition is a common application of entity classification, which involves identifying entities with specific categories from text.
[0031] For example, entity recognition based on BiLSTM-CRF, which uses a model combining a bidirectional long short-term memory network (BiLSTM) and a conditional random field (CRF), performs well in NER tasks. BiLSTM can capture long-distance dependencies in text, while the CRF layer can use the constraint relationship between adjacent tags to improve the accuracy of annotation.
[0032] In order to facilitate the processing of graph data, the preset graph database stores a relationship network diagram, which is a graph generated by extracting entity features from the preset graph data. It is specifically generated based on a preset model, which is a pre-trained neural network model. The neural network model adopts the existing graph neural network, which will not be elaborated here.
[0033] Here, entity recognition is performed on the preset graph database, mainly using traditional entity recognition methods, which rely on a large number of rules and dictionaries, but this method often does not work well when processing complex texts. With the development of deep learning, neural network-based methods have become mainstream. For example, the model of bidirectional long short-term memory network (BiLSTM) combined with conditional random field (CRF) performs well in NER tasks.
[0034] It should be noted here that the data stored in the preset graph database also includes the relationship between entities and the attributes corresponding to each entity, that is, it is used to describe and supplement the entity. The amount of data stored in the acquired data set to be labeled is roughly the same, mainly to facilitate the processing of data by the smart terminal.
[0035] The smart terminal mentioned here processes and stores data on the entire preset graph database, making it easier for subsequent clients to quickly obtain corresponding information based on the network reasoning method based on the knowledge graph.
[0036] When the preset graph database is huge, containing hundreds of millions of nodes and relationships, the traditional way of obtaining data is to search for information based on the knowledge graph generated by the preset graph database. This requires traversing the entire graph, which results in a long retrieval time. At this time, we divide the entire knowledge graph into thousands of small tiles, each of which has a corresponding multi-level index table to support fast positioning.
[0037] Combination Figure 2 , dividing the preset graph database based on the preset entity categories to obtain several sets of data sets to be labeled, including the following steps: S110, obtaining data distribution characteristics of a preset atlas database based on statistical analysis, and identifying high-density areas according to the data distribution characteristics.
[0038] S120, determining whether there are multiple entity categories in the high-density area.
[0039] S130: If there are multiple entity categories in the high-density area, the preset atlas database is divided according to the entity categories to obtain several groups of data sets to be labeled.
[0040] S140: If there are no multiple entity categories in the high-density area, entity features and corresponding attribute features are obtained based on the high-density area, and low-dimensional embedding vectors are generated from the entity features and the corresponding attribute features based on a preset network model.
[0041] Among them, the data distribution feature represents the data distribution displayed in the relationship network diagram formed in the preset graph database. The high-density area represents the data area with multiple entity features. The entity category represents that the entity belongs to different types.
[0042] What needs to be said here is that we first determine the data distribution characteristics of the knowledge graph through statistical analysis, and then design a two-stage partitioning algorithm. In the first stage, a clustering algorithm is used to identify high-density areas and sparse areas to form preliminary partitioning suggestions. In the second stage, combined with the results of the preliminary partitioning, the genetic algorithm is used for iterative optimization to find the best partitioning scheme. In this way, even in the case of extremely unbalanced data distribution, a good partitioning effect can be maintained, ensuring that the load of each tile is roughly equal, maximizing retrieval efficiency.
[0043] S200, obtaining a corresponding set of entities to be labeled according to a set of data to be labeled.
[0044] The to-be-labeled entity set includes several to-be-labeled entities, and the to-be-labeled entities represent entities that need to be subsequently labeled. The specific method of obtaining the corresponding to-be-labeled entity set based on the to-be-labeled data set is through a pre-trained language model.
[0045] The preset training language model can be BERT (Bidirectional Encoder Representations from Transformers), which can more accurately identify entities by understanding contextual semantics. BERT and other models learn rich language features through pre-training on a large amount of unlabeled text, so that they can be effectively applied to entity recognition tasks.
[0046] For example, specific characters and places can be identified in fairy tales. For example, by analyzing a fairy tale of Little Red Riding Hood, the entity recognition system can identify entities such as the protagonist, story name, and important events mentioned in the text. This is of great significance to fields such as information aggregation, information retrieval, and analysis.
[0047] S300: acquiring related entities corresponding to the entity to be marked based on attribute features, and acquiring related frequencies according to the entity to be marked and the related entities.
[0048] Among them, attribute features characterize the relationship between entities, thereby confirming whether entities are associated with each other. The specific attribute features are traditional methods that rely on rules and pattern matching, while attribute extraction methods based on machine learning can automatically identify attributes by learning patterns in data. Deep learning, especially methods based on pre-trained models such as RNN (recurrent neural network) and BERT, perform well in attribute extraction. These models can capture contextual information to more accurately identify and classify attributes. With the development of deep learning technology, especially the emergence of pre-trained language models (such as BERT), the accuracy and efficiency of attribute extraction have been significantly improved. These models can understand complex contextual information and extract relevant attributes more accurately.
[0049] Reference Figure 3 , the association frequency represents the degree of association between the entity to be marked and the associated entity. The association frequency is obtained according to the entity to be marked and the associated entity, including the following steps: S310, based on the preset atlas data, the regional positions corresponding to all entities to be marked are obtained, and the data to be scanned are divided in the regional positions according to the preset range.
[0050] S320: Determine whether the data to be scanned includes an associated entity.
[0051] S330: If the data to be scanned includes associated entities, frequency marking is performed on the regional position.
[0052] S340: Obtain the associated frequency according to the frequency tag and the total number of entities to be tagged.
[0053] The regional position represents the position obtained in the preset atlas data, which is the area where the entity to be marked is located. The data to be scanned represents the data obtained according to the preset range, and the preset range is a range set in advance, which can be set according to a sentence or a paragraph in the preset atlas data.
[0054] Determine whether the data to be scanned includes associated entities. If the data to be scanned includes associated entities, it indicates that the associated entities and the entities to be marked appear synchronously, so it is necessary to perform frequency marking on the area position. After all the entities to be marked in the preset atlas data have been subjected to the above steps S310 to S330, all frequency markings and the total number of all entities to be marked are compared according to step S340 to obtain the corresponding associated frequency.
[0055] S400: Index and mark the entity to be marked and the related entity according to the association frequency to obtain a feature marked entity.
[0056] The feature-tagged entity represents an entity that has a search index. The index tag includes a category tag and an association tag. The index tagging of the tagged entity and the associated entity is performed according to the association frequency, including the following steps: S410: Determine whether the associated frequency is a first preset frequency.
[0057] S420: If the association frequency is the first preset frequency, the entity to be marked and the associated entity are associated and marked.
[0058] S430: If the associated frequency is not the first preset frequency, determine whether the associated frequency is in the first preset frequency set.
[0059] S440: If the associated frequency is within the first preset frequency set, classify and associate the entity to be marked and the associated entity.
[0060] S450: If the associated frequency is not in the first preset frequency set, determine whether the associated frequency is in the second preset frequency set.
[0061] S460: If the associated frequency is within the second preset frequency set, the entity to be marked and the associated entity are categorized. It should be noted that the index mark may be a hash index.
[0062] Among them, the first preset frequency represents that the entity to be marked and the associated entity are completely related, that is, the associated entity will appear where the entity to be marked appears, so the first preset frequency can be set to 1. The first preset frequency set represents that there is a correlation between the entity to be marked and the associated entity. The first preset frequency set can be set to 0.4~1, and the first preset frequency set is a set that is closed at the beginning and open at the end. The second preset frequency set represents that there is an association between the entity to be marked and the associated entity, but the correlation is not high. Therefore, the association between the entity to be marked and the associated entity can be ignored. The second preset frequency set can be set to 0~0.4, and the second preset set is a set that is closed at the beginning and open at the end.
[0063] It should be noted here that the association mark corresponding to the characteristic mark entity is at the first-level mark, and the category mark is at the second-level mark. When searching, the search will be performed based on marks of different levels.
[0064] Specifically, in step S410-step S460, it is described to determine whether the associated frequency is the first preset frequency. If the associated frequency is the first preset frequency, the entity to be marked and the associated entity are associated. If the associated frequency is not the first preset frequency, it is determined whether the associated frequency is in the first preset frequency set. If the associated frequency is in the first preset frequency set, the entity to be marked and the associated entity are category marked and associated marked. If the associated frequency is not in the first preset frequency set, it is determined whether the associated frequency is in the second preset frequency set. If the associated frequency is in the second preset frequency set, the entity to be marked and the associated entity are category marked.
[0065] S500, generating a marked entity set corresponding to the to-be-marked entity set based on the feature marked entity, and acquiring a network audit model corresponding to the marked entity set based on the index mark.
[0066] Among them, the marked entity set represents that there are multiple marked entities in the set. Therefore, based on the existing database, the marked entity set is used to generate a network audit model. The network audit model specifically includes a knowledge graph generated based on preset graph data, which is convenient for subsequent client users to obtain corresponding information based on the network audit model. The index established according to the present technical solution can reduce the time cost of customer retrieval.
[0067] In one embodiment, after the high-density area is identified according to the data distribution characteristics, the following steps are also included: S600: Obtain the corresponding data entity quantity according to the high-density area, and compare the data entity quantity with a preset compression quantity.
[0068] S610: If the number of data entities is greater than the preset compression number, a low-dimensional embedding vector is generated based on a preset network model for the high-density area and the corresponding attribute features.
[0069] Among them, the number of data entities represents the number of data entities existing in the high-density area, and the preset compression number represents the minimum data standard for determining whether the data in the high-density area needs to be compressed. If the number of data entities is not greater than the preset compression number, the high-density area does not need to be compressed, and the intelligent network generation end can perform corresponding processing. When the number of data entities is greater than the preset compression number, a low-dimensional embedding vector is generated for the high-density area and the corresponding attribute features based on the preset network model. Specifically, a lightweight embedding representation strategy is adopted to compress the vector representation of entities and relationships and reduce the GPU memory burden. For example, the vector dimensions of entities and relationships are compressed to less than 64 or 32 dimensions, and a low-dimensional embedding vector is generated through a neural network model.
[0070] In one embodiment, before obtaining the associated frequency according to the frequency tag and the total number of entities to be tagged, the following steps are also included: S700: Acquire information features between the entity to be marked and the associated entity according to the data to be scanned.
[0071] S710, analyzing the information feature to obtain an analysis result, and determining whether the analysis result is a valid result.
[0072] S720: If the analysis result is a valid result, the frequency mark of the entity to be marked is removed.
[0073] Among them, the information feature represents the data information between the entity to be marked and the associated entity. If the analysis result is a valid result, it means that other entities exist in the information feature, and the corresponding relationship between the entity to be marked and the associated entity does not exist, so the frequency mark needs to be removed.
[0074] An embodiment of the present application also discloses a network generation system based on a knowledge graph.
[0075] Reference Figure 4 The network reasoning system based on the knowledge graph includes a data partitioning module 10, an entity acquisition module 20 connected to the data partitioning module 10 in a network, a frequency acquisition module 30 connected to the entity acquisition module 20 in a network, a tag generation module 40 connected to the frequency acquisition module 30 in a network, and a model generation module 50 connected to the tag generation module 40 in a network.
[0076] The data partitioning module 10 is used to partition the preset graph database based on preset entity categories to obtain several groups of data sets to be labeled. The entity acquisition module 20 is used to obtain the corresponding entity set to be labeled based on the data set to be labeled, and the entity set to be labeled includes several entities to be labeled. The frequency acquisition module 30 is used to obtain the associated entity corresponding to the entity to be labeled based on the attribute features corresponding to the entity to be labeled, and obtain the associated frequency based on the entity to be labeled and the associated entity. The tag generation module 40 is used to disclose a network generation system based on the knowledge graph according to the embodiment of the present application.
[0077] The association frequency indexes the entity to be marked and the associated entity to obtain the characteristic marked entity. The model generation module 50 is used to generate a marked entity set corresponding to the entity set to be marked according to the characteristic marked entity, and obtain a network audit model corresponding to the marked entity set based on the index mark.
[0078] The other functions performed in the above-mentioned data partitioning module 10, entity acquisition module 20, frequency acquisition module 30, tag generation module 40 and model generation module 50 as well as the technical details of each function are the same as or similar to the corresponding features in the network generation method based on the knowledge graph described above, so they will not be repeated here.
[0079] The embodiment of the present application also discloses a smart terminal.
[0080] Reference Figure 5 The intelligent terminal includes a processor 61 and a memory 62 coupled to each other, and the memory 62 stores a computer program 63 that can be run on the processor 61. When the computer program 63 is executed by the processor 61, a generation method based on a knowledge graph is implemented.
[0081] The memory 62 may be a ROM or other type of static storage device capable of storing static information and instructions, a random access memory, or other type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable read-only memory, a read-only optical disc or other optical disc storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program codes in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto. The memory 62 may be an internal storage unit in some embodiments.
[0082] In addition, the processor 61 may be a central processing unit, a general purpose processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic device, a transistor logic device, a hardware component or any combination thereof, for running the program code stored in the memory 62 or processing data.
[0083] Specifically, the processor 61 and the memory 62 are connected via a bus. The bus may include a path to transmit information between the above components. The bus may be a peripheral component interconnect standard bus or an extended industrial standard architecture bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0084] Figure 5 Only the intelligent terminal having the memory 62, the processor 61 and the bus is shown, and it can be understood by those skilled in the art that Figure 5 The structure shown does not constitute a limitation on the intelligent terminal, which can be a bus structure or a star structure. The intelligent terminal can also include more or fewer components than shown in the figure, or combine certain components, or deploy different components. Other existing or future electronic devices may be applicable, should also be included in the scope of protection, and are included here by reference.
[0085] The implementation principle is: First, the intelligent terminal obtains the data distribution characteristics of the preset graph database based on statistical analysis based on the data processing end, identifies the high-density area based on the data distribution characteristics, and judges whether there are multiple entity categories in the high-density area. If there are multiple entity categories in the high-density area, the preset graph database is divided according to the entity category to obtain several groups of data sets to be labeled. If there are no multiple entity categories in the high-density area, the entity features and the corresponding attribute features are obtained based on the high-density area, and the entity features and the corresponding attribute features are generated into a low-dimensional embedding vector based on the preset network model.
[0086] Next, the intelligent terminal obtains a corresponding set of entities to be marked according to the set of data to be marked, obtains associated entities corresponding to the entities to be marked based on attribute features, and obtains associated frequencies according to the entities to be marked and the associated entities.
[0087] Finally, the intelligent terminal indexes and marks the entities to be marked and the associated entities according to the association frequency to obtain the characteristic marked entities, generates a marked entity set corresponding to the set of entities to be marked based on the characteristic marked entities, and obtains the network audit model corresponding to the marked entity set based on the index mark.
[0088] Reference Figure 6 The present application also discloses a network reasoning method based on a knowledge graph. The network audit model generated by the network generation method based on the knowledge graph provided in the above embodiment includes the following steps: S800, obtaining information to be identified, and obtaining main entity features and corresponding main attribute features based on the information to be identified.
[0089] S810, obtaining a corresponding entity category according to the main entity feature, and obtaining main network information corresponding to the main entity feature in the network audit model according to the entity category.
[0090] S820, obtaining an estimated association frequency according to the main entity feature and the main attribute feature, and obtaining estimated network data with a corresponding index mark in the main network information according to the estimated association frequency.
[0091] S830, screening is performed in the estimated network data according to the main attribute characteristics to obtain tail entity information corresponding to the information to be identified.
[0092] The information to be identified represents the information inputted by the client, including the target information that the client needs to obtain. The information to be identified can be the information directly obtained by the intelligent terminal from the client. The main entity feature represents the main feature that needs to be obtained for the feature of the information to be identified. The feature extraction of the information to be identified is the same or similar to the description in the above embodiment, so it will not be described in detail here.
[0093] It should be noted that since the information to be identified includes several entity features, it is necessary to analyze several entity features to obtain the main entity feature. The main attribute feature represents the attribute corresponding to the main entity feature, and the main attribute feature here may include multiple ones. Therefore, the smart terminal can perform corresponding processing according to the information to be identified.
[0094] The main network information represents the entity category obtained based on the main entity characteristics, and the information obtained in the same entity category of the network audit model based on the entity category. The estimated association frequency represents the estimated frequency obtained by the smart terminal based on the main entity characteristics and the main attribute characteristics. This estimated association frequency can be obtained through the search history data, and the main entity characteristics and the main attribute characteristics are matched in the search history data. If a match is found, the location of the existing estimated network data is directly located, and a more specific search is performed based on the main entity characteristics.
[0095] In one embodiment, after obtaining the main network information corresponding to the main entity feature in the network audit model according to the entity category, the following steps are also included: S910, obtaining a tail entity type according to the main attribute feature, and determining whether the tail entity type belongs to a common attribute.
[0096] S920: If the tail entity type belongs to a common attribute, the common marking data is obtained in the main network information based on the main attribute feature.
[0097] S930, screening is performed in the common tag data according to the main attribute feature to obtain tail entity information corresponding to the information to be identified.
[0098] Among them, the shared attributes mainly include attributes existing in the public area, such as country, location, etc. The main attribute feature is to ask where a person is from. By analyzing the main attribute feature, it can be obtained that the tail entity type is country, and country is a shared attribute. Therefore, the shared tag data with country tags can be directly obtained, and the country can be screened in the country to screen out the country with the same tag as the main entity feature as the tail entity information corresponding to the main entity feature.
[0099] It should be noted here that the network reasoning method based on the knowledge graph is specifically applied to the client end corresponding to the smart terminal, and then the network reasoning method based on the knowledge graph can realize the network reasoning based on the knowledge graph, which is convenient for obtaining tail entity information.
[0100] Combination Figure 5 It should be noted that the intelligent terminal includes a processor 61 and a memory 62 coupled to each other, the processor 61 is also used to execute the above steps S800-S830 and execute steps S900-S830, and the memory 62 stores a computer program 63 that can be run on the processor 61. When the computer program 63 is executed by the processor 61, a network reasoning method based on a knowledge graph is used.
[0101] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the instructions of the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise clearly stated in this document, the execution of these steps is not strictly limited in order and can be performed in other orders.
[0102] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, any equivalent changes made according to the structure, shape, and principle of the present application should be included in the protection scope of the present application.
Claims
1. A network generation method based on knowledge graph, characterized in that: The following steps are involved: Divide the preset graph database based on preset entity categories to obtain several groups of data sets to be labeled; Acquire a corresponding entity set to be labeled according to the data set to be labeled, wherein the entity set to be labeled includes a plurality of entities to be labeled; Acquire an associated entity corresponding to the entity to be marked based on the attribute features corresponding to the entity to be marked, and acquire an associated frequency according to the entity to be marked and the associated entity; Indexing and marking the entity to be marked and the associated entity according to the association frequency to obtain a characteristic marked entity; A marked entity set corresponding to the to-be-marked entity set is generated according to the characteristic marked entity, and a network audit model corresponding to the marked entity set is obtained based on the index mark.
2. The network generation method based on knowledge graph according to claim 1 is characterized in that: The index mark includes a category mark and an association mark, and the index mark of the to-be-marked entity and the associated entity according to the association frequency includes the following steps: Determining whether the associated frequency is a first preset frequency; If the association frequency is the first preset frequency, performing the association marking on the entity to be marked and the associated entity; If the associated frequency is not the first preset frequency, determining whether the associated frequency is within the first preset frequency set; If the associated frequency is within the first preset frequency set, performing the category marking and the association marking on the entity to be marked and the associated entity; If the associated frequency is not in the first preset frequency set, determining whether the associated frequency is in a second preset frequency set; If the associated frequency is within the second preset frequency set, the category marking is performed on the entity to be marked and the associated entity.
3. The network generation method based on knowledge graph according to claim 1 is characterized in that: The method of dividing the preset graph database based on the preset entity categories to obtain a plurality of sets of data to be labeled includes the following steps: Acquire data distribution characteristics of the preset atlas database based on statistical analysis, and identify high-density areas according to the data distribution characteristics; Determining whether there are multiple entity categories in the high-density area; If there are multiple entity categories in the high-density area, the preset atlas database is divided according to the entity categories to obtain several groups of data sets to be marked; If there are no multiple entity categories in the high-density area, entity features and corresponding attribute features are obtained based on the high-density area, and low-dimensional embedding vectors are generated from the entity features and the corresponding attribute features based on a preset network model.
4. The network generation method based on knowledge graph according to claim 3 is characterized in that: After the high-density area is identified according to the data distribution characteristics, the following steps are also included: Acquire the corresponding data entity quantity according to the high-density area, and compare the data entity quantity with a preset compression quantity; If the number of data entities is greater than the preset compression number, a low-dimensional embedding vector is generated for the high-density area and the corresponding attribute features based on a preset network model.
5. The network generation method based on knowledge graph according to claim 1, characterized in that: The obtaining of the association frequency according to the entity to be marked and the associated entity comprises the following steps: Based on the preset atlas data, the regional positions corresponding to all the entities to be marked are acquired, and the data to be scanned are divided in the regional positions according to a preset range; Determining whether the data to be scanned includes the associated entity; If the data to be scanned includes the associated entity, frequency marking the area position; The associated frequency is obtained according to the frequency tag and the total number of the entities to be marked.
6. The network generation method based on knowledge graph according to claim 5 is characterized in that: Before obtaining the associated frequency according to the frequency mark and the total number of the entities to be marked, the following steps are also included: Acquire information features between the entity to be marked and the associated entity according to the data to be scanned; Analyze the information characteristics to obtain analysis results, and determine whether the analysis results are valid results; If the analysis result is a valid result, the frequency mark of the entity to be marked is removed.
7. A network reasoning method based on knowledge graph, characterized in that: The network audit model generated by the network generation method based on the knowledge graph according to any one of claims 1 to 6 above comprises the following steps: Acquire information to be identified, and acquire a main entity feature and a corresponding main attribute feature based on the information to be identified; Acquire a corresponding entity category according to the main entity feature, and acquire main network information corresponding to the main entity feature in the network audit model according to the entity category; Acquire an estimated association frequency according to the main entity feature and the main attribute feature, and acquire estimated network data having a corresponding index mark in the main network information according to the estimated association frequency; The estimated network data is screened according to the main attribute feature to obtain tail entity information corresponding to the information to be identified.
8. The network reasoning method based on knowledge graph according to claim 7 is characterized in that: After acquiring the main network information corresponding to the main entity feature in the network audit model according to the entity category, the following steps are also included: Acquire a tail entity type according to the main attribute feature, and determine whether the tail entity type belongs to a common attribute; If the tail entity type belongs to a common attribute, obtaining common marking data in the main network information based on the main attribute feature; The shared tag data is screened according to the main attribute feature to obtain tail entity information corresponding to the information to be identified.
9. A network generation system based on knowledge graph, characterized in that: include: A data partitioning module (10), the data partitioning module (10) being used to partition a preset graph database based on preset entity categories to obtain a plurality of sets of data to be labeled; An entity acquisition module (20), the entity acquisition module (20) being used to acquire a corresponding entity set to be labeled based on the data set to be labeled, the entity set to be labeled comprising a plurality of entities to be labeled; A frequency acquisition module (30), the frequency acquisition module (30) being used to acquire an associated entity corresponding to the entity to be marked based on the attribute characteristics corresponding to the entity to be marked, and to acquire an associated frequency based on the entity to be marked and the associated entity; A tag generation module (40), the tag generation module (40) being used to index and tag the entity to be tagged and the associated entity according to the association frequency, so as to obtain a characteristic tagged entity; A model generation module (50), the model generation module (50) is used to generate a marked entity set corresponding to the to-be-marked entity set based on the feature marked entity, and to obtain a network audit model corresponding to the marked entity set based on the index mark.
10. An intelligent terminal, characterized in that: The intelligent terminal comprises a processor (61) and a memory (62) coupled to each other, wherein the memory (62) stores a computer program (63) that can be run on the processor (61); When the computer program (63) is executed by the processor (61), the method for generating a network based on a knowledge graph as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method and device for creating knowledge graph
CN107665252A
Establishment method and device of entity intention system, equipment and medium
CN111091006A
Data standard generation and automatic mapping method based on knowledge graph technology
CN115374108A
User question extension method and system and electronic equipment
CN117171313A
Knowledge graph construction method in knowledge field
CN117390201A