Data retrieval method and system based on knowledge graph, terminal and storage medium
By constructing a knowledge graph, the problem of low information retrieval efficiency is solved, enabling convenient information search and accurate recommendations, allowing users to quickly and accurately find the best target information.
Patent Information
- Application Number
- CN202511019001.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-07
Smart Images

Figure CN120910116A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data retrieval, in particular to a data retrieval method and system based on a knowledge graph, a terminal and a computer readable storage medium. BACKGROUND
[0002] The information data of various information platforms is complex and diverse, and lacks perfect information integration forms. At the same time, the information data of various information platforms is scattered in different systems, and the data in a single channel is not comprehensive, and structured API (Application Programming Interface) is not provided, so there is a lack of cross-platform information integration and information island problems. In addition, information data is also dynamically changing, and many information platforms have information lag problems.
[0003] At present, in the corresponding data retrieval, due to the lack of unified and perfect data acquisition channel, when users find suitable data, they mainly use manual network search to obtain related information, which has the problems of low information retrieval efficiency and incomplete information search. And because these information platforms are still based on traditional relational databases when integrating information data, they have not established the dynamic relationship between information data, and cannot perform correlation analysis, so users cannot effectively integrate and use information resources to assist in selecting target data.
[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY
[0005] The main purpose of the present application is to provide a data retrieval method and system based on a knowledge graph, a terminal and a storage medium, which aims to solve the problems of low information retrieval efficiency, incomplete information search and inability to perform correlation analysis in the prior art for information data query, so that users cannot quickly and accurately obtain useful information.
[0006] To achieve the above purpose, the present application provides a data retrieval method based on a knowledge graph, which comprises the following steps:
[0007] Obtain information data of a target user, pre-process the information data to obtain target information data;
[0008] Construct a knowledge expression framework of the target user, perform data extraction on the target information data based on the knowledge expression framework to obtain a plurality of knowledge expression data;
[0009] Perform multi-source information knowledge fusion on all the knowledge expression data to obtain a plurality of fusion data, and construct a knowledge graph according to all the fusion data to obtain a corresponding knowledge graph;
[0010] The knowledge graph is dynamically updated to obtain a target knowledge graph, and data retrieval is performed according to the target knowledge graph to obtain a target retrieval result.
[0011] Optionally, the knowledge graph-based data retrieval method, wherein the information data of the target user is obtained, and the information data is preprocessed to obtain target information data, specifically including:
[0012] The information data of the target user is collected and processed to obtain information data, and the information data is data cleaned, wherein the data cleaning includes de-duplication processing, missing value processing and abnormal value detection processing;
[0013] The information data after cleaning is subjected to format standardization processing, and the information data after format standardization is subjected to text enhancement processing to obtain target information data, wherein the format standardization processing includes uniform time format, standard address and uniform case, and the text enhancement includes special symbol processing and abbreviation expansion.
[0014] Optionally, the knowledge graph-based data retrieval method, wherein the knowledge expression framework of the target user is constructed, specifically including:
[0015] The basic concept of the target user is defined;
[0016] The static abstract concept of the basic concept is defined to obtain a static feature concept;
[0017] The dynamic abstract concept of the basic concept is defined to obtain a dynamic feature concept;
[0018] The spatial feature concept of the target user is defined;
[0019] The framework concept layer is constructed according to the basic concept, the static feature concept, the dynamic feature concept and the spatial feature concept;
[0020] The static relationship dimension between the entity of the basic concept and the entity of the static feature concept is set, the dynamic relationship dimension between the entity of the basic concept and the entity of the dynamic feature concept is set, and the framework relationship layer is constructed according to the static relationship dimension and the dynamic relationship dimension;
[0021] The feature concept entity attribute and the corresponding time attribute between the basic concept, the static feature concept and the dynamic feature concept are set, and the relationship feature attribute of the feature concept entity attribute is set, and the framework attribute layer is constructed according to the feature concept entity attribute, the time attribute and the relationship feature attribute;
[0022] The knowledge expression framework is constructed according to the framework concept layer, the framework relationship layer and the framework attribute layer.
[0023] Optionally, the knowledge graph-based data retrieval method, wherein the multi-source information knowledge fusion on all the knowledge expression data is performed to obtain a plurality of fusion data, and specifically includes:
[0024] The knowledge expression data is subjected to coarse block processing to obtain a plurality of data coarse blocks, all the data coarse blocks are subjected to encoding processing to obtain a plurality of data vectors, and all the data vectors are subjected to clustering block processing to obtain a candidate entity matching library;
[0025] The associated features of all the candidate matching entities in the candidate entity matching library are obtained, all the candidate matching entities and the corresponding associated features are subjected to splicing processing to obtain a plurality of entity description information, and vector calculation is performed on all the entity description information to obtain a plurality of embedding vector representations;
[0026] All the embedding vector representations are subjected to pairwise matching to obtain a plurality of matching results, and the entity similarity distance of all the matching results is calculated to obtain a target calculation result;
[0027] The calculation result with a cosine similarity greater than a preset threshold value in the target calculation result is obtained, and all the calculation results corresponding to two entities are subjected to merging processing to obtain a plurality of fusion data.
[0028] Optionally, the knowledge graph-based data retrieval method, wherein the encoding processing on all the data coarse blocks is specifically:
[0029] H1=WEM(N)=[h1,…,h d ];
[0030] The vector calculation on all the entity description information is specifically:
[0031] H2=SEM(D)=[h1,…,h d ];
[0032] The calculation on the entity similarity distance of all the matching results is specifically:
[0033]
[0034] wherein H1 is a data vector of an entity name, WEM is a word-level embedding model, N is an entity name, h1 is a data vector of a first dimension, d is a dimension of a data vector, h dis a data vector of the d-th dimension, H2 is an embedding vector representation of entity description information, SEM is a sentence-level embedding model, D is the description information of an entity, Similarity(A, B) is the cosine similarity between embedding vector A and embedding vector B, A is an embedding vector of a first entity, B is an embedding vector of a second entity, ‖A‖ is the norm of embedding vector A, and ‖B‖ is the norm of embedding vector B.
[0035] Optionally, the knowledge graph-based data retrieval method, wherein the dynamic updating of the knowledge graph to obtain a target knowledge graph specifically comprises:
[0036] setting a timing task, collecting new information data of the target user according to the timing task, and performing data confirmation on the new information data;
[0037] if the new information data has update data, comparing the new information data with the information data to obtain new data;
[0038] determining an entity type of the new data, and dynamically updating the knowledge graph according to an update mode corresponding to the entity type to obtain a target knowledge graph, wherein the entity type is any one of a basic concept type, a dynamic feature concept type, and a static feature concept type.
[0039] Optionally, the knowledge graph-based data retrieval method, wherein the dynamic updating of the knowledge graph according to the update mode corresponding to the entity type to obtain a target knowledge graph specifically comprises:
[0040] if the entity type is the basic concept type, performing entity embedding and similarity calculation on the new data to obtain a first processing result, and determining whether the new data is an existing entity node according to the first processing result;
[0041] if the new data is a non-existing entity node, adding an entity node corresponding to the new data in the knowledge graph, and establishing an association relationship of the entity node to obtain a target knowledge graph;
[0042] if the entity type is the dynamic feature concept type, adding a dynamic feature node corresponding to the new data in the knowledge graph according to an update period of the timing task, and establishing a dynamic association relationship of the dynamic feature node to obtain a target knowledge graph;
[0043] if the entity type is the static feature concept type, performing entity embedding and similarity calculation on the new data to obtain a second processing result;
[0044] According to the second processing result, it is judged whether the new data is an existing entity node, if the new data is an existing entity node, a static feature node corresponding to the new data is added in the knowledge graph, and the association relationship of the static feature node is established, and a target knowledge graph is obtained.
[0045] Optionally, the knowledge graph-based data retrieval method, wherein the knowledge graph-based data retrieval system comprises:
[0046] The data processing module is configured to acquire information data of a target user, pre-process the information data, and obtain target information data.
[0047] The data extraction module is configured to construct a knowledge expression framework of the target user, extract data from the target information data based on the knowledge expression framework, and obtain a plurality of knowledge expression data.
[0048] The graph construction module is configured to perform multi-source information knowledge fusion on all the knowledge expression data, obtain a plurality of fusion data, construct a knowledge graph based on all the fusion data, and obtain a corresponding knowledge graph.
[0049] The graph updating module is configured to dynamically update the knowledge graph, obtain a target knowledge graph, perform data retrieval based on the target knowledge graph, and obtain a target retrieval result.
[0050] In addition, to achieve the above object, the present application further provides a terminal, wherein the terminal comprises a memory, a processor, and a knowledge graph-based data retrieval program stored in the memory and executable on the processor, and the knowledge graph-based data retrieval program implements the steps of the knowledge graph-based data retrieval method when executed by the processor.
[0051] In addition, to achieve the above object, the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores a knowledge graph-based data retrieval program, and the knowledge graph-based data retrieval program implements the steps of the knowledge graph-based data retrieval method when executed by a processor.
[0052] In the present application, the information data of the target user is acquired, the information data is preprocessed to obtain target information data, the knowledge expression framework of the target user is constructed, the target information data is data-extracted based on the knowledge expression framework to obtain a plurality of knowledge expression data, all the knowledge expression data is subjected to multi-source information knowledge fusion to obtain a plurality of fusion data, and the corresponding knowledge graph is constructed according to all the fusion data; the knowledge graph is dynamically updated to obtain a target knowledge graph, and data retrieval is performed according to the target knowledge graph to obtain a target retrieval result. The present application integrates complex and multi-element related information in a highly structured manner by constructing a knowledge graph, thereby providing convenient information search and accurate recommendation services for users, so that the users can quickly and accurately find the best target information. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 is a flowchart of a preferred embodiment of the data retrieval method based on the knowledge graph of the present application;
[0054] Figure 2 is a schematic diagram of the overall architecture of the knowledge graph construction method in the preferred embodiment of the present application;
[0055] Figure 3 is a structure diagram of a preferred embodiment of the data retrieval system based on the knowledge graph of the present application;
[0056] Figure 4 is a structure diagram of a preferred embodiment of the terminal of the present application. DETAILED DESCRIPTION
[0057] To make the objectives, technical solutions and advantages of the present application clearer and more explicit, the present application is further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0058] It should be noted that if the present application examples involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement condition, etc. between the components in a certain specific posture (as shown in the drawings), and if the specific posture changes, the directional indications also change accordingly.
[0059] In addition, if the description of "first", "second" and the like is involved in the embodiments of the present application, the description of "first", "second" and the like is only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can be explicitly or implicitly included at least one of the features. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the realization of ordinary skilled in the art, when the combination of technical solutions appears contradictory or cannot be realized, it should be considered that the combination of technical solutions does not exist, also not within the protection scope required by the present application.
[0060] The data retrieval method based on knowledge graph according to the preferred embodiment of the present application, as shown in Figure 1 The data retrieval method based on knowledge graph includes the following steps:
[0061] Step S10, obtaining the information data of the target user, preprocessing the information data to obtain the target information data.
[0062] Specifically, in the embodiments of the present application, the best college tutor is searched when the overseas student applies for studying abroad, for example, in order to solve the problems of complex and multiple international college tutor information, scattered sources and lack of comprehensive information sources, resulting in low efficiency and incomplete information of the applicant, the present application provides a data retrieval method based on knowledge graph, which integrates the tutor information by using the advantage of high structure of knowledge graph, and provides convenient information search and accurate recommendation service for the applicant, the specific processing process is as shown in Figure 2 Firstly, the information data (i.e. tutor data) of the target user (for example, a tutor of an international college) needs to be obtained, specifically, a crawler program is customized and written for different data sources from academic databases (for example, Google Scholar database), social platforms (for example, LinkedIn) and open data sources (for example, OpenAlex), the relevant information of the tutor is collected under the condition of legality and compliance, it should be noted that the information (including user equipment information, user personal information, etc.), data (including user analysis data, user stored data, user displayed data, etc.) and signals involved in the present application are all information, data and signals authorized by the user or fully authorized by all parties; and the collection, use and processing of the relevant information, data and signals comply with the relevant national and regional laws, regulations and standards.
[0063] After obtaining the information data of the target user, the information data needs to be preprocessed, including data cleaning, format standardization and text enhancement. Specifically, the information data is cleaned, wherein the data cleaning includes duplicate removal processing, missing value processing and outlier detection processing; the cleaned information data is subjected to format standardization processing, and the format standardized information data is subjected to text enhancement processing to obtain target information data, wherein the format standardization processing includes uniform time format, standardized address and uniform case, and the text enhancement includes special symbol processing and abbreviation expansion, thereby improving the data quality of the information data and being more beneficial to subsequent data analysis.
[0064] Step S20, constructing a knowledge expression framework of the target user, and performing data extraction on the target information data based on the knowledge expression framework to obtain a plurality of knowledge expression data.
[0065] Specifically, as shown in the figure, Figure 2 After completing the preprocessing of the information data, a corresponding knowledge expression framework (i.e. a tutor knowledge expression framework) needs to be constructed. The specific construction process is as follows: first, design a framework concept layer of the knowledge expression framework, which is divided into four dimensions, namely basic concept dimension, static feature dimension, dynamic feature dimension and spatial feature concept dimension, wherein the basic concept dimension is described by the other three dimension concepts. The concept layer of the knowledge expression framework corresponds to the expression:
[0066] C = (B, (St, Dc, Sp));
[0067] Wherein, C is the concept layer of the knowledge representation framework, B is the basic concept, St is the static characteristic concept, Dc is the dynamic characteristic concept, and Sp is the spatial characteristic concept. Then, the basic concept of the target user is defined, wherein the basic concept includes the user class (i.e. the supervisor class), the paper class, the journal class, and the conference class. The static abstract concept of the basic concept is defined to obtain the static characteristic concept, and the characteristic expressed by the entity corresponding to the static characteristic concept does not change with time or changes with time in a long period, and in general cases, it can be considered to belong to static, wherein the static characteristic concept includes the homepage class of the supervisor, the Email class, the college class, and the affiliation class. The dynamic abstract concept of the basic concept is defined to obtain the dynamic characteristic concept, and the characteristic expressed by the entity corresponding to the dynamic characteristic concept changes with time and the interval is short, and in the updating process of the knowledge graph, the dynamic change process needs to be considered, wherein the dynamic characteristic concept includes the h-index class, the target user citation quantity (i.e. the supervisor_citations class), and the paper_citations class. The spatial characteristic concept of the target user is defined, wherein the spatial characteristic concept includes the zone class, the country class, and the city class.
[0068] Secondly, the framework relationship layer of the design knowledge expression framework is designed, which is divided into two dimensions, namely static relationship dimension and dynamic relationship dimension. Specifically, the static relationship dimension between the entity of the basic concept and the entity of the static characteristic concept is set. The static relationship represents the static relationship between the concept entities, which does not have obvious time variation characteristics. The relationship included in this static relationship dimension mainly includes the relationship between the basic concept entity and the basic concept entity, and the relationship between the basic concept entity and the static characteristic concept entity. Among them, the relationship between the basic concept and the basic concept includes: "publish" (the concept of the head entity: supervisor, the concept of the tail entity: paper), "be_published_in" (the concept of the head entity: paper, the concept of the tail entity: journal), "conference as" (the concept of the head entity: paper, the concept of the tail entity: conference). The relationship between the basic concept and the static characteristic concept includes: "research_fields_are" (the concept of the head entity: supervisor, the concept of the tail entity: research_fields), "work_college_is" (the concept of the head entity: supervisor, the concept of the tail entity: college), "work_affiliation_is" (the concept of the head entity: supervisor, the concept of the tail entity: affiliation).
[0069] The dynamic relationship dimension between the entity of the basic concept and the entity of the dynamic characteristic concept is set. The dynamic relationship represents the dynamic relationship between the concept entities, which has variability with time. The relationship included in this dynamic relationship dimension mainly includes the relationship between the basic concept entity and the dynamic characteristic concept entity, including: "h-index_is" (the concept of the head entity: supervisor, the concept of the tail entity: h-index), "total_citations_is" (the concept of the head entity: supervisor, the concept of the tail entity: supervisor_paper_citations), "paper_citations_is" (the concept of the head entity: paper, the concept of the tail entity: paper_citations).
[0070] Subsequently, the framework attribute layer of the knowledge representation framework is set. Specifically, the feature concept entity attributes between the basic concepts, the static feature concepts, and the dynamic feature concepts are set. For example, the feature attributes of the school class include: "school name", "establishment time", "student size", etc. The corresponding time attributes are set to express the time corresponding to the concept entity. For dynamic feature concepts, such as citations, in addition to the feature attribute value, it also includes the "update_time" attribute corresponding to the citation count. For basic concept entities, such as paper, it has a clear publication time, so it has the time attribute "publish_time". The relation feature attributes of the feature concept entity attributes are set to express the attributes of the relationship. In addition to the basic attribute "name", the relationship has other feature attributes to further describe it. For example, the relation "publish" (the concept of the head entity: supervisor, the concept of the tail entity: paper) has feature attributes such as: "rank", "weather_corresponding author", etc.
[0071] Finally, a knowledge representation framework is constructed based on the framework concept layer, framework relationship layer, and framework attribute layer, and the target information data is refined based on the knowledge representation framework to obtain multiple knowledge representation data.
[0072] Step S30: Perform multi-source information knowledge fusion on all the knowledge representation data to obtain multiple fused data, and construct a knowledge graph based on all the fused data to obtain a knowledge graph.
[0073] Specifically, such as Figure 2 As shown in this embodiment of the invention, after obtaining multiple knowledge representation data, it is necessary to fuse all the knowledge representation data. Specifically, firstly, a corresponding candidate entity matching library is constructed, which consists of two main steps: rule-based coarse segmentation and clustering-based sub-segmentation. For rule-based coarse segmentation, the knowledge representation data is coarsely segmented according to the knowledge representation framework and the country where the supervisor works, resulting in multiple coarse data blocks. The purpose of coarse segmentation is to perform a rough segmentation before sub-segmentation, as segmentation requires a certain amount of computation. Here, coarse segmentation is based on rule classification, which requires less computation and can reduce the workload of subsequent coarse segmentation. For clustering-based sub-segmentation, a word-level pre-trained embedding model is used to encode all the coarse data blocks, resulting in multiple data vectors. Specifically, the encoding of all the coarse data blocks involves:
[0074] H1 = WEM(N) = [h1,…,h d ];
[0075] wherein, H1 is a data vector of entity name, WEM is a word-level embedding model, N is an entity name, h1 is a data vector of the first dimension, d is a dimension of the data vector, h d is a data vector of the dth dimension, and the purpose of the subdivision block is to finely block, and the purpose of the whole block is to reduce the amount of subsequent entity matching calculation, because the amount of entity matching calculation is large, the candidate matching entities need to be blocked first to reduce the candidate matching entities. All the data vectors are clustered and blocked to obtain a candidate entity matching library. The entities in the same cluster block are close entities and need to be matched subsequently.
[0076] Secondly, the associated feature entity embedding is performed. Specifically, the associated features of all the candidate matching entities in the candidate entity matching library are obtained, all the candidate matching entities and the corresponding associated features are spliced to obtain a plurality of entity description information, and the corresponding form is:
[0077] descriptor = f"name:{name} affiliation:{affiliation}. Research fields:{','.join(research_fields)}. Papers:{','.join(publications[:3])}……";
[0078] The sentence-level pre-trained embedding model is used for vector calculation of all the entity description information to obtain a plurality of embedding vector representations, and the specific process is as follows:
[0079] H2 = SEM(D) = [h1,..., h d ];
[0080] wherein, H2 is an embedding vector representation of entity description information, SEM is a sentence-level embedding model, and D is an entity description information.
[0081] Finally, the entity matching and merging are performed. Specifically, all the embedding vector representations are matched two by two to obtain a plurality of matching results, the entity similarity distance of all the matching results is calculated to obtain a target calculation result, and the corresponding calculation process is as follows:
[0082]
[0083] Similarity(A, B) is the cosine similarity between embedding vector A and embedding vector B, A is the embedding vector of the first entity, B is the embedding vector of the second entity, ‖A‖ is the norm of embedding vector A, and ‖B‖ is the norm of embedding vector B. The calculation result with a cosine similarity greater than a preset threshold (for example, 0.8) in the target calculation result is obtained, because when the cosine similarity is greater than 0.8, two entities are considered to be the same entity, and then the two entities and the associated feature entities need to be merged, that is, the two entities corresponding to all calculation results are respectively merged and processed to obtain a plurality of fusion data, for example, the name of a person in one data is A, and the name of a person in another data is also A. After matching, it is found that the two pieces of information are very similar, and there is a high probability that they are the same person, so they are merged into one piece of data. Otherwise, the two pieces of data are actually two pieces of data with the same name but different people, and cannot be merged. If the cosine similarity is less than or equal to 0.8, it is determined that the data is conflict data, and the data with a more recent update time is used as a reference.
[0084] Step S40, dynamically updating the knowledge graph to obtain a target knowledge graph, and performing data retrieval according to the target knowledge graph to obtain a target retrieval result.
[0085] Specifically, as shown in the figure, Figure 3 In the embodiment of the application, the generated knowledge graph needs to be dynamically updated. Specifically, based on the written crawler program, a timing task is set, the target user is re-collected according to the corresponding period in the timing task, new information data is obtained, and the new information data is data confirmed. If the new information data has updated data, it needs to be updated to the knowledge graph. The updating process is to compare and process the new information data and the information data to obtain new data. The entity type of the new data is determined, and the knowledge graph is dynamically updated according to the update mode corresponding to the entity type to obtain a target knowledge graph, wherein the entity type is any one of a basic concept type, a dynamic feature concept type and a static feature concept type.
[0086] If the entity type is the basic concept type, the new data is subjected to entity embedding and similarity calculation to obtain a first processing result, and it is judged according to the first processing result whether the new data is an existing entity node. If the new data is an existing entity node, the knowledge graph is not updated. If the new data is a non-existing entity node, the entity node corresponding to the new data is added in the knowledge graph, and the association relationship of the entity node is established, for example, a school hires a tutor, the tutor publishes a new article, etc., to obtain a target knowledge graph.
[0087] If the entity type is the dynamic feature concept type, a dynamic feature node corresponding to the new data is added in the knowledge graph according to an update period of the timing task, and a dynamic association relationship of the dynamic feature node is established, so as to obtain a target knowledge graph; for example, the h-index index and the citation amount data of each tutor in Google Scholar are captured in a period of years, the new information data is formed into a new entity node, the time attribute value of the node is the update time, the numerical attribute value of the node is the updated data, a new relationship of "h-index index" and "citation amount" is established between the new entity node and the tutor entity node, and the update time is recorded in the time attribute value of the relationship.
[0088] If the entity type is the static feature concept type, entity embedding and similarity calculation are performed on the new data to obtain a second processing result; whether the new data is an existing entity node is judged according to the second processing result, if the new data is an existing entity node, a static feature node corresponding to the new data is added in the knowledge graph, and an association relationship of the static feature node is established, so as to obtain a target knowledge graph, and data retrieval is performed according to the target knowledge graph to obtain a target retrieval result.
[0089] The data retrieval method based on the knowledge graph provided by the application provides technical details of the construction method of the knowledge graph for the study application from the aspects of data collection, model construction, graph fusion and graph updating. Through the technical method provided by the application, the knowledge graph for the study application can be gradually constructed from scratch, and complex and multiple tutor-related knowledge can be integrated in a highly structured manner. Using the knowledge graph, applicants can be provided with convenient information search and accurate recommendation services to support applicants to find suitable tutor information more quickly and accurately. The knowledge graph can also support knowledge services such as knowledge-based question answering and knowledge-based intelligent recommendation, thereby effectively improving the intelligent level of the study application service industry.
[0090] Further, as shown in Figure 4 Based on the above-mentioned data retrieval method based on the knowledge graph, the application further correspondingly provides a data retrieval system based on the knowledge graph, wherein the data retrieval system based on the knowledge graph comprises:
[0091] The data processing module 51 is configured to obtain information data of a target user, and pre-process the information data to obtain target information data.
[0092] The data extraction module 52 is configured to construct a knowledge expression framework, and extract data from the target information data based on the knowledge expression framework to obtain a plurality of knowledge expression data.
[0093] The knowledge graph construction module 53 is used to perform multi-source information knowledge fusion on all the knowledge representation data to obtain multiple fused data, and to construct a knowledge graph based on all the fused data to obtain a knowledge graph.
[0094] The knowledge graph update module 54 is used to dynamically update the knowledge graph to obtain the target knowledge graph.
[0095] Furthermore, such as Figure 4 As shown, based on the above-mentioned knowledge graph-based data retrieval method, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0096] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a knowledge graph-based data retrieval program 40, which can be executed by the processor 10 to implement the knowledge graph-based data retrieval method of this application.
[0097] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the knowledge graph-based data retrieval method.
[0098] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0099] In an embodiment, the following steps are implemented when the processor 10 executes the knowledge graph-based data retrieval program 40 in the memory 20:
[0100] Obtain information data of a target user, pre-process the information data to obtain target information data;
[0101] Construct a knowledge expression framework of the target user, perform data extraction on the target information data based on the knowledge expression framework to obtain a plurality of knowledge expression data;
[0102] Perform multi-source information knowledge fusion on all the knowledge expression data to obtain a plurality of fusion data, and perform knowledge graph construction according to all the fusion data to obtain a corresponding knowledge graph;
[0103] Perform dynamic updating on the knowledge graph to obtain a target knowledge graph, and perform data retrieval according to the target knowledge graph to obtain a target retrieval result.
[0104] The obtaining of the information data of the target user and the pre-processing of the information data to obtain the target information data specifically includes:
[0105] Collect and process related information of the target user to obtain information data, and perform data cleaning on the information data, wherein the data cleaning includes de-duplication processing, missing value processing, and outlier detection processing;
[0106] Perform format standardization processing on the cleaned information data, and perform text enhancement processing on the format-standardized information data to obtain target information data, wherein the format standardization processing includes unifying time format, standardizing address, and unifying case, and the text enhancement includes special symbol processing and abbreviation expansion.
[0107] The construction of the knowledge expression framework of the target user specifically includes:
[0108] Define basic concepts of the target user;
[0109] Define static abstract concepts of the basic concepts to obtain static feature concepts;
[0110] Define dynamic abstract concepts of the basic concepts to obtain dynamic feature concepts;
[0111] Define spatial feature concepts of the target user;
[0112] Construct a framework concept layer according to the basic concepts, the static feature concepts, the dynamic feature concepts, and the spatial feature concepts;
[0113] set a static relationship dimension between the entity of the basic concept and the entity of the static feature concept, set a dynamic relationship dimension between the entity of the basic concept and the entity of the dynamic feature concept, and construct a framework relationship layer according to the static relationship dimension and the dynamic relationship dimension;
[0114] set a feature concept entity attribute and a corresponding time attribute between the basic concept, the static feature concept and the dynamic feature concept, and set a relationship feature attribute of the feature concept entity attribute, and construct a framework attribute layer according to the feature concept entity attribute, the time attribute and the relationship feature attribute;
[0115] construct a knowledge expression framework according to the framework concept layer, the framework relationship layer and the framework attribute layer.
[0116] The multi-source information knowledge fusion of all the knowledge expression data obtains a plurality of fusion data, specifically including:
[0117] perform coarse block processing on the knowledge expression data to obtain a plurality of data coarse blocks, perform encoding processing on all the data coarse blocks to obtain a plurality of data vectors, and perform clustering block processing on all the data vectors to obtain a candidate entity matching library;
[0118] obtain the associated features of all the candidate matching entities in the candidate entity matching library, perform splicing processing on all the candidate matching entities and the corresponding associated features to obtain a plurality of entity description information, and perform vector calculation on all the entity description information to obtain a plurality of embedded vector representations;
[0119] match all the embedded vector representations two by two to obtain a plurality of matching results, and calculate the entity similarity distance of all the matching results to obtain a target calculation result;
[0120] obtain the calculation results with a cosine similarity greater than a preset threshold in the target calculation result, and perform merging processing on the two entities corresponding to all the calculation results respectively to obtain a plurality of fusion data.
[0121] The encoding processing on all the data coarse blocks is specifically:
[0122] H1=WEM(N)=[h1,…,h d ];
[0123] The vector calculation on all the entity description information is specifically:
[0124] H2=SEM(D)=[h1,…,h d ];
[0125] The entity similarity distance of all the matching results is calculated, specifically:
[0126]
[0127] wherein H1 is a data vector of the entity name, WEM is a word-level embedding model, N is the entity name, h1 is a data vector of the first dimension, d is the dimension of the data vector, h d is a data vector of the dth dimension, H2 is an embedding vector representation of the entity description information, SEM is a sentence-level embedding model, D is the description information of the entity, Similarity(A,B) is the cosine similarity between embedding vector A and embedding vector B, A is the embedding vector of the first entity, B is the embedding vector of the second entity, ‖A‖ is the norm of embedding vector A, and ‖B‖ is the norm of embedding vector B.
[0128] The knowledge graph is dynamically updated to obtain a target knowledge graph, specifically including:
[0129] A timing task is set, information of the target user is collected according to the timing task to obtain new information data, and the new information data is subjected to data confirmation;
[0130] If the new information data has update data, the new information data is compared with the information data to obtain new data;
[0131] The entity type of the new data is determined, the knowledge graph is dynamically updated according to the update mode corresponding to the entity type to obtain a target knowledge graph, and the entity type is any one of a basic concept type, a dynamic feature concept type and a static feature concept type.
[0132] The knowledge graph is dynamically updated according to the update mode corresponding to the entity type to obtain a target knowledge graph, specifically including:
[0133] If the entity type is the basic concept type, entity embedding and similarity calculation are performed on the new data to obtain a first processing result, and it is judged whether the new data is an existing entity node according to the first processing result;
[0134] If the new data is a non-existing entity node, an entity node corresponding to the new data is added in the knowledge graph, and an association relationship of the entity node is established to obtain a target knowledge graph;
[0135] If the entity type is the dynamic feature concept type, a dynamic feature node corresponding to the new data is added in the knowledge graph according to an update period of the timing task, and a dynamic association relationship of the dynamic feature node is established, so as to obtain a target knowledge graph.
[0136] If the entity type is the static feature concept type, entity embedding and similarity calculation are performed on the new data, so as to obtain a second processing result.
[0137] It is judged according to the second processing result whether the new data is an existing entity node, if the new data is a non-existing entity node, a static feature node corresponding to the new data is added in the knowledge graph, and an association relationship of the static feature node is established, so as to obtain a target knowledge graph.
[0138] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a knowledge graph based data retrieval program, and the knowledge graph based data retrieval program is executed by a processor to realize the steps of the knowledge graph based data retrieval method.
[0139] In summary, the application provides a knowledge graph based data retrieval method, system, terminal and storage medium, the method comprising: obtaining information data of a target user, preprocessing the information data to obtain target information data; constructing a knowledge expression framework of the target user, performing data extraction on the target information data based on the knowledge expression framework to obtain a plurality of knowledge expression data; performing multi-source information knowledge fusion on all the knowledge expression data to obtain a plurality of fusion data, and constructing a knowledge graph based on all the fusion data to obtain a corresponding knowledge graph; dynamically updating the knowledge graph to obtain a target knowledge graph, and performing data retrieval based on the target knowledge graph to obtain a target retrieval result. The application integrates complex and multi-element related information in a highly structured manner by constructing a knowledge graph, thereby providing convenient information search and accurate recommendation services for users, so that users can quickly and accurately find the best target information.
[0140] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, methods, articles or terminals including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles or terminals. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or terminal including the element.
[0141] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable computer-readable storage medium, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0142] It should be understood that the application is not limited to the above examples, and those skilled in the art can make improvements or changes according to the above description, and all these improvements and changes shall belong to the protection scope of the appended claims of the application.
Claims
1. A knowledge graph based data retrieval method, characterized in that, The knowledge graph-based data retrieval method comprises the following steps: Obtain information data of a target user, preprocess the information data to obtain target information data; Construct a knowledge expression framework of the target user, perform data extraction on the target information data based on the knowledge expression framework to obtain a plurality of knowledge expression data; Fuse all the knowledge expression data to obtain a plurality of fusion data, and construct a knowledge graph based on all the fusion data to obtain a corresponding knowledge graph; Dynamically update the knowledge graph to obtain a target knowledge graph, and perform data retrieval based on the target knowledge graph to obtain a target retrieval result. 2.The knowledge graph based data retrieval method according to claim 1, characterized in that, The information data of the target user is obtained, and the information data is preprocessed to obtain target information data, specifically including: Collect and process the relevant information of the target user to obtain information data, and perform data cleaning on the information data, wherein the data cleaning includes de-duplication processing, missing value processing and outlier detection processing; Perform format standardization processing on the cleaned information data, and perform text enhancement processing on the format-standardized information data to obtain target information data, wherein the format standardization processing includes unified time format, standardized address and unified case, and the text enhancement includes special symbol processing and abbreviation expansion. 3.The knowledge graph based data retrieval method of claim 1, wherein, The knowledge expression framework of the target user is constructed, specifically including: Define the basic concept of the target user; Define the static abstract concept of the basic concept to obtain a static feature concept; Define the dynamic abstract concept of the basic concept to obtain a dynamic feature concept; Define the spatial feature concept of the target user; Construct a framework concept layer according to the basic concept, the static feature concept, the dynamic feature concept and the spatial feature concept; Set the static relationship dimension between the entity of the basic concept and the entity of the static feature concept, set the dynamic relationship dimension between the entity of the basic concept and the entity of the dynamic feature concept, and construct a framework relationship layer according to the static relationship dimension and the dynamic relationship dimension; Set the feature concept entity attribute, the corresponding time attribute and the relationship feature attribute between the basic concept, the static feature concept and the dynamic feature concept, and construct a framework attribute layer according to the feature concept entity attribute, the time attribute and the relationship feature attribute; Construct a knowledge expression framework according to the framework concept layer, the framework relationship layer and the framework attribute layer. 4.The knowledge graph based data retrieval method of claim 1, wherein, The knowledge expression data is subjected to rough block processing to obtain a plurality of data rough blocks, all the data rough blocks are subjected to encoding processing to obtain a plurality of data vectors, and all the data vectors are subjected to clustering block to obtain a candidate entity matching library; Obtaining the associated features of all candidate matching entities in the candidate entity matching library, splicing all the candidate matching entities and the corresponding associated features to obtain a plurality of entity description information, and performing vector calculation on all the entity description information to obtain a plurality of embedding vector representations; Matching all the embedding vector representations two by two to obtain a plurality of matching results, and calculating the entity similarity distance of all the matching results to obtain a target calculation result; Obtaining the calculation result with a cosine similarity greater than a preset threshold in the target calculation result, and merging the two entities corresponding to all the calculation results to obtain a plurality of fusion data. 5.The knowledge graph based data retrieval method of claim 4, wherein, The encoding processing of all the data coarse blocks is specifically: H1=WEM(N)=[h1,…,h d ]; The vector calculation of all the entity description information is specifically: H2 = SEM(D) = [h1,..., h d ]; The calculation of the entity similarity distance of all the matching results is specifically: wherein H1 is a data vector of entity name, WEM is a word-level embedding model, N is an entity name, h1 is a data vector of the first dimension, d is a dimension of the data vector, h d is a data vector of the dth dimension, H2 is an embedding vector representation of entity description information, SEM is a sentence-level embedding model, D is a description information of an entity, Similarity(A, B) is a cosine similarity between embedding vector A and embedding vector B, A is an embedding vector of a first entity, B is an embedding vector of a second entity, ‖A‖ is a norm of embedding vector A, and ‖B‖ is a norm of embedding vector B. 6.The knowledge graph based data retrieval method according to claim 1, characterized in that, The dynamic updating of the knowledge graph to obtain a target knowledge graph specifically includes: Setting a timing task, collecting information of the target user according to the timing task to obtain new information data, and performing data confirmation on the new information data; If the new information data has update data, compare the new information data with the information data to obtain new data; Determine the entity type of the new data, and dynamically update the knowledge graph according to the update mode corresponding to the entity type to obtain a target knowledge graph, wherein the entity type is any one of a basic concept type, a dynamic feature concept type and a static feature concept type. 7.The knowledge graph based data retrieval method of claim 6, wherein, The dynamic updating of the knowledge graph according to the update mode corresponding to the entity type to obtain a target knowledge graph specifically includes: If the entity type is the basic concept type, perform entity embedding and similarity calculation on the new data to obtain a first processing result, and determine whether the new data is an existing entity node according to the first processing result; If the new data is a non-existing entity node, add an entity node corresponding to the new data in the knowledge graph, and establish the association relationship of the entity node to obtain a target knowledge graph; If the entity type is the dynamic feature concept type, add a dynamic feature node corresponding to the new data in the knowledge graph according to the update period of the timing task, and establish the dynamic association relationship of the dynamic feature node to obtain a target knowledge graph; If the entity type is the static feature concept type, perform entity embedding and similarity calculation on the new data to obtain a second processing result; Determine whether the new data is an existing entity node according to the second processing result, and if the new data is a non-existing entity node, add a static feature node corresponding to the new data in the knowledge graph, and establish the association relationship of the static feature node to obtain a target knowledge graph. 8.A knowledge graph based data retrieval system, characterized in that, The data retrieval system based on the knowledge graph includes: A data processing module is configured to obtain information data of a target user, and pre-process the information data to obtain target information data; The data extraction module is configured to construct a knowledge expression framework of the target user, perform data extraction on the target information data based on the knowledge expression framework, and obtain a plurality of knowledge expression data. The graph construction module is configured to perform multi-source information knowledge fusion on all the knowledge expression data, obtain a plurality of fusion data, and construct a knowledge graph according to all the fusion data to obtain a corresponding knowledge graph. The graph updating module is configured to perform dynamic updating on the knowledge graph to obtain a target knowledge graph, perform data retrieval according to the target knowledge graph, and obtain a target retrieval result.
9. A terminal, characterized by comprising: The terminal comprises a memory, a processor, and a knowledge graph-based data retrieval program stored on the memory and executable on the processor. When the knowledge graph-based data retrieval program is executed by the processor, the steps of the knowledge graph-based data retrieval method according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a knowledge graph-based data retrieval program. When the knowledge graph-based data retrieval program is executed by the processor, the steps of the knowledge graph-based data retrieval method according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Data processing method based on knowledge map
CN109255031A
Text matching method and device based on slot similarity, equipment and storage medium
CN111400493A
Implementation method for entity matching and partitioning
CN116028596A
Traffic accident report time series knowledge graph modeling method based on deep learning
CN119761479A