Data processing method and device based on knowledge graph and computer equipment

Through the data processing method based on knowledge graph, the entity recognition and semantic matching model are used to solve the limitations of the prior art when dealing with relationship generalization and entity positioning, and the precise identification and efficient retrieval of requested entities and relationships are realized.

CN120196765APending Publication Date: 2025-06-24JIANGNAN INST OF COMPUTING TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510290305.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has limitations when dealing with the problem of relationship generalization and cannot flexibly deal with the problem of user diversity. In addition, information retrieval methods have ambiguity and generalization when locating entities and relationships, making it difficult to accurately identify the entities and relationships involved in the problem.

Method used

Through the data processing method based on knowledge graph, the preset entity recognition model is used to identify the request information, and the entity searches are performed based on the preset database and search thresholds, and the post-retrieval mapping relationship is obtained, and the semantic matching model is used for matching analysis to accurately identify entities and relationships.

Benefits of technology

It improves the search efficiency of the precise search and mapping relationship of the requesting entity, can flexibly respond to user diversification problems, and enhances the accuracy and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196765A_ABST
    Figure CN120196765A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing method and device based on a knowledge graph and computer equipment. The method comprises the steps of obtaining a current data request and determining current request information based on the current data request; performing entity recognition on the current request information according to a preset entity recognition model to obtain a current entity recognition result; under the condition that the current entity recognition result is that the current request entity exists, performing entity retrieval on the current request entity based on a preset database and a preset retrieval threshold value to obtain a current post-retrieval mapping relation; and performing matching analysis on the current post-retrieval mapping relationship according to the current request information, a preset semantic matching model and a preset matching threshold to obtain current target data. By adopting the method, the accuracy of searching the current request entity can be improved, and the searching efficiency of accurately searching the mapping relation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing and analysis, and particularly to a data processing method, apparatus, and computer device based on a knowledge graph. Background Art

[0002] In today's information age, enterprises and organizations have accumulated a vast amount of structured data, including various types such as equipment, people, and locations. With the continuous increase in the amount of data, how to effectively manage and utilize this data has become a crucial issue. A knowledge graph, as a powerful knowledge representation tool, can formalize the description of structured knowledge and construct a networked knowledge structure through the relationships between entities, which greatly promotes the organization and acquisition of knowledge. Especially in intelligent question-answering systems, the application of knowledge graphs gradually shows advantages such as fast retrieval efficiency and high answer accuracy.

[0003] In the related technologies of knowledge graph data processing, it is mainly divided into two methods: semantic parsing-based and information retrieval-based. The semantic parsing-based method usually relies on machine learning models, combines dictionaries and rules, and retrieves answers by identifying entities and relationships in the question. The information retrieval-based method first identifies the entities in the question, then obtains the candidate relationships related to them, and then uses a semantic similarity matching model to find the relationship closest to the question.

[0004] However, in the current related technologies, although the semantic parsing-based method performs well in specific scenarios, it has obvious limitations in dealing with the problem of relationship generalization and often cannot flexibly handle the diverse questions of users. The information retrieval-based method often has a certain degree of ambiguity and generalization for users' questions, making it difficult to directly locate entities and relationships based on the question alone. Summary of the Invention

[0005] Based on this, it is necessary to provide a data processing method, apparatus, computer device, computer-readable storage medium, and computer program product based on a knowledge graph that can accurately identify the entities involved in the question and improve the retrieval efficiency of entity candidate relationships for the above technical problems.

[0006] In a first aspect, this application provides a data processing method based on a knowledge graph. The method includes:

[0007] Obtain a current data request and determine current request information based on the current data request;

[0008] Perform entity recognition on the current request information according to a preset entity recognition model to obtain a current entity recognition result;

[0009] When the current entity recognition result indicates the existence of the current requested entity, entity retrieval is performed on the current requested entity based on a preset database and a preset retrieval threshold to obtain the current retrieved mapping relationship;

[0010] Match analysis is performed on the current retrieved mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain the current target data.

[0011] In one embodiment, when the current entity recognition result indicates the existence of the current requested entity, entity retrieval is performed on the current requested entity based on a preset database and a preset retrieval threshold to obtain the current retrieved mapping relationship, including:

[0012] Retrieve and compare the current requested entity based on a preset retrieval lower threshold and a preset entity to obtain a first type of entity and a second type of entity;

[0013] When the number of the first type of entity is greater than zero, retrieval is performed based on a preset graph and the first type of entity to obtain the current retrieved mapping relationship.

[0014] In one embodiment, when the number of the first type of entity is zero, obtain the next data request.

[0015] In one embodiment, when the number of the first type of entity is greater than zero, retrieval is performed based on a preset graph and the first type of entity to obtain the current retrieved mapping relationship, including:

[0016] Divide the first type of entity based on a preset retrieval upper threshold to obtain a third type of entity and a fourth type of entity;

[0017] When the number of the third type of entity is greater than zero, perform mapping retrieval based on the third type of entity and a preset graph to obtain the current retrieved mapping relationship;

[0018] When the number of the third type of entity is zero, screen the fourth type of entity to obtain the screened entity;

[0019] Perform mapping retrieval based on the screened entity and a preset graph to obtain the current retrieved mapping relationship.

[0020] In one embodiment, match analysis is performed on the current retrieved mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain the current target data, including:

[0021] Perform data splicing on the current request information and the current retrieved mapping relationship to obtain the spliced data;

[0022] Calculate the spliced data based on a preset similarity algorithm to obtain a similarity value;

[0023] Determine the target mapping relationship based on the similarity value and the preset matching threshold;

[0024] Determine the current target data based on the target mapping relationship and the current request information.

[0025] In one embodiment, determining the target mapping relationship based on the similarity value and the preset matching threshold includes:

[0026] Divide the current retrieved mapping relationship based on the similarity value and the preset matching threshold to obtain a first mapping relationship and a second mapping relationship;

[0027] When the number of the first mapping relationships is greater than zero, set the first mapping relationship as the target mapping relationship;

[0028] When the number of the first mapping relationships is zero, screen the second mapping relationship to obtain the screened mapping relationship and set the screened mapping relationship as the target mapping relationship.

[0029] In one embodiment, after performing entity recognition on the current request information according to the preset entity recognition model to obtain the current entity recognition result, the method further includes:

[0030] When the current entity recognition result indicates that the current request entity does not exist, determine the previous request information and the previous retrieved mapping relationship based on the historical query database;

[0031] Perform request analysis on the current request information based on the previous request information and the preset interval threshold to obtain a request analysis result;

[0032] When the request analysis result indicates that the preset interval threshold is not exceeded, perform matching analysis on the previous retrieved mapping relationship according to the current request information, the preset semantic matching model, and the preset matching threshold to obtain the current target data.

[0033] In one embodiment, performing request analysis on the current request information based on the previous request information and the preset interval threshold to obtain a request analysis result includes:

[0034] Perform data analysis on the previous request information and the current request information to obtain a data analysis result and a request interval;

[0035] When the data analysis result is the same user identifier, perform request analysis on the request interval and the preset interval threshold to obtain a request analysis result.

[0036] In a second aspect, the present application further provides a data processing device based on a knowledge graph. The device includes:

[0037] An acquisition module, configured to acquire a current data request and determine current request information based on the current data request;

[0038] An identification module, configured to perform entity identification on the current request information according to a preset entity identification model to obtain a current entity identification result;

[0039] A retrieval module, configured to perform entity retrieval on the current request entity based on a preset database and a preset retrieval threshold when the current entity identification result indicates the existence of the current request entity, to obtain a current post-retrieval mapping relationship;

[0040] A matching module, configured to perform matching analysis on the current post-retrieval mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain current target data.

[0041] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0042] Acquire a current data request and determine current request information based on the current data request;

[0043] Perform entity identification on the current request information according to a preset entity identification model to obtain a current entity identification result;

[0044] When the current entity identification result indicates the existence of the current request entity, perform entity retrieval on the current request entity based on a preset database and a preset retrieval threshold to obtain a current post-retrieval mapping relationship;

[0045] Perform matching analysis on the current post-retrieval mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain current target data.

[0046] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0047] Acquire a current data request and determine current request information based on the current data request;

[0048] Perform entity identification on the current request information according to a preset entity identification model to obtain a current entity identification result;

[0049] When the current entity identification result indicates the existence of the current request entity, perform entity retrieval on the current request entity based on a preset database and a preset retrieval threshold to obtain a current post-retrieval mapping relationship;

[0050] Perform matching analysis on the currently retrieved mapping relationship according to the current request information, the preset semantic matching model, and the preset matching threshold to obtain the current target data.

[0051] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps:

[0052] Obtain the current data request and determine the current request information based on the current data request;

[0053] Perform entity recognition on the current request information according to the preset entity recognition model to obtain the current entity recognition result;

[0054] In the case where the current entity recognition result indicates the existence of the current request entity, perform entity retrieval on the current request entity based on the preset database and the preset retrieval threshold to obtain the currently retrieved mapping relationship;

[0055] Perform matching analysis on the currently retrieved mapping relationship according to the current request information, the preset semantic matching model, and the preset matching threshold to obtain the current target data.

[0056] The above data processing method, device, computer device, storage medium, and computer program product based on the knowledge graph, after obtaining the problem, that is, the current data request, use the pre-trained entity recognition model, that is, the preset entity recognition model, to find entities in the current request information, and after finding the current request entity, retrieve similar preset entities in the database storing the preset entity information, and then obtain the entity mapping relationship corresponding to the similar preset entity in the database storing the entity mapping relationship, that is, the currently retrieved mapping relationship. Then, use the preset semantic matching model to find the mapping relationship with a similar semantics corresponding to the current request information and find the current target data in the mapping relationship, improving the accuracy of the retrieval of the current request entity and the retrieval efficiency of accurately retrieving the mapping relationship. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic flowchart of the data processing method based on the knowledge graph in an embodiment;

[0058] Figure 2 It is a schematic structural diagram of the preset entity recognition model in an embodiment;

[0059] Figure 3 It is a schematic structural diagram of the preset semantic matching model in an embodiment;

[0060] Figure 4 It is a schematic flowchart of the matching analysis step in an embodiment;

[0061] Figure 5Schematic structural diagram of a knowledge graph in an embodiment;

[0062] Figure 6 Flowchart of a data processing method based on a knowledge graph in an embodiment;

[0063] Figure 7 Block diagram of the structure of a data processing device based on a knowledge graph in an embodiment;

[0064] Figure 8 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0065] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0066] In one embodiment, as Figure 1 shown, a data processing method based on a knowledge graph is provided. In this embodiment, the method is exemplified by being applied to a terminal. It can be understood that the method can also be applied to a server and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0067] Step 102: Obtain a current data request and determine current request information based on the current data request.

[0068] Exemplarily, obtain the question input by the terminal user as the current data request, and perform data preprocessing on the current data request to obtain relevant data request information, that is, the current request information. Among them, the current request information includes information such as question data, user identification, Chinese date, and date standardization of fields representing time significance.

[0069] Step 104: Perform entity recognition on the current request information according to a preset entity recognition model to obtain a current entity recognition result.

[0070] Among them, the preset entity recognition model is a Bert-BiLSTM-CRF (Bidirectional Encoder Representations from Transformers - Long Short-Term Memory - Conditional Random Fields) model, and the preset entity recognition model is as Figure 2 shown.

[0071] Exemplarily, when using the Bert-BiLSTM-CRF model to perform entity recognition on the current request information, the specific steps are as follows:

[0072] The BERT layer encodes the original data to generate context-related encoding vectors for the sequence; then, the BiLSTM layer can effectively capture long-distance text dependencies and rich context semantic information; subsequently, the CRF layer adds constraints to the output of the BiLSTM, reduces the probability of unreasonable sequences through optimized calculations, and thus trains to obtain the optimal label sequence, and finally outputs the prediction result, that is, the current recognition result.

[0073] The current recognition result includes the problem input by the end user, that is, the current data request has a current request entity, and the problem input by the end user, that is, the current data request does not have a current request entity.

[0074] If the problem input by the end user, that is, the current data request does not have a current request entity, then it is necessary to obtain the problem re-entered by the end user, that is, the next data request, and then perform entity recognition operations on the next data request.

[0075] Step 106, in the case where the current entity recognition result is that there is a current request entity, perform entity retrieval on the current request entity based on a preset database and a preset retrieval threshold to obtain the current retrieved mapping relationship.

[0076] Among them, the preset database includes preset entities and a preset graph. The preset entities include other entity key information such as the full entity name, entity abbreviation, and entity alias. The preset retrieval threshold includes a preset retrieval lower limit threshold.

[0077] Exemplarily, before making a data request, an auxiliary Elasticsearch database for relevant entity nodes is established according to the knowledge graph database, that is, the knowledge graph database and the auxiliary Elasticsearch database form the preset database.

[0078] In the case where the problem input by the end user, that is, the current data request has a current request entity, calculate the similarity between the preset entities in the auxiliary Elasticsearch database and the current request entity, and use the preset retrieval lower limit threshold in the preset retrieval threshold to combine the similarity calculation results to determine whether there are similar preset entities.

[0079] If there is no similarity result greater than or equal to the preset retrieval lower limit threshold, the user is reminded to re-enter the problem; otherwise, the full entity name of the similar entity is output.

[0080] Then, in the knowledge graph database, search for all mapping relationships of the output entity full name according to the output entity full name, and set them as the current retrieved mapping relationship.

[0081] Step 108, perform matching analysis on the current retrieved mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain the current target data.

[0082] Among them, the preset semantic matching model is the Bert (Bidirectional Encoder Representations from Transformers) semantic similarity matching model. The current retrieved mapping relationship includes multiple mapping relationships, and the preset semantic matching model is as Figure 3 shown.

[0083] Exemplarily, before making a data request, a model is constructed and trained by combining the powerful language representation ability of BERT, the sequence modeling ability of BiLSTM, and the attention mechanism (Attention) to obtain the Bert semantic similarity matching model.

[0084] Input the current request information and the current retrieved mapping relationship into the Bert semantic similarity matching model, that is, the preset semantic matching model, calculate the similarity between the current request information and each mapping relationship in the current retrieved mapping relationship, and output the probability value of the current request information and each mapping relationship. Among them, the probability value is a numerical value between 0 and 1, and the closer the probability value is to 1, the more semantically similar the current request information is to the mapping relationship corresponding to the probability value.

[0085] Then determine the current target data based on the current request entity in the current request information and the mapping relationship corresponding to the maximum probability value.

[0086] In the above data processing method based on the knowledge graph, after obtaining the question, that is, the current data request, use the pre-trained entity recognition model, that is, the preset entity recognition model, to find entities in the current request information, and after finding the current request entity, retrieve similar preset entities in the database storing the preset entity information, and then obtain the entity mapping relationship corresponding to the similar preset entity in the database storing the entity mapping relationship, that is, the current retrieved mapping relationship. Then use the preset semantic matching model to find the mapping relationship semantically similar to the current request information and find the current target data in the mapping relationship, which improves the accuracy of retrieving the current request entity and the retrieval efficiency of accurately retrieving the mapping relationship.

[0087] In an exemplary embodiment, when the current entity recognition result is that there is a current request entity, perform entity retrieval on the current request entity based on a preset database and a preset retrieval threshold to obtain the current retrieved mapping relationship, including:

[0088] Retrieve and compare the current requested entity based on a preset retrieval lower threshold and preset entities to obtain a first type of entity and a second type of entity; when the number of the first type of entity is greater than zero, retrieve based on a preset knowledge graph and the first type of entity to obtain the current retrieved mapping relationship.

[0089] Among them, the preset database includes preset entities and a preset knowledge graph. The preset entities include other entity key information such as the full entity name, entity abbreviation, and entity alias. The preset retrieval threshold includes a preset retrieval lower threshold.

[0090] Exemplarily, after obtaining the current requested entity, retrieve and match the preset entities in the auxiliary Elasticsearch database with the current requested entity to obtain the retrieval matching values of the preset entities. Then compare all the retrieval matching values with the preset retrieval lower threshold, and divide the preset entities according to the comparison results, specifically:

[0091] Set the preset entities corresponding to the retrieval matching values greater than or equal to the retrieval lower threshold as the first type of entity;

[0092] Set the preset entities corresponding to the retrieval matching values less than the retrieval lower threshold as the second type of entity.

[0093] After that, perform the next step according to the number of the first type of entity, specifically including:

[0094] If the number of the first type of entity is greater than zero, generate a mapping relationship according to the preset knowledge graph corresponding to the first type of entity in the knowledge graph database, and set this mapping relationship as the current retrieved mapping relationship;

[0095] If the number of the first type of entity is zero, it is necessary to obtain the problem re-entered by the terminal user, that is, the next data request.

[0096] In an exemplary embodiment, retrieving based on a preset knowledge graph and the first type of entity to obtain the current retrieved mapping relationship includes:

[0097] Divide the first type of entity based on a preset retrieval upper threshold to obtain a third type of entity and a fourth type of entity; when the number of the third type of entity is greater than zero, perform a mapping retrieval based on the third type of entity and the preset knowledge graph to obtain the current retrieved mapping relationship; when the number of the third type of entity is zero, screen the fourth type of entity to obtain the screened entity; perform a mapping retrieval based on the screened entity and the preset knowledge graph to obtain the current retrieved mapping relationship.

[0098] Among them, the preset retrieval threshold further includes a preset retrieval upper threshold.

[0099] Exemplarily, after obtaining the first type of entities, compare the retrieval matching values corresponding to all the first type of entities with a preset retrieval upper limit threshold, and classify the first type of entities according to the comparison results. Specifically:

[0100] Set the preset entities corresponding to the retrieval matching values greater than or equal to the retrieval upper limit threshold as the third type of entities;

[0101] Set the preset entities corresponding to the retrieval matching values less than the retrieval upper limit threshold as the fourth type of entities.

[0102] After that, perform the next step according to the number of the third type of entities, which specifically includes:

[0103] If the number of the third type of entities is greater than zero, generate a mapping relationship according to the preset graph corresponding to the third type of entities in the knowledge graph database, and set this mapping relationship as the current post-retrieval mapping relationship;

[0104] If the number of the first type of entities is zero, generate a mapping relationship according to the preset graph corresponding to the fourth type of entities in the knowledge graph database, output several preset entities with larger retrieval matching values corresponding to the fourth type of entities, and then let the user select. After that, set the mapping relationship corresponding to the preset entity selected by the user as the current post-retrieval mapping relationship.

[0105] In an exemplary embodiment, as Figure 4 shown, perform matching analysis on the current post-retrieval mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain the current target data, including:

[0106] Step 402, splice the current request information and the current post-retrieval mapping relationship to obtain the spliced data.

[0107] Exemplarily, after obtaining the current post-retrieval mapping relationship through retrieval, splice the current request information and the current post-retrieval mapping relationship. Specifically, add the "[CLS]" identifier at the start of the current data request corresponding to the current request information and add the "[SEP]" identifier at the end, then sequentially splice a mapping relationship and add the "[SEP]" identifier at the tail, thereby obtaining the spliced data.

[0108] For example, if the current data request corresponding to the current request information is P = {P1,..., P n}, and the current post-screening mapping relationship is Q = {Q1,..., Q n}, then the spliced data is:

[0109] X = {[CLS], P1,..., P n , [SEP], Q1,..., Q n , [SEP]}.

[0110] Step 404: Calculate the spliced data based on a preset similarity algorithm to obtain a similarity value.

[0111] Exemplarily, after obtaining the spliced data, use the Bert encoding layer of the preset semantic matching model to encode the spliced data to obtain corresponding encoding vectors, specifically:

[0112] Bert(X)=L={l1, l2, …, l m}

[0113] where l ∈ R m*d , m is the length of the input X, and l i is the representation vector of the i-th character.

[0114] Then, connect the encoding vector output by Bert and the information obtained by Attention on the aggregation layer of the preset semantic matching model, and input it into a bidirectional BiLSTM layer. The BiLSTM layer can better train and learn the long-distance text dependency relationship and the semantic information of the context. Finally, obtain a fixed-length vector through pooling and convert it into a probability value, that is, the similarity value, as shown below:

[0115] P = Softmax[w * r + b]

[0116] where r is the text vector output after pooling, P is the predicted similarity value, and w and b are the weight parameter and the bias term parameter, respectively.

[0117] Step 406: Determine the target mapping relationship based on the similarity value and a preset matching threshold.

[0118] Exemplarily, compare the calculated similarity value with the preset matching threshold and determine the target mapping relationship according to the comparison result, specifically:

[0119] If the maximum similarity threshold is greater than or equal to the preset matching threshold, set the mapping relationship corresponding to the maximum similarity threshold as the target mapping relationship;

[0120] If the maximum similarity threshold is less than the preset matching threshold, select a suitable mapping relationship from several with larger similarity thresholds and set it as the target mapping relationship.

[0121] Step 408: Determine the current target data based on the target mapping relationship and the current request information.

[0122] Exemplarily, search for the target data corresponding to the current data request in the target mapping relationship according to the current request information.

[0123] For example, the current data request, i.e., the question is "Which country manufactures the F16?", and the target mapping relationship, i.e., the connection relationship of the R & D companies in the knowledge graph, is obtained through data processing.

[0124] In an exemplary embodiment, determining the target mapping relationship based on the similarity value and the preset matching threshold includes:

[0125] Dividing the currently retrieved mapping relationship based on the similarity value and the preset matching threshold to obtain a first mapping relationship and a second mapping relationship; when the number of the first mapping relationship is greater than zero, setting the first mapping relationship as the target mapping relationship; when the number of the first mapping relationship is zero, screening the second mapping relationship to obtain the screened mapping relationship and setting the screened mapping relationship as the target mapping relationship.

[0126] Exemplarily, when determining the target mapping relationship, it is necessary to compare the calculated similarity value with the preset matching threshold. The specific steps are as follows:

[0127] After calculating the similarity values of all candidate mapping relationships, find the largest similarity value among them. If this largest similarity value is greater than or equal to the preset matching threshold, then the mapping relationship corresponding to this similarity value can be set as the target mapping relationship; otherwise, several relatively high candidate mapping relationships need to be selected from all the calculated similarity values.

[0128] In an exemplary embodiment, after performing entity recognition on the current request information according to the preset entity recognition model to obtain the current entity recognition result, the method further includes:

[0129] When the current entity recognition result indicates that the current request entity does not exist, determining the previous request information and the previous retrieved mapping relationship based on the historical query database; performing request analysis on the current request information based on the previous request information and the preset interval threshold to obtain a request analysis result; when the request analysis result does not exceed the preset interval threshold, performing matching analysis on the previous retrieved mapping relationship according to the current request information, the preset semantic matching model, and the preset matching threshold to obtain the current target data.

[0130] Among them, the historical query database stores data such as the request information, the retrieved mapping relationship, and the request time of the historical data requests.

[0131] Exemplarily, when the current entity recognition result indicates the non-existence of the current requested entity, the previous request information and the mapping relationship after the previous retrieval are searched for in the historical query database, and the previous request information includes the user identifier. Then, the user identifier in the previous request information is compared with the user identifier in the current request information to determine whether they are the same user. At the same time, the time interval between the two requests is compared with the preset interval threshold to determine whether the continuous request time has been exceeded.

[0132] If the same user continuously makes data requests within the preset interval threshold, the preset semantic matching model is used to perform the matching analysis operation as shown in Figure 4 to obtain the current target data.

[0133] In an exemplary embodiment, request analysis is performed on the current request information based on the previous request information and the preset interval threshold to obtain a request analysis result, including:

[0134] Data analysis is performed based on the previous request information and the current request information to obtain a data analysis result and a request interval; when the data analysis result is the same user identifier, the request interval is subjected to request analysis with the preset interval threshold to obtain a request analysis result.

[0135] Exemplarily, the time interval between the two requests is calculated according to the request time of the previous request information and the request time of the current request information, and it is compared with the preset interval threshold to determine whether the continuous request time has been exceeded.

[0136] In an exemplary embodiment, a data processing method based on a knowledge graph is provided, and the method includes the following steps:

[0137] Obtain the current data request and determine the current request information based on the current data request.

[0138] Perform entity recognition on the current request information according to the preset entity recognition model to obtain the current entity recognition result.

[0139] When the current entity recognition result indicates the existence of the current requested entity, based on the preset retrieval lower limit threshold and the preset entity, the current requested entity is retrieved and compared to obtain a first type of entity and a second type of entity.

[0140] When the number of the first type of entity is zero, obtain the next data request.

[0141] When the number of the first type of entity is greater than zero, divide the first type of entity based on the preset retrieval upper limit threshold to obtain a third type of entity and a fourth type of entity.

[0142] When the number of entities of the third type is greater than zero, perform mapping retrieval based on the entities of the third type and a preset knowledge graph to obtain the mapping relationship after the current retrieval.

[0143] When the number of entities of the third type is zero, screen the entities of the fourth type to obtain the screened entities.

[0144] Perform mapping retrieval based on the screened entities and a preset knowledge graph to obtain the mapping relationship after the current retrieval.

[0145] Perform data splicing on the current request information and the mapping relationship after the current retrieval to obtain the spliced data;

[0146] Calculate the spliced data based on a preset similarity algorithm to obtain a similarity value.

[0147] Divide the mapping relationship after the current retrieval based on the similarity value and a preset matching threshold to obtain a first mapping relationship and a second mapping relationship.

[0148] When the number of the first mapping relationship is greater than zero, set the first mapping relationship as the target mapping relationship.

[0149] When the number of the first mapping relationship is zero, screen the second mapping relationship to obtain the screened mapping relationship and set the screened mapping relationship as the target mapping relationship.

[0150] Determine the current target data based on the target mapping relationship and the current request information.

[0151] When the current entity recognition result is that the current request entity does not exist, determine the previous request information and the mapping relationship after the previous retrieval based on the historical query database.

[0152] Perform data analysis based on the previous request information and the current request information to obtain a data analysis result and a request interval.

[0153] When the data analysis result is the same user identifier, perform request analysis on the request interval and a preset interval threshold to obtain a request analysis result.

[0154] When the request analysis result is that the preset interval threshold is not exceeded, perform matching analysis on the mapping relationship after the previous retrieval according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain the current target data.

[0155] The flowchart of the above steps is as Figure 6 shown.

[0156] In an exemplary embodiment, based on Figure 6 the flowchart shown and taking the request "What is the country of origin of F16?" as an example, and the knowledge graph is as Figure 5as shown

[0157] See Figure 2 is a schematic structural diagram of a preset entity recognition model. Based on the trained named entity recognition model, entities in the question are recognized and obtained, and its input form is: [CLS user question SEP]. The entity recognition result in the question is: "F16".

[0158] Based on the question entities recognized by the named entity recognition model, through the constructed entity Elasticsearch database, the full name of the corresponding entity in the graph is obtained. If the Es retrieval threshold is greater than the preset threshold, the corresponding entity full name is directly returned; otherwise, the top k entities are recommended for the user to select.

[0159] Since there are multiple similar entities in the Es database and the Es retrieval result of the entity "F16" in the question is less than the set threshold, the entities "F-16 "Fighting Falcon" fighter" and "F-16V fighter" are recommended for the user to select. Suppose the user selects "F-16 "Fighting Falcon" fighter".

[0160] According to the question entity, accurately obtain the candidate relationships of this entity:

[0161] The associated relationships of "F-16 "Fighting Falcon" fighter" are "maximum speed", "flight speed", "country of origin", "length of the aircraft", "height of the aircraft", and "maximum takeoff weight".

[0162] See Figure 3 is a schematic structural diagram of a preset semantic matching model. Based on the above-obtained candidate relationships, the question and the candidate relationships are concatenated and input into the semantic similarity matching model, and its input form is:

[0163] [CLSWhat is the country of origin of F16?SEPMaximum speedSEPFlight speedSEPCountry of originSEPLength of the aircraftSEPHeight of the aircraftSEPMaximum takeoff weightSEP]

[0164] If the probability of the model output result is greater than the preset threshold, it is determined that this candidate relationship and the question are the most matched. Since the relationship mentioned in the question does not exist in the candidate relationship, the entity relationships ["maximum speed", "flight speed", "country of origin", "length of the aircraft", "height of the aircraft", "maximum takeoff weight"] are recommended for the user to select.

[0165] If the user selects "country of origin", then according to F-16 "Fighting Falcon" fighter -> country of origin -> answer: United States, the answer to the question can be obtained as the United States

[0166] If later, take "What is its takeoff weight?" as an example.

[0167] See Figure 2It is a schematic structural diagram of a preset entity recognition model. Based on the trained named entity recognition model, entities in the question are recognized and obtained, and its input form is: [CLS User Question SEP]. If the entity recognition result in the question is "empty", it is determined that the question is a continuous question of the user.

[0168] It is determined that the question meets the continuous question setting rule. Based on the historical database Mongodb, the previous question entity name and entity relationship are obtained as: {Entity name: "F-16 Fighting Falcon" fighter; Entity relationship: ["Maximum speed", "Flight speed", "Country of origin", "Length of the aircraft", "Height of the aircraft", "Maximum takeoff weight"]}

[0169] See Figure 3 It is a schematic structural diagram of a preset semantic matching model. For the candidate relationships obtained above, the question and the candidate relationships are concatenated and input into the semantic similarity matching model, and its input form is:

[0170] [CLS What is its takeoff weight? SEP Maximum speed SEP Flight speed SEP Country of origin SEP Length of the aircraft SEP Height of the aircraft SEP Maximum takeoff weight SEP]

[0171] The model output result is: "Maximum takeoff weight"

[0172] According to F-16 Fighting Falcon -> Maximum takeoff weight -> answer: 19,190 kg, the answer to the question can be obtained as 19,190 kg.

[0173] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0174] Based on the same inventive concept, the embodiments of the present application also provide a knowledge graph-based data processing device for implementing the above-mentioned knowledge graph-based data processing method. The implementation solutions provided by this device for solving problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the following knowledge graph-based data processing device can refer to the limitations on the knowledge graph-based data processing method in the above text, and will not be repeated here.

[0175] In one embodiment, as Figure 7 shown, a data processing apparatus based on a knowledge graph is provided, including: an acquisition module 702, an identification module 704, a retrieval module 706, and a matching module 708, where:

[0176] The acquisition module 702 is configured to acquire a current data request and determine current request information based on the current data request;

[0177] The identification module 704 is configured to perform entity identification on the current request information according to a preset entity identification model to obtain a current entity identification result;

[0178] The retrieval module 706 is configured to, when the current entity identification result indicates the existence of a current request entity, perform entity retrieval on the current request entity based on a preset database and a preset retrieval threshold to obtain a current retrieved mapping relationship;

[0179] The matching module 708 is configured to perform matching analysis on the current retrieved mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain current target data.

[0180] In an exemplary embodiment, the retrieval module 706 is further configured to perform retrieval comparison on the current request entity based on a preset retrieval lower limit threshold and a preset entity to obtain a first type of entity and a second type of entity; when the number of the first type of entities is greater than zero, perform retrieval based on a preset graph and the first type of entity to obtain a current retrieved mapping relationship.

[0181] In an exemplary embodiment, the retrieval module 706 is further configured to divide the first type of entity based on a preset retrieval upper limit threshold to obtain a third type of entity and a fourth type of entity; when the number of the third type of entities is greater than zero, perform mapping retrieval based on the third type of entity and the preset graph to obtain a current retrieved mapping relationship; when the number of the third type of entities is zero, screen the fourth type of entity to obtain a screened entity; perform mapping retrieval based on the screened entity and the preset graph to obtain a current retrieved mapping relationship.

[0182] In an exemplary embodiment, the acquisition module 702 is further configured to acquire a next data request when the number of the first type of entities is zero.

[0183] In an exemplary embodiment, the matching module 708 is further configured to splice the current request information and the current retrieved mapping relationship to obtain spliced data; calculate the spliced data based on the preset similarity algorithm to obtain a similarity value; determine a target mapping relationship based on the similarity value and a preset matching threshold; and determine current target data based on the target mapping relationship and the current request information.

[0184] In an exemplary embodiment, the matching module 708 is further configured to divide the current retrieved mapping relationship based on the similarity value and a preset matching threshold to obtain a first mapping relationship and a second mapping relationship; in a case where the number of the first mapping relationships is greater than zero, set the first mapping relationship as the target mapping relationship; and in a case where the number of the first mapping relationships is zero, screen the second mapping relationship to obtain a screened mapping relationship and set the screened mapping relationship as the target mapping relationship.

[0185] In an exemplary embodiment, a data processing device based on a knowledge graph is configured to, in a case where the current entity recognition result indicates that there is no current request entity, determine a previous request information and a previous retrieved mapping relationship based on a historical query database; perform request analysis on the current request information based on the previous request information and a preset interval threshold to obtain a request analysis result; and in a case where the request analysis result does not exceed the preset interval threshold, perform matching analysis on the previous retrieved mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain current target data.

[0186] In an exemplary embodiment, a data processing device based on a knowledge graph is configured to perform data analysis on the previous request information and the current request information to obtain a data analysis result and a request interval; and in a case where the data analysis result is the same user identifier, perform request analysis on the request interval and a preset interval threshold to obtain a request analysis result.

[0187] Each module in the above data processing device based on a knowledge graph can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute operations corresponding to the above respective modules.

[0188] In an embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store request information data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a data processing method based on a knowledge graph.

[0189] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0190] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0191] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0192] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0193] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0194] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0195] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0196] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A data processing method based on knowledge graph, characterized in that: The method comprises: Obtaining a current data request and determining current request information based on the current data request; Perform entity recognition on the current request information according to a preset entity recognition model to obtain a current entity recognition result; When the current entity recognition result indicates that the current request entity exists, performing entity search on the current request entity based on a preset database and a preset search threshold to obtain a current post-search mapping relationship; The current post-retrieval mapping relationship is matched and analyzed according to the current request information, a preset semantic matching model and a preset matching threshold to obtain current target data.

2. The data processing method based on knowledge graph according to claim 1 is characterized in that: The preset database includes preset entities and preset graphs; the preset search threshold includes a preset search lower limit threshold; when the current entity recognition result is that the current request entity exists, the current request entity is searched for based on the preset database and the preset search threshold to obtain the current post-search mapping relationship, including: Based on a preset search lower limit threshold and a preset entity, the current request entity is searched and compared to obtain a first category entity and a second category entity; When the number of the first-category entities is greater than zero, a search is performed based on the preset graph and the first-category entities to obtain the current post-search mapping relationship.

3. The data processing method based on knowledge graph according to claim 2 is characterized in that: The preset search threshold also includes a preset search upper threshold; when the number of the first type of entities is greater than zero, searching based on the preset graph and the first type of entities to obtain the current post-search mapping relationship includes: Dividing the first category of entities based on a preset retrieval upper threshold to obtain third category entities and fourth category entities; When the number of the third-category entities is greater than zero, a mapping search is performed based on the third-category entities and the preset graph to obtain a current post-search mapping relationship; When the number of entities of the third category is zero, screening the entities of the fourth category to obtain screened entities; A mapping search is performed based on the screened entities and the preset graph to obtain a current post-search mapping relationship.

4. The data processing method based on knowledge graph according to claim 2 is characterized in that: After performing a search and comparison on the current request entity based on the preset search lower limit threshold and the preset entity to obtain the first category entity and the second category entity, the method further includes: When the number of entities of the first category is zero, a next data request is obtained.

5. The data processing method based on knowledge graph according to claim 1 is characterized in that: The preset semantic matching model includes a preset similarity algorithm; Performing a matching analysis on the current post-retrieval mapping relationship according to the current request information, a preset semantic matching model, and a preset matching threshold to obtain current target data includes: Performing data splicing on the current request information and the current post-retrieval mapping relationship to obtain spliced ​​data; Calculating the spliced ​​data based on the preset similarity algorithm to obtain a similarity value; Determine a target mapping relationship based on the similarity value and a preset matching threshold; The current target data is determined based on the target mapping relationship and the current request information.

6. The data processing method based on knowledge graph according to claim 5 is characterized in that: The determining the target mapping relationship based on the similarity value and a preset matching threshold comprises: Dividing the current post-retrieval mapping relationship based on the similarity value and a preset matching threshold to obtain a first mapping relationship and a second mapping relationship; When the number of the first mapping relationships is greater than zero, setting the first mapping relationship as a target mapping relationship; When the number of the first mapping relationships is zero, the second mapping relationships are screened to obtain screened mapping relationships and the screened mapping relationships are set as target mapping relationships.

7. The data processing method based on knowledge graph according to claim 1 is characterized in that: After performing entity recognition on the current request information according to the preset entity recognition model to obtain a current entity recognition result, the method further includes: When the current entity recognition result is that the current requested entity does not exist, determining the last requested information and the last post-retrieval mapping relationship based on the historical query database; Performing request analysis on the current request information based on the previous request information and a preset interval threshold to obtain a request analysis result; When the request analysis result does not exceed the preset interval threshold, a matching analysis is performed on the last searched mapping relationship according to the current request information, the preset semantic matching model and the preset matching threshold to obtain the current target data.

8. The data processing method based on knowledge graph according to claim 7 is characterized in that: The performing request analysis on the current request information based on the previous request information and the preset interval threshold to obtain a request analysis result includes: Perform data analysis based on the previous request information and the current request information to obtain data analysis results and request intervals; In the case that the data analysis result is the same user identification, the request interval is subjected to request analysis with a preset interval threshold to obtain a request analysis result.

9. A data processing device based on knowledge graph, characterized in that: The device comprises: An acquisition module, used to acquire a current data request and determine current request information based on the current data request; An identification module is used to perform entity identification on the current request information according to a preset entity identification model to obtain a current entity identification result; A retrieval module, configured to, when the current entity recognition result indicates that the current request entity exists, perform entity retrieval on the current request entity based on a preset database and a preset retrieval threshold to obtain a current post-retrieval mapping relationship; The matching module is used to perform matching analysis on the current post-retrieval mapping relationship according to the current request information, a preset semantic matching model and a preset matching threshold to obtain the current target data.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.