Request response method and device, electronic equipment and storage medium

By finding named entities that match the pending request or node describing named entities in the node relationship diagram, and using the large language model to generate response information, the analysis perspective problem caused by structured and unstructured data separation processing is solved, and a more accurate user request response is achieved.

CN120407755AActive Publication Date: 2025-08-01BEIJING DIPEAK TECHNOLOGY CO LTD
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510919706.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

In the prior art, the separation of structured data and unstructured data processing and storage leads to a one-sided analysis perspective, making it difficult to achieve efficient information integration and utilization, resulting in inaccurate responses to user requests.

Method used

By obtaining a node relationship diagram containing multiple nodes representing named entities, finding a named entity that matches the pending request or a node describing the named entity, and sending it to a large language model for processing, generating response information.

Benefits of technology

Improve the accuracy of response to user requests and improve the user's user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407755A_ABST
    Figure CN120407755A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a request response method and device, electronic equipment and a storage medium, and relates to the technical field of computers. The method comprises the following steps: acquiring a to-be-processed request and a first node relation graph; for each first type node, determining a first matching degree between the first type node and each other node based on the to-be-processed request; and for each first matching degree greater than a first preset threshold, taking the named entity represented by the first type node corresponding to the first matching degree and the text content represented by the corresponding other nodes as prompt information corresponding to the to-be-processed request, and obtaining response information corresponding to the to-be-processed request based on the prompt information through a large language model. Through the method, the user request can be answered from the perspective of multiple named entities when the user request is answered, the answering accuracy of the user request can be effectively improved, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology. Specifically, this application relates to a request-response method, apparatus, electronic device, and storage medium. Background Art

[0002] In actual scenarios, a user's request often involves multiple requirements. There may be sub-questions related to document content and sub-questions related to numerical calculations, or it may be necessary to infer the dimensions and metrics involved in numerical calculations from document content. Under the traditional data management mode, structured data is efficiently processed through relational databases. However, unstructured data has long faced problems such as scattered storage, difficult retrieval, and information silos due to its complex format and lack of a unified model. That is to say, structured data and unstructured data are processed and stored separately, resulting in a one-sided analysis perspective and difficulty in achieving efficient information integration and utilization, thereby leading to inaccurate responses to user requests. Summary of the Invention

[0003] The purpose of this application aims to solve at least one of the above technical defects. The technical solutions provided by the embodiments of this application are as follows: In a first aspect, an embodiment of this application provides a request-response method, including: Obtain a request to be processed and a first node relationship graph. The first node relationship graph includes multiple nodes and multiple directed edges. Each node is a first type of node or a second type of node. The first type of node represents a named entity, and the second type of node represents the text content describing the named entity. Among them, a directed edge pointing from one node to another node indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to one node; For each first type of node, determine a first matching degree between the first type of node and other nodes in the first node relationship graph except the first type of node based on the request to be processed; where the first matching degree represents the matching degree between the named entity represented by the first type of node and the request to be processed under the condition of the text content represented by other nodes; For each first matching degree greater than a first preset threshold, use the named entity represented by the first type of node corresponding to the first matching degree and the text content represented by the corresponding other nodes as the prompt information corresponding to the request to be processed, and obtain the response information corresponding to the request to be processed through a large language model based on the prompt information.

[0004] In a second aspect, an embodiment of this application provides a request-response apparatus, including: A request acquisition module is used to acquire a request to be processed and a first node relationship graph. The first node relationship graph includes multiple nodes and multiple directed edges. Each node is a first type of node or a second type of node. The first type of node represents a named entity, and the second type of node represents the text content describing the named entity. Among them, a directed edge pointing from one node to another indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to one node; A node matching module is used to, for each first type of node, determine the first matching degree between the first type of node and other nodes in the first node relationship graph except the first type of node based on the request to be processed. Among them, the first matching degree represents the matching degree between the named entity represented by the first type of node and the request to be processed under the condition of the text content represented by other nodes; A request response module is used to, for each first matching degree greater than the first preset threshold, use the named entity represented by the first type of node corresponding to the first matching degree and the text content represented by the corresponding other nodes as the prompt information corresponding to the request to be processed, and obtain the response information corresponding to the request to be processed through a large language model based on the prompt information.

[0005] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory; The processor executes the computer program to implement the method provided in the embodiment of the first aspect or any optional embodiment of the first aspect.

[0006] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method provided in the embodiment of the first aspect or any optional embodiment of the first aspect.

[0007] The beneficial effects brought by the technical solution provided by the embodiment of the present application are: The solution provided by the embodiment of the present application obtains a node relationship graph including multiple nodes representing named entities and representing the attribution relationship between each named entity, and finds nodes representing named entities or nodes describing named entities that match the request to be processed from the node relationship graph, and then sends the named entity represented by the matched node or the text content describing the named entity to the large language model for processing to obtain the response information, so that when answering the user's request, it can answer from the perspective of multiple named entities, which can effectively improve the accuracy of answering the user's request and enhance the user's usage experience. Description of the Drawings

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0009] Figure 1 It is a schematic flow chart of a request - response method provided by an embodiment of the present application; Figure 2 It is a schematic flow chart of integrating data to be integrated in an example of an embodiment of the present application; Figure 3 It is a schematic flow chart of extracting a first node relationship graph from preset relationship nodes and generating response information in an example of an embodiment of the present application; Figure 4 It is a structural block diagram of a request - response device provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0010] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0011] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations, etc. supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0012] To make the purpose, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0013] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps, etc. in different embodiments, they will not be described repeatedly.

[0014] Figure 1 The following is a schematic flowchart of a request-response method provided by an embodiment of the present application. The execution subject of this method may be a terminal (such as a computer, a mobile phone, etc.). As Figure 1 shown, this method may include: Step S101: Obtain a request to be processed and a first node relationship graph. The first node relationship graph includes multiple nodes and multiple directed edges. Each node is a first type of node or a second type of node. The first type of node represents a named entity, and the second type of node represents the text content describing the named entity. Among them, a directed edge pointing from one node to another node indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to one node.

[0015] In the embodiments of the present application, the request to be processed may be sent by a user. Generally speaking, it may be in the form of natural language. The request may be a query request, such as "Please help me query the revenue of XX Company last year", etc., or it may be a question raised, such as "May I ask where XX Company is?", etc. The embodiments of the present application do not make any limitations here. The first node relationship graph may be a directed graph associated with the request to be processed. Each node in the graph represents a named entity that may be related to the field to which the semantics of the request to be processed belongs. For example, when the question raised by the user is "Please help me query the revenue of XX Company last year", the semantics to which the semantics of this question belongs may be the financial field. Therefore, each node included in the obtained first node relationship graph is also related to the financial field. The nodes in the present application may be divided into a first type of node and a second type of node. Among them, the first type of node may generally be a word of a named entity, such as "Shanghai", etc., and the second type of node may be a text describing the named entity, such as "Shanghai is a municipality directly under the Central Government", etc. The embodiments of the present application do not make any limitations here.

[0016] Specifically, after the terminal obtains the request to be processed sent by the user, it may perform semantic analysis according to the request to be processed, and then determine the first node relationship graph related to the semantics of the request to be processed according to the semantic analysis result.

[0017] Step S102: For each first type of node, determine the first matching degree between the first type of node and other nodes in the first node relationship graph except the first type of node based on the request to be processed; among them, the first matching degree represents the matching degree between the named entity represented by the first type of node and the request to be processed under the condition of the text content represented by other nodes.

[0018] In the embodiments of the present application, the first matching degree may be obtained by calculating the cosine similarity of vectors after converting the text content represented by each node and the text content of the request to be processed into corresponding vector forms.

[0019] Specifically, after obtaining the first node relationship graph, since each node in the first node relationship graph is related to the field to which the request to be processed belongs, it is necessary to further screen out nodes from the first node relationship graph according to the semantics of the request to be processed. The screening method can be to screen according to the semantics of the text content of the request to be processed, and the semantics can be represented by converting it into a corresponding vector. The semantics of each node in each first node relationship graph can also be represented by converting it into a corresponding vector. Then, the cosine similarity can be calculated based on the vectors converted from each text content to determine the relevance between each semantics, thereby completing the screening of nodes.

[0020] Specifically, after obtaining the first node relationship graph, the text content of the request to be processed and the named entities represented by each first type node in the first node relationship graph can be first converted into vector forms, and then the ComplEx (Complex Embedding) model is used to represent each vector in the form of complex numbers. The specific form can be as follows:

[0021] Among them, S represents the complex vector form of the named entity of the first type node, R represents the complex vector form of the text content of the request to be processed, s represents the original vector form of the named entity of the first type node, r represents the original vector form of the text content of the request to be processed, Re() represents the real part, and Im() represents the imaginary part. For example, for the vector [1, 2], the converted complex form vector can be .

[0022] After that, the text content represented by each other node can also be converted into the form of complex vectors in the same way as above, and then the first matching degree between each first type node and each other node is determined through the following formula:

[0023] Among them, f(S, R, E) is the first matching degree, E is the complex vector form of the other node, is the conjugate complex vector of E, represents element-wise multiplication, and Re() represents taking the real part to ensure that the first matching degree is a real number. represents the k-th dimension of the real part of the complex form vector S (for example, for the complex form vector , it can also be regarded as , then the first dimension is , and the second dimension is ), represents the k-th dimension of the real part of the complex form vector R, represents the k-th dimension of the real part of the complex form vector E; Denotes the k-th dimension of the imaginary part of the complex vector S, Denotes the k-th dimension of the imaginary part of the complex vector R, Denotes the k-th dimension of the real part of the complex vector E.

[0024] By using the above formula, the matching degree is calculated for each first-type node and each other node respectively, and the first matching degree between each first-type node and each other node can be obtained.

[0025] Optionally, the process of calculating the first matching degree can also be completed by a trained model. For example, the text content of various requests to be processed that the user may give and the named entities that may appear can be used as the training samples of the initial model in advance, and the text content that may describe the named entities can be used as the training labels to train the initial model until the training loss of the model is less than the threshold, and then the training ends to obtain the trained model. The model loss function of the model can be as follows:

[0026] Where, is the square of the Euclidean norm of the error between the model prediction value and the true value, that is, the sum of the squares of the components of the vector. γ is the margin hyperparameter (usually taken as 1-2), and λ is the L2 regularization coefficient. Through optimization, the model learns a complex vector representation that can reflect complex semantics and capture anti-symmetric relationships such as "father-son" and "son-father".

[0027] Step S103, for each first matching degree greater than the first preset threshold, the named entity represented by the first-type node corresponding to the first matching degree and the text content represented by the corresponding other node are used as the prompt information corresponding to the request to be processed, and the response information corresponding to the request to be processed is obtained through the large language model based on the prompt information.

[0028] In an embodiment of the present application, the first preset threshold can be freely set. A first matching degree exceeding this threshold indicates that, under the condition of the text content represented by the corresponding other nodes, the named entity represented by the corresponding first-type node matches the to-be-processed request. The prompt information can include multiple named entities or multiple segments of text content for describing named entities. Specifically, the first-type node corresponding to any first matching degree greater than the first preset threshold can be used as the starting node, and the corresponding other nodes can be used as the ending nodes. Then, according to the attribution relationship represented by the directed edges in the first node relationship graph, the text content they represent is integrated to obtain the prompt information; alternatively, the named entity or the text content describing the named entity can be directly used as the prompt information and input into the large language model respectively. The large language model will integrate the input named entities or text content by itself and output the corresponding response information according to the integration result. The embodiments of the present application do not make limitations here.

[0029] It should be noted that when the to-be-processed request is a query request, the large language model in the embodiments of the present application can also first generate a corresponding SQL (Structured Query Language) query statement according to the prompt information, query relevant results from the preset database through this SQL statement, and then integrate the results into a natural language query result to output as the response information.

[0030] The solution provided by the embodiments of the present application is to obtain a node relationship graph including multiple nodes representing named entities and nodes representing the attribution relationship between each named entity, find the nodes of the named entity or the nodes of the text content describing the named entity that match the to-be-processed request from the node relationship graph, and then send the named entity or the text content describing the named entity represented by the matched nodes to the large language model for processing to obtain the response information. This enables answering the user's request from the perspective of multiple named entities, effectively improving the accuracy of answering the user's request and enhancing the user's experience.

[0031] Based on the above various embodiments, as an alternative embodiment, the first node relationship graph is obtained in the following manner: Obtain a preset node relationship graph, and obtain at least one candidate node and at least one candidate directed edge from the preset node relationship graph, where the semantics of the text content represented by the candidate node match the semantics of each named entity included in the to-be-processed request; the preset node relationship graph includes multiple nodes and multiple directed edges. Each node includes a first-type node and a second-type node. The first-type node represents a named entity, and the second-type node represents the text content describing the named entity. Among them, a directed edge pointing from one node to another node indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to one node; Construct an initial node relationship graph based on at least one candidate node and at least one candidate directed edge; Perform a traversal operation on each candidate node. If a corresponding first node is determined through the traversal operation, perform the deletion operation corresponding to the candidate node; Take the initial node relationship graph after performing the deletion operations corresponding to each candidate node as the first node relationship graph; Among them, for each candidate node, the traversal operation on the candidate node includes: Determine the second node currently being traversed, and the first node traversed is the candidate node; Take the node in the preset node relationship graph that is connected to the second node and pointed to by a directed edge as the associated node corresponding to the current traversal operation; Determine the first similarity between the semantics of the attribution relationship represented by the directed edge between the second node and the associated node and the semantics represented by the request to be processed. If the first similarity is greater than the second preset threshold, take the associated node as the second node for the next traversal; If the first similarity is not greater than the second preset threshold, take the associated node as one of the first nodes corresponding to the candidate node; The deletion operation includes: Delete each first node and the directed edges starting from each first node from the preset node relationship graph.

[0032] In the embodiments of the present application, the preset node relationship graph is pre-constructed and stored in a graph database (such as a neo4j database). The first node relationship graph in the embodiments of the present application can be regarded as a sub-graph (i.e., a part of the preset node relationship graph) in the preset node relationship graph. The candidate node can be a node selected from the preset node relationship graph for constructing the first node relationship graph. The associated node can be a node directly connected to the candidate node by a directed edge.

[0033] Specifically, the preset node relationship graph in the present application stores nodes representing various named entities from different fields and the text content describing the named entities. If directly looking for nodes related to the request to be processed from the preset node relationship graph containing various fields, it may affect the accuracy of the response information due to the mutual interference between different fields. Therefore, it is necessary to first perform a preliminary screening on each node in the preset node relationship graph, and then determine the response information after forming the first node relationship graph from the selected nodes.

[0034] Specifically, the embodiments of the present application provide an adaptive pruning multi-level sub-graph search method, such as Figure 3As shown, in this method, first, multiple candidate nodes and multiple matching candidate directed edges that match the semantics of the request to be processed are determined based on the semantics of each node. Then, an initial node relationship graph is constructed based on the multiple candidate nodes and multiple directed edges obtained by the matching. Next, traversal operations are performed on each candidate node in the initial node relationship graph. The method will be described by way of example below. For example, for candidate node A1, the nodes directly connected to A1 in the preset node relationship graph are B1, B2, and B3. Then, B1, B2, and B3 are respectively determined as associated nodes. Then, the first similarity between B1, B2, and B3 and A1 is respectively determined in combination with the semantics of the attribution relationship between the nodes. At this time, the first similarity of B1 is obtained to be greater than the second preset threshold, while the first similarities of B2 and B3 are less than the second preset threshold. Then, B2 and B3 can be determined as the first nodes, while B1 is determined as the second node. Then, starting from B1, other nodes directly connected to B1 in the preset node relationship graph are determined, and so on, until when starting from a certain node, all the directly connected nodes determined are the first nodes or there are no other directly connected nodes.

[0035] After the traversal of all candidate nodes is completed, all the first nodes and the directed edges starting from each first node are deleted from the preset node relationship graph, and the other nodes pointed to by the directed edges starting from the deleted nodes can also be deleted (for example, for the first node B2, there are the following relationships B2→C1→D1 and B2→C2→D2 in the preset node relationship graph. Then, the nodes C1, C2, D1, D2 and the directed edges B2→C1, C1→D1, B2→C2, and C2→D2 are also deleted), so as to obtain the first node relationship graph.

[0036] Based on the above various embodiments, as an optional embodiment, determining the first similarity between the semantics of the attribution relationship represented by the directed edge between the second node and the associated node and the semantics represented by the request to be processed specifically includes: Obtain the first embedding vector corresponding to the text content of the request to be processed, and obtain the second embedding vector corresponding to the text content of the attribution relationship, and determine the second similarity between the first embedding vector and the second embedding vector; Determine the first number of identical characters that appear in the text content of the request to be processed and the total number of characters, and determine the proportion of the first number of identical characters in the total number of characters; the total number of characters is the sum of the number of characters in the text content represented by the second node, the number of characters in the text content represented by the associated node, and the number of characters in the text content of the attribution relationship; Perform a weighted sum of the second similarity and the proportion to obtain the first similarity.

[0037] In the embodiments of the present application, the embedding vectors corresponding to each text content can be obtained by inputting the text content into a pre-trained text vector conversion model (such as the bge-m3 vector model) and then output by the model.

[0038] Specifically, since the similarity between two text contents is affected by multiple different aspects, such as the number of words, semantics, etc., in order to improve the accuracy of the traversal process, the embodiments of the present application will consider the similarity between the text content of the attribution relationship and the text content of the pending request from multiple aspects respectively. Specifically, the embodiments of the present application provide the following formula to calculate the first similarity:

[0039] Wherein, E query is the first embedding vector, E path is the second embedding vector, cos(E query ,E path ) is the second similarity, overlap(query,path) is the number of first identical characters, len(path) is the total number of characters (for example, if the pending request is "What honors has the Shanghai Branch obtained", the text content represented by the second node, the text content represented by the associated node, and the text content of the attribution relationship is "XX Company Headquarters - Won - Technology Pioneer", then the repeated characters are "obtained" and "won", so overlap(query,path) = 2, indicating that there are 2 overlapping characters. At this time, the total number of characters len(path) is 10), β is the harmonic coefficient between the two, which is used to weight the second similarity and the ratio, and can be freely set according to actual needs.

[0040] Through the above formula, the similarity between the text content of the attribution relationship and the text content of the pending request can be considered from multiple aspects, thereby improving the accuracy of the traversal process.

[0041] Based on the above various embodiments, as an optional embodiment, the preset node relationship graph is obtained in the following manner: Obtain multiple pieces of data to be integrated; wherein, the data to be integrated is tabular data or non-tabular data; For each piece of data to be integrated, if the data to be integrated is tabular data, then each cell in the tabular data representing the first text content is respectively determined as a named entity, and an attribution relationship is established between the named entity represented by the cell at the table header in each column of the table and the named entity represented by the cell not at the table header in the corresponding column; For each piece of data to be integrated, if the data to be integrated is non-table type data, obtain the second text content corresponding to the non-table type data, split the second text content into multiple text blocks containing complete semantics, and extract the named entities contained in each text block and the attribution relationships between the named entities through a large language model; Take each named entity and each text block as a node respectively, and take the attribution relationships between each named entity and the attribution relationships between each text block and the named entity as directed edges between the corresponding two nodes to construct a second node relationship graph; Perform a deduplication operation on each node in the second node relationship graph to obtain a preset node relationship graph.

[0042] In the embodiments of the present application, the non-table type data can be picture type data, text type data, etc. For non-text type data in the non-table type data, these data can be first converted into corresponding text type data. For example, for picture type data, technologies such as layout analysis and ocr (Optical Character Recognition) can be used to convert it into text form. The embodiments of the present application do not make limitations here.

[0043] Specifically, in the embodiments of the present application, since the preset node relationship graph needs to contain data materials in various fields, it is necessary to first collect the data materials in various fields. The collected data materials can be table type data or non-table type data. Usually, the table type data and the non-table type data cannot be directly integrated. Therefore, in the embodiments of the present application, these data can be processed separately and then the processed data can be integrated. Specifically, as Figure 2 shown, for table type data, the first text content in each cell in the table can be used as a named entity. At the same time, there is an attribution relationship between the named entity of the first text content in the header cell in this column and the named entities of the first text content in other cells. This relationship can be represented by adding a directed edge in the preset relationship node graph. For example, as shown in Table 1 below:

[0044] Table 1 In Table 1, the cells in the table have "City", "Branch Name", and "Interest Income". Then, "City", "Branch Name", "Interest Income", "Shanghai", "Beijing", "Guangzhou", "Shanghai Branch", "Beijing Branch", "Guangzhou Branch", "1 million", "1.5 million", and "800,000" will be determined as the first text contents to be determined respectively. And in this table, there is an attribution relationship between "City" and "Shanghai", "Beijing", "Guangzhou" respectively (that is, "City" includes "Shanghai", "Beijing", "Guangzhou"), an attribution relationship between "Branch Name" and "Shanghai Branch", "Beijing Branch", "Guangzhou Branch" respectively, and an attribution relationship between "Interest Income" and "1 million", "1.5 million", "800,000" respectively. These attribution relationships can be represented by directed edges in the preset node relationship diagram. For example, add a directed edge between the node of "City" and the node of "Shanghai" and the edge points to "City", and the directed edge contains the text content of "attribution" to indicate that "Shanghai" belongs to a "City".

[0045] It should be noted that the named entities in the cells at the head of each column in the table can also be called "dimensions", and the named entities in the cells of each column other than the column head can also be called "code values". In some application scenarios, some data do not appear in the table but outside the table. These data can be called "indicators". "Indicators" can be calculated from the data in the table and the relationships between the data, or there can be an attribution relationship with "dimensions" or "code values". And like "dimensions" and "code values", "indicators" can also be used as a named entity and a separate node can be set in the preset node relationship diagram.

[0046] For non-table type data, such as Figure 2 shown, first, these data need to be uniformly converted into text form (if it is text type data, there is no need to convert), to obtain each second text content, then each second text content is split into multiple text blocks containing complete semantics, and then each text block is respectively input into the large language model, and the named entities contained in each text block and the attribution relationships of each named entity are extracted by the way of knowledge extraction through the large language model. For example, for a text block "Shanghai and Beijing are both cities", three named entities "Shanghai", "Beijing", and "city" can be extracted by the way of knowledge extraction through the large language model, and the attribution relationship that both "Shanghai" and "Beijing" belong to "city" can be extracted.

[0047] It should be noted that in the embodiments of the present application, in order to make the preset node relationship diagram contain more detailed data information, the text block can also be used as one of the nodes, and the named entities contained in the text block will also establish corresponding attribution relationships with the text block.

[0048] After extracting each named entity and the attribution relationships between each named entity, each named entity can be used as a node respectively, and each attribution relationship can be used as a directed edge respectively to construct a second node relationship graph. It can be understood that since the sources of the data to be integrated are different, there will inevitably be the same or similar named entities among the extracted named entities. Therefore, it is necessary to further perform a deduplication operation on these named entities. After the deduplication operation is completed, a preset node relationship graph can be obtained.

[0049] Optionally, as Figure 2 shown, each text block, each named entity, and each attribution relationship can be input into a pre-trained bge-m3 (Beijing General Embedding M3) vectorization model to be converted into corresponding embedding vectors and then used as each node in the preset node relationship graph.

[0050] Based on the above various embodiments, as an optional embodiment, a deduplication operation is performed on each node in the second node relationship graph to obtain a preset node relationship graph, which specifically includes: For each node in the second node relationship graph, obtain each adjacent node connected to the node through an edge; For any two nodes in the second node relationship graph, determine the number of identical adjacent nodes in the two nodes, and determine the sum of the number of adjacent nodes of each node in the two nodes. Obtain the ratio of the number of identical adjacent nodes to the sum of the number of adjacent nodes. If the ratio is not less than the third preset threshold, then delete any one of the two nodes and the directed edge connected to the one node from the second node relationship graph; Use the second node relationship graph obtained after the deletion operation as the preset node relationship graph.

[0051] Specifically, for each node in the second node relationship graph, first, each node with exactly the same named entity represented by the node can be obtained. For these nodes, only one of them needs to be retained, and the other nodes can be directly deleted, and the directed edges of the other nodes can be merged into the retained node. Then, similar nodes can be further retrieved. For each remaining node, each adjacent node connected to the node and the total number of adjacent nodes can be obtained. Then, each pair of nodes can be compared one by one, comparing the number of adjacent nodes and the number of identical adjacent nodes of the two nodes. Specifically, it can be represented by the following formula:

[0052] where A and B represent any two nodes, is the ratio, represents the number of identical adjacent nodes of the two nodes, Represents the total number of adjacent nodes of two nodes.

[0053] When the ratio is not less than the third preset threshold, it can be considered that the text contents represented by these two nodes are the same. Then, any one of these two nodes and the directed edge connected to any one of the nodes can be deleted from the second node relationship graph, that is, the duplicate removal operation of the second node relationship graph is realized. After the duplicate removal operation is completed, the preset relationship node graph is obtained.

[0054] Optionally, as Figure 2 shown, in the embodiments of the present application, duplicate removal operations can also be performed on the attribution relationships represented by the directed edges between nodes. The duplicate removal method is similar to the method of removing duplicates from nodes, and the embodiments of the present application will not elaborate here.

[0055] Based on the above various embodiments, as an optional embodiment, obtaining at least one candidate node and at least one candidate directed edge that match each named entity included in the to-be-processed request from the preset node relationship graph specifically includes: Performing word segmentation processing on the text content of the to-be-processed request to obtain each first field, obtaining the node corresponding to the named entity that is the same as any one of the first fields from the preset node relationship graph as a candidate node, and obtaining the directed edge corresponding to the attribution relationship that is the same as any one of the first fields from the preset node relationship graph as a candidate directed edge; Obtaining a first embedding vector corresponding to the text content of the to-be-processed request, and obtaining a third embedding vector corresponding to the text content represented by each node of the second node type in the preset node relationship graph; For each node of the second node type, obtaining the third similarity between the third embedding vector and the first embedding vector; Sorting the third similarities in descending order, and determining the nodes of the second node type corresponding to the preset number of the third similarities with the top ranking results as candidate nodes.

[0056] Specifically, before executing the multi-level subgraph search method of adaptive pruning, an initial node relationship graph needs to be constructed first, and each node and directed edge in the initial node relationship graph need to be screened from the preset node relationship graph first. In the embodiments of the present application, these candidate nodes and candidate directed edges can be obtained from the preset node relationship graph by means of retrieval. Specifically, as Figure 3As shown, first, the text content of the request to be processed can be segmented to obtain multiple first fields. The segmentation method can be to use the segmentation and part of speech of the jieba tool, which is not limited in this embodiment of the present application; then, the nodes and directed edges representing the same text content as any first field are found from the preset node relationship graph by exact matching. These nodes can be used as candidate nodes, and these directed edges can be used as candidate directed edges; at the same time, in order to enable the constructed first node relationship graph to provide richer information, a vector similarity matching method can also be adopted, that is, the first embedding vector corresponding to the text content of the request to be processed and the third embedding vector of the text content represented by each node of the second node type (that is, a text content describing the named entity) are similarly calculated. The specific formula can be as follows:

[0057] in, Score similarity is the third similarity, E query is the first embedding vector, E chunk is the third embedding vector, | E query | is the modulus of the first embedding vector, | E chunk | is the modulus of the third embedding vector.

[0058] After calculating the third similarity between each node and the text content of the request to be processed, a preset number of nodes ranked top in descending order can be determined as candidate nodes. It is understood that if a node ranked top has already been determined as a candidate node, that node can be skipped and other nodes ranked lower can be determined as candidate nodes.

[0059] Based on the above embodiments, as an optional embodiment, if there is a first field that is not the same as the named entity represented by any node in the preset node relationship graph; After obtaining a node corresponding to a named entity having the same first field as any one of the candidate nodes from the preset node relationship graph, the method further includes: Obtaining at least one second field from each first field, where the second field is different from the named entity represented by any node in the preset node relationship graph; For each second field, determining an edit distance and a second number of identical characters between the second field and the text content represented by each node in a preset node relationship graph; for each node, performing a weighted summation of the edit distance and the second number of identical characters of the node to obtain a second degree of matching between the node and the second field; and determining nodes having a second degree of matching greater than a fourth preset threshold as candidate nodes; The edit distance is the minimum number of character-based edit operations in the process of adjusting the second field to the text content represented by the node.

[0060] In the embodiments of the present application, the edit distance may be the number of edits required to modify one piece of text content into another. For example, to modify "Shanghai is a city" into "Beijing is a city", it is necessary to modify "Shang" and "Hai" into "Bei" and "Jing" respectively, and a total of two modifications are required. Therefore, the edit distance is 2.

[0061] Specifically, as Figure 3 shown, when there is a first field in the to-be-processed request that fails to match a corresponding node in the preset node relationship graph, these fields can be used as the second field, and then further fuzzy matching is performed in the preset node relationship graph. Specifically, for each second field, it can be determined to compare the second field with the text content represented by each node in the preset node relationship graph. The comparison content includes the edit distance and the number of identical characters, and then the above two comparison contents are weighted and summed respectively to obtain the second matching degree. The specific formula is as follows:

[0062] Wherein, Score fuzzy is the second matching degree, d edit (keyword,note) is the edit distance, overlap (keyword,note) is the number of second identical characters, len(note) is the number of characters of the text content represented by the node, and together with the harmonic coefficient β, they form the weighting parameter.

[0063] After calculating each second matching degree, the nodes with each second matching degree greater than the fourth preset threshold can be determined as candidate nodes.

[0064] Figure 4 is a structural block diagram of a request-response device provided by an embodiment of the present application. As Figure 4 shown, the request-response device 400 may include: a request acquisition module 401, a node matching module 402, and a request-response module 403. Among them, The request acquisition module 401 is used to acquire a to-be-processed request and a first node relationship graph. The first node relationship graph includes multiple nodes and multiple directed edges. Each node is a first type of node or a second type of node. The first type of node represents a named entity, and the second type of node represents the text content describing the named entity. Among them, a directed edge pointing from one node to another node indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to one node; The node matching module 402 is used to, for each first-type node, determine the first matching degree between the first-type node and other nodes in the first node relationship graph except the first-type node based on the request to be processed; wherein, the first matching degree represents the matching degree between the named entity represented by the first-type node and the request to be processed under the condition of the text content represented by other nodes. The request response module 403 is used to, for each first matching degree greater than the first preset threshold, use the named entity represented by the first-type node corresponding to the first matching degree and the text content represented by the corresponding other nodes as the prompt information corresponding to the request to be processed, and obtain the response information corresponding to the request to be processed through the large language model based on the prompt information.

[0065] The solution provided by the embodiment of the present application, by obtaining a node relationship graph including multiple nodes representing named entities and nodes representing the attribution relationships between the named entities, and finding nodes of named entities that match the request to be processed or nodes describing named entities from the node relationship graph, and then sending the named entities represented by the matched nodes or the text content describing the named entities to the large language model for processing to obtain the response information, enables answering the user's request from the perspective of multiple named entities, can effectively improve the accuracy of answering the user's request, and enhance the user's experience.

[0066] Based on the above various embodiments, as an optional embodiment, the device further includes a node graph acquisition module, specifically used for: Obtain a preset node relationship graph, and obtain at least one candidate node and at least one candidate directed edge whose semantic meanings of the represented text contents match the semantic meanings of the named entities included in the request to be processed from the preset node relationship graph; the preset node relationship graph includes multiple nodes and multiple directed edges, each node includes a first-type node and a second-type node, the first-type node represents a named entity, the second-type node represents the text content describing the named entity, wherein, a directed edge pointing from one node to another node indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to one node; Construct an initial node relationship graph based on at least one candidate node and at least one candidate directed edge; Perform a traversal operation on each candidate node, and if a corresponding first node is determined through the traversal operation, perform a deletion operation corresponding to the candidate node; Use the initial node relationship graph after performing the deletion operations corresponding to each candidate node as the first node relationship graph; Wherein, for each candidate node, the traversal operation on the candidate node includes: Determine the second node currently being traversed, and the first node traversed is the candidate node; Use the nodes in the preset node relationship graph that are connected to the second node and pointed to by directed edges as the associated nodes corresponding to the current traversal operation; Determine the first similarity between the semantics of the attribution relationship represented by the directed edge between the second node and the associated node and the semantics represented by the request to be processed. If the first similarity is greater than the second preset threshold, use the associated node as the second node for the next traversal; If the first similarity is not greater than the second preset threshold, use the associated node as one of the first nodes corresponding to the candidate nodes; The deletion operation includes: Delete each first node and the directed edges starting from each first node from the preset node relationship graph.

[0067] Based on the above various embodiments, as an alternative embodiment, the node graph acquisition module is further configured to: Obtain the first embedding vector corresponding to the text content of the request to be processed, and obtain the second embedding vector corresponding to the text content of the attribution relationship, and determine the second similarity between the first embedding vector and the second embedding vector; Determine the first number of identical characters that appear in the text content of the request to be processed and the total number of characters, and determine the proportion of the first number of identical characters to the total number of characters; the total number of characters is the sum of the number of characters in the text content represented by the second node, the number of characters in the text content represented by the associated node, and the number of characters in the text content of the attribution relationship; Perform a weighted sum of the second similarity and the proportion to obtain the first similarity.

[0068] Based on the above various embodiments, as an alternative embodiment, the device further includes a node graph construction module, which is specifically configured to: Obtain multiple pieces of data to be integrated; wherein, the data to be integrated is tabular data or non-tabular data; For each piece of data to be integrated, if the data to be integrated is tabular data, determine the first text content represented by each cell in the tabular data as a named entity, and establish an attribution relationship between the named entity represented by the cell at the table header in each column of the table and the named entity represented by the cell not at the table header in the corresponding column; For each piece of data to be integrated, if the data to be integrated is non-tabular data, obtain the second text content corresponding to the non-tabular data, split the second text content into multiple text blocks containing complete semantics, and extract the named entities contained in each text block and the attribution relationship between the named entities through a large language model; Taking each named entity and each text block as a node respectively, and taking the attribution relationship between each named entity and the attribution relationship between each text block and the named entity as the directed edges between the corresponding two nodes, a second node relationship graph is constructed; Perform a duplicate removal operation on each node in the second node relationship graph to obtain a preset node relationship graph.

[0069] Based on the above various embodiments, as an alternative embodiment, the apparatus further includes a node duplicate removal module, specifically used for: For each node in the second node relationship graph, obtain each adjacent node connected to the node through an edge; For any two nodes in the second node relationship graph, determine the number of identical adjacent nodes in the two nodes, and determine the sum of the number of adjacent nodes of each node in the two nodes. Obtain the ratio of the number of identical adjacent nodes to the sum of the number of adjacent nodes. If the ratio is not less than a third preset threshold, then delete any one of the two nodes and the directed edge connected to the node from the second node relationship graph; Take the second node relationship graph obtained after the deletion operation as the preset node relationship graph.

[0070] Based on the above various embodiments, as an alternative embodiment, the node graph acquisition module can also be used for: Perform word segmentation on the text content of the to-be-processed request to obtain each first field. Obtain the nodes corresponding to the named entities identical to any first field from the preset node relationship graph as candidate nodes, and obtain the directed edges corresponding to the attribution relationships identical to any first field from the preset node relationship graph as candidate directed edges; Obtain the first embedding vector corresponding to the text content of the to-be-processed request, and obtain the third embedding vector corresponding to the text content represented by each node of the second node type in the preset node relationship graph; For each node of the second node type, obtain the third similarity between the third embedding vector and the first embedding vector; Sort the third similarities in descending order, and determine the nodes of the second node type corresponding to the preset number of the third similarities with the top ranking results as candidate nodes.

[0071] Based on the above various embodiments, as an alternative embodiment, the apparatus further includes a node fuzzy matching module, specifically used for: Obtain at least one second field from each first field, and the second field is not the same as the named entity represented by any node in the preset node relationship graph; For each second field, determine the edit distance and the number of identical characters of the second field with the text content represented by each node in the preset node relationship graph. For each node, perform a weighted sum of the edit distance of the node and the number of identical characters of the second field of the node to obtain the second matching degree between the node and the second field. Determine the nodes with the second matching degree greater than the fourth preset threshold as candidate nodes; Wherein, the edit distance is the minimum number of character-based edit operations in the process of adjusting the second field to the text content represented by the node.

[0072] The following refers to Figure 5 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server that executes the method shown in Figure 1 ) 500 suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0073] The electronic device includes: a memory and a processor. The memory is used to store programs for executing the methods described in the above various method embodiments; the processor is configured to execute the programs stored in the memory. Among them, the processor here can be referred to as the processing device 501 described below, and the memory may include at least one of the read-only memory (ROM) 502, random access memory (RAM) 503, and storage device 508 described below, as specifically shown below: As Figure 5 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 502 or the programs loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, ROM 502, and RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0074] Typically, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 508 including, for example, magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or wireline to exchange data. Although Figure 5 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0075] In particular, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present application are executed.

[0076] It should be noted that the above computer-readable storage medium in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0077] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0078] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0079] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: Obtain a request to be processed and a first node relationship graph. The first node relationship graph includes multiple nodes and multiple directed edges. Each node is a first type of node or a second type of node. The first type of node represents a named entity, and the second type of node represents the text content describing the named entity. Among them, a directed edge from one node to another indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to one node. For each first type of node, determine the first matching degree between the first type of node and each second type of node based on the request to be processed. The first matching degree represents the matching degree between the named entity represented by the first type of node and the request to be processed under the condition of the text content represented by the second type of node. For each first matching degree greater than the first preset threshold, use the named entity represented by the first type of node corresponding to the first matching degree and the text content represented by the corresponding second type of node as the prompt information corresponding to the request to be processed, and obtain the response information corresponding to the request to be processed through a large language model based on the prompt information.

[0080] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0081] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0082] The modules or units involved in the embodiments described in the present application can be implemented in software or in hardware. Among them, the name of the module or unit does not, in some cases, constitute a limitation on the unit itself. For example, the first constraint acquisition module can also be described as "the module for acquiring the first constraint".

[0083] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0084] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0085] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0086] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A request-response method, characterized in that, Including: Obtain a request to be processed and a first node relationship graph, where the first node relationship graph includes multiple nodes and multiple directed edges. Each node is a first type of node or a second type of node. The first type of node represents a named entity, and the second type of node represents the text content describing the named entity. Among them, a directed edge pointing from one node to another indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to the one node; For each first type of node, determine the first matching degree between the first type of node and other nodes in the first node relationship graph except the first type of node based on the request to be processed; where the first matching degree represents the matching degree between the named entity represented by the first type of node and the request to be processed under the condition of the text content represented by the other nodes; For each first matching degree greater than a first preset threshold, use the named entity represented by the first type of node corresponding to the first matching degree and the text content represented by the corresponding other nodes as the prompt information corresponding to the request to be processed, and obtain the response information corresponding to the request to be processed through a large language model based on the prompt information.

2. The method according to claim 1, wherein The first node relationship graph is obtained through the following method: Obtain a preset node relationship graph, and obtain at least one candidate node and at least one candidate directed edge whose semantic meaning of the represented text content matches the semantic meaning of each named entity included in the request to be processed from the preset node relationship graph; The preset node relationship graph includes multiple nodes and multiple directed edges. Each node includes a first type of node and a second type of node. The first type of node represents a named entity, and the second type of node represents the text content describing the named entity. Among them, a directed edge pointing from one node to another indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to the one node; Construct an initial node relationship graph based on the at least one candidate node and the at least one candidate directed edge; Perform a traversal operation on each candidate node. If a corresponding first node is determined through the traversal operation, perform the deletion operation corresponding to the candidate node; Use the initial node relationship graph after performing the deletion operations corresponding to each candidate node as the first node relationship graph; Among them, for each candidate node, the traversal operation on the candidate node includes: Determine the second node currently being traversed, and the first node being traversed is the candidate node; Use the node connected to the second node and pointed to by a directed edge in the preset node relationship graph as the associated node corresponding to the current traversal operation; Determine the first similarity between the semantic meaning of the attribution relationship represented by the directed edge between the second node and the associated node and the semantic meaning represented by the request to be processed. If the first similarity is greater than a second preset threshold, use the associated node as the second node for the next traversal; If the first similarity is not greater than the second preset threshold, use the associated node as a first node corresponding to the candidate node; The deletion operation includes: Delete each first node and the directed edges starting from each first node from the initial node relationship graph.

3. The method according to claim 2, wherein The first similarity between the semantics of the directed edge representation between the second node and the associated node and the semantics of the to-be-processed request includes: Obtain a first embedding vector corresponding to the text content of the to-be-processed request, and obtain a second embedding vector corresponding to the text content of the attribution relationship, and determine the second similarity between the first embedding vector and the second embedding vector; Determine the first number of identical characters that appear in the text content of the to-be-processed request and the total number of characters, and determine the proportion of the first number of identical characters to the total number of characters; the total number of characters is the sum of the number of characters in the text content represented by the second node, the number of characters in the text content represented by the associated node, and the number of characters in the text content of the attribution relationship; Perform a weighted sum of the second similarity and the proportion to obtain the first similarity.

4. The method according to claim 2, wherein The preset node relationship graph is obtained in the following manner: Obtain multiple pieces of data to be integrated; wherein, the data to be integrated is tabular data or non-tabular data; For each piece of data to be integrated, if the data to be integrated is tabular data, then determine each first text content represented by each cell in the tabular data as a named entity, and establish an attribution relationship between the named entity represented by the cell at the table header in each column of the table and the named entity represented by the cell not at the table header in the corresponding column; For each piece of data to be integrated, if the data to be integrated is non-tabular data, then obtain the second text content corresponding to the non-tabular data, split the second text content into multiple text blocks containing complete semantics, and extract the named entities contained in each text block and the attribution relationships between the named entities through the large language model; Construct a second node relationship graph with each named entity and each text block as a node respectively, and with the attribution relationships between each named entity and the attribution relationships between each text block and the named entity as directed edges between the corresponding two nodes; Perform a deduplication operation on the nodes in the second node relationship graph to obtain the preset node relationship graph.

5. The method according to claim 4, characterized in that, The operation of performing a deduplication operation on the nodes in the second node relationship graph to obtain the preset node relationship graph includes: For each node in the second node relationship graph, obtain each adjacent node connected to the node through an edge; For any two nodes in the second node relationship graph, determine the number of identical adjacent nodes in the any two nodes, and determine the sum of the number of adjacent nodes of each node in the any two nodes, obtain the proportion of the number of identical adjacent nodes to the sum of the number of adjacent nodes, and if the proportion is not less than a third preset threshold, then delete any one of the any two nodes and the directed edge connected to the any one node from the second node relationship graph; Use the second node relationship graph obtained after the deletion operation as the preset node relationship graph.

6. The method according to claim 2, wherein Obtaining at least one candidate node and at least one candidate directed edge that match each named entity included in the to-be-processed request from the preset node relationship graph includes: Performing word segmentation on the text content of the to-be-processed request to obtain each first field, obtaining, from the preset node relationship graph, a node corresponding to the named entity that is the same as any one of the first fields as the candidate node, and obtaining, from the preset node relationship graph, a directed edge corresponding to the attribution relationship that is the same as any one of the first fields as the candidate directed edge; Obtaining a first embedding vector corresponding to the text content of the to-be-processed request, and obtaining a third embedding vector corresponding to the text content represented by each node of the second node type in the preset node relationship graph; For each node of the second node type, obtaining a third similarity between the third embedding vector and the first embedding vector; Sorting the third similarities in descending order, and determining the nodes of the second node type corresponding to a preset number of the third similarities with the top ranking results as the candidate nodes.

7. The method according to claim 6, wherein If there is a first field that is not different from any named entity represented by any node in the preset node relationship graph; After obtaining, from the preset node relationship graph, a node corresponding to the named entity that is the same as any one of the first fields as the candidate node, it further includes: Obtaining at least one second field from each of the first fields, where the second field is not the same as any named entity represented by any node in the preset node relationship graph; For each second field, determining the edit distance and the number of second identical characters between the second field and the text content represented by each node in the preset node relationship graph. For each node, performing a weighted sum of the edit distance of the node and the number of second identical characters of the node to obtain a second matching degree between the node and the second field, and determining the node with the second matching degree greater than a fourth preset threshold as the candidate node; Wherein, the edit distance is the minimum number of character-based edit operations in the process of adjusting the second field to the text content represented by the node.

8. A request-response device, characterized in that, It includes: A request acquisition module, configured to acquire a to-be-processed request and a first node relationship graph, where the first node relationship graph includes multiple nodes and multiple directed edges, each node is a first type node or a second type node, the first type node represents a named entity, and the second type node represents the text content describing the named entity. Wherein, a directed edge pointing from one node to another node indicates that the named entity corresponding to the other node belongs to the category represented by the named entity corresponding to the one node; A node matching module, configured to, for each first type node, determine a first matching degree between the first type node and other nodes in the first node relationship graph except the first type node based on the to-be-processed request; wherein, the first matching degree represents the matching degree between the named entity represented by the first type node and the to-be-processed request under the condition of the text content represented by the other nodes; The request-response module is configured to, for each first matching degree greater than a first preset threshold, use the named entity represented by the first type of node corresponding to the first matching degree and the text content represented by the corresponding other nodes as the prompt information corresponding to the to-be-processed request, and obtain the response information corresponding to the to-be-processed request through a large language model based on the prompt information.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Intelligent customer service voice response method and device, computer equipment and storage medium

    CN112653798A

  • Knowledge question and answer method and device, computer readable medium and electronic equipment

    CN115292457A

  • Question and answer processing method and device, equipment and storage medium

    CN117556012A

  • Knowledge graph retrieval method and device, medium and product

    CN118227668A

  • Intelligent question and answer method, system and device, computer equipment and readable storage medium

    CN118377881A