Conversation processing method, device and computer program product
By matching dialogue text with document block entities in the knowledge graph, the problem of inaccurate matching in the existing dialogue system is solved, and higher accuracy and comprehensiveness are achieved, and the quality of reply is improved.
Patent Information
- Application Number
- CN202510032807.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
AI Technical Summary
The existing RAG-based dialogue system still has room for improvement in accuracy, especially since document blocks may contain noisy content, resulting in inaccurate matching between dialogue text and document blocks, and there are problems such as ‘inaccurate’ and ‘incomplete’.
By matching dialogue text with document block entities in the knowledge graph, rather than matching traditional dialogue text with document blocks, the document block associated with the target document block entity obtained by the matching is used as the reference document block to improve the accuracy and comprehensiveness of the match.
This method can better reflect the key information of the conversation text, avoid noise influence, improve the accuracy and comprehensiveness of the reference document block, and thus improve the accuracy of the final reply.
Smart Images

Figure CN119988544A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing and artificial intelligence technology, and in particular to a dialogue processing method, device and computer program product. Background Art
[0002] With the continuous development of artificial intelligence technology, intelligent dialogue systems can automatically generate responses to dialogue texts when users input dialogue texts and feed them back to users. The dialogue system based on RAG (Retrieval Augmented Generation) technology has greatly improved the representation learning ability of contextual language, and can show better performance in various common text tasks, which has strong practical value in a wide range of industrial applications.
[0003] However, the accuracy of current RAG-based dialogue systems still needs to be improved. Summary of the invention
[0004] In view of this, the present application provides a dialogue processing method, device and computer program product to improve the accuracy of the RAG-based dialogue system.
[0005] This application provides the following solutions:
[0006] In a first aspect, a dialog processing method is provided, the method comprising:
[0007] Get the conversation text;
[0008] The conversation text is used to search in the knowledge graph to obtain a document block entity matching the conversation text as a target document block entity; wherein the knowledge graph at least includes an association relationship between a document block and a document block entity, and the document block entity is obtained by performing content analysis on its associated document block;
[0009] Searching for a document block associated with the target document block entity in the knowledge graph to obtain a reference document block;
[0010] constructing prompt instructions using the dialogue text and the reference document block;
[0011] The prompt instruction is provided to the dialogue model to obtain a reply to the prompt instruction.
[0012] In a second aspect, a dialog processing device is provided, the device comprising:
[0013] A question acquisition unit, configured to acquire a dialogue text;
[0014] An entity retrieval unit is configured to use the conversation text to search in the knowledge graph to obtain a document block entity matching the conversation text as a target document block entity; wherein the knowledge graph at least includes an association relationship between a document block and a document block entity, and the document block entity is obtained by performing content analysis on its associated document block;
[0015] A document retrieval unit, configured to search the knowledge graph for a document block associated with the target document block entity to obtain a reference document block;
[0016] The reply acquisition unit is configured to construct a prompt instruction using the dialogue text and the reference document block, provide the prompt instruction to the dialogue model, and obtain a reply to the prompt instruction.
[0017] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the first aspect are implemented.
[0018] In a fourth aspect, an electronic device is provided, including:
[0019] one or more processors; and
[0020] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the first aspects above.
[0021] In a fifth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.
[0022] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0023] In the RAG-based solution proposed in the present application, the knowledge graph includes the association relationship between document blocks and document block entities, and replaces the traditional method of matching dialogue text and document blocks with the matching of dialogue text and document block entities. The target document block entity matched in this way better reflects the key information of the dialogue text. Therefore, the document block associated with the target document block entity is used as the reference document block, which can make the reference document block contain the key information of the dialogue text, and avoid the problems of "inaccurate search" and "incomplete search" caused by the traditional method of matching dialogue text and document blocks due to the noise contained in the document block, thereby improving the accuracy and comprehensiveness of the reference document block, and thereby improving the accuracy of the answer obtained based on the reference document block. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0025] Figure 1 is a system architecture diagram applicable to the embodiments of the present application;
[0026] Figure 2 A flow chart of a conversation processing method provided in an embodiment of the present application;
[0027] Figure 3 A flow chart of a method for constructing a knowledge graph provided in an embodiment of the present application;
[0028] Figure 4 A schematic diagram of a document segmentation method provided in an embodiment of the present application;
[0029] Figure 5 A schematic diagram of a method for extracting document block entities provided in an embodiment of the present application;
[0030] Figure 6 A schematic diagram of a method for extracting document entities provided in an embodiment of the present application;
[0031] Figure 7 A schematic diagram of a knowledge graph provided for an embodiment of the present application;
[0032] Figure 8 A schematic block diagram of a conversation processing device provided in an embodiment of the present application;
[0033] Fig. 9 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0034] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0035] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0036] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0037] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0038] The solution currently adopted by the traditional RAG-based dialogue system is mainly to split each document in the document set into document blocks, convert the document blocks into vectors and store them in the database. In response to the dialogue text input by the user, the vector corresponding to the dialogue text is matched with the vector corresponding to the document block in the database, and the matched document block is used as the reference document block. The prompt instruction is constructed using the dialogue text and the reference document block and input into the LLM (Large Language Model), and then the reply output by the LLM according to the prompt instruction is obtained, and the reply is returned to the user. However, in the traditional method, because the document block may contain noise content, the text block obtained by matching the dialogue text with the vector corresponding to the document block is often inaccurate, and often lacks key information related to the query requirements of the dialogue text, that is, there are defects of "inaccurate query" and "incomplete query", which makes the reply finally returned to the user inaccurate and does not meet the user's query requirements.
[0039] In view of this, the present application provides a new idea. In order to facilitate the understanding of the present application, the system architecture on which the present application is based is first described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in , the system architecture may include: user equipment and a dialogue processing device located at the server end.
[0040] The user can input a query (conversation text) through the user equipment, and the user equipment sends the conversation text to the conversation processing device on the server side.
[0041] User devices may include, but are not limited to, smart mobile terminals, smart home devices, wearable devices, PCs (Personal Computers), etc. Smart mobile devices may include mobile phones, tablet computers, laptops, PDAs (Personal Digital Assistants), Internet cars, etc. Smart home devices may include smart TVs, smart refrigerators, etc. Wearable devices may include smart watches, smart glasses, virtual reality devices, augmented reality devices, mixed reality devices (i.e., devices that can support virtual reality and augmented reality), etc.
[0042] The dialogue processing device can use the method provided in the embodiment of the present application to generate a reply for the dialogue text. The process of generating a reply by the dialogue processing device involves the use of a dialogue model, which is described in the subsequent embodiments for details.
[0043] The conversation processing device can be set up as an independent server, or in a server group, or in a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a host product in the cloud computing service system to solve the problems of difficult management and weak service scalability in traditional physical hosts and virtual private servers (VPS). Figure 1 In addition to the shown architecture, the question-answer processing device can also be set in a computer terminal with strong computing power.
[0044] As one possible implementation method, a user can input a conversation text through a user device, and the user device sends the conversation text to a conversation processing device on the server side through a network. After the conversation processing device generates a reply for the conversation text, it returns the reply to the user device through the network. The user device displays the reply to the user.
[0045] It should be understood that Figure 1 The user equipment, dialogue processing device and dialogue model in the figure are only illustrative. According to the implementation requirements, there can be any number of user equipment, dialogue processing device and dialogue model.
[0046] Figure 2 A flow chart of a method for processing a conversation provided in an embodiment of the present application. The method can be performed by Figure 1 The dialogue processing device in the system shown in the figure is executed. Figure 2 As shown in , the method may include the following steps:
[0047] Step 201: Get the conversation text.
[0048] Step 203: Use the conversation text to search in the knowledge graph and obtain the document block entity that matches the conversation text as the target document block entity; wherein the knowledge graph at least includes the association relationship between the document block and the document block entity, and the document block entity is obtained by performing content analysis on its associated document block.
[0049] Step 205: Search the knowledge graph for a document block associated with the target document block entity to obtain a reference document block.
[0050] Step 207: construct a prompt instruction using the dialogue text and the reference document block, provide the prompt instruction to the dialogue model, and obtain a response to the prompt instruction.
[0051] It can be seen from the above process that in the RAG-based solution proposed in the present application, the knowledge graph includes the association relationship between document blocks and document block entities, and replaces the traditional way of matching dialogue text and document blocks with the matching of dialogue text and document block entities. The target document block entity matched in this way better reflects the key information of the dialogue text. Therefore, the document block associated with the target document block entity is used as the reference document block, which can make the reference document block contain the key information of the dialogue text, and avoid the problems of "inaccurate search" and "incomplete search" caused by the traditional way of matching dialogue text and document blocks due to the noise contained in the document block, thereby improving the accuracy and comprehensiveness of the reference document block, and thus improving the accuracy of the answer obtained based on the reference document block.
[0052] The following is a detailed description of each step in the above process and the effects that can be further produced in conjunction with the embodiments. It should be noted that the limitations such as "first" and "second" involved in the present disclosure do not have limitations in terms of size, order and quantity, and are only used to distinguish between the names. For example, "first document block entity", "second document block entity", etc. are only used to distinguish document block entities in name. For another example, "first document entity", "second document entity", etc. are only used to distinguish document entities in name.
[0053] First, the above step 201, namely "obtaining the conversation text", is described in detail in conjunction with the embodiment.
[0054] In the embodiment of the present application, the dialogue text may be input by the user and sent to the server through the user device. For example, a dialog box, a web page form or a specific page component may be provided to the user, and the user enters the dialogue text through the dialog box or the specific page component. The dialogue text usually contains the user's question or contains content describing the user's needs.
[0055] In addition, in addition to the form of users directly inputting text, voice input can also be used. Users can input voice through specific page components, and the voice is sent to the server through the user device. The server recognizes the voice and converts it into text to obtain the conversation text.
[0056] The above step 203, i.e., "using the conversation text to search in the knowledge graph to obtain the document block entity matching the conversation text as the target document block entity" is described in detail below in conjunction with the embodiments.
[0057] In the embodiment of the present application, the conversation text is not matched with the document blocks obtained after the document segmentation in the document library, but a knowledge graph is constructed for the document blocks, and the knowledge graph at least includes the association relationship between the document blocks and the document block entities, wherein the document block entities are extracted by content analysis of the associated document blocks. In order to facilitate the understanding of the present application, the construction process of the knowledge graph is first described. Figure 3 A flow chart of a method for constructing a knowledge graph provided in an embodiment of the present application, such as Figure 3 As shown in , the method may include the following steps:
[0058] Step 301: Split the documents in the document library into multiple document blocks.
[0059] The so-called document block refers to a smaller text unit obtained by segmenting a document. Depending on the segmentation method, the size of the obtained document block may also be different.
[0060] As one of the feasible methods, the documents in the document library can be segmented according to a fixed document block size to obtain document blocks. The sizes of the document blocks obtained by this segmentation method are relatively uniform, but this segmentation method does not consider the structure and semantics of the document at all, which easily leads to "structural disorder" and "semantic discontinuity" in the document blocks obtained after segmentation.
[0061] Therefore, the present application embodiment provides another implementation method, such as Figure 4 As shown in , each document is first segmented according to the document structure, for example, according to the title in the document, to obtain multiple preliminary document blocks. In addition, before the document is segmented, the document format can be first converted into a document format that can be understood by the large language model. Since the obtained preliminary document blocks are segmented based on the document structure, for example, each subheading is segmented as a unit, some subheadings have a lot of content, and some subheadings have very little content. Accordingly, the obtained preliminary document blocks may be very long or very short. Therefore, the preliminary document blocks can be further optimized using the large language model to finally obtain multiple document blocks.
[0062] It should be noted here that a large language model, or LLM (Large Language Model), refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. It is characterized by its large scale and huge number of parameters (usually reaching more than 10 billion levels), and is usually based on a deep learning architecture such as the Transformer architecture. The difference between LLM and ordinary pre-trained language models lies in the parameter scale. When the parameter scale exceeds a certain level, the model achieves significant performance improvement and exhibits capabilities that do not exist in small models, such as in-context learning capabilities, which can learn complex patterns in language and perform a wide range of tasks, including text summarization, translation, sentiment analysis, multi-round dialogue, etc. Therefore, in order to distinguish it from traditional pre-trained language models, such models with a parameter scale exceeding a certain level are called LLMs. In general, it can be considered that a language model with a certain degree of parameter scale based on a deep learning architecture is called a large language model.
[0063] Based on the powerful semantic understanding and text processing capabilities of the large language model, the preliminary document block can be input into the large language model and an optimization rule can be provided so that the large language model can optimize the preliminary document block according to the optimization rule. The optimization can include: segmenting the preliminary document block whose length exceeds a first threshold based on semantics, and / or merging multiple preliminary document blocks whose length is less than a second threshold based on semantics. The first threshold and the second threshold can be defined in the optimization rule, and the size range of the text block obtained after segmentation can also be defined.
[0064] In this way, the large language model can merge "small" preliminary document blocks with similar semantics based on semantics, and reasonably divide the "large" preliminary document blocks based on semantics. The document blocks thus obtained can ensure a certain degree of integrity in structure and semantics, and are relatively uniform in size, which can be used as a preferred implementation method.
[0065] Step 303: Perform content analysis on each document block to extract document block entities.
[0066] In an embodiment of the present application, document block entities can be extracted from each document block based on content analysis, and are usually semantically critical entities involved in each document block. Taking a technical document block as an example, the document block entities extracted therefrom are some key technical terms involved in the document block, such as "method", "process", "API", "component", "model", etc. In addition, the document block entity may include descriptive information such as entity name, entity type, entity content, etc. Among them, which entity types can be pre-defined, that is, a pre-defined entity type set, and the entity type of the document block entity is selected from the entity type set. The entity content can be extracted from the document block, which can usually be a description related to the document block entity in the document block.
[0067] As one of the feasible ways, this step can be implemented based on a large language model. Each document block and document block entity extraction rule obtained in step 301 are provided to the large language model respectively, and the large language model extracts document block entities from each document block based on the extraction rule, such as Figure 5 as shown in .
[0068] After extracting each document block entity, each document block entity can be deduplicated, and then an association is established between the document block and the document block entity in the knowledge graph. The association relationship corresponds to the above-mentioned extracted relationship, that is, if a document block entity is extracted from a document block, there is an association relationship between the two. Figure 7 In the knowledge graph shown in , the document block entities "method" and "model" are extracted from document block 1, and the document block entities "method" and "model" are respectively associated with document block 1. The document block entity "model" is extracted from both document block 3 and document block 1, and the document block entity "model" is associated with both document block 3 and document block 1.
[0069] It can be seen that there are already two layers in the knowledge graph. One layer includes each document block as a node, and the other layer includes each document block entity as a node.
[0070] Furthermore, the constructed knowledge graph may further include the association relationship between the document block entity and the document entity, and may further include the association relationship between the document entity and the field. Accordingly, the construction process of the knowledge graph may further perform the following steps:
[0071] Step 305: Perform attribute analysis on each document to extract document entities and relationships between document entities.
[0072] In the embodiment of the present application, the document entity is obtained by analyzing the attributes of the document, and the extracted document entity represents the attributes of the document, such as the document name, document type, document keywords, document summary, document path, etc. In addition, the relationship between the document entities can be further extracted.
[0073] As one possible implementation, the type of relationship between two document entities can be represented by the domain to which the two document entities belong. If the two document entities do not have a common domain, there is no relationship between the two document entities, and there is no edge between the two document entities in the corresponding knowledge graph. Furthermore, in addition to the relationship type, the relationship between two document entities can also include relationship strength, which represents the degree to which the two document entities correspond to the relationship type.
[0074] In addition to characterizing the type of relationship between document entities by the domain to which two text entities belong, other methods can also be used, which are not listed here one by one.
[0075] It should be noted that the "field" involved in the embodiments of the present application can be any granularity representing the scope of industry, academic, social activities, etc. The field type can be pre-defined to form a field space, and the document entity can be mapped to the corresponding field space to obtain the field type to which the document entity belongs, and then the relationship between the two document entities can be obtained.
[0076] For example, the document entities "music player" and "video editor" both belong to the multimedia field, and the relationship between the document entities "music player" and "video editor" includes: "relationship type: multimedia, strength: 5".
[0077] As one of the feasible ways, the extraction of the relationship between the document entities can be implemented based on a large language model. Each document entity and the relationship extraction rule are provided to the large language model respectively, and the large language model extracts the relationship between the document entities from each document entity based on the extraction rule, such as Figure 6 As shown in . The extraction rules may include all relationship types, i.e., domain types, and may also include evaluation criteria for relationship strength to instruct the large language model to extract the relationships between document entities from these relationship types, including relationship types and relationship strengths.
[0078] After extracting the document entity, the document entity is deduplicated and an association is established between the document block entity and the document entity in the knowledge graph. The document block entity and the document entity from the same document can be associated. Figure 7 As shown in , document entity 1 is obtained by performing attribute analysis and extraction on document 1, and document block entities "method" and "process" are obtained by performing content analysis on document block 1. Document block 1 is obtained by segmenting document 1, which means that document entity 1 and document block entities "method" and "process" all originate from document 1. Therefore, document entity 1 is associated with document block entities "method" and "process" respectively.
[0079] Furthermore, in the knowledge graph, the relationship between document entities can also be reflected by the edges between document entities. In addition to the relationship type, the relationship between document entities can also include information such as relationship strength and relationship description.
[0080] So far, there are three layers in the knowledge graph. In addition to document blocks and document block entities, there is also a document entity layer, which includes each document entity as a node.
[0081] Step 307: Clustering and domain analysis are performed on multiple document entities according to the relationships between the document entities to obtain multiple domains.
[0082] After obtaining each document entity, the document entities can be clustered according to the relationship between them, and then each clustering result can be analyzed for domains to obtain multiple domains. Taking technical documents as an example, after clustering and domain analysis of each document entity, we can obtain "multimedia domain", "network request domain", "file operation domain", etc.
[0083] Step 309: Construct a knowledge graph using document blocks, document block entities, document entities, and domains.
[0084] After obtaining multiple fields, we can further establish associations between fields and document entities. For example, if document entity 1 and document entity 2 belong to the same cluster, and the cluster corresponds to the multimedia field, we can establish associations between the multimedia field and document entity 1 and document entity 2, such as Figure 7 As shown in . So far, there are four layers in the knowledge graph, including the bottom layer where the document block is located, the second to last layer where the document block entity is located, the second layer where the document entity is located, and the top layer where the domain is located. There are associations between the nodes in each layer.
[0085] When executing step 203, i.e., using the conversation text to search in the knowledge graph and obtaining the document block entity matching the conversation text as the target document block entity, a local search can be performed, i.e., matching the conversation text with the document block entity in the knowledge graph, i.e., calculating the matching degree between the conversation text and the document block entity, and determining the document block entity whose matching degree meets the first requirement as the first document block entity; and obtaining the target document block entity using the first document block entity.
[0086] Among them, when calculating the matching degree between the dialogue text and the document block entity, a variety of methods can be used, such as calculating the similarity between the vector corresponding to the dialogue text and the vector corresponding to the text block entity, calculating the Jaccard distance between the dialogue text and the document block entity (that is, the matching degree calculated based on the number of common words in the dialogue text and the document block entity), calculating the edit distance between the dialogue text and the document block entity, using a neural network model to determine the similarity between the dialogue text and the document block entity, and so on.
[0087] The first requirement may be that the matching degree ranks in the top preset number, the matching degree is greater than or equal to the preset first matching degree threshold, etc. For example, the document block entity that ranks in the top N1 (Top N1) matching degrees with the dialogue text may be taken as the first document block entity, where N1 is a preset positive integer.
[0088] Furthermore, when performing local search to match the dialogue text with the document block entity, each type of document block entity can be searched separately to obtain the top several document block entities in each type with the dialogue text, and then the matching degrees corresponding to these document entities are weighted, and the corresponding weight is determined by the type of the document block entity, that is, each document block entity type corresponds to its own weight. Then, the retrieved document block entities are sorted according to the weighted matching degree, and the document block entities with the top N1 matching degrees after weighted matching are selected.
[0089] As one possible implementation method, the first document block entity may be used as the target document block entity.
[0090] As another feasible method, the global search method can be further combined on the basis of local search. When performing global search, the document entity associated with the first document block entity in the knowledge graph can be queried first as the first document entity, and then the document block entity associated with the first document entity in the knowledge graph can be queried, and the queried document block entity can be deduplicated based on the first document block entity to obtain the second document block entity; based on the matching degree between the dialogue text and the second document block entity, the document block entity whose matching degree meets the second requirement is determined as the third document block entity; and the target document block entity is obtained using the first document block entity and the third document block entity.
[0091] That is to say, first, the first document block entity retrieved locally is searched upward in the knowledge graph for the associated document entity. For the convenience of description, these document entities are called first document entities. Then, the document block entities associated with the first document entity are searched downward. Since these queried document block entities may be repeated with the first document block entity retrieved locally, deduplication processing is performed to remove the duplicate content of the queried document block entity with the first document block entity retrieved locally, and the remaining document block entities are called second document block entities. The matching degree of these second document block entities with the dialogue text can be calculated. The matching degree calculation method can be referred to the relevant records in the previous embodiment and will not be repeated. Then take the second document block entity ranked in the top N2, N2 is a preset positive integer, or take the second document entity whose corresponding matching degree is greater than or equal to the preset second matching degree threshold. For example, take the second document block entity with the matching degree ranked in the top N2 with the dialogue text as the third document block entity.
[0092] For example, suppose that the conversation text input by the user is retrieved locally and the first document block entity includes "API", "process" and "framework". In the global search, the document entity associated with "API", "process" and "framework" is queried upward as the first document entity, including "document entity 2" and "document entity 4". Then the document block entity associated with "document entity 2" and "document entity 4" is queried downward, including "process", "API", "framework" and "component". After deduplication, the second document block entity including "component" is obtained. The matching degree between "component" and the conversation text is calculated respectively. Assuming that "component" meets the second requirement, "component" is used as the third document block entity.
[0093] As one achievable manner, the target document block entity may be obtained by using the first document block entity and the third document block entity.
[0094] Furthermore, during global retrieval, the knowledge graph can be further queried upward for fields associated with the first document entity, and the target field can be determined from these associated fields; then, a downward query can be started from the target field to query the document entities associated with the target field, and the document entities associated with the target field can be deduplicated based on the first document entity to obtain the second document entity; the knowledge graph can be queried for the document block entity associated with the second document entity to obtain the fourth document block entity; based on the matching degree between the dialogue text and the fourth document block entity, the document block entity whose matching degree meets the third requirement can be determined as the fifth document block entity; the first document block entity, the third document block entity and the fifth document block entity can be merged and deduplicated to obtain the target document block entity.
[0095] That is to say, the first document entity is continued to be queried upward for the associated fields. When determining the target field from the associated fields, since multiple first document entities may be associated with the same field, one first document entity hits the field once, and multiple first document entities hit the field multiple times. Therefore, the fields associated with the first document entity can be sorted according to the number of times the field is hit by the query, and according to the sorting result, the first M fields are selected as the target fields, where M is a positive integer.
[0096] Continuing with the above example, after obtaining the first document entity "document entity 2" and "document entity 4", continue to query upward to obtain the associated fields "multimedia field" and "file operation field". Assuming that the first two fields are selected based on the number of hits, the "multimedia field" and "file operation field" are used as the target fields to continue to query downward. "Document entity 1", "Document entity 2", "Document entity 3" and "Document entity 4" are queried. After deduplication based on the first document entity, "Document entity 1" and "Document entity 3" are obtained as the second document entity. Continue to query downward the document block entities associated with the second document entity, namely "method", "model" and "component" as the fourth document block entity. Assuming that the matching degree between the "component" and the dialogue text meets the third requirement, "component" is used as the fifth document block entity. After merging and deduplication of the first document block entities "API", "process" and "framework", the third document block entity "component" and the fifth document block entity "component", "API", "process", "framework" and "component" are finally obtained as the target document block entities.
[0097] After obtaining the target document block entity, step 205 is executed, that is, searching the knowledge graph for document blocks associated with the target document block entity to obtain a reference document block. Continuing with the above example, searching the knowledge graph for document blocks associated with the target document block entities "API", "process", "framework" and "component" to obtain document block 2, document block 4 and document 5 as reference document blocks.
[0098] It should be noted here that Figure 7 The knowledge graph shown is only a schematic diagram for easy understanding. In actual scenarios, the number of nodes at each layer in the knowledge graph may be more than Figure 7 As shown, the implementation principle can refer to the examples given above.
[0099] The above step 207, namely "using the dialogue text and the reference document block to construct a prompt instruction, providing the prompt instruction to the dialogue model, and obtaining a reply to the prompt instruction" is described in detail below in conjunction with the embodiments.
[0100] In the embodiment of the present application, the dialogue model can be implemented based on a large language model or a common pre-trained language model. A prompt can be constructed using the dialogue text and the reference document block to instruct the dialogue model to use the content of the reference document block as a reference to reason about the dialogue text to obtain a reply.
[0101] The generated reply may be text, or other modal content such as images, tables, videos, etc. The generated reply may be returned to the user terminal through a dialog box, a web form, or a specific page component on the web page, and displayed to the user by the user terminal. In addition, the reply text may be converted into speech, and the corresponding speech may be returned to the user terminal, and played to the user by the user terminal.
[0102] It can be seen from the above embodiments that the technical solution provided by the present application can also have the following advantages:
[0103] 1) The present application replaces the traditional method of matching dialogue text and document blocks with the matching of dialogue text and document block entities. The target document block entity matched in this way better reflects the key information of the dialogue text. Therefore, the document block associated with the target document block entity is used as the reference document block, which can make the reference document block contain the key information of the dialogue text, and avoid the problems of "inaccurate search" and "incomplete search" caused by the traditional method of matching dialogue text and document blocks due to the noise contained in the document block, thereby improving the accuracy and comprehensiveness of the reference document block, and further improving the accuracy of the answer obtained based on the reference document block.
[0104] 2) Based on the matched first document block entity, the present application can first query upwards to obtain the associated document entity to obtain the first document entity, then retrieve downwards to obtain the associated document block entity and perform deduplication processing to obtain the second document block entity, and further select a third document block entity from the second document block entity to expand the target document block entity based on the matching degree between the conversation text and the second document block entity, thereby being able to expand the target document block entity based on document attribute matching, further improving the accuracy and comprehensiveness of the reference document block, and further improving the accuracy of the reply obtained based on the reference document block.
[0105] 3) The present application can further query upward based on the first document entity obtained to obtain the associated fields, and after determining the target field therefrom, retrieve downward to obtain the associated document block entity and perform deduplication processing to obtain the fifth document block entity to expand the target document entity, thereby being able to expand the target document block entity based on the field, further improving the accuracy and comprehensiveness of the reference document block, and further improving the accuracy of the answer obtained based on the reference document block.
[0106] 4) When determining the target field from the queried fields, the present application sorts the queried fields according to the number of times the field is hit by the query, and selects the target field based on the sorting result. This method can determine the field that is most relevant to the conversation text on the one hand, and can effectively improve the efficiency and resource consumption of downward queries on the other hand.
[0107] 5) This application provides a new method for constructing a knowledge graph, establishing a layer-by-layer association between document blocks, document block entities, document entities and domains, thereby providing a basis for document block retrieval in question-and-answer processing methods.
[0108] 6) In this application, the document is first segmented according to the document structure, and then the powerful semantic understanding ability of the large language model is used to segment or merge the preliminary document blocks obtained by segmentation. The document blocks thus obtained can ensure a certain degree of integrity in structure and semantics, and the size is relatively uniform.
[0109] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time.
[0110] The above method can be applied to dialogue systems in various scenarios, such as online customer service, intelligent question-and-answer functions of smart speakers, dialogue robots based on large models, etc. It can also be applied to various fields, such as intelligent question-and-answer for technical documents, intelligent question-and-answer for medical documents, etc.
[0111] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0112] According to an embodiment of another aspect, a dialog processing device is provided. Figure 8 This is a schematic block diagram of a dialog processing device provided in an embodiment of the present application. Figure 8 As shown, the device 800 includes: a question acquisition unit 801, an entity retrieval unit 802, a document retrieval unit 803 and a reply acquisition unit 804, and may further include a reply output unit 805 and a graph construction unit 806. The main functions of each component unit are as follows:
[0113] The question acquisition unit 801 is configured to acquire the dialogue text.
[0114] The entity retrieval unit 802 is configured to use the conversation text to search in the knowledge graph and obtain the document block entity matching the conversation text as the target document block entity; wherein the knowledge graph at least includes the association relationship between the document block and the document block entity, and the document block entity is obtained by performing content analysis on its associated document block.
[0115] The document retrieval unit 803 is configured to search for a document block associated with a target document block entity in the knowledge graph to obtain a reference document block.
[0116] The reply acquisition unit 804 is configured to construct a prompt instruction using the dialogue text and the reference document block, provide the prompt instruction to the dialogue model, and obtain a reply to the prompt instruction.
[0117] Furthermore, the reply output unit 805 is configured to output a reply.
[0118] As one of the achievable methods, the entity retrieval unit 802 can be specifically configured to match the conversation text with the document block entity in the knowledge graph, determine the document block entity whose matching degree meets the first requirement as the first document block entity; and use the first document block entity to obtain the target document block entity.
[0119] Furthermore, the knowledge graph also includes the association between document chunk entities and document entities. Document entities are extracted by analyzing the attributes of documents, and document chunks are obtained by segmenting documents.
[0120] When the entity retrieval unit 802 uses the first document block entity to obtain the target document block entity, it can be configured as follows: query the document entity associated with the first document block entity in the knowledge graph as the first document entity; query the document block entity associated with the first document entity in the knowledge graph, and deduplicate the queried document block entity based on the first document block entity to obtain the second document block entity; determine the document block entity whose matching degree meets the second requirement as the third document block entity based on the matching degree between the dialogue text and the second document block entity; and obtain the target document block entity using the first document block entity and the third document block entity.
[0121] Furthermore, the knowledge graph also includes the association relationship between the document entity and the domain. When the entity retrieval unit 802 obtains the target document block entity by using the first document block entity and the third document block entity, it can be specifically configured as follows:
[0122] Querying the domain associated with the first document entity in the knowledge graph, and determining the target domain from the domain associated with the first document entity;
[0123] Query the document entities associated with the target field in the knowledge graph, and remove duplicate document entities associated with the target field based on the first document entity to obtain a second document entity;
[0124] Query the document block entity associated with the second document entity in the knowledge graph to obtain a fourth document block entity;
[0125] According to the matching degree between the dialogue text and the fourth document block entity, determining the document block entity whose matching degree meets the third requirement as the fifth document block entity;
[0126] The first document block entity, the third document block entity and the fifth document block entity are merged and deduplicated to obtain a target document block entity.
[0127] As one of the achievable methods, when the entity retrieval unit 802 determines the target domain from the domains associated with the first document entity, it can be specifically configured as follows: sorting the domains associated with the first document entity according to the number of times the domains are hit by the query; and based on the sorting result, selecting the top M domains as the target domains, where M is a positive integer.
[0128] Furthermore, the graph construction unit 806 can be configured to: pre-segment the documents in the document library to obtain multiple document blocks; perform content analysis on each document block to extract document block entities; perform attribute analysis on each document to extract document entities and the relationships between document entities; perform clustering and domain analysis on multiple document entities based on the relationships between document entities to obtain multiple domains; and construct a knowledge graph using document blocks, document block entities, document entities, and domains.
[0129] As one of the achievable ways, when the graph construction unit 806 divides the documents in the document library to obtain multiple document blocks, it can be specifically configured as follows: dividing the documents in the document library according to the document structure to obtain multiple preliminary document blocks; optimizing the multiple preliminary document blocks using a large language model to obtain multiple document blocks; wherein the optimization includes: dividing the preliminary document blocks whose length exceeds a first threshold based on semantics, and / or merging multiple preliminary document blocks whose length is less than a second threshold based on semantics, and the first threshold is greater than the second threshold.
[0130] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0132] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0133] And an electronic device, comprising:
[0134] one or more processors; and
[0135] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.
[0136] The present application also provides a computer program product, including a computer program, which implements the steps of any one of the methods in the aforementioned method embodiments when executed by a processor.
[0137] in, Fig. 9The architecture of the electronic device is shown as an example, which may include a processor 910, a video display adapter 911, a disk drive 912, an input / output interface 913, a network interface 914, and a memory 920. The processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920 may be communicatively connected via a communication bus 930.
[0138] Among them, the processor 910 can be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solution provided in this application.
[0139] The memory 920 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 920 can store an operating system 921 for controlling the operation of the electronic device 900, and a basic input and output system (BIOS) 922 for controlling the low-level operation of the electronic device 900. In addition, a web browser 923, a data storage management system 924, and a dialogue processing device 925, etc. can also be stored. The above-mentioned dialogue processing device 925 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 920 and is called and executed by the processor 910.
[0140] The input / output interface 913 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0141] The network interface 914 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0142] The bus 930 comprises a pathway for transmitting information between the various components of the device (eg, the processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920).
[0143] It should be noted that, although the above device only shows a processor 910, a video display adapter 911, a disk drive 912, an input / output interface 913, a network interface 914, a memory 920, a bus 930, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.
[0144] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0145] The technical solution provided by the present application is described in detail above. The principle and implementation method of the present application are described in detail using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A method for processing a conversation, characterized in that: The method comprises: Get the conversation text; The conversation text is used to search in the knowledge graph to obtain a document block entity matching the conversation text as a target document block entity; wherein the knowledge graph at least includes an association relationship between a document block and a document block entity, and the document block entity is obtained by performing content analysis on its associated document block; Searching for a document block associated with the target document block entity in the knowledge graph to obtain a reference document block; constructing prompt instructions using the dialogue text and the reference document block; The prompt instruction is provided to the dialogue model to obtain a reply to the prompt instruction.
2. The method according to claim 1, characterized in that Using the conversation text to search in the knowledge graph, obtaining a document block entity matching the conversation text as a target document block entity includes: Matching the conversation text with the document block entity in the knowledge graph, and determining the document block entity whose matching degree meets the first requirement as the first document block entity; The target document block entity is obtained by using the first document block entity.
3. The method according to claim 2, characterized in that The knowledge graph also includes an association relationship between a document block entity and a document entity, wherein the document entity is obtained by performing attribute analysis on the document, and the document block is obtained by segmenting the document; Obtaining the target document block entity using the first document block entity includes: Querying the knowledge graph for a document entity associated with the first document block entity as the first document entity; Querying the document block entity associated with the first document entity in the knowledge graph, and deduplicating the queried document block entity based on the first document block entity to obtain a second document block entity; According to the matching degree between the dialogue text and the second document block entity, determining a document block entity whose matching degree meets a second requirement as a third document block entity; The target document block entity is obtained by using the first document block entity and the third document block entity.
4. The method according to claim 3, characterized in that The knowledge graph also includes the association relationship between document entities and domains; Using the first document block entity and the third document block entity, obtaining the target document block entity includes: Querying the fields associated with the first document entity in the knowledge graph, and determining a target field from the fields associated with the first document entity; Querying the knowledge graph for document entities associated with the target domain, and removing duplicate document entities associated with the target domain based on the first document entity to obtain a second document entity; Querying the document block entity associated with the second document entity in the knowledge graph to obtain a fourth document block entity; According to the matching degree between the dialogue text and the fourth document block entity, determining a document block entity whose matching degree meets a third requirement as a fifth document block entity; The first document block entity, the third document block entity and the fifth document block entity are merged and deduplicated to obtain the target document block entity.
5. The method according to claim 4, characterized in that Determining the target domain from the domain associated with the first document entity includes: Sorting the fields associated with the first document entity according to the number of times the fields are hit by the query; According to the ranking result, the first M fields are selected as target fields, where M is a positive integer.
6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Pre-segment the documents in the document library to obtain multiple document blocks; Performing content analysis on each document block to extract document block entities; Perform attribute analysis on each document to extract document entities and relationships between document entities; Performing clustering and domain analysis on the multiple document entities according to the relationships between the document entities to obtain multiple domains; The knowledge graph is constructed using the document block, the document block entity, the document entity and the field.
7. The method according to claim 6, characterized in that The pre-segmentation of the documents in the document library to obtain multiple document blocks includes: The documents in the document library are segmented according to the document structure to obtain multiple preliminary document blocks; The multiple preliminary document blocks are optimized using a large language model to obtain the multiple document blocks; wherein the optimization includes: segmenting the preliminary document blocks whose length exceeds a first threshold based on semantics, and / or merging the multiple preliminary document blocks whose length is less than a second threshold based on semantics, and the first threshold is greater than the second threshold.
8. The method according to claim 6, characterized in that The relationship between the document entities includes a relationship type and a relationship strength, and the relationship type represents the field to which the two document entities belong.
9. A dialogue processing device, characterized in that: The device comprises: A question acquisition unit, configured to acquire a dialogue text; An entity retrieval unit is configured to use the conversation text to search in the knowledge graph to obtain a document block entity matching the conversation text as a target document block entity; wherein the knowledge graph at least includes an association relationship between a document block and a document block entity, and the document block entity is obtained by performing content analysis on its associated document block; A document retrieval unit, configured to search the knowledge graph for a document block associated with the target document block entity to obtain a reference document block; The reply acquisition unit is configured to construct a prompt instruction using the dialogue text and the reference document block, provide the prompt instruction to the dialogue model, and obtain a reply to the prompt instruction.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.