Data processing method and device and electronic equipment
By constructing text block collection and tree index structure, using a large model to summarize historical Q&A information from multiple granularity, the problem of low reply accuracy of Q&A system is solved, efficient and accurate replies and dialogue coherence are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510539525.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-22
AI Technical Summary
When answering new questions, the existing Q&A system has low accuracy in replying, making it difficult to effectively use historical Q&A information to respond efficiently.
By constructing a collection of text blocks, using a large model to summarize historical question-and-answer information from multiple granularity, obtain target text blocks related to the questions to be answered, and generate replies based on these text blocks, including slicing, embedding, clustering and spawning tree-like index structures, achieving efficient retrieval and replies.
It improves the accuracy of the reply and dialogue coherence of the Q&A system, enhances the user experience, and improves the quality and efficiency of the reply through multi-grained generalization and efficient retrieval technology.
Smart Images

Figure CN120354947A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and particularly to a data processing method, apparatus, and electronic device. Background Art
[0002] In existing question-and-answer systems, a new question raised by a user and historical question-and-answer information are usually used as inputs to a question-and-answer model to answer the new question. Summary of the Invention
[0003] In view of this, this application provides a data processing method, apparatus, and electronic device as follows:
[0004] A data processing method includes:
[0005] Obtain a question to be answered of a target user;
[0006] Obtain at least one target text block related to the question to be answered from a set of text blocks corresponding to the target user, where the set of text blocks includes multiple summary text blocks, and different summary text blocks summarize the historical question-and-answer information corresponding to the target user from different granularities;
[0007] Based on the question to be answered and the target text block, guide a large model to generate an answer corresponding to the question to be answered.
[0008] Preferably, the above method further includes:
[0009] Segment the historical question-and-answer information to obtain multiple original text blocks;
[0010] Perform multiple iterations until an iteration termination condition is reached to obtain the set of text blocks;
[0011] Wherein, any one iteration includes:
[0012] Cluster multiple embedded text blocks to obtain at least one set of text block combinations, where the embedded text blocks in the first iteration are the original text blocks, and the embedded text blocks in any iteration other than the first iteration are the summary text blocks obtained in the previous iteration;
[0013] Generate summary text blocks for each of the text block combinations and add them to the set of text blocks.
[0014] Preferably, in the above method, the target text block is a summary text block that meets a relevance condition, and the relevance condition includes: the text similarity with the question to be answered is greater than a similarity threshold.
[0015] Preferably, the above method further includes:
[0016] Taking each of the original text blocks as leaf nodes, and using the summary text block of the text block combination in each iteration as the parent node of the corresponding embedded text block, a tree-like index structure is generated.
[0017] In the above method, preferably, at least one target text block related to the question to be answered is obtained from the set of text blocks corresponding to the target user, including:
[0018] Starting from the root node of the tree-like index structure, perform layer by layer: calculate the text similarity between the question to be answered and each first node; select second nodes from the current layer based on the text similarity, where the second nodes include the first nodes with a text similarity greater than the similarity threshold to the question to be answered;
[0019] Wherein, the first node of the first layer is the root node of the tree-like index structure, and the first nodes of the nth layer include the child nodes of the second nodes of the n - 1th layer; n is greater than 1 and less than N, and N is the maximum traversal layer;
[0020] From each layer of the tree-like index structure, select a first target node, which has the maximum text similarity to the question to be answered;
[0021] Obtain the target text block according to the summary text block corresponding to the first target node.
[0022] In the above method, preferably, at least one target text block related to the question to be answered is obtained from the set of text blocks corresponding to the target user, including:
[0023] Calculate the text similarity between each node in the tree-like index structure and the question to be answered respectively;
[0024] In the tree-like index structure, select a second target node, which has a text similarity greater than the second similarity threshold to the question to be answered;
[0025] Obtain the target text block according to the summary text block corresponding to the second target node.
[0026] In the above method, preferably, calculating the text similarity between the target node and the question to be answered includes:
[0027] Calculate the cosine similarity between the semantic vector of the question to be answered and the semantic vector of the summary text block corresponding to the target node as the text similarity, where the target node is any node.
[0028] In the above method, preferably, based on the question to be answered and the target text block, guiding the large model to generate a reply corresponding to the question to be answered includes:
[0029] Input the question to be answered and the target text block into the large model to obtain the answer to the question to be answered generated by the large model using the knowledge base. The knowledge base includes a private knowledge base and a public knowledge base, and the private knowledge base is the topic knowledge base corresponding to the question to be answered.
[0030] A data processing device includes:
[0031] A question acquisition unit for acquiring a question to be answered of a target user;
[0032] A question summarization unit for acquiring at least one target text block related to the question to be answered from the set of text blocks corresponding to the target user. The set of text blocks includes multiple summary text blocks, and different summary text blocks summarize the historical Q&A information corresponding to the target user at different granularities;
[0033] A question answering unit for guiding the large model to generate an answer corresponding to the question to be answered based on the question to be answered and the target text block.
[0034] An electronic device includes: a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above data processing method.
[0035] As can be seen from the above technical solutions, a data processing method, device, and electronic device disclosed in this application acquire a question to be answered of a target user. At least one target text block related to the question to be answered is acquired from the set of text blocks corresponding to the target user, and based on the question to be answered and the target text block, the large model is guided to generate an answer corresponding to the question to be answered. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 It is a flowchart of a data processing method disclosed in this application;
[0038] Figure 2 It is a flowchart of a method for constructing a set of text blocks disclosed in this application;
[0039] Figure 3 It is a structural diagram of a tree-like index structure disclosed in this application;
[0040] Figure 4Flow diagram of another data processing method disclosed in this application;
[0041] Figure 5 Flow diagram of another data processing method disclosed in this application;
[0042] Figure 6 Flow diagram of another data processing method disclosed in this application;
[0043] Figure 7 Structural diagram of a data processing device disclosed in this application;
[0044] Figure 8 Structural diagram of an electronic device provided in an embodiment of this application. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0046] To solve the problem of low accuracy of question answers in existing question-and-answer systems, an embodiment of this application provides a data processing method. Next, the data processing method in the embodiment of this application will be introduced in detail in conjunction with the accompanying drawings.
[0047] Refer to Figure 1 , Figure 1 which is a flow diagram of a data processing method provided in an embodiment of this application. As Figure 1 shown, the data processing method provided in the embodiment of this application includes steps 101 to 103. Next, these steps will be described in detail respectively.
[0048] Step 101: Obtain the question to be answered of the target user.
[0049] In this embodiment, the target user continuously inputs multiple questions through the interaction interface configured by the question-and-answer system. The question-and-answer system answers the questions one by one according to the input order. The question to be answered is the new question to be replied currently.
[0050] For example, after starting a round of question and answer, the target user has successively input questions 1 to question n-1 according to the requirements. The question-and-answer system has answered questions 1 to question n-1 in sequence. In response to receiving the new question input by the target user, that is, question n, this question n is used as the current question to be answered.
[0051] Step 102: Obtain at least one target text block related to the question to be answered from the text block set corresponding to the target user.
[0052] In this embodiment, the text block set includes multiple summary text blocks, and different summary text blocks summarize the historical question-and-answer information corresponding to the target user from different granularities.
[0053] In this embodiment, the historical question-and-answer information corresponding to the target user at least includes historical questions. Optionally, it may also include historical answers corresponding to the historical questions. For example, the historical question-and-answer information corresponding to the target user is the text obtained by splicing questions 1 to n - 1. Another example is that the historical question-and-answer information corresponding to the target user is the text obtained by splicing questions 1 to n - 1 and the corresponding answers 1 to n - 1.
[0054] In this embodiment, the multiple summary text blocks included in the text block set can summarize the historical question-and-answer information from the macro granularity, meso granularity, and micro granularity respectively. For example, the first type of summary text block is obtained by refining the core theme, main idea, or core conclusion of the historical question-and-answer information and is used to summarize the historical question-and-answer information from the macro granularity. The second type of summary text block is obtained by extracting the structural framework and key supporting points of the historical question-and-answer information and is used to summarize the historical question-and-answer information from the meso granularity. The third type of summary text block is obtained by extracting the specific details or data of the historical question-and-answer information and is used to summarize the historical question-and-answer information from the micro granularity.
[0055] In an alternative embodiment, when obtaining at least one target text block related to the question to be answered from the text block set corresponding to the target user in this step, select the summary text block that meets the relevance condition as the target text block. Optionally, the relevance condition includes: the text similarity with the question to be answered is greater than the similarity threshold, where the text similarity is an index to measure the similarity between two texts and is used to evaluate the similarity of different texts at multiple levels such as semantics, grammar, and vocabulary.
[0056] Step 103: Based on the question to be answered and the target text block, guide the large model to generate an answer corresponding to the question to be answered.
[0057] In this embodiment, the large model is a pre-trained question-and-answer model, which is used to deeply understand the question to be answered and the target text block and generate an answer to the question to be answered.
[0058] As can be seen from the above technical solution, a data processing method disclosed in this application obtains the questions to be answered of the target user. At least one target text block related to the question to be answered is obtained from the text block set corresponding to the target user, and based on the question to be answered and the target text block, the large model is guided to generate a reply corresponding to the question to be answered. Since the text block set includes multiple summary text blocks, different summary text blocks summarize the historical question-and-answer information corresponding to the target user from different granularities. Therefore, under the limited data processing capacity of the large model, by selecting the target text blocks related to the question to be answered from the multi-granularity summarization of the historical question-and-answer information, while streamlining the input information volume of the large model, guiding the large model to make a reply to the question to be answered based on richer historical question-and-answer information, and improving the accuracy of the reply.
[0059] In an alternative embodiment, before step 102 of the data processing method provided in the embodiment of this application, it further includes a construction process of the text block set. Refer to Figure 2 , Figure 2 which is a schematic flow chart of a method for constructing a text block set provided in the embodiment of this application. As Figure 2 shown, this method includes steps 201 to 205, and the following will describe these steps in detail.
[0060] Step 201: Split the historical question-and-answer information corresponding to the target user to obtain multiple original text blocks.
[0061] In this embodiment, the historical question-and-answer information is split into multiple original text blocks by combining a splitting method based on semantic logic and a splitting method based on content.
[0062] Step 202: Embed the multiple original text blocks through an encoder to obtain initial embedded text blocks, and use each initial embedded text block as a leaf node of the tree-like index structure.
[0063] In this embodiment, the encoder is an encoder based on BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture). Each original text block is embedded through the encoder to obtain an initial embedded text block represented by a semantic vector.
[0064] Next, perform multiple iterations until the iteration termination condition is reached to obtain a text block set represented by a tree-like index structure. Among them, any iteration includes steps 203 to 206, as follows:
[0065] Step 203: Cluster the multiple embedded text blocks to obtain at least one group of text block combinations.
[0066] In this embodiment, the embedded text block in the first iteration is the initial embedded text block corresponding to the original text block, and the embedded text block in any iteration other than the first iteration is the summary text block obtained in the previous iteration.
[0067] Optionally, in this step, a soft clustering method based on the Gaussian mixture model is used to group similar embedded text blocks to obtain at least one group of text block combinations.
[0068] Step 204: Generate summary text blocks for each group of text block combinations respectively, and use the summary text blocks as the parent nodes of each embedded text block in the corresponding group of text block combinations.
[0069] In this embodiment, a summary algorithm is used to summarize each group of text block combinations to obtain summary text blocks, and the summary text blocks are used to summarize all the embedded text blocks within the text block combinations.
[0070] Furthermore, the summary text blocks are used as the parent nodes of each embedded text block in the corresponding group of text block combinations and added to the tree-like index structure.
[0071] Step 205: Determine whether the iteration stop condition is reached. If the iteration stop condition is not reached, return to step 203 to perform the next iteration.
[0072] In this embodiment, the iteration stop condition includes reaching the maximum number of iterations or the number of text block combinations reaching a quantity threshold. For example, the quantity threshold is 2, and when the number of text block combinations is 2, the iteration ends.
[0073] Step 206: If the iteration stop condition is reached, stop the iteration and generate a tree-like index structure.
[0074] In this embodiment, the tree-like index structure is used for structuring the text block set. In the tree-like index structure after reaching the iteration stop condition, except for the leaf nodes, different nodes are indexed by semantic vectors, representing different summary text blocks.
[0075] Exemplarily, Figure 3 is a schematic diagram of a tree-like index structure provided by an embodiment of the present application. As Figure 3 shown, text blocks 1 to 5 are the initial embedded text blocks obtained by dividing historical Q&A information, and text blocks 1 to 5 are used as the leaf nodes of the tree-like index structure.
[0076] In the first iteration, the first text block combination obtained through clustering includes text block 1 and text block 2, the second text block combination includes text block 2, text block 3, and text block 4, the third text block combination includes text block 4 and text block 5. The summary text block of the first text block combination obtained through summarization is text block 6, the summary text block of the second text block combination is text block 7, and the summary text block of the third text block combination is text block 8. Further, text block 6 is added as the parent node of text block 1 and text block 2 to the tree-like index structure, text block 7 is added as the parent node of text block 2, text block 3, and text block 4 to the tree-like index structure, and text block 8 is added as the parent node of text block 4 and text block 5 to the tree-like index structure.
[0077] In the second iteration, the fourth text block combination obtained through clustering includes text block 6 and text block 7, the fifth text block combination includes text block 7 and text block 8. The summary text block of the fourth text block combination obtained through summarization is text block 9, and the summary text block of the fifth text block combination is text block 10. Further, text block 9 is added as the parent node of text block 6 and text block 7 to the tree-like index structure, and text block 10 is added as the parent node of text block 7 and text block 8 to the tree-like index structure.
[0078] Since the number of text block combinations obtained through clustering in the second iteration is 2, reaching the quantity threshold, iteration is no longer performed, and the final tree-like index structure is obtained.
[0079] It should be noted that when adding a text block as a node to the tree-like index structure, the semantic vector of the text block is used as the index of the corresponding node.
[0080] In summary, in a data processing method provided by an embodiment of the present application, multiple iterations are performed until an iteration termination condition is reached to obtain a text block set. In one iteration, multiple embedded text blocks are clustered to obtain at least one group of text block combinations, summary text blocks of each text block combination are generated respectively and added to the text block set. It can be seen that different summary text blocks in the text block set summarize the historical question-and-answer information corresponding to the target user from different granularities. Therefore, by using the original text blocks obtained by splitting the historical question-and-answer information as the initial embedded text blocks, clustering and summarization are iteratively performed to realize multi-granularity summarization of the historical question-and-answer information. Based on multiple iterations, different summary text blocks in the text block set summarize the historical question-and-answer information from different granularities, so as to realize multi-granularity summary and generalization of the historical question-and-answer information of the target user.
[0081] Furthermore, in the present application, each original text block is used as a leaf node, and in each iteration, the summary text block of the text block combination is used as the parent node of the corresponding embedded text block to generate a tree-like index structure. Thus, through recursive embedding, clustering, and summarization operations, a tree-like structure with different abstraction levels from bottom to top is constructed, and the set of text blocks is structured using the tree-like index structure to capture high-level and low-level details of historical question-and-answer information and support efficient retrieval methods.
[0082] Based on the above embodiment, step 101, that is, obtaining at least one target text block related to the question to be answered from the set of text blocks corresponding to the target user, can be implemented by retrieving the target text block that meets the relevance condition from the tree-like index structure through a tree traversal retrieval method. Refer to Figure 4 the optional flowchart of a data processing method shown in FIG. The specific implementation process of step 101 includes steps 401 to 404, and these steps will be described in detail below.
[0083] Step 401: Starting from the root node of the tree-like index structure, traverse layer by layer to calculate the text similarity between the question to be answered and each first node.
[0084] In this embodiment, the cosine similarity between the semantic vector of the question to be answered and the semantic vector of the summary text block corresponding to any node, that is, the target node, is calculated as the text similarity.
[0085] In this embodiment, the first node in the first layer is the root node of the tree-like index structure, that is, the topmost node, and the first node in the nth layer includes the children of the second node in the n - 1th layer. Where n is greater than 1 and less than N, and N is the maximum traversal layer.
[0086] Step 402: Select a second node from the current layer based on the text similarity.
[0087] In this embodiment, the second node includes the first node whose text similarity to the question to be answered is greater than the similarity threshold.
[0088] Step 403: Select a first target node from the second nodes in the current layer.
[0089] In this embodiment, the first target node has the maximum text similarity to the question to be answered.
[0090] With Figure 3Taking the example of the tree - like index structure, the first node of the first layer of the tree - like index structure includes text block 9 and text block 10. By calculating the cosine similarity, the text similarity s9 between text block 9 and the text of the question to be answered is obtained, and the text similarity s10 between text block 10 and the text of the question to be answered is obtained. And, let y1 represent the first similarity threshold, and s9 > s10 > y1. Then it is determined that the second node of the first layer includes text block 9 and text block 10, and the first target text block is text block 9.
[0091] The first node of the second layer of the tree - like index structure includes the child nodes of text block 9 and the child nodes of text block 10, that is, text block 6, text block 7, and text block 8. By calculating the cosine similarity, the text similarities between text block 6, text block 7, and text block 8 and the text of the question to be answered are obtained respectively, denoted as s6, s7, and s8 respectively, and s6 > s7 > y1 > s8. Then it is determined that the second node of the second layer includes text block 6 and text block 7, and the first target text block is text block 6.
[0092] Assume that the maximum traversal layer is 3, then the obtained first target text blocks include text block 6 and text block 9, that is, the text blocks with the highest text similarity to the question to be answered are selected as the target text blocks from different layers.
[0093] It should be noted that in other alternative embodiments, the second nodes can also be sorted in descending order according to the text similarity to the question to be answered, and the text blocks corresponding to the first i second nodes are selected as the first target text blocks, where i and N are configured according to the data - processing capabilities of the large model.
[0094] Step 404: Determine whether the traversal end condition is reached. If not, return to execute step 401.
[0095] In this embodiment, the traversal end condition includes that the number of second nodes is 0 or the maximum traversal layer is reached.
[0096] Step 405: Obtain the target text blocks according to the summary text blocks corresponding to the first target nodes of each layer in the tree - like index structure.
[0097] In this embodiment, all the first target text blocks are used as the target text blocks or the first target text blocks are screened to obtain multiple target text blocks that meet the quantity requirements.
[0098] In summary, this solution quickly locates the nodes related to the question to be answered by gradually delving into the tree - like index structure, while retaining the context information, so as to achieve accurate retrieval in complex tasks. Benefiting from the high efficiency of tree traversal, multi - scale understanding ability, and context - awareness characteristics, when the data volume of historical question - answering information is large, by adjusting the first similarity threshold and the traversal end condition, the retrieval calculation amount is reduced, and at the same time, the retrieval efficiency and the flexibility of the retrieval granularity are improved.
[0099] Based on the above embodiments, refer to Figure 5 the flowchart of the data processing method shown in FIG. In step 101, that is, obtaining at least one target text block related to the question to be answered, it can also be implemented by retrieving the target text block that meets the relevance condition from the tree index structure through the retrieval method of folding tree traversal. Refer to Figure 5 the optional flowchart of a data processing method shown in FIG. The specific implementation process of step 101 includes steps 501 to 503, and these steps will be described in detail below.
[0100] Step 501: Calculate the text similarity between each node in the tree index structure and the question to be answered.
[0101] In this embodiment, for any node, that is, the target node, calculate the cosine similarity between the semantic vector of the question to be answered and the semantic vector of the abstract text block corresponding to the target node as the text similarity.
[0102] Step 502: Based on the text similarity, select the second target node in the tree index structure.
[0103] In this embodiment, the text similarity between the second target node and the question to be answered is greater than the second similarity threshold.
[0104] Still taking the Figure 3 example tree index structure as an example, based on the cosine similarity, calculate the text similarity between text blocks 1 to 10 and the question to be answered as s1 to s10. Since s1, s6, s7, and s9 are greater than the second similarity threshold y2, and s2, s3, s4, s5, s8, and s10 are all less than y2, therefore, select text blocks 1, text block 6, text block 7, and text block 9 as the second target nodes.
[0105] Step 503: Obtain the target text block according to the abstract text block corresponding to the second target node.
[0106] In this embodiment, all the second target text blocks are used as the target text blocks or the second target text blocks are screened to obtain multiple target text blocks that meet the quantity requirements.
[0107] In summary, this solution constructs the retrieval condition based on the text similarity according to the folding tree traversal strategy, "flattens" the tree index structure into a single-layer list, and then retrieves the second target text block from it. By simplifying the tree structure, the retrieval efficiency is improved.
[0108] Based on the above respective embodiments, step 103, that is, based on the question to be answered and the target text block, the specific implementation of guiding the large model to generate a reply corresponding to the question to be answered is: inputting the question to be answered and the target text block into the large model to obtain the reply to the question to be answered generated by the large model using the knowledge base.
[0109] In this embodiment, the knowledge base includes a private knowledge base and a public knowledge base. Among them, the private knowledge base is the topic knowledge base corresponding to the question to be answered. For example, if the topic of the question to be answered is identified as medicine, the private knowledge base is the knowledge base in the medical field.
[0110] In this embodiment, when the large model generates a reply to the question to be answered using the knowledge base, it uses the context features of the question to be answered indicated by the target text block and the question features indicated by the question to be answered, and uses the private knowledge base and the public knowledge base to generate a reply to the question to be answered, thereby realizing combining historical question-and-answer information and new questions to enhance the reply quality and dialogue coherence of the question-and-answer system.
[0111] In an optional embodiment, this solution can also support the configuration of trigger conditions. For example, the trigger condition for configuring the data processing method is that the number of continuously input questions reaches the question number threshold m, and the trigger condition for configuring the construction method of the text block set is that the number of continuously input questions reaches m - 1.
[0112] Based on this, Figure 6 An example of a question-and-answer scenario schematic diagram is as Figure 6 shown. The target user has asked m - 1 historical questions in a round of question-and-answer, namely questions 1 to question m - 1. Then, based on the historical questions, a tree-like index structure is generated and saved with reference to the Figure 2 process shown.
[0113] After the user inputs the latest question m, the question-and-answer system uses question m as the question to be answered. Based on the tree-like index structure, with reference to the Figure 4 or Figure 5 process shown, k target text blocks are retrieved, namely text blocks x1 to xk.
[0114] The question m and the k target text blocks are input into the large model to obtain the reply generated by the large model based on the private knowledge base and the public knowledge base. Among them, in order to further improve the personalization and specialization of the private knowledge base, the private knowledge base is the private knowledge base determined by combining the user type of the target user and the question topic of question m.
[0115] After the reply to question m is completed, question m is added to the historical questions, and the tree-like index structure is updated based on the current historical questions.
[0116] By Figure 6As can be seen from the example Q&A scenario, the data processing method provided by the embodiments of the present application utilizes historical Q&A information to enhance the answer quality of the Q&A system, ensure the coherence of the conversation, and improve the user experience. Through the tree-like index structure, the efficiency and accuracy of target text block retrieval are improved.
[0117] An embodiment of the present application also provides a data processing device. Refer to Figure 7 As shown, the data processing device includes:
[0118] A question acquisition unit 701 for acquiring the question to be answered of the target user;
[0119] A question summary unit 702 for obtaining at least one target text block related to the question to be answered from the set of text blocks corresponding to the target user, where the set of text blocks includes multiple summary text blocks, and different summary text blocks summarize the historical Q&A information corresponding to the target user from different granularities;
[0120] A question answering unit 703 for guiding a large model to generate an answer corresponding to the question to be answered based on the question to be answered and the target text block.
[0121] Optionally, the data processing device further includes an information summary unit for:
[0122] Segmenting the historical Q&A information to obtain multiple original text blocks;
[0123] Performing multiple iterations until an iteration termination condition is reached to obtain the set of text blocks;
[0124] Wherein, any one iteration includes:
[0125] Clustering multiple embedded text blocks to obtain at least one set of text block combinations, wherein the embedded text blocks in the first iteration are the original text blocks, and the embedded text blocks in any iteration other than the first iteration are the summary text blocks obtained in the previous iteration;
[0126] Generating summary text blocks for each of the text block combinations and adding them to the set of text blocks.
[0127] Optionally, the target text block is a summary text block that meets the relevance condition, and the relevance condition includes: the text similarity with the question to be answered is greater than a similarity threshold.
[0128] Optionally, the data processing device further includes an information structuring unit for:
[0129] Generating a tree-like index structure with each of the original text blocks as leaf nodes and the summary text blocks of the text block combinations in each iteration as the parent nodes corresponding to the embedded text blocks.
[0130] Optionally, when the question summarization unit is used to obtain at least one target text block related to the question to be answered from the set of text blocks corresponding to the target user, it is specifically used for:
[0131] Starting from the root node of the tree-like index structure, perform layer by layer: calculate the text similarity between the question to be answered and each first node; select second nodes from the current layer, where the second nodes include first nodes with a text similarity greater than the similarity threshold to the question to be answered;
[0132] Wherein, the first nodes of the first layer are the root nodes of the tree-like index structure, and the first nodes of the nth layer include the child nodes of the second nodes of the n - 1th layer; n is greater than 1 and less than N, and N is the maximum traversal layer;
[0133] Select a first target node from each layer of the tree-like index structure, where the first target node has the maximum text similarity to the question to be answered;
[0134] Obtain the target text block according to the summary text block corresponding to the first target node.
[0135] Optionally, when the question summarization unit is used to obtain at least one target text block related to the question to be answered from the set of text blocks corresponding to the target user, it is specifically used for:
[0136] Calculate the text similarity between each node in the tree-like index structure and the question to be answered respectively;
[0137] In the tree-like index structure, select a second target node, where the second target node has a text similarity greater than the second similarity threshold to the question to be answered;
[0138] Obtain the target text block according to the summary text block corresponding to the second target node.
[0139] Optionally, when the question summarization unit is used to calculate the text similarity between the target node and the question to be answered, it is specifically used for:
[0140] Calculate the cosine similarity between the semantic vector of the question to be answered and the semantic vector of the summary text block corresponding to the target node as the text similarity, where the target node is any node.
[0141] Optionally, when the question answering unit is used to guide the large model to generate a reply corresponding to the question to be answered based on the question to be answered and the target text block, it is specifically used for:
[0142] Input the question to be answered and the target text block into the large model to obtain the answer to the question to be answered generated by the large model using the knowledge base. The knowledge base includes a private knowledge base and a public knowledge base, and the private knowledge base is the topic knowledge base corresponding to the question to be answered.
[0143] An electronic device is also provided in an embodiment of the present application. Refer to Figure 8 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, and the like. Figure 8 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiment of the present application.
[0144] As Figure 8 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0145] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a memory card, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0146] An embodiment of the present application also provides a computer program product, including computer-readable instructions, which, when running on an electronic device, cause the electronic device to implement each step of any data processing method provided in the embodiment of the present application.
[0147] In an embodiment of the present application, a computer-readable storage medium is further provided. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement each step of any data processing method provided in the embodiment of the present application.
[0148] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for the present application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0150] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0151] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0152] The embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0153] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0154] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the technical field.
[0155] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data processing method, comprising: Obtaining a question to be answered of a target user; Obtaining at least one target text block related to the question to be answered from a set of text blocks corresponding to the target user, where the set of text blocks includes multiple summary text blocks, and different summary text blocks summarize the historical Q&A information corresponding to the target user at different granularities; Guiding a large model to generate a reply corresponding to the question to be answered based on the question to be answered and the target text blocks.
2. The data processing method according to claim 1, further comprising: Segmenting the historical Q&A information to obtain multiple original text blocks; Performing multiple iterations until an iteration termination condition is reached to obtain the set of text blocks; Wherein, any one iteration includes: Clustering multiple embedded text blocks to obtain at least one set of text block combinations, where the embedded text blocks in the first iteration are the original text blocks, and the embedded text blocks in any iteration other than the first iteration are the summary text blocks obtained in the previous iteration; Generating summary text blocks for each of the text block combinations and adding them to the set of text blocks.
3. The data processing method according to claim 2, wherein the target text block is an abstract text block that meets the relevance condition, and the relevance condition includes: The text similarity with the question to be answered is greater than a similarity threshold.
4. The data processing method according to claim 2, further comprising: Generating a tree-like index structure with each of the original text blocks as leaf nodes and the summary text blocks of the text block combinations in each iteration as the parent nodes corresponding to the embedded text blocks.
5. The data processing method according to claim 4, obtaining at least one target text block related to the question to be answered from the set of text blocks corresponding to the target user, comprising: Starting from the root node of the tree-like index structure, performing layer by layer: calculating the text similarity between the question to be answered and each first node; Selecting second nodes from the current layer based on the text similarity, where the second nodes include first nodes with a text similarity with the question to be answered greater than the similarity threshold; Wherein, the first nodes in the first layer are the root nodes of the tree-like index structure, and the first nodes in the nth layer include the child nodes of the second nodes in the (n - 1)th layer; n is greater than 1 and less than N, and N is the maximum traversal layer; Selecting a first target node from each layer of the tree-like index structure, where the first target node has the maximum text similarity with the question to be answered; Obtaining the target text block according to the summary text block corresponding to the first target node.
6. The data processing method according to claim 4, obtaining at least one target text block related to the question to be answered from the set of text blocks corresponding to the target user, comprising: Calculating the text similarity between each node in the tree-like index structure and the question to be answered respectively; Selecting a second target node in the tree-like index structure, where the second target node has a text similarity with the question to be answered greater than a second similarity threshold; Obtaining the target text block according to the summary text block corresponding to the second target node.
7. The data processing method according to claim 5 or 6, calculating the text similarity between a target node and the question to be answered, comprising: Calculate the cosine similarity between the semantic vector of the question to be answered and the semantic vector of the summary text block corresponding to the target node as the text similarity, where the target node is any node.
8. The method according to claim 1, based on the question to be answered and the target text block, guiding the large model to generate a reply corresponding to the question to be answered, including: Inputting the question to be answered and the target text block into the large model to obtain a reply to the question to be answered generated by the large model using a knowledge base, where the knowledge base includes a private knowledge base and a public knowledge base, and the private knowledge base is a topic knowledge base corresponding to the question to be answered.
9. A data processing device, including: A question acquisition unit for acquiring a question to be answered of a target user; A question summarization unit for acquiring at least one target text block related to the question to be answered from a set of text blocks corresponding to the target user, where the set of text blocks includes multiple summary text blocks, and different summary text blocks summarize the historical Q&A information corresponding to the target user from different granularities; A question reply unit for guiding the large model to generate a reply corresponding to the question to be answered based on the question to be answered and the target text block.
10. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the data processing method according to claim 1.