Data processing method and device, equipment, storage medium and computer program product
By segmenting and integrating contents of the input content of the large model, the problem of token exceeding the limit in complex dialogue tasks is solved, the lightweight and semantic integrity of the input content of the model is achieved, and the effect and accuracy of task processing are improved.
Patent Information
- Application Number
- CN202510322760.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-11
AI Technical Summary
When large models deal with complex dialogue tasks such as role deduction and long text induction, the number of tokens exceeds the limit due to the long input content, which affects the task effect and accuracy. Although the existing technologies such as improving token limits and cropping have been alleviated, there are problems such as missing information or slowing interface response speed.
By segmenting multiple rounds of dialogue, dialogue shards are generated, and each shard is fused with the remaining shards according to the target task information to generate dialogue information to ensure the semantic integrity of task execution and the lightweighting of model input content.
While compressing the scale of input content of the model, the complete semantics of multiple rounds of conversations are retained, the effectiveness and accuracy of task processing are improved, and information loss and response speed are avoided.
Smart Images

Figure CN120296117A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as large models, natural language processing, and deep learning. In particular, it relates to a data processing method, apparatus, device, storage medium, and computer program product. Background Art
[0002] In the field of natural language processing, although large models show remarkable levels of intelligence, their inherent context window limitations (such as the 4k - 100k tokens thresholds of common models) can significantly affect task performance when dealing with complex dialogue tasks such as role-playing and long text summarization. Summary of the Invention
[0003] The present disclosure provides a data processing method, apparatus, device, storage medium, and computer program product.
[0004] According to a first aspect of the present disclosure, there is provided a data processing method, including:
[0005] Obtaining model input content, where the model input content includes multi-round conversations and target task information;
[0006] Segmenting the multi-round conversations according to the conversation segment length to obtain a plurality of conversation segments; wherein, any one of the conversation segments includes at least one round of conversation;
[0007] For any target conversation segment among the plurality of conversation segments, according to the target task information, fusing the content of the target conversation segment with the remaining conversation segments to generate conversation information corresponding to the target conversation segment; wherein, the remaining conversation segments are at least one of the plurality of conversation segments other than the target conversation segment;
[0008] Performing a target task on the conversation information corresponding to at least one of the conversation segments according to the target task information to obtain a task execution result of the model input content.
[0009] According to a second aspect of the present disclosure, there is provided a data processing apparatus, including:
[0010] An obtaining module, configured to obtain model input content, where the model input content includes multi-round conversations and target task information;
[0011] A segmentation module, configured to segment the multi-round conversations according to the conversation segment length to obtain a plurality of conversation segments; wherein, any one of the conversation segments includes at least one round of conversation;
[0012] A fusion module, configured to, for any target dialogue fragment among a plurality of dialogue fragments, fuse the content of the target dialogue fragment with the remaining dialogue fragments according to the target task information to generate dialogue information corresponding to the target dialogue fragment; wherein, the remaining dialogue fragments are at least one dialogue fragment among the plurality of dialogue fragments other than the target dialogue fragment;
[0013] A processing module, configured to perform a target task on the dialogue information corresponding to at least one of the dialogue fragments according to the target task information to obtain a task execution result of the model input content.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method as described in the first aspect.
[0018] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the data processing method as described in the first aspect.
[0019] According to a fifth aspect of the present disclosure, there is provided a computer program product, including computer instructions, and the computer instructions, when executed by a processor, implement the steps of the data processing method as described in the first aspect.
[0020] A data processing method, apparatus, device, storage medium, and computer program product provided by the present disclosure have the following beneficial effects:
[0021] Obtain the model input content, where the model input content includes multi-turn conversations and target task information; according to the conversation shard length, slice the multi-turn conversations to obtain multiple conversation shards; wherein, any conversation shard includes at least one turn of conversation; for any target conversation shard among the multiple conversation shards, according to the target task information, fuse the content of the target conversation shard with the remaining conversation shards to generate the conversation information corresponding to the target conversation shard; wherein, the remaining conversation shards are at least one conversation shard among the multiple conversation shards other than the target conversation shard; according to the target task information, perform the target task on the conversation information corresponding to at least one conversation shard to obtain the task execution result of the model input content. In the present disclosure, by slicing the multi-turn conversations in the model input content and fusing the content of each conversation shard with other conversation shards according to the task information, and performing the task based on the generated conversation information, it is possible to effectively compress the scale of the model input content while ensuring that the conversation information used to perform the task retains the complete semantics of the multi-turn conversations, achieving a dual optimization of the light weight and semantic integrity of the model input content, and improving the effect and accuracy of task processing.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0023] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0024] Figure 1 is a schematic flowchart of a data processing method according to the first embodiment of the present disclosure;
[0025] Figure 2 is a schematic flowchart of a data processing method according to the second embodiment of the present disclosure;
[0026] Figure 3 is a schematic flowchart of a data processing method according to the third embodiment of the present disclosure;
[0027] Figure 4 is a schematic flowchart of a data processing method according to the fourth embodiment of the present disclosure;
[0028] Figure 5 is a schematic flowchart of a data processing method according to the fifth embodiment of the present disclosure;
[0029] Figure 6 is a schematic structural diagram of a data processing device according to the sixth embodiment of the present disclosure;
[0030] Figure 7A schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure is shown. Detailed implementation
[0031] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0032] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processing are all carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0033] Large models have demonstrated powerful capabilities in natural language processing tasks. However, when dealing with complex dialogue tasks such as role-playing and long text induction, large models often face the problem of exceeding the token limit due to the excessive length of the input content. Although the maximum input lengths supported by different large models vary, and there are also large models that are good at processing extremely long texts, in actual application scenarios, it is still common for the model input length to exceed its processing capacity limit due to factors such as long role dialogues, large reference document lengths, or excessive numbers of documents.
[0034] In related technologies, for the processing of overly long model input content, there are mainly the following processing solutions:
[0035] I. Increase the input token limit of the large model.
[0036] The input token limit of the model has gradually expanded from 5k in the early stage to 128k or even higher, which significantly alleviates the input token limit problem in related scenarios and enables the model to process larger-scale input content.
[0037] II. Trim the previous context.
[0038] Among them, the previous context refers to the content mentioned or discussed previously in a dialogue or text relative to the current discussion point or question.
[0039] Trimming the previous context specifically means trimming the less relevant content in the previous context to avoid the input token limit. For example, in a multi-turn dialogue system, if the user has asked about multiple topics such as weather and news in the previous few turns of dialogue, and only asks for travel advice in the current turn, then trimming the previous context may be to remove the previous dialogue content about weather and news and only retain the dialogue history directly related to travel advice.
[0040] However, although the maximum number of tokens that the large model can process is gradually increasing, as the number of processed tokens increases, the interface response speed slows down, the answer accuracy decreases, and at the same time, the call cost also rises significantly. Moreover, with the accumulation of the previous context, there is still a possibility of exceeding the token limit. In addition, the large model is prone to missing parts of the overly long previous context, making it difficult to guarantee the quality of task processing.
[0041] Although trimming the previous context can avoid the input token limit, it may cause information loss, thereby affecting the effect and accuracy of task execution.
[0042] In view of the above problems, the present disclosure provides a data processing method, apparatus, device, storage medium, and computer program product. By splitting multi-turn conversations in the model input content and fusing the content of each conversation slice with other conversation slices according to task information, and performing tasks based on the generated conversation information, it is possible to effectively compress the scale of the model input content while ensuring that the conversation information used to perform tasks retains the complete semantics of multi-turn conversations, achieving a dual optimization of the lightweight of the model input content and semantic integrity, and improving the effect and accuracy of task processing.
[0043] The following describes the data processing method, apparatus, device, storage medium, and computer program product according to the embodiments of the present disclosure with reference to the accompanying drawings.
[0044] It should be noted that the execution subject of the data processing method in this embodiment is a data processing device, and the data processing device can be implemented in software and / or hardware and can be configured in an electronic device.
[0045] The embodiments of the present disclosure are mainly applied to the situation where the input content of the large model is extremely long (exceeding the input limit of the large model), including but not limited to the following scenarios:
[0046] Role-playing type: The user has a large number of historical conversation turns with the large model, and the accumulated content is large;
[0047] Document processing type: For example, document summarization, comparison, extraction, statistics, etc., where the number of documents to be used as the previous context is large, or the document content is extremely long, or both.
[0048] Figure 1 It is a schematic flowchart of the data processing method according to the first embodiment of the present disclosure.
[0049] As Figure 1 shown, the data processing method includes:
[0050] Step 101, obtain model input content, where the model input content includes multi-turn conversations and target task information.
[0051] In the embodiments of the present disclosure, a multi-turn conversation refers to a continuous multi-turn communication between a user and a model (or system) during an interaction. Each turn of the conversation may include the user's question, the model's answer, and possible user feedback.
[0052] In the embodiments of the present disclosure, target task information refers to information related to a target task. For example, the target task information may include the explicit requirements input by the user, i.e., the target task (such as scoring a multi-turn conversation, extracting certain words or phrases from a multi-turn conversation, summarizing a multi-turn conversation, etc.), the task type (such as scoring type, extraction type, summary type, etc.), and any information related to the execution of the target task (such as the scoring criteria when scoring a multi-turn conversation, the specified words or phrases when extracting certain words or phrases from a multi-turn conversation, etc.).
[0053] By way of example and not limitation, the model input content may be:
[0054] User: "Hello, I want to inquire about the status of my order."
[0055] Intelligent customer service: "Hello, your order has been shipped and is expected to arrive tomorrow."
[0056] User: "Oh, then I want to ask again, what is your product return and exchange policy?"
[0057] Intelligent customer service: "Our return and exchange policy is... (detailed answer omitted here)"
[0058] User: "Okay, then I also want to know what your company's after-sales service is like?"
[0059] Intelligent customer service: "Our after-sales service is... (detailed answer omitted here)"
[0060] ...
[0061] The above multi-turn conversation is scored according to the following scoring criteria:
[0062] 1. Skills (full score 10 points)
[0063] Whether the question asked by the user is answered [if answered, +5 points]
[0064] Whether the user repeats the same question after getting a reply [if not, +5 points]
[0065] 2. Coherence (full score 10 points)
[0066] Whether the reply is fluent [if satisfied, +5 points]
[0067] Whether the transition is reasonable and coherent [if satisfied, +5 points]
[0068] 3. Objection Handling [If the objection handling issue is resolved well, or the user has no objection handling related issues, +5 points]
[0069] When the user has common doubts, be able to answer well and relieve the user's doubts.
[0070] Examples of objection handling are as follows:
[0071] User: Why is the processing time for returns so long?
[0072] Intelligent Customer Service: We fully understand your concern about the return processing time. To improve the processing efficiency, we have optimized the after-sales service process and strengthened the personnel allocation. At the same time, we have also accelerated the speed of commodity inspection and confirmation and simplified the financial processing process. Nevertheless, since the return processing involves the collaboration of multiple links and departments, it may still take some time. To provide you with a better shopping experience, we promise to do our best to shorten the return processing time and ensure close communication with you during the process.
[0073] 4. Result Judgment (Full score: 5 points)
[0074] Whether the user is satisfied with the reply [If satisfied, +5 points]
[0075] Step 102: According to the length of the multi-round dialogue slice, slice the multi-round dialogue to obtain multiple dialogue slices.
[0076] Among them, any dialogue slice includes at least one round of dialogue.
[0077] In the embodiments of the present disclosure, each dialogue slice must be a complete dialogue turn and cannot be split within one round of dialogue.
[0078] In an optional embodiment, the length of the dialogue slice can be determined based on the number of tokens (character units) of the multi-round dialogue, the number of tokens of the model input upper limit, and the number of tokens occupied by the Prompt template. Thus, for long dialogue content, the dialogue content can be flexibly sliced by combining the model input upper limit to ensure that each dialogue slice has sufficient coherence and integrity and avoid excessive fragmentation.
[0079] In the embodiments of the present disclosure, the length of the dialogue slice is only a reference for slicing the multi-round dialogue. In actual slicing of the multi-round dialogue, since each dialogue slice must be a complete dialogue turn, the length of each dialogue slice obtained by slicing is not necessarily the same as the length of the dialogue slice.
[0080] In an alternative embodiment, there is at least one round of repeated conversation between adjacent dialogue segments. Thus, it can be avoided that different rounds of conversation on the same topic are split into adjacent two dialogue segments during segmentation, resulting in poor task execution results when performing the target task on the dialogue information corresponding to any dialogue segment based on the target task information.
[0081] By way of example and not limitation, when segmenting multi-round conversations, 1 to 2 rounds of overlap can be retained between adjacent dialogue segments, that is, there is 1 to 2 rounds of repeated conversation between adjacent dialogue segments.
[0082] Step 103, for any target dialogue segment among multiple dialogue segments, according to the target task information, fuse the target dialogue segment with the remaining dialogue segments to generate the dialogue information corresponding to the target dialogue segment.
[0083] Wherein, the remaining dialogue segments are at least one dialogue segment among multiple dialogue segments other than the target dialogue segment.
[0084] To avoid the problem of token overrun caused by too long input for the model, in an alternative embodiment, for any target dialogue segment among multiple dialogue segments, the content related to the target task information in the target dialogue segment can be fused with the content of the remaining dialogue segments to obtain the dialogue information corresponding to the target dialogue segment.
[0085] In another alternative embodiment, for any target dialogue segment among multiple dialogue segments, the content of the target dialogue segment can be fused with the content related to the target task information in the remaining dialogue segments to obtain the dialogue information corresponding to the target dialogue segment.
[0086] In another alternative embodiment, for any target dialogue segment among multiple dialogue segments, the content related to the target task information in the target dialogue segment can be fused with the content related to the target task information in the remaining dialogue segments to obtain the dialogue information corresponding to the target dialogue segment.
[0087] Wherein, the content related to the target task information in any dialogue segment can be obtained by summarizing the original content of the dialogue segment according to the target task information.
[0088] Thus, by fusing the content related to the task information in the dialogue segments, not only the amount of dialogue content is effectively reduced, but also a certain integrity of the multi-round conversation is retained as a whole.
[0089] By way of example and not limitation, assuming the model input content is as described in the example of step 101, if the user: "Hello, I want to query the status of my order."
[0090] Intelligent Customer Service: "Hello, your order has been shipped and is expected to arrive tomorrow." This is a dialogue segment (Dialogue Segment 1);
[0091] User: "Oh, then I would like to ask, what is your product's return and exchange policy?"
[0092] Intelligent Customer Service: "Our return and exchange policy is... (detailed answer omitted here)" This is a dialogue segment (Dialogue Segment 2);
[0093] User: "Okay, then I also want to know what your company's after-sales service is like?"
[0094] Intelligent Customer Service: "Our after-sales service is... (detailed answer omitted here)" This is a dialogue segment (Dialogue Segment 3);
[0095] ...
[0096] Then the content related to the task information in Dialogue Segment 1 is
[0097] User: Query the order status.
[0098] Intelligent Customer Service: The user's order has been shipped and is expected to arrive tomorrow.
[0099] The content related to the task information in Dialogue Segment 2 is
[0100] User: Consult the return and exchange policy.
[0101] Intelligent Customer Service: The return and exchange policy is... (summary of the detailed answer on the return and exchange policy).
[0102] The content related to the task information in Dialogue Segment 3 is
[0103] User: Consult the after-sales service.
[0104] Intelligent Customer Service: The after-sales service is... (summary of the detailed answer on the after-sales service).
[0105] ...
[0106] Furthermore, for Dialogue Segment 1, the content related to the task information in Dialogue Segment 1 can be integrated with the content of Dialogue Segment 2, Dialogue Segment 3... to obtain the following dialogue information corresponding to Dialogue Segment 1:
[0107] User: Query the order status.
[0108] Intelligent Customer Service: The user's order has been shipped and is expected to arrive tomorrow.
[0109] User: "Oh, then I would like to ask, what is your product's return and exchange policy?"
[0110] Intelligent Customer Service: "Our return and exchange policy is... (detailed answer omitted here)"
[0111] User: "Okay, then I also want to know what your company's after-sales service is like?"
[0112] Intelligent Customer Service: "Our after-sales service is... (detailed answer omitted here)"
[0113] ……
[0114] Or, fuse the content of dialogue segment 1 with the content related to the task information in dialogue segment 2, the content related to the task information in dialogue segment 3... to obtain the dialogue information corresponding to the following dialogue segment 1:
[0115] User: "Hello, I want to check the status of my order."
[0116] Intelligent Customer Service: "Hello, your order has been shipped and is expected to arrive tomorrow."
[0117] User: Inquire about the return and exchange policy.
[0118] Intelligent Customer Service: The return and exchange policy is... (summary of the detailed answer on the return and exchange policy).
[0119] User: Inquire about after-sales service.
[0120] Intelligent Customer Service: The after-sales service is... (summary of the detailed answer on the after-sales service).
[0121] ……
[0122] Or, fuse the content related to the task information in dialogue segment 1 with the content related to the task information in dialogue segment 2, the content related to the task information in dialogue segment 3... to obtain the dialogue information corresponding to the following dialogue segment 1:
[0123] User: Check the order status.
[0124] Intelligent Customer Service: The user's order has been shipped and is expected to arrive tomorrow.
[0125] User: Inquire about the return and exchange policy.
[0126] Intelligent Customer Service: The return and exchange policy is... (summary of the detailed answer on the return and exchange policy).
[0127] User: Inquire about after-sales service.
[0128] Intelligent Customer Service: The after-sales service is... (summary of the detailed answer on the after-sales service).
[0129] ……
[0130] Step 104: Perform a target task on the dialogue information corresponding to at least one dialogue segment according to the target task information, and obtain the task execution result of the model input content.
[0131] In the embodiments of the present disclosure, for each dialogue segment, corresponding dialogue information will be generated, and the dialogue information corresponding to each dialogue segment retains a certain integrity for the overall multi-round dialogue. Therefore, a target task can be performed on the dialogue information corresponding to at least one dialogue segment according to the target task information to obtain the task execution result of the dialogue information corresponding to at least one dialogue segment. Furthermore, based on the obtained task execution result, the task execution result of the model input content can be determined.
[0132] In an optional embodiment, one task execution result can be selected from the task execution results of the dialogue information corresponding to at least one dialogue segment as the task execution result of the model input content.
[0133] In another optional embodiment, the aggregated result of the task execution results of the dialogue information corresponding to multiple dialogue segments can be used as the task execution result of the model input content.
[0134] In the embodiments of the present disclosure, the model input content is obtained. The model input content includes multi-round dialogue and target task information. According to the dialogue segment length, the multi-round dialogue is segmented to obtain multiple dialogue segments. Any one of the dialogue segments includes at least one round of dialogue. For any target dialogue segment among the multiple dialogue segments, according to the target task information, the target dialogue segment is content-fused with the remaining dialogue segments to generate the dialogue information corresponding to the target dialogue segment. The remaining dialogue segments are at least one dialogue segment other than the target dialogue segment among the multiple dialogue segments. A target task is performed on the dialogue information corresponding to at least one dialogue segment according to the target task information to obtain the task execution result of the model input content. In the present disclosure, by segmenting the multi-round dialogue in the model input content and content-fusing each dialogue segment with other dialogue segments based on the task information, and performing tasks based on the generated dialogue information, it is possible to effectively compress the scale of the model input content while ensuring that the dialogue information used to perform tasks retains the complete semantics of the multi-round dialogue, achieving a dual optimization of the light weight and semantic integrity of the model input content, and improving the effect and accuracy of task processing.
[0135] To clearly illustrate how the dialogue information corresponding to the target dialogue segment is generated in the present disclosure, another data processing method is provided in this embodiment. Figure 2 It is a schematic flowchart of the data processing method provided according to the second embodiment of the present disclosure.
[0136] As Figure 2 shown, the data processing method includes:
[0137] Step 201: Obtain the model input content, which includes multi-turn conversations and target task information.
[0138] Step 202: According to the conversation shard length, split the multi-turn conversations to obtain multiple conversation shards.
[0139] Step 203: For any one of the multiple conversation shards, summarize the content of the original corpus within the conversation shard according to the target task information to obtain the summary content of the conversation shard.
[0140] In the embodiments of the present disclosure, a conversation shard is a conversation segment split from multi-turn conversations and having relatively independent meaning. The original corpus within the conversation shard is the unprocessed original text content within the conversation shard, which contains the communication information between two or more parties of the conversation. The target task information is used to guide the content summary of the original corpus within the conversation shard.
[0141] In an alternative embodiment, when summarizing the content of the original corpus within the conversation shard according to the target task information, the following conditions should be met:
[0142] Accuracy: Ensure that the summary content of the obtained conversation shard accurately reflects the content within the conversation shard;
[0143] Relevance: Ensure that the summary content of the obtained conversation shard only contains content related to the target task information and avoid introducing irrelevant content;
[0144] Conciseness: Try to use short language to express the content related to the target task information within the conversation shard;
[0145] Coherence: Ensure that the summary content of the obtained conversation shard is logically coherent.
[0146] Thus, by summarizing the content of the original corpus within the conversation shard according to the target task information, it can be ensured that the summary content of the obtained conversation shard closely revolves around the target task information and can express the content related to the target task information within the conversation shard in a relatively short length.
[0147] In an alternative embodiment, for any one of the multiple conversation shards, summarize the content of the original corpus related to the target task information within the conversation shard to obtain the summary content of the conversation shard. Thus, only summarizing the original corpus related to the target task information within the conversation shard can accurately focus on the key information in the conversation and avoid the interference of irrelevant information.
[0148] Step 204: For any target dialogue slice, fuse the original corpus within the target dialogue slice with the summary content of other dialogue slices to generate the dialogue information corresponding to the target dialogue slice, or fuse the summary content of the target dialogue slice with the original corpus within other dialogue slices to generate the dialogue information corresponding to the target dialogue slice.
[0149] In the embodiments of the present disclosure, the original corpus within the target dialogue slice is the unprocessed and original dialogue content in the target dialogue slice, which contains the direct communication information between the dialogue participants. By summarizing the content of the original corpus within the target dialogue slice according to the target task information, the summary content of the target dialogue slice can be obtained.
[0150] Among them, the remaining dialogue slices are at least one dialogue slice among the multiple dialogue slices other than the target dialogue slice.
[0151] The summary content of other dialogue slices and the original corpus within other dialogue slices are both for each dialogue slice among the multiple dialogue slices other than the target dialogue slice. The summary content of other dialogue slices includes the summary content of each dialogue slice among the multiple dialogue slices other than the target dialogue slice, and the original corpus within other dialogue slices includes the original corpus within each dialogue slice among the multiple dialogue slices other than the target dialogue slice.
[0152] The summary content of each dialogue slice among the multiple dialogue slices other than the target dialogue slice is obtained by summarizing the content of the original corpus within each dialogue slice among the multiple dialogue slices other than the target dialogue slice according to the target task information.
[0153] The original corpus within each dialogue slice among the multiple dialogue slices other than the target dialogue slice is the unprocessed and original dialogue content in each dialogue slice among the multiple dialogue slices other than the target dialogue slice.
[0154] In the embodiments of the present disclosure, for any target dialogue slice, two different content fusion strategies can be adopted to generate the dialogue information corresponding to the target dialogue slice:
[0155] Strategy 1: Original corpus within the target dialogue slice + Summary content of other dialogue slices
[0156] As an example rather than a limitation, assume that there are N dialogue slices in total. The process of generating dialogue information is as follows:
[0157] for(0 < i <= N)
[0158] {
[0159] The i-th dialogue slice retains the original content;
[0160] The other N-1 dialogue segments except the i-th dialogue segment use the summary content;
[0161] Combine them into a new dialogue content (i.e., the dialogue information corresponding to the i-th dialogue segment);
[0162] i = i + 1;
[0163] }
[0164] Strategy 2: The summary content of the target dialogue segment + the original corpus in other dialogue segments
[0165] As an example rather than a limitation, assume that there are N dialogue segments in total, and the process of generating dialogue information is as follows:
[0166] for(0 < i <= N)
[0167] {
[0168] The i-th dialogue segment uses the summary content;
[0169] The other N-1 dialogue segments except the i-th dialogue segment retain the original content;
[0170] Combine them into a new dialogue content (i.e., the dialogue information corresponding to the i-th dialogue segment);
[0171] i = i + 1;
[0172] }
[0173] Step 205: According to the target task information, perform the target task on the dialogue information corresponding to at least one dialogue segment to obtain the task execution result of the model input content.
[0174] It should be noted that the explanations of Step 201, Step 202, and Step 205 can refer to the relevant descriptions in any embodiment of the present disclosure, and will not be elaborated here.
[0175] In the embodiments of the present disclosure, for any one of a plurality of dialogue segments, according to the target task information, the original corpus within the dialogue segment is summarized to obtain the summary content of the dialogue segment; for any target dialogue segment, the original corpus within the target dialogue segment and the summary contents of other dialogue segments are content-fused to generate the dialogue information corresponding to the target dialogue segment, or the summary content of the target dialogue segment and the original corpus within other dialogue segments are content-fused to generate the dialogue information corresponding to the target dialogue segment. By summarizing the original corpus within the dialogue segment according to the target task information, it can be ensured that the summary content of the obtained dialogue segment closely revolves around the target task information, and at the same time, the content related to the target task information in the dialogue segment can be expressed in a relatively short length. By performing content fusion based on the original corpus within each dialogue segment and the summary content of each dialogue segment, not only the amount of dialogue content is effectively reduced, but also a certain integrity of the multi-round dialogue as a whole is retained.
[0176] To clearly illustrate how to determine the task execution result of the model input content in the present disclosure, another data processing method is provided in this embodiment. Figure 3 It is a schematic flowchart of the data processing method provided in the third embodiment of the present disclosure.
[0177] As Figure 3 shown, the data processing method includes:
[0178] Step 301, obtain the model input content, where the model input content includes multi-round dialogues and target task information.
[0179] Step 302, segment the multi-round dialogues according to the dialogue segment length to obtain a plurality of dialogue segments.
[0180] Step 303, for any target dialogue segment among the plurality of dialogue segments, according to the target task information, content-fuse the target dialogue segment with the remaining dialogue segments to generate the dialogue information corresponding to the target dialogue segment.
[0181] Step 304, for the dialogue information corresponding to any dialogue segment, perform the target task according to the target task information to obtain the task execution result of the dialogue information.
[0182] As an example rather than a limitation, assume that the model input content is as described in the example in step 101, where the target task information includes:
[0183] Score the above multi-round dialogues, and the scoring criteria are as follows:
[0184] 1. Skills (full score 10 points)
[0185] Whether the question asked by the user is replied to [if there is a reply, +5 points]
[0186] Whether the user repeats the same question when getting a response [if not, +5 points]
[0187] 2. Coherence (full score: 10 points)
[0188] Is the response fluent? [if satisfied, +5 points]
[0189] Is the transition from beginning to middle, then to turning point and ending reasonable and coherent? [if satisfied, +5 points]
[0190] 3. Objection handling [if the objection handling problem is solved well, or the user has no problem related to objection handling, +5 points]
[0191] When the user has common doubts, it can answer well and relieve the user's doubts.
[0192] Examples of objection handling are as follows:
[0193] User: Why is the return processing time so long?
[0194] Intelligent customer service: We fully understand your concern about the return processing time. To improve the processing efficiency, we have optimized the after-sales service process and strengthened the personnel allocation. At the same time, we have also accelerated the speed of commodity inspection and confirmation and simplified the financial processing process. Nevertheless, since the return processing involves the cooperation of multiple links and departments, it may still take some time. To bring you a better shopping experience, we promise to do our best to shorten the return processing time and ensure to keep in close communication with you during the process.
[0195] 4. Result judgment (full score: 5 points)
[0196] Is the user satisfied with the response? [if satisfied, +5 points]
[0197] After obtaining the dialogue information corresponding to each dialogue slice, for the dialogue information corresponding to any dialogue slice, based on the above scoring criteria, the dialogue information can be scored to obtain the task execution result of the dialogue information.
[0198] Step 305: Determine the task execution result of the model input content according to the type of the target task and the task execution results of at least one dialogue information.
[0199] In the embodiments of the present disclosure, when the type of the target task is different, the method for determining the task execution result of the model input content based on the task execution results of at least one dialogue information is different.
[0200] In an alternative embodiment, when the target task is a scoring task, the target task execution result is selected from the task execution results of at least one conversation message as the task execution result of the model input content; when the target task is an extraction task, the task execution results of multiple conversation messages are aggregated to obtain the task execution result of the model input content.
[0201] Thus, for a scoring task, selecting the target task execution result (such as the highest value, or average value, or median value, or a value selected according to specific rules) from the task execution results of at least one conversation message as the task execution result of the model input content can directly reflect the overall evaluation situation, improving the effectiveness and accuracy of task processing; for an extraction task, by aggregating the task execution results of multiple conversation messages, the accuracy of task processing can be improved and the robustness of task processing can be enhanced.
[0202] As an example but not limitation, assume there are N conversation shards, then N conversation messages will be generated, and thus N task execution results will be obtained. If the target task is a scoring task, the highest value among the N task execution results can be used as the task execution result of the model input content at this time; if the target task is an extraction task, the union of the N task execution results can be used as the task execution result of the model input content.
[0203] It should be noted that the explanations of steps 301 to 303 can be referred to the relevant descriptions in any embodiment of the present disclosure, and will not be elaborated here.
[0204] In the embodiments of the present disclosure, for the conversation message corresponding to any conversation shard, according to the target task information, the target task is executed on the conversation message to obtain the task execution result of the conversation message; based on the type of the target task and the task execution results of at least one conversation message, the task execution result of the model input content is determined. By executing the target task on the conversation message corresponding to each conversation shard, it provides a certain basis for subsequent determination of the task execution result of the model input content. By determining the task execution result of the model input content according to the type of the target task and combining the task execution results of at least one conversation message, it not only realizes targeted determination of the task result, but also significantly improves the precision of task processing and enhances its robustness.
[0205] To clearly illustrate how multiple conversation shards are obtained in the present disclosure, another data processing method is provided in this embodiment. Figure 4 It is a schematic flowchart of the data processing method provided in the fourth embodiment of the present disclosure.
[0206] As Figure 4 shown, the data processing method includes:
[0207] Step 401: Obtain the model input content, which includes multi-turn conversations and target task information.
[0208] Step 402: Determine the semantic similarity between adjacent conversations in the multi-turn conversations based on the original corpus within the multi-turn conversations.
[0209] In an optional embodiment, for any adjacent conversation in the multi-turn conversations, the semantic similarity of the adjacent conversation can be determined according to the original corpus within the adjacent conversation.
[0210] In the embodiments of the present disclosure, any method for determining semantic similarity can be used to determine the semantic similarity between adjacent conversations in the multi-turn conversations, including but not limited to:
[0211] Calculate the similarity between adjacent conversation texts using text matching algorithms (such as cosine similarity, Jaccard similarity, etc.);
[0212] Apply natural language processing (NLP) techniques (such as named entity recognition, part-of-speech tagging, syntactic analysis, etc.) to preprocess and analyze the original corpus within the adjacent conversations, extract key information (such as entities, relationships, emotions, etc.), and calculate the semantic similarity between adjacent conversations based on this information;
[0213] Use a deep learning model (such as a convolutional neural network (CNN), a recurrent neural network (RNN) and its variants: long short-term memory network (LSTM), etc.) to encode and represent the original corpus within the adjacent conversations, and measure the semantic similarity between adjacent conversations by calculating the distance or similarity between the encoded vectors.
[0214] Step 403: Segment the multi-turn conversations based on the conversation segment length and the semantic similarity between adjacent conversations in the multi-turn conversations to obtain multiple conversation segments.
[0215] In the embodiments of the present disclosure, in order to avoid different turns of conversations on the same topic being segmented into adjacent two conversation segments during segmentation, resulting in poor task execution results when performing the target task on the conversation information corresponding to the conversation segment based on the target task information, in addition to limiting that there is at least one round of repeated conversation between adjacent conversation segments, the semantic similarity between adjacent conversations in the multi-turn conversations can also be added to the reference for segmenting the multi-turn conversations, that is, the multi-turn conversations can be segmented based on two references, namely the conversation segment length and the semantic similarity between adjacent conversations in the multi-turn conversations, to obtain multiple conversation segments.
[0216] In an alternative embodiment, multiple candidate dialogue segments can be obtained by segmenting a multi-turn dialogue according to the length of the dialogue segments; in the case where the semantic similarity between any two adjacent dialogues in the multi-turn dialogue is higher than a set threshold, it is determined that the adjacent dialogues belong to the same dialogue segment; based on the adjacent dialogues determined to belong to the same dialogue segment, the multiple candidate dialogue segments are adjusted to obtain multiple dialogue segments. Thus, if two adjacent dialogue turns are highly similar semantically (i.e., their semantic similarity is higher than the set threshold) but are wrongly segmented into different dialogue segments due to length limitations, they can be adjusted back to the same dialogue segment, thereby effectively maintaining the semantic integrity and coherence of the dialogue.
[0217] Step 404: For any target dialogue segment among the multiple dialogue segments, according to the target task information, fuse the content of the target dialogue segment with the content of the remaining dialogue segments to generate the dialogue information corresponding to the target dialogue segment.
[0218] Step 405: According to the target task information, perform the target task on the dialogue information corresponding to at least one dialogue segment to obtain the task execution result of the model input content.
[0219] It should be noted that the explanations of Step 401, Step 404, and Step 405 can be found in the relevant descriptions in any embodiment of the present disclosure, and will not be elaborated here.
[0220] In the embodiments of the present disclosure, the semantic similarity between adjacent dialogues in a multi-turn dialogue is determined based on the original corpus within the multi-turn dialogue; according to the length of the dialogue segments and the semantic similarity between adjacent dialogues in the multi-turn dialogue, the multi-turn dialogue is segmented to obtain multiple dialogue segments. By combining the two factors of the length of the dialogue segments and the semantic similarity between adjacent dialogues in the multi-turn dialogue to segment the multi-turn dialogue, it can be ensured that the obtained dialogue segments not only meet the length limitations but also maintain semantic integrity.
[0221] To clearly illustrate how the length of the dialogue segments is determined in the present disclosure, another data processing method is provided in this embodiment. Figure 5 It is a schematic flowchart of the data processing method provided in the fifth embodiment of the present disclosure.
[0222] As Figure 5 shown, the data processing method includes:
[0223] Step 501: Obtain the model input content, where the model input content includes a multi-turn dialogue and target task information.
[0224] Step 502: Determine the target number of character units based on the difference between the upper limit of the number of character units of the model input and the number of character units occupied by the prompt template.
[0225] Among them, the number of target character units is the upper limit of the number of character units occupied by any dialogue segment.
[0226] In the embodiments of the present disclosure, the upper limit of the number of tokens (character units) occupied by a dialogue segment can be determined according to the upper limit of the number of tokens input to the model and the number of tokens occupied by the Prompt (prompt) template. Specifically, the difference between the upper limit of the number of tokens input to the model and the number of tokens occupied by the Prompt template can be determined as the upper limit of the number of tokens occupied by the dialogue segment.
[0227] As an example rather than a limitation, assume that the total upper limit of the number of tokens input to the model (note that the output is not included) is L, and the number of tokens occupied by the Prompt template (note that the content of each Prompt is different) is P. Then, the available number of tokens U for the dialogue segment = L - P.
[0228] Step 503, determine the length of the dialogue segment according to the restriction condition of the dialogue segment, the first relationship between the number of target character units and the length of the dialogue segment, and the second relationship between the number of character units of the multi-round dialogue and the length of the dialogue segment.
[0229] Among them, the restriction condition of the dialogue segment is that any dialogue segment includes at least one round of dialogue. Thus, it is ensured that the same round of dialogue will not be split into two dialogue segments, ensuring sufficient coherence and integrity for each dialogue segment.
[0230] Among them, the first relationship is that the ratio of the number of target character units to the length of the dialogue segment is a first set value; the second relationship is that the ratio of the number of character units of the multi-round dialogue to the length of the dialogue segment is a second set value. Thus, through the first relationship, it is ensured that the number of dialogue segments obtained by splitting is moderate, neither too many nor too few. Through the second relationship, it is ensured that the ratio of the length of the dialogue segment to the total number of tokens of the multi-round dialogue is moderate, improving the rationality of dialogue segmentation.
[0231] In the embodiments of the present disclosure, the first set value and the second set value can be the same or different, and no limitation is imposed thereon.
[0232] In an alternative embodiment, the first value range of the length of the dialogue segment can be determined according to the first relationship; the second value range of the length of the dialogue segment can be determined according to the second relationship; the length of the dialogue segment can be determined according to the restriction condition of the dialogue segment, the first value range, and the second value range; wherein, the length of the dialogue segment satisfies the restriction condition of the dialogue segment and is within the first value range and the second value range. Thus, by comprehensively considering various factors to determine the length of the dialogue segment, the dialogue segmentation can be made more reasonable.
[0233] By way of example and not limitation, assume that the total number of tokens in a multi-turn conversation is D, the total number of tokens that the model input can reach (note that this does not include the output) is L, the number of tokens occupied by the Prompt template (note that the content of each Prompt is different) is P, the number of available tokens for a conversation shard is U, and the length of a conversation shard is S. Then, the limiting conditions for the length of a conversation shard (arranged in order of priority) are as follows:
[0234] 1. Each conversation shard must contain a complete conversation turn and cannot be split within one turn. This is the primary limiting condition to ensure that the integrity and coherence of the conversation shard are not broken.
[0235] 2. U / S = 2 - 3
[0236] This limiting condition requires that the number of conversation shards be moderate, neither too many nor too few. If S is too large, the number of conversation shards may be too small, resulting in some conversation shards being too long and potentially exceeding the model's processing capacity; if S is too small, the number of conversation shards may be too many, increasing the complexity of processing.
[0237] 3. D / S = 2 - 5
[0238] This limiting condition requires that the ratio of the length of a conversation shard to the total number of tokens in a multi-turn conversation be moderate. If S is too small, the number of conversation shards will be very large, potentially resulting in overly fragmented conversation shards; if S is too large, the number of conversation shards may be too small to effectively disperse the conversation content.
[0239] From this, a specific length of a conversation shard can be obtained. For example, assume D = 3000 (total number of tokens in a multi-turn conversation), L = 4096 (total number of tokens that the model input can reach), and P = 500 (number of tokens occupied by the Prompt template). Then:
[0240] U = L - P = 4096 - 500 = 3596
[0241] Make a preliminary estimate of the range of S: According to U / S = 2 - 3, the possible range of S is between 1199 □ 1798.
[0242] Consider the integrity of conversation turns: Assume that there are multiple turns in the conversation and the number of tokens in each turn is approximately equal. By examining the conversation content, determine a suitable value of S such that each shard contains a complete conversation turn.
[0243] Comprehensively consider D / S = 2 - 5: In this example, if S takes the lower limit of the preliminary estimated range (such as 1200), then D / S ≈ 2.5, which meets the range requirements; if S takes the upper limit (such as 1800), then D / S ≈ 1.67, slightly lower than the lower limit of the range but still within the acceptable range (adjust according to the specific situation).
[0244] Final determination: Suppose that after comprehensive consideration, it is determined that S = 1500. This value can ensure the integrity of the conversation turns, and also make the number of shards moderate (N = U / S ≈ 2.4, rounded up to 3 shards), and the ratio of the shard length to the total number of characters in the historical conversation is reasonable (D / S ≈ 2).
[0245] It should be noted that if there is at least one round of repeated conversation between adjacent conversation shards, it is necessary to consider that the total number of slices is expected to increase, and the upper limit of D / S needs to be relaxed.
[0246] Step 504: According to the conversation shard length, segment the multi-round conversation to obtain multiple conversation shards.
[0247] Step 505: For any target conversation shard among the multiple conversation shards, according to the target task information, fuse the content of the target conversation shard with the other conversation shards to generate the conversation information corresponding to the target conversation shard.
[0248] Step 506: According to the target task information, perform the target task on the conversation information corresponding to at least one conversation shard to obtain the task execution result of the model input content.
[0249] It should be noted that the explanations of Step 501, Step 504 to Step 506 can be referred to the relevant descriptions in any embodiment of the present disclosure, and will not be elaborated here.
[0250] In the embodiments of the present disclosure, the target character unit quantity is determined according to the difference between the upper limit quantity of the character units input by the model and the character unit quantity occupied by the prompt template; according to the restriction condition of the conversation shard, the first relationship between the target character unit quantity and the conversation shard length, and the second relationship between the character unit quantity of the multi-round conversation and the conversation shard length, the conversation shard length is determined. Thus, for long conversation content, the conversation content can be flexibly segmented by combining the model input upper limit to ensure that each conversation shard has sufficient coherence and integrity, and avoid excessive fragmentation.
[0251] In an alternative embodiment, the model input content includes multi-round conversation and target task information, where the total number of tokens of the multi-round conversation is D, the total number of tokens of the model input upper limit (note that the output is not included) is L, the number of tokens occupied by the Prompt template (note that the content of each Prompt is different) is P, the available number of tokens of the conversation shard is U, and the conversation shard length is S. Then, the restriction conditions (arranged in priority) of the conversation shard length are as follows:
[0252] 1. Each conversation shard must be a complete conversation turn and cannot be split within 1 conversation turn;
[0253] 2. U / S = 2 - 3;
[0254] 3. D / S = 2 - 5.
[0255] Thus, a specific dialogue slice length can be obtained, and based on this length, the multi-turn dialogue can be segmented into N dialogue slices.
[0256] After that, the following processing is performed on the segmented dialogue slices:
[0257] According to the target task information, summarize the N dialogue slices respectively (only summarize the content related to the target task information);
[0258] for (0 < i <= N)
[0259] {
[0260] Keep the original content for the i-th dialogue slice;
[0261] For the other N - 1 dialogue slices except the i-th one, use the summary content;
[0262] Combine them into new dialogue content (i.e., the dialogue information corresponding to the i-th dialogue slice);
[0263] According to the target task information, perform the target task on the new dialogue content according to the Prompt template to obtain the i-th task execution result;
[0264] i = i + 1;
[0265] }
[0266] Or,
[0267] for (0 < i <= N)
[0268] {
[0269] Use the summary content for the i-th dialogue slice;
[0270] Keep the original content for the other N - 1 dialogue slices except the i-th one;
[0271] Combine them into new dialogue content (i.e., the dialogue information corresponding to the i-th dialogue slice);
[0272] According to the target task information, perform the target task on the new dialogue content according to the Prompt template to obtain the i-th task execution result;
[0273] i = i + 1;
[0274] }
[0275] Finally, organize the task execution results: For the N task execution results, determine the task execution results of the model input content according to the type of the target task. Among them, for the scoring-type task, take a specific value (such as the highest value, or the average value, or the median value, or the value selected according to specific rules) of the N task execution results as the task execution results of the model input content; for the extraction-type task, take the union of the N task execution results as the task execution results of the model input content.
[0276] Thus, the scale of the model input content is effectively compressed, and at the same time, a certain integrity of the overall conversation in the model input content is retained, the impact of segmentation on the task results is reduced, and the number of requests to the model is only 2N times, significantly improving the efficiency of task processing.
[0277] In addition, to avoid different rounds of conversations on the same topic being segmented into two adjacent dialogue segments during segmentation, resulting in poor task execution results when performing the target task on the dialogue information corresponding to the dialogue segment based on the target task information, the following optimization methods can be used to further optimize the above solution:
[0278] Optimization method 1: When segmenting, retain 1 to 2 rounds of overlap in adjacent segments. In this way, it is necessary to consider that the total number of segments is expected to increase, and the upper limit of D / S needs to be relaxed.
[0279] Optimization method 2: When segmenting, calculate the semantic similarity of each round of conversation. Try to choose adjacent conversations with large semantic changes for segmentation, and avoid adjacent conversations with similar semantics being segmented.
[0280] To implement the above embodiments, the present disclosure also provides a data processing device.
[0281] Figure 6 It is a schematic structural diagram of a data processing device provided according to the sixth embodiment of the present disclosure.
[0282] As Figure 6 shown, the data processing device includes:
[0283] An acquisition module 601, configured to acquire model input content, where the model input content includes multiple rounds of conversations and target task information;
[0284] A segmentation module 602, configured to segment the multiple rounds of conversations according to the dialogue segment length to obtain multiple dialogue segments; where any dialogue segment includes at least one round of conversation;
[0285] The fusion module 603 is configured to, for any target dialogue slice among a plurality of dialogue slices, fuse the content of the target dialogue slice with the content of the remaining dialogue slices according to the target task information to generate dialogue information corresponding to the target dialogue slice, where the remaining dialogue slices are at least one dialogue slice among the plurality of dialogue slices other than the target dialogue slice;
[0286] The processing module 604 is configured to execute a target task on the dialogue information corresponding to at least one dialogue slice according to the target task information to obtain a task execution result of the model input content.
[0287] As a possible implementation manner of an embodiment of the present disclosure, the fusion module 603 includes:
[0288] A summarization unit is configured to, for any dialogue slice among a plurality of dialogue slices, summarize the content of the original corpus in the dialogue slice according to the target task information to obtain a summary content of the dialogue slice;
[0289] A fusion unit is configured to, for any target dialogue slice, fuse the original corpus in the target dialogue slice with the summary content of other dialogue slices to generate dialogue information corresponding to the target dialogue slice, or fuse the summary content of the target dialogue slice with the original corpus in other dialogue slices to generate dialogue information corresponding to the target dialogue slice.
[0290] As a possible implementation manner of an embodiment of the present disclosure, the summarization unit is further configured to:
[0291] For any dialogue slice, summarize the original corpus related to the target task information in the original corpus in the dialogue slice to obtain a summary content of the dialogue slice.
[0292] As a possible implementation manner of an embodiment of the present disclosure, the processing module 604 includes:
[0293] A processing unit is configured to, for the dialogue information corresponding to any dialogue slice, execute a target task on the dialogue information according to the target task information to obtain a task execution result of the dialogue information;
[0294] A first determination unit is configured to determine a task execution result of the model input content based on the type of the target task and the task execution results of at least one dialogue information.
[0295] As a possible implementation manner of an embodiment of the present disclosure, the first determination unit is further configured to:
[0296] In the case where the target task is a scoring task, select a target task execution result from the task execution results of at least one dialogue information as the task execution result of the model input content;
[0297] In the case where the target task is an extraction task, the task execution results of multiple conversation messages are aggregated to obtain the task execution result of the model input content.
[0298] As a possible implementation manner of the embodiments of the present disclosure, the splitting module 602 includes:
[0299] A second determination unit, configured to determine the semantic similarity of adjacent conversations in the multi-round conversation according to the original corpus in the multi-round conversation;
[0300] A splitting unit, configured to split the multi-round conversation according to the conversation slice length and the semantic similarity of adjacent conversations in the multi-round conversation to obtain multiple conversation slices.
[0301] As a possible implementation manner of the embodiments of the present disclosure, the splitting unit is further configured to:
[0302] Split the multi-round conversation according to the conversation slice length to obtain multiple candidate conversation slices;
[0303] In the case where the semantic similarity of any adjacent conversations in the multi-round conversation is higher than a set threshold, determine that the adjacent conversations belong to the same conversation slice;
[0304] Adjust the multiple candidate conversation slices according to the determined adjacent conversations belonging to the same conversation slice to obtain multiple conversation slices.
[0305] As a possible implementation manner of the embodiments of the present disclosure, the above device further includes:
[0306] A first determination module, configured to determine a target character unit quantity according to the difference between the upper limit quantity of character units input to the model and the quantity of character units occupied by the prompt template; wherein, the target character unit quantity is the upper limit quantity of character units occupied by any conversation slice;
[0307] A second determination module, configured to determine the conversation slice length according to the restriction condition of the conversation slice, the first relationship between the target character unit quantity and the conversation slice length, and the second relationship between the quantity of character units of the multi-round conversation and the conversation slice length.
[0308] As a possible implementation manner of the embodiments of the present disclosure, the second determination module is further configured to:
[0309] Determine a first value range of the conversation slice length according to the first relationship;
[0310] Determine a second value range of the conversation slice length according to the second relationship;
[0311] Determine the dialogue segment length according to the limiting conditions of dialogue segmentation, the first value range, and the second value range; wherein, the dialogue segment length meets the limiting conditions of dialogue segmentation and is within the first value range and the second value range.
[0312] As a possible implementation manner of the embodiment of the present disclosure, the limiting condition of dialogue segmentation is that any dialogue segment includes at least one round of dialogue.
[0313] As a possible implementation manner of the embodiment of the present disclosure, the first relationship is that the ratio of the number of target character units to the dialogue segment length is a first set value;
[0314] The second relationship is that the ratio of the number of character units in multiple rounds of dialogue to the dialogue segment length is a second set value.
[0315] As a possible implementation manner of the embodiment of the present disclosure, there is at least one round of repeated dialogue between adjacent dialogue segments.
[0316] It should be noted that the foregoing explanation of the data processing method also applies to the data processing device of this embodiment, and will not be elaborated here.
[0317] In the embodiment of the present disclosure, obtain the model input content, where the model input content includes multiple rounds of dialogue and target task information; according to the dialogue segment length, segment the multiple rounds of dialogue to obtain multiple dialogue segments; wherein, any dialogue segment includes at least one round of dialogue; for any target dialogue segment among the multiple dialogue segments, according to the target task information, fuse the content of the target dialogue segment with the remaining dialogue segments to generate dialogue information corresponding to the target dialogue segment; wherein, the remaining dialogue segments are at least one dialogue segment other than the target dialogue segment among the multiple dialogue segments; according to the target task information, perform the target task on the dialogue information corresponding to at least one dialogue segment to obtain the task execution result of the model input content. In the present disclosure, by segmenting the multiple rounds of dialogue in the model input content, and fusing the content of each dialogue segment with other dialogue segments according to the task information, and performing the task based on the generated dialogue information, it is possible to effectively compress the scale of the model input content while ensuring that the dialogue information used to perform the task retains the complete semantics of multiple rounds of dialogue, realizing the dual optimization of the lightweight and semantic integrity of the model input content, and improving the effect and accuracy of task processing.
[0318] According to the embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0319] Figure 7FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device 700 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0320] As Figure 7 shown, the electronic device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 702 or a computer program loaded from a storage unit 708 into a RAM (Random Access Memory) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An I / O (Input / Output) interface 705 is also connected to the bus 704.
[0321] A plurality of components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0322] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the data processing method in any other suitable way (e.g., by means of firmware).
[0323] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0324] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0325] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0326] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0327] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0328] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0329] It should be noted that artificial intelligence is a discipline that studies enabling a computer to simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0330] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0331] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A data processing method, comprising: Obtaining model input content, where the model input content includes multi-turn conversations and target task information; Segmenting the multi-turn conversations according to the conversation segmentation length to obtain multiple conversation segments; wherein, any one of the conversation segments includes at least one turn of conversation; For any target conversation segment among the multiple conversation segments, according to the target task information, fusing the content of the target conversation segment with the content of the remaining conversation segments to generate conversation information corresponding to the target conversation segment; wherein, the remaining conversation segments are at least one conversation segment other than the target conversation segment among the multiple conversation segments; Performing a target task on the conversation information corresponding to at least one of the conversation segments according to the target task information to obtain a task execution result of the model input content.
2. The method according to claim 1, wherein, The step of, for any target conversation segment among the multiple conversation segments, according to the target task information, fusing the content of the target conversation segment with the content of the remaining conversation segments to generate conversation information corresponding to the target conversation segment, includes: For any one of the multiple conversation segments, summarizing the content of the original corpus within the conversation segment according to the target task information to obtain the summary content of the conversation segment; For any one of the target conversation segments, fusing the original corpus within the target conversation segment with the summary content of the other conversation segments to generate conversation information corresponding to the target conversation segment, or fusing the summary content of the target conversation segment with the original corpus within the other conversation segments to generate conversation information corresponding to the target conversation segment.
3. The method according to claim 2, wherein, The step of, for any one of the multiple conversation segments, summarizing the content of the original corpus within the conversation segment according to the target task information to obtain the summary content of the conversation segment, includes: For any one of the conversation segments, summarizing the original corpus related to the target task information within the conversation segment to obtain the summary content of the conversation segment.
4. The method according to claim 1, wherein, The step of, according to the target task information, performing a target task on the conversation information corresponding to at least one of the conversation segments to obtain a task execution result of the model input content, includes: For the conversation information corresponding to any one of the conversation segments, performing a target task on the conversation information according to the target task information to obtain a task execution result of the conversation information; Determining the task execution result of the model input content based on the type of the target task and the task execution results of at least one of the conversation information.
5. The method according to claim 4, wherein The step of determining the task execution result of the model input content based on the type of the target task and the task execution results of at least one of the conversation information, includes: In the case where the target task is a scoring task, selecting a target task execution result from the task execution results of at least one of the conversation information as the task execution result of the model input content; In the case where the target task is an extraction task, aggregating the task execution results of the multiple conversation information to obtain the task execution result of the model input content.
6. The method according to claim 1, wherein, Segmenting the multi-turn dialogue according to the dialogue segment length to obtain multiple dialogue segments, including: Determining the semantic similarity of adjacent dialogues in the multi-turn dialogue according to the original corpus in the multi-turn dialogue; Segmenting the multi-turn dialogue according to the dialogue segment length and the semantic similarity of adjacent dialogues in the multi-turn dialogue to obtain multiple dialogue segments.
7. The method according to claim 6, wherein, The segmenting the multi-turn dialogue according to the dialogue segment length and the semantic similarity of adjacent dialogues in the multi-turn dialogue to obtain multiple dialogue segments, including: Segmenting the multi-turn dialogue according to the dialogue segment length to obtain multiple candidate dialogue segments; When the semantic similarity of any adjacent dialogues in the multi-turn dialogue is higher than a set threshold, determining that the adjacent dialogues belong to the same dialogue segment; Adjusting the multiple candidate dialogue segments according to the adjacent dialogues determined to belong to the same dialogue segment to obtain multiple dialogue segments.
8. The method according to any one of claims 1-7, wherein The method further includes: Determining a target number of character units according to the difference between the upper limit of the number of character units input to the model and the number of character units occupied by the prompt template; wherein, the target number of character units is the upper limit of the number of character units occupied by any one of the dialogue segments; Determining the dialogue segment length according to the restriction condition of the dialogue segment, the first relationship between the target number of character units and the dialogue segment length, and the second relationship between the number of character units of the multi-turn dialogue and the dialogue segment length.
9. The method according to claim 8, wherein, The determining the dialogue segment length according to the restriction condition of the dialogue segment, the first relationship between the target number of character units and the dialogue segment length, and the second relationship between the number of character units of the multi-turn dialogue and the dialogue segment length, includes: Determining a first value range of the dialogue segment length according to the first relationship; Determining a second value range of the dialogue segment length according to the second relationship; Determining the dialogue segment length according to the restriction condition of the dialogue segment, the first value range and the second value range; wherein, the dialogue segment length satisfies the restriction condition of the dialogue segment and is within the first value range and the second value range.
10. The method according to claim 8, wherein, The restriction condition of the dialogue segment is that any one of the dialogue segments includes at least one turn of dialogue.
11. The method according to claim 8, wherein The first relationship is that the ratio of the target number of character units to the dialogue segment length is a first set value; The second relationship is that the ratio of the number of character units of the multi-turn dialogue to the dialogue segment length is a second set value.
12. The method according to any one of claims 1-7, wherein There is at least one turn of repeated dialogue between adjacent dialogue segments.
13. A data processing device, including: An acquisition module, configured to acquire model input content, where the model input content includes multi-turn dialogue and target task information; A segmentation module, configured to segment the multi-turn dialogue according to the dialogue segment length to obtain multiple dialogue segments; wherein, any one of the dialogue segments includes at least one turn of dialogue; A fusion module, configured to, for any target dialogue segment among a plurality of dialogue segments, fuse the content of the target dialogue segment with the remaining dialogue segments according to the target task information to generate dialogue information corresponding to the target dialogue segment; wherein the remaining dialogue segments are at least one dialogue segment among the plurality of dialogue segments other than the target dialogue segment. A processing module, configured to perform a target task on the dialogue information corresponding to at least one of the dialogue segments according to the target task information to obtain a task execution result of the model input content.
14. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-12.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.
16. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-12.