Information processing method and equipment

By performing multiple rounds of segmentation on related texts and filtering target text segments based on their relevance, the problem of low accuracy in related text filtering in existing technologies is solved, thereby improving the accuracy and efficiency of generating response information.

CN120910213APending Publication Date: 2025-11-07LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511074754.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing technologies, methods for filtering related text based on query information have low accuracy, resulting in insufficient efficiency and accuracy in generating response information.

Method used

By performing multiple rounds of segmentation on the related text, the target text segment is determined based on the relevance of the text segments, and the response information is generated.

Benefits of technology

It improves the accuracy and efficiency of generating response information, retains content that is helpful in generating response information, and reduces redundancy and omissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910213A_ABST
    Figure CN120910213A_ABST
Patent Text Reader

Abstract

The invention discloses an information processing method and equipment, and the method comprises the steps: obtaining query information and a corresponding associated text; performing multi-round segmentation processing on the associated text to obtain a plurality of target text segments; wherein the (K + 1) th round of segmentation processing process comprises the following steps: carrying out segmentation processing on a plurality of alternative Kth text segments to obtain a plurality of (K + 1) th text segments; according to the relevancy of the multiple (K + 1) th text segments, at least one target Kth text segment or alternative (K + 1) th text segment or target (K + 1) th text segment is determined, and the relevancy represents the relevancy of the text segments and the query information; wherein the plurality of first text segments are obtained by segmenting the associated text in the first round; the multiple target text segments are determined based on a target Kth text segment, and K is a positive integer; and reply information corresponding to the query information is generated based on the target text segment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information processing, and particularly relates to an information processing method and device. BACKGROUND

[0002] A common use of a data processing model based on deep learning technology is that a specific query information is input into the model, and the model generates corresponding reply information based on the query information. Generally, in order to obtain more accurate and comprehensive reply information, at least one text related to the query information can be obtained and queried from multiple channels, and the reply information is generated based on the text.

[0003] For the purpose of improving efficiency, after obtaining multiple texts, the texts can be screened based on the relevance of the texts and the query information, and only the texts with high relevance are retained for generating the reply information. The accuracy of this screening method is low. SUMMARY

[0004] Therefore, the present application provides an information processing method and device.

[0005] The present application provides an information processing method, comprising:

[0006] obtaining query information and corresponding associated text;

[0007] performing segmentation processing on the associated text for multiple rounds to obtain multiple target text segments;

[0008] The segmentation processing procedure of the K+1th round comprises:

[0009] performing segmentation processing on the multiple candidate Kth text segments to obtain multiple K+1th text segments;

[0010] determining at least one target Kth text segment, candidate K+1th text segment or target K+1th text segment according to the relevance of the multiple K+1th text segments, wherein the relevance represents the relevance of the text segment and the query information;

[0011] The multiple first text segments are obtained by performing segmentation on the associated text in the first round; and the multiple target text segments are determined based on the target Kth text segment, wherein K is a positive integer;

[0012] generating reply information corresponding to the query information based on the target text segment.

[0013] Optionally, the determining at least one target Kth text segment, candidate K+1th text segment or target K+1th text segment according to the relevance of the multiple K+1th text segments comprises:

[0014] If the relevance of each of the plurality of K+1 text segments obtained by segmenting the same K text segment does not meet the screening condition, the K text segment to which the plurality of K+1 text segments belong is determined as a target K text segment.

[0015] Optionally, the text length of the K+1 text segment is greater than the minimum segmentation length corresponding to the associated text.

[0016] The determining of the at least one target K text segment or the alternative K+1 text segment or the target K+1 text segment according to the relevance of the plurality of K+1 text segments comprises:

[0017] The K+1 text segment with the relevance meeting the screening condition in the plurality of K+1 text segments is determined as an alternative K+1 text segment.

[0018] Optionally, the text length of the K+1 text segment is the minimum segmentation length corresponding to the associated text.

[0019] The determining of the at least one target K text segment or the alternative K+1 text segment or the target K+1 text segment according to the relevance of the plurality of K+1 text segments comprises:

[0020] The K+1 text segment with the relevance meeting the screening condition in the plurality of K+1 text segments is determined as a target K+1 text segment.

[0021] Optionally, the screening condition comprises at least one of the following:

[0022] The relevance of the K+1 text segment is greater than or equal to a first threshold value.

[0023] The relevance difference of the K+1 text segment is less than or equal to the relevance corresponding to the K+1 text segment.

[0024] The relevance difference of the K+1 text segment is less than or equal to a second threshold value corresponding to the K+1 text segment.

[0025] The relevance difference is the difference between the relevance of the K+1 text segment and the relevance of the K text segment to which the K+1 text segment belongs.

[0026] The second threshold value is negatively correlated with the length of the corresponding K+1 text segment.

[0027] The determining of the at least one target K text segment or the alternative K+1 text segment or the target K+1 text segment according to the relevance of the plurality of K+1 text segments comprises:

[0028] If the relevance difference of each of the plurality of K+1 text segments obtained by the same K text segment is greater than the corresponding relevance of the K+1 text segment or the second threshold, the K text segment to which the plurality of K+1 text segments belong is determined as the target K text segment.

[0029] If the relevance difference of at least one of the plurality of K+1 text segments obtained by the same K text segment is less than or equal to the corresponding relevance of the K+1 text segment or the second threshold, the K+1 text segment whose relevance meets the screening condition is determined as the target K+1 text segment.

[0030] Optionally, the multiple rounds of segmentation processing on the associated text to obtain a plurality of target text segments comprises:

[0031] The associated text is segmented in multiple rounds according to a plurality of length parameters corresponding to the associated text.

[0032] Each round of segmentation processing corresponds to one of the length parameters, and the length of each text segment obtained by each round of segmentation processing is less than or equal to the corresponding length parameter.

[0033] Optionally, the method further comprises:

[0034] According to at least one of the length of the associated text and the length of the query information, the number of length parameters corresponding to the associated text and the numerical value of each length parameter are determined.

[0035] Optionally, the generating of the reply information corresponding to the query information based on the target text segment comprises:

[0036] The target text segments are sequentially spliced and fused according to the order of appearance of the target text segments in the associated text.

[0037] The reply information corresponding to the query information is generated based on the target text.

[0038] The present application also provides an electronic device comprising a memory and one or more processors.

[0039] The memory is used to store a computer program.

[0040] The processor is used to execute the computer program to perform:

[0041] Obtaining query information and corresponding associated text;

[0042] Segmenting the associated text in multiple rounds to obtain a plurality of target text segments;

[0043] The segmentation processing of the K+1th round includes:

[0044] The plurality of candidate Kth text segments are segmented to obtain a plurality of K+1th text segments;

[0045] At least one target Kth text segment, a candidate K+1th text segment, or a target K+1th text segment is determined according to the relevance of the plurality of K+1th text segments, and the relevance represents the relevance of the text segment and the query information.

[0046] The plurality of first text segments are obtained by segmenting the associated text in the first round; and the plurality of target text segments are determined based on the target Kth text segment, and K is a positive integer.

[0047] The reply information corresponding to the query information is generated based on the target text segment. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0049] Figure 1 is a flowchart of an information processing method disclosed by the present application;

[0050] Figure 2 is an example diagram of the multi-round segmentation processing disclosed by the present application;

[0051] Figure 3 is a structural block diagram of an electronic device disclosed by the present application. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0053] In the related art, a commonly used method for screening text is that after obtaining a plurality of associated texts related to the query information based on the query information, the relevance of each associated text and the query information is determined respectively, and screening is performed according to the relevance, and only the associated texts with high relevance are retained, and the reply information is generated based on the retained associated texts.

[0054] The problem of this method is that:

[0055] On the one hand, there may be a considerable amount of redundant text in the retained associated text after screening, which is less relevant to the query information and does not help to generate reply information. Directly generating reply information based on these retained associated text will reduce the efficiency of generating reply information.

[0056] On the other hand, there may be some text in the associated text with low relevance that is screened out, which is highly relevant to the query information and is important for generating reply information. Therefore, directly screening out the associated text with low relevance and not using it to generate reply information may cause the process of generating reply information to miss these important text, making the generated reply information less accurate.

[0057] In view of the above two problems, it can be seen that the above-mentioned text screening method has the problem of low accuracy.

[0058] In order to solve the above-mentioned problems, a more accurate text screening method is provided, please refer to Figure 1 A flowchart of a text processing method provided for the embodiment, which can include the following steps.

[0059] S101, obtaining query information and corresponding associated text.

[0060] In the process of implementing step S101, the query information is obtained, and at least one associated text related to the query information is searched according to the query information.

[0061] In one example, the text content or text segment related to the query information is searched from the knowledge base or database, thereby obtaining at least one associated text related to the query information.

[0062] For example, in the large model knowledge base question and answer scene, the query information can be Query, and the text content related to Query is searched from the knowledge base through the Retrieval Augmented Generation (RAG) module, thereby obtaining at least one associated text related to Query.

[0063] The source of the associated text includes but is not limited to various published papers, various articles, blogs and the like published by media or individuals on the network.

[0064] S102, performing multi-round segmentation processing on the associated text to obtain a plurality of target text segments.

[0065] The process of the K+1th round of segmentation processing in step S102 can include:

[0066] S1011, perform segmentation processing on the plurality of candidate K-th text segments to obtain a plurality of K+1-th text segments.

[0067] In the K+1-th round of segmentation processing performed in step S1011, segmentation processing can be performed on each candidate K-th text segment to obtain a plurality of K+1-th text segments. Each candidate K-th text segment is segmented into at least two K+1-th text segments.

[0068] K is a positive integer, that is, K is greater than or equal to 1. The candidate K-th text segment is determined in the K-th round of segmentation processing. The plurality of first text segments are obtained by segmenting the associated text in the first round.

[0069] In one example, when K=1, the associated text is segmented in the first round of segmentation processing to obtain a plurality of first text segments, and the plurality of candidate first text segments are determined according to the relevance of the first text segments.

[0070] The relevance refers to the relevance of the text segment to the query information. One way to obtain the relevance is to input the text segment and the query information into the rerank model to obtain the relevance of the text segment to the query information. The relevance can be in the form of a score.

[0071] For example, for a certain text segment, the text segment and the query information are input into the rerank model to obtain the relevance of the text segment to the query information.

[0072] S1012, determine at least one target K-th text segment or candidate K+1-th text segment or target K+1-th text segment according to the relevance of the plurality of K+1-th text segments.

[0073] In the implementation of step S1012, the relevance of each K+1-th text segment is obtained, and at least one target K-th text segment or candidate K+1-th text segment or target K+1-th text segment is determined according to the relevance of the plurality of K+1-th text segments.

[0074] One way to obtain the relevance of each K+1-th text segment is to input the K+1-th text segment and the query information into the rerank model to obtain the relevance of the K+1-th text segment.

[0075] S103, generate reply information corresponding to the query information based on the target text segment, and the plurality of target text segments are determined based on the target K-th text segment.

[0076] The target K-th text segment herein can include the target text segment determined in all rounds of segmentation processing.

[0077] As an example, assuming that the method of the embodiment accumulates 4 rounds of segmentation processing on one associated text, and determines at least one target first text segment, at least one target second text segment, at least one target third text segment, and at least one target fourth text segment in sequence according to the above method, the "target Kth text segment" in S103 can include the above target first text segment, target second text segment, target third text segment, and target fourth text segment.

[0078] Further, in the case of multiple associated texts, multiple target text segments corresponding to each associated text can be obtained according to the above method, and the target text segments used to generate the reply information can include each target text segment corresponding to each associated text.

[0079] The embodiment has the beneficial effect that all associated texts corresponding to the query information are subjected to multiple rounds of segmentation processing to obtain multiple target text segments. The reply information corresponding to the query information is generated based on the target text segments. In the process of generating the reply information, the content of each associated text that is helpful to generating the reply information is retained, so that the generated reply information is more accurate.

[0080] Optionally, in an embodiment, determining the at least one target Kth text segment or the alternative K+1th text segment or the target K+1th text segment according to the relevance of the multiple K+1th text segments includes:

[0081] If the relevance of the multiple K+1th text segments segmented from the same alternative Kth text segment does not meet the screening condition, the Kth text segment to which the multiple K+1th text segments belong is determined as the target Kth text segment.

[0082] Specifically, each alternative Kth text segment can be segmented to obtain multiple K+1th text segments.

[0083] For "the multiple K+1th text segments segmented from the same alternative Kth text segment", if the relevance of the multiple K+1th text segments segmented from the same alternative Kth text segment does not meet the preset screening condition, the Kth text segment to which the multiple K+1th text segments belong is determined as the target Kth text segment, and the multiple K+1th text segments can be deleted.

[0084] For example: for a certain alternative Kth text segment, the alternative Kth text segment is segmented to obtain 2 K+1th text segments, and if the relevance of the 2 K+1th text segments does not meet the screening condition, the alternative Kth text segment is determined as the target Kth text segment.

[0085] For example, assuming K=2, the alternative second text segment is the second text segment c; in the third round of segmentation processing, the second text segment c is segmented into the third text segment g and the third text segment h; if the relevance of the third text segment g and the relevance of the third text segment h do not meet the screening condition, the second text segment c to which the third text segment g and the third text segment h belong is determined as the target second text segment, and at this time the third text segment g and the third text segment h can be deleted.

[0086] Optionally, in an embodiment, the text length of the K+1th text segment is greater than the minimum segmentation length corresponding to the associated text.

[0087] According to the relevance of the plurality of K+1th text segments, at least one target Kth text segment or alternative K+1th text segment or target K+1th text segment is determined, comprising:

[0088] The K+1th text segment with the relevance meeting the screening condition in the plurality of K+1th text segments is determined as the alternative K+1th text segment.

[0089] It should be noted that in the process of multiple rounds of segmentation processing, segmentation can be performed according to different segmentation lengths corresponding to the associated text. As the segmentation round increases, the segmentation length used will become smaller and smaller until the minimum segmentation length is reached.

[0090] In the case where the text length of the K+1th text segment is greater than the minimum segmentation length corresponding to the associated text, the K+1th round of segmentation processing is not the last round of segmentation processing. In this case, if there are K+1th text segments with relevance meeting the screening condition in the plurality of K+1th text segments, these K+1th text segments with relevance meeting the screening condition can be determined as alternative K+1th text segments, i.e., the alternative K+1th text segment is the K+1th text segment with relevance meeting the screening condition.

[0091] Correspondingly, if there are K+1th text segments with relevance not meeting the screening condition in the plurality of K+1th text segments, these K+1th text segments with relevance not meeting the screening condition can be deleted.

[0092] For example, assuming K=1, the alternative first text segment includes the first text segment a and the first text segment b; in the second round of segmentation processing, the first text segment a is segmented into the second text segment c and the second text segment d, and the first text segment b is segmented into the second text segment e and the second text segment f.

[0093] The second round of segmentation processing is not the last round of segmentation processing, and the text lengths of the second text segment c, the second text segment d, the second text segment e, and the second text segment f are all greater than the minimum segmentation length.

[0094] The relevance of the second text segment c and the relevance of the second text segment d both satisfy the screening condition, the relevance of the second text segment e satisfies the screening condition, and the relevance of the second text segment f does not satisfy the screening condition.

[0095] Therefore, the second text segment c, the second text segment d, and the second text segment e are all determined as candidate second text segments, and the second text segment f can be deleted.

[0096] Optionally, in an embodiment, the text length of the K+1th text segment is the minimum segmentation length corresponding to the associated text.

[0097] According to the relevance of the plurality of K+1th text segments, at least one target Kth text segment or candidate K+1th text segment or target K+1th text segment is determined, comprising:

[0098] The K+1th text segment with a relevance meeting the screening condition in the plurality of K+1th text segments is determined as a target K+1th text segment.

[0099] In the case where the text length of the K+1th text segment is the minimum segmentation length corresponding to the associated text, the K+1th round of segmentation processing is the last round of segmentation processing. In this case, if there are K+1th text segments with a relevance meeting the screening condition in the plurality of K+1th text segments, these K+1th text segments with a relevance meeting the screening condition can be determined as target K+1th text segments, i.e., the "target K+1th text segment" is the "K+1th text segment with a relevance meeting the screening condition".

[0100] Correspondingly, if there are K+1th text segments with a relevance not meeting the screening condition in the plurality of K+1th text segments, these K+1th text segments with a relevance not meeting the screening condition can be deleted.

[0101] For example: assuming K=2, and assuming that there are a total of 3 rounds of segmentation processing, and the candidate second text segment is the second text segment d; in the third round of segmentation processing, the second text segment d is segmented to obtain the third text segment i and the third text segment k.

[0102] The text length of the third text segment i and the third text segment k is the minimum segmentation length corresponding to the associated text, and the third round of segmentation processing is the last round of segmentation processing.

[0103] The relevance of the third text segment i meets the screening condition, and the relevance of the third text segment k does not meet the screening condition, so the third text segment i is determined as a target third text segment, and the third text segment k can be deleted.

[0104] Optionally, in an embodiment, the above-mentioned screening condition at least includes one of the following:

[0105] The relevance of the K+1th text segment is greater than or equal to a first threshold value, which is denoted as Θ.

[0106] the relevance difference of the K+1th text segment is less than or equal to the relevance corresponding to the K+1th text segment;

[0107] the relevance difference of the K+1th text segment is less than or equal to a second threshold corresponding to the K+1th text segment, which is denoted as Δ (also referred to as relevance decay bias);

[0108] The relevance difference is the difference between the relevance of the K+1th text segment and the relevance of the Kth text segment to which the K+1th text segment belongs. Specifically, the relevance difference is the difference between the relevance of the K+1th text segment and the relevance of the Kth text segment to which the K+1th text segment belongs.

[0109] The second threshold is negatively correlated with the length of the corresponding K+1th text segment.

[0110] In one example, the first threshold Θ can be set to -1, and the initial value of the second threshold Δ can be set to 1.

[0111] The second threshold is negatively correlated with the length of the corresponding K+1th text segment, which means that the longer the K+1th text segment, the smaller the second threshold used to determine whether the relevance of the K+1th text segment meets the screening condition; the shorter the K+1th text segment, the larger the second threshold used to determine whether the relevance of the K+1th text segment meets the screening condition.

[0112] In other words, as the rounds of segmentation processing proceed, the second threshold used to determine whether the K+1th text segment meets the screening condition will gradually increase; that is, the second threshold used in each round of segmentation processing will be larger than the second threshold used in the previous round. This is because even if the text segment is strongly related to the query information, the smaller the text segment is divided, the relevance will show a slight downward trend, so the second threshold usually does not use a fixed threshold.

[0113] The way in which the second threshold increases can be set according to actual conditions and is not limited. In one example, the second threshold increases by 0.1 each time a round of segmentation processing is performed.

[0114] For example, the second threshold used in the second round of segmentation processing is 1, the second threshold used in the third round of segmentation processing is 1.1, and the second threshold used in the fourth round of segmentation processing is 1.2.

[0115] It is worth noting that the screening condition includes at least one of the following: the relevance of the K+1 text segment is greater than or equal to the first threshold value; the relevance difference of the K+1 text segment is less than or equal to the relevance corresponding to the K+1 text segment; the relevance difference of the K+1 text segment is less than or equal to the second threshold value corresponding to the K+1 text segment, that is, the screening condition can be any one of the above three or a combination of any number of them.

[0116] Therefore, "meeting the screening condition" means that the relevance of the K+1 text segment is greater than or equal to the first threshold value, and / or the relevance difference of the K+1 text segment is less than or equal to the relevance corresponding to the K+1 text segment, and / or the relevance difference of the K+1 text segment is less than or equal to the second threshold value corresponding to the K+1 text segment.

[0117] Further, "the relevance of the K+1 text segment is greater than or equal to the first threshold value" can be represented by S K+1 (i) >= Θ, S K+1 (i) represents the relevance of the i-th K+1 text segment obtained after the K+1 round of segmentation.

[0118] "The relevance difference of the K+1 text segment is less than or equal to the second threshold value corresponding to the K+1 text segment" can be represented by S K (i) < S K+1 (i) + Δ, wherein S K (i) is the relevance of the K text segment to which the K+1 text segment belongs.

[0119] Optionally, in an embodiment, determining at least one target K text segment or alternative K+1 text segment or target K+1 text segment according to the relevance of a plurality of K+1 text segments includes:

[0120] If the relevance difference of a plurality of K+1 text segments obtained by the same alternative K text segment is greater than the relevance corresponding to the K+1 text segment or the second threshold value, the K text segment to which the plurality of K+1 text segments belong is determined as the target K text segment.

[0121] If the relevance difference of at least one K+1 text segment obtained by the same alternative K text segment is less than or equal to the relevance corresponding to the K+1 text segment or the second threshold value, the K+1 text segment with relevance meeting the screening condition is determined as the alternative K+1 text segment or the target K+1 text segment.

[0122] It is worth noting that since the relevance output by the rearrangement model is easily affected by the size of the text segment, the following individual cases may occur: a text segment with high relevance is cut and the relevance of the cut text segment is greatly reduced due to the destruction of the context semantics.

[0123] For the above individual cases, the present application provides a processing mode of "if the relevance difference of multiple K+1 text segments obtained by the same candidate K text segment is greater than the relevance corresponding to the K+1 text segment or the second threshold value, determining the K text segment to which the multiple K+1 text segments belong as the target K text segment".

[0124] Specifically, for multiple K+1 text segments obtained by the same candidate K text segment, if the relevance difference of the multiple K+1 text segments is greater than the relevance corresponding to the K+1 text segment, or if the relevance difference of the multiple K+1 text segments is greater than the second threshold value corresponding to the K+1 text segment, the K text segment to which the multiple K+1 text segments belong is determined as the target K text segment, and the multiple K+1 text segments can be deleted.

[0125] Further, "the relevance difference of multiple K+1 text segments is greater than the second threshold value corresponding to the K+1 text segment" can be represented by S K (i) > max [S K+1 (i) | i = 1 ~ n] + Δ, where S K (i) is the relevance of the i-th candidate K text segment, max[] represents taking the maximum value, n is the total number of K+1 text segments obtained by performing K+1 round of segmentation processing on the i-th candidate K text segment, that is, n K+1 text segments are obtained by performing K+1 round of segmentation on the i-th candidate K text segment, S K+1 (i) represents the relevance of the i-th K+1 text segment among the n K+1 text segments, max[SK+1(i) | i = 1 ~ n] represents taking the maximum value among the relevances of the n K+1 text segments, and Δ is the second threshold value corresponding to the K+1 text segment.

[0126] For example: assuming K = 2, the candidate 2 text segment is the 2 text segment d; in the third round of segmentation processing, the 2 text segment d is segmented into the 3 text segment i and the 3 text segment k. The relevance of the 3 text segment i meets the screening condition, and the relevance of the 3 text segment k does not meet the screening condition; however, the relevance difference of the 3 text segment i is greater than the second threshold value corresponding to the 3 text segment, and the relevance difference of the 3 text segment k is also greater than the second threshold value, that is, the 3 text segment i and k meet the condition of "the relevance difference of multiple K+1 text segments is greater than the second threshold value corresponding to the K+1 text segment", so the 2 text segment d is determined as the target 2 text segment, and the 3 text segment i and the 3 text segment k can be deleted.

[0127] If the relevance difference of at least one of the plurality of K+1 text segments is less than or equal to the relevance corresponding to the K+1 text segment, or if the relevance difference of at least one of the plurality of K+1 text segments is less than or equal to the second threshold corresponding to the K+1 text segment, the K+1 text segment whose relevance meets the screening condition is determined as a candidate K+1 text segment or a target K+1 text segment, and the K+1 text segment that does not meet the screening condition can be deleted.

[0128] For example, assuming K=2 and assuming that there are 3 rounds of segmentation processing in total, the candidate 2 text segment is the 2 text segment e; in the 3rd round of segmentation processing (the last round of segmentation processing), the 2 text segment e is segmented into the 3 text segment m and the 3 text segment n.

[0129] The relevance of the 3 text segment m is less than the first threshold, and the relevance difference of the 3 text segment m is greater than the second threshold corresponding to the 3 text segment; the relevance of the 3 text segment n is less than the first threshold, but the relevance difference of the 3 text segment n is less than the second threshold; therefore, the relevance of the 3 text segment m does not meet the screening condition, the relevance of the 3 text segment n meets the screening condition, and the 3 text segments m and n do not meet the condition that the relevance difference of the plurality of K+1 text segments is greater than the second threshold corresponding to the K+1 text segment, so it can be determined that the 3 text segment n is the target 3 text segment, and the 3 text segment m can be directly deleted.

[0130] In addition, in the above example, if the 3rd round of segmentation processing is not the last round of segmentation processing for the associated text, that is, the length of the 3 text segment is greater than the minimum segmentation length corresponding to the associated text, then after the above determination, the 3 text segment n is determined as a candidate 3 text segment, and then the 4th round of segmentation processing is performed.

[0131] Optionally, in an embodiment, the associated text is subjected to multiple rounds of segmentation processing to obtain a plurality of target text segments, including:

[0132] The associated text is subjected to multiple rounds of segmentation processing according to a plurality of length parameters corresponding to the associated text;

[0133] Each round of segmentation processing corresponds to a length parameter (i.e., a segmentation length), and the length of the plurality of text segments obtained by each round of segmentation processing is less than or equal to the corresponding length parameter.

[0134] Specifically, the associated text is processed by multiple rounds of segmentation processing, and the associated text corresponds to multiple length parameters. In each round of segmentation processing, a corresponding length parameter is applied, and the length of the text segment obtained in each round of segmentation processing is less than or equal to the length parameter applied in the round.

[0135] For example, assuming that there are 4 rounds of segmentation processing in total, and the associated text corresponds to multiple length parameters (1024, 512, 256, 128) in turn. In the first round of segmentation processing, the associated text is segmented based on the segmentation length of 1024, and multiple text segments with a length less than or equal to 1024 are obtained as the first text segment.

[0136] In the second round of segmentation processing, for each determined candidate first text segment, the candidate first text segment is segmented into two text segments with a length of 512 as the second text segment.

[0137] In the third round of segmentation processing, for each determined candidate second text segment, the candidate second text segment is segmented to obtain two text segments with a length of 256 as the third text segment.

[0138] In the fourth round of segmentation processing, for each determined candidate third text segment, the candidate third text segment is segmented to obtain two text segments with a length of 128 as the fourth text segment.

[0139] Optionally, in an embodiment, the method further comprises:

[0140] The number of length parameters corresponding to the associated text and the value of each length parameter are determined according to at least one of the length of the associated text and the length of the query information.

[0141] Specifically, for the length parameters applied in each round of segmentation processing, the number of length parameters corresponding to the associated text and the value of each length parameter can be determined according to at least one of the length of the associated text and the length of the query information.

[0142] The number of length parameters corresponds to the number of rounds of segmentation processing, that is, the associated text corresponds to several length parameters, and the associated text is processed by several rounds of segmentation processing. For example, in the foregoing example, the associated text corresponds to 4 length parameters, and the associated text is processed by 4 rounds of segmentation processing.

[0143] The way of determining the number and value of the length parameters can be: determining a maximum value of the length parameter according to the length of the associated text, which can be less than the length of the associated text; determining a minimum value of the length parameter according to the length of the query information, which can be greater than the length of the query information; then adding at least one length parameter at a certain interval between the maximum value of the length parameter and the minimum value of the length parameter, thereby obtaining a plurality of length parameters corresponding to the associated text. The minimum value of the length parameter corresponds to the minimum segmentation length in the foregoing embodiments.

[0144] For example, the length of the associated text obtained by searching is 1900, and the length of the query information obtained is 50, so the maximum value of the length parameter can be determined as 1024, the minimum value of the length parameter can be determined as 128, and then the plurality of length parameters corresponding to the associated text are determined in turn as (1024, 512, 256, 128).

[0145] For another example, the length of the associated text obtained by searching is 3066, and the length of the query information obtained is 150, so the maximum value of the length parameter can be determined as 2048, the minimum value of the length parameter can be determined as 256, and then the plurality of length parameters corresponding to the associated text are determined in turn as (2048, 1024, 512, 256).

[0146] The length of the associated text can be represented by the number of characters contained in the associated text, or represented in other ways, without limitation.

[0147] Among the plurality of associated texts obtained based on a query information, different associated texts can correspond to different numbers of length parameters, and the values of the corresponding length parameters can be different. For example, the length parameters corresponding to the associated text 1 can be (1024, 256, 128), and the length parameters corresponding to the associated text 2 can be (2048, 1024, 512, 128).

[0148] Optionally, in an embodiment, the reply information corresponding to the query information is generated based on the target text segment, comprising:

[0149] According to the appearance order of the target text segment in the associated text, the target text is spliced and fused; the reply information corresponding to the query information is generated based on the target text.

[0150] Specifically, after all the target text segments are obtained through the above various embodiments, all the target text segments are spliced in turn according to the appearance order of each target text segment in the associated text to fuse the target text; the reply information corresponding to the query information is generated based on the target text, thereby realizing the extraction and compression of the associated text.

[0151] For example, the number of characters of the associated text searched from the knowledge base and related to the query information is 4201, and the associated text is processed by multiple rounds of segmentation in the manner given in the above embodiments to obtain all target text segments, which are sequentially spliced according to the order of appearance in the associated text to fuse the target text. The number of characters of the reply information corresponding to the query information generated based on the target text is 1229.

[0152] To better understand the information processing logic of the present application, the following will be illustrated by a specific application example, in which there are a total of 3 rounds of segmentation processing. The example content is as follows: Figure 2

[0153] The first round of segmentation processing is performed on the associated text to obtain two first text segments, which are denoted as first text segment a and first text segment b. The relevance of the first text segment a and the relevance of the first text segment b both satisfy the screening condition, so the first text segment a and the first text segment b are both determined as candidate first text segments.

[0154] The second round of segmentation processing is performed on each candidate first text segment, wherein the first text segment a is segmented into second text segment c and second text segment d, and the first text segment b is segmented into second text segment e and second text segment f.

[0155] The relevance of the second text segment c and the relevance of the second text segment d both satisfy the screening condition, the relevance of the second text segment e satisfies the screening condition, and the relevance of the second text segment f does not satisfy the screening condition. The second text segment c, the second text segment d, and the second text segment f are all determined as candidate second text segments, and the second text segment f can be directly deleted.

[0156] The third round of segmentation processing is performed on each candidate second text segment, wherein the second text segment c is segmented into third text segment g and third text segment h, the second text segment d is segmented into third text segment i and third text segment k, and the second text segment e is segmented into third text segment m and third text segment n.

[0157] The text length of the third text segment is equal to the minimum segmentation length corresponding to the associated text, indicating that the third round of segmentation processing is the last round of segmentation.

[0158] Case one: the relevance of the third text segment g and the relevance of the third text segment h both do not satisfy the screening condition, so the second text segment c to which the third text segment g and the third text segment h belong is determined as a target second text segment, and the third text segment g and the third text segment h can be deleted.

[0159] ​The relevance of the third text segment i meets the screening condition, the relevance of the third text segment k does not meet the screening condition, so the third text segment i is determined as the target third text segment, and the third text segment k can be deleted.

[0160] The relevance of the third text segment m does not meet the screening condition, the relevance of the third text segment n meets the screening condition, and the third text segment n is determined as the target third text segment, and the third text segment m can be deleted.

[0161] Case two: The relevance of the third text segment i meets the screening condition, the relevance of the third text segment k does not meet the screening condition, but the relevance difference of the third text segment i (the difference between the relevance of the third text segment i and the second text segment d) is greater than the second threshold corresponding to the third text segment, and the relevance difference of the third text segment k is also greater than the second threshold, so the second text segment d is determined as the target second text segment, and the third text segment i and the third text segment k can be deleted.

[0162] The relevance of the third text segment m is less than the first threshold, and the relevance difference of the third text segment m is greater than the second threshold; the relevance of the third text segment n is less than the first threshold, but the relevance difference of the third text segment n is less than the second threshold; therefore, the relevance of the third text segment m does not meet the screening condition, the relevance of the third text segment n meets the screening condition, the third text segment n is determined as the target third text segment, and the third text segment m can be deleted.

[0163] It is worth noting that the third round in the above example is the last round of segmentation, but if the text length of the third text segment is greater than the minimum segmentation length corresponding to the associated text (i.e., if the third round is not the last round of segmentation), then "determining the third text segment n as the target third text segment" is changed to "determining the third text segment n as the candidate third text segment", and then the fourth round of segmentation processing is performed on the candidate third text segment (i.e., the third text segment n). Among them, for the third text segment i, although the relevance of the third text segment i meets the screening condition, the second text segment d to which the third text segment i belongs has been determined as the target second text segment, so the third text segment i does not need to be segmented in subsequent rounds.

[0164] It should be noted that in the above example, the target Kth text segment is determined only after the second round of segmentation processing, and only the candidate first text segment is determined after the first round of segmentation processing, without determining the target Kth text segment and the target K+1th text segment.

[0165] The comparison of various indicators of the related technology for compressing text in the present application is shown in Table 1 below.

[0166] Table 1

[0167]

[0168] The related technology in Table 1 for comparison can be a compression technology based on a large language model to compress text, and the performance of the related technology is different according to the used model. The number of characters represents the number of characters of the text before compression, which is equivalent to the length of the associated text in the foregoing embodiment. The time refers to the time used to compress one associated text.

[0169] The number of characters after compression refers to the length of the text after compression. For the method of the application, the number of characters after compression can be understood as the sum of the number of characters of all target text segments corresponding to one associated text.

[0170] The recall rate refers to the accuracy of the reply information generated based on the compressed text. For the method of the application, the recall rate refers to the accuracy of the reply information generated based on the obtained target text segment.

[0171] Among them, model 1 can be Qwen2.5-70B, model 2 can be Qwen2.5-vl, and the rearrangement model can be Bge-m3-reranker.

[0172] As can be seen from the comparison given in Table 1 above, compared with the related technology, the application has a significant advantage in time. Specifically, the compression time of the related technology is affected by the model size and the length of the text after compression, which will exacerbate the waiting time of the compression process, and thus affect the efficiency of generating reply information subsequently. The method of the application has a recall rate close to that of the related technology, and the time is only 11.3% of the time of the related technology. It can be seen that the application can greatly shorten the time used for compression.

[0173] Compared with the model containing fewer parameters, the compression quality of the application is more stable, and will not miss part of the related text information due to hallucination when compressing long text.

[0174] Compared with the additional deployment of the compression model, the application can reuse the rearrangement model to maximize the utilization of computing resources and alleviate the hardware resource shortage of the local service.

[0175] Corresponding to the method embodiment described above, the application embodiment discloses an electronic device, the structure of the electronic device is as shown in Figure 3 It can include a memory 301 and one or more processors 302.

[0176] The memory 301 is used to store a computer program;

[0177] The processor 302 is used to execute the computer program to perform:

[0178] Obtain query information and corresponding associated text;

[0179] The associated text is segmented in multiple rounds to obtain multiple target text segments;

[0180] The segmentation process of the K+1th round includes:

[0181] The multiple candidate Kth text segments are segmented to obtain multiple K+1th text segments;

[0182] According to the relevance of the multiple K+1th text segments, at least one target Kth text segment, candidate K+1th text segment, or target K+1th text segment is determined, and the relevance represents the relevance of the text segment and the query information;

[0183] The multiple first text segments are obtained by segmenting the associated text in the first round; and the multiple target text segments are determined based on the target Kth text segment, and K is a positive integer;

[0184] The reply information corresponding to the query information is generated based on the target text segment.

[0185] Optionally, the processor 302 determines at least one target Kth text segment, candidate K+1th text segment, or target K+1th text segment according to the relevance of the multiple K+1th text segments, including:

[0186] If the relevance of the multiple K+1th text segments segmented from the same candidate Kth text segment does not meet the screening condition, the Kth text segment to which the multiple K+1th text segments belong is determined as the target Kth text segment.

[0187] Optionally, the text length of the K+1th text segment is greater than the minimum segmentation length corresponding to the associated text;

[0188] The processor 302 determines at least one target Kth text segment, candidate K+1th text segment, or target K+1th text segment according to the relevance of the multiple K+1th text segments, including:

[0189] The K+1th text segment with relevance meeting the screening condition in the multiple K+1th text segments is determined as the candidate K+1th text segment.

[0190] Optionally, the text length of the K+1th text segment is the minimum segmentation length corresponding to the associated text;

[0191] The processor 302 determines at least one target Kth text segment, candidate K+1th text segment, or target K+1th text segment according to the relevance of the multiple K+1th text segments, including:

[0192] The K+1th text segment with relevance meeting the screening condition in the multiple K+1th text segments is determined as the target K+1th text segment.

[0193] Optionally, the screening condition comprises at least one of the following:

[0194] the relevance of the K+1th text segment is greater than or equal to a first threshold value;

[0195] the difference value of the relevance of the K+1th text segment is less than or equal to the relevance corresponding to the K+1th text segment;

[0196] the difference value of the relevance of the K+1th text segment is less than or equal to a second threshold value corresponding to the K+1th text segment;

[0197] wherein the difference value is the difference between the relevance of the K+1th text segment and the relevance of the Kth text segment to which the K+1th text segment belongs;

[0198] the second threshold value is negatively correlated with the length of the corresponding K+1th text segment.

[0199] Optionally, the processor 302 determines at least one target Kth text segment, or an alternative K+1th text segment, or a target K+1th text segment, according to the relevance of the plurality of K+1th text segments, comprising:

[0200] if the difference value of the relevance of the plurality of K+1th text segments obtained by the same alternative Kth text segment is greater than the relevance corresponding to the K+1th text segment or the second threshold value, the Kth text segment to which the plurality of K+1th text segments belong is determined as the target Kth text segment;

[0201] if the difference value of the relevance of at least one K+1th text segment obtained by the same alternative Kth text segment is less than or equal to the relevance corresponding to the K+1th text segment or the second threshold value, the K+1th text segment whose relevance meets the screening condition is determined as the alternative K+1th text segment or the target K+1th text segment.

[0202] Optionally, the processor 302 performs multiple rounds of segmentation processing on the associated text to obtain a plurality of target text segments, comprising:

[0203] performing multiple rounds of segmentation processing on the associated text according to a plurality of length parameters corresponding to the associated text;

[0204] wherein each round of segmentation processing corresponds to a length parameter, and the length of the plurality of text segments obtained by each round of segmentation processing is less than or equal to the corresponding length parameter.

[0205] Optionally, the processor 302 can also be configured to:

[0206] determine the number of length parameters and the value of each length parameter corresponding to the associated text according to at least one of the length of the associated text and the length of the query information.

[0207] Optionally, the processor 302 generates reply information corresponding to the query information based on the target text segment, including:

[0208] According to the appearance order of the target text segment in the associated text, the target text is spliced and fused in sequence;

[0209] Generate reply information corresponding to the query information based on the target text.

[0210] The working principle of the electronic device of the embodiment can be referred to the related steps in the information processing method of the foregoing embodiments, and will not be described herein.

[0211] It should be noted that each embodiment in the present specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts between embodiments can be referred to each other.

[0212] For the convenience of description, the above system or device is described as various modules or units in terms of functions. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware in the implementation of the present application.

[0213] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and necessary general hardware platforms. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0214] Finally, it should be noted that in this paper, relationship terms such as first, second, third and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0215] The above merely describes the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. An information processing method comprising: obtaining query information and corresponding associated text; performing a plurality of rounds of segmentation processing on the associated text to obtain a plurality of target text segments; wherein a K+1th round of segmentation processing comprises: performing segmentation processing on a plurality of candidate Kth text segments to obtain a plurality of K+1th text segments; determining at least one target Kth text segment or candidate K+1th text segment or target K+1th text segment according to a relevance of the plurality of K+1th text segments, the relevance representing a relevance of the text segment to the query information; wherein a plurality of 1st text segments are obtained by performing segmentation on the associated text in a 1st round; the plurality of target text segments are determined based on the target Kth text segment, K being a positive integer; generating reply information corresponding to the query information based on the target text segments.

2. The method of claim 1, wherein the determining at least one target Kth text segment or candidate K+1th text segment or target K+1th text segment according to a relevance of the plurality of K+1th text segments comprises: if the relevance of a plurality of K+1th text segments obtained from the same candidate Kth text segment does not meet a screening condition, determining a Kth text segment to which the plurality of K+1th text segments belong as a target Kth text segment.

3. The method of claim 1, wherein a text length of the K+1th text segment is greater than a minimum segmentation length corresponding to the associated text; the determining at least one target Kth text segment or candidate K+1th text segment or target K+1th text segment according to a relevance of the plurality of K+1th text segments comprises: determining a K+1th text segment having a relevance meeting a screening condition from the plurality of K+1th text segments as a candidate K+1th text segment.

4. The method of claim 1, wherein a text length of the K+1th text segment is a minimum segmentation length corresponding to the associated text; the determining at least one target Kth text segment or candidate K+1th text segment or target K+1th text segment according to a relevance of the plurality of K+1th text segments comprises: determining a K+1th text segment having a relevance meeting a screening condition from the plurality of K+1th text segments as a target K+1th text segment.

5. The method of any one of claims 2 to 4, wherein the screening condition comprises at least one of: the relevance of the K+1th text segment is greater than or equal to a first threshold value; a relevance difference of the K+1th text segment is less than or equal to a relevance corresponding to the K+1th text segment; the relevance difference of the K+1th text segment is less than or equal to a second threshold value corresponding to the K+1th text segment; wherein the relevance difference is a difference between the relevance of the K+1th text segment and a relevance of a Kth text segment to which the K+1th text segment belongs; the second threshold value and a length of the corresponding K+1th text segment are negatively correlated.

6. The method of claim 5, wherein the determining at least one target Kth text segment or candidate K+1th text segment or target K+1th text segment according to a relevance of the plurality of K+1th text segments comprises: If the relevance difference of each of the plurality of K+1 text segments obtained by the same K text segment is greater than the corresponding relevance of the K+1 text segment or the second threshold value, the K text segment to which the plurality of K+1 text segments belong is determined as the target K text segment; If the relevance difference of at least one of the plurality of K+1 text segments obtained by the same K text segment is less than or equal to the corresponding relevance of the K+1 text segment or the second threshold value, the K+1 text segment whose relevance meets the screening condition is determined as the target K+1 text segment.

7. The method of claim 1, wherein the plurality of target text segments are obtained by performing a plurality of rounds of segmentation on the associated text, comprising: performing a plurality of rounds of segmentation on the associated text according to a plurality of length parameters corresponding to the associated text; wherein each round of segmentation corresponds to a length parameter, and the length of each text segment obtained by each round of segmentation is less than or equal to the corresponding length parameter.

8. The method of claim 7, further comprising: determining the number of length parameters and the value of each length parameter according to at least one of the length of the associated text and the length of the query information.

9. The method of claim 1, wherein the reply information corresponding to the query information is generated based on the target text segments, comprising: sequentially concatenating and fusing the target text segments according to their appearance order in the associated text; and generating the reply information corresponding to the query information based on the target text.

10. An electronic device, comprising a memory and one or more processors; the memory is configured to store a computer program; the processor is configured to execute the computer program to perform the following steps: obtaining query information and corresponding associated text; performing a plurality of rounds of segmentation on the associated text to obtain a plurality of target text segments; wherein the K+1 round of segmentation process comprises: segmenting a plurality of candidate K text segments to obtain a plurality of K+1 text segments; determining at least one target K text segment, candidate K+1 text segment, or target K+1 text segment based on the relevance of the plurality of K+1 text segments, wherein the relevance represents the relevance of the text segment to the query information; wherein the plurality of first text segments are obtained by performing the first round of segmentation on the associated text, and the plurality of target text segments are determined based on the target K text segment, K being a positive integer; generating reply information corresponding to the query information based on the target text segments. ​ ​ ​ ​ ​ ​