A data processing method and related apparatus
Patent Information
- Application Number
- CN202510325425.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]然而目前的输出文本的生成过程存在精度不高和耗时较多的问题
[0091] As can be seen from the above technical solution, the retrieved text can be obtained based on the prompt text. The retrieved text is related to the prompt text and has more information. It can serve as knowledge text related to the domain of the prompt text, which helps improve the accuracy of text prediction. Segmenting the retrieved text yields multiple word units. The probabilities of these word units are used to indicate their information entropy. The higher the information entropy, the greater the information content of the word unit. Based on the order of these word units in the retrieved text, multiple initial word groups can be determined. Each initial word group includes at least one word unit. The probability of the initial word group is determined by the probabilities of the included word units. Therefore, the probability of the initial word group reflects the information entropy of the word units in the initial word group, and thus reflects the total information entropy of the initial word group. Based on the probabilities of the initial word groups, target word groups can be selected from them. That is, the total information entropy of the target word group meets certain selection criteria, and it has a large amount of information. It can serve as a summary of the retrieved text, reflecting the key information of the retrieved text. Compared to the retrieved text, the target word group has better redundant information and more key information, which is equivalent to compressing the retrieved text to obtain the target word group. Based on the prompt text and the target phrase, the output text corresponding to the prompt text can be determined. Since the target phrase contains the key information of the search text and has less data compared to the search text, the output text is related to both the prompt text and the search text. This realizes the domain expansion based on the search text during the output text generation process, ensuring high output accuracy. Moreover, the key information of the search text can be used without processing all the content of the search text. While ensuring high output accuracy, it also reduces the time consumption of text generation.
Smart Images

Figure CN122779072A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a data processing method and related apparatus. Background Technology
[0002] Currently, it is possible to provide users with output text related to the prompts, thereby achieving natural language interaction. For example, output text can be generated using a Large Language Model (LLM). LLM is an artificial intelligence model based on deep learning technology, which has powerful language understanding and generation capabilities and has achieved significant results in the field of natural language processing.
[0003] However, the current process for generating output text suffers from low accuracy and excessive time consumption. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides a data processing method and related apparatus. The output text is related to both the prompt text and the search text, enabling domain expansion based on the search text during the output text generation process. This ensures high output accuracy and allows the use of key information from the search text without processing all of its content. While maintaining high output accuracy, this also reduces the time required for text generation.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] On the one hand, this application provides a data processing method, the method comprising:
[0007] The search text corresponding to the prompt text is obtained based on the prompt text search;
[0008] The retrieved text is segmented into multiple word units, and the probabilities of the multiple word units are used to indicate the information entropy of the multiple word units.
[0009] Based on the order of the multiple word units in the search text, multiple initial word groups are determined according to the multiple word units, each initial word group including at least one word unit, and the probability of the initial word group is determined by the probability of the included word units;
[0010] Based on the probability of the initial word groups, a target word group is selected from the plurality of initial word groups;
[0011] Based on the prompt text and the target phrase, determine the output text corresponding to the prompt text.
[0012] Optionally, the step of selecting the target word group from the plurality of initial word groups based on the probability of the initial word group includes:
[0013] Based on the target compression rate of the retrieved text and the probability of the initial word group, determine the probability threshold of the initial word group;
[0014] Based on the probability threshold and the probability of the initial word group, the target word group is selected from the plurality of initial word groups.
[0015] Optionally, before determining the probability threshold of the initial word group based on the target compression rate of the retrieved text and the probability of the initial word group, the method further includes:
[0016] The target compression ratio of the retrieved text is determined based on the text length of the retrieved text, and the target compression ratio is inversely correlated with the text length.
[0017] Optionally, the method further includes:
[0018] The initial model is optimized to obtain a compressed model. The optimization operation includes at least one of model distillation, model quantization, model pruning, and operator fusion.
[0019] The probability of the multiple word units is calculated using the compression model.
[0020] Optionally, the method further includes:
[0021] Determine the text structure information corresponding to the prompt text;
[0022] Determine keywords corresponding to the text structure information from the plurality of word units;
[0023] Set the probability of the keyword to a preset value.
[0024] Optionally, the step of segmenting the retrieved text into multiple word units includes:
[0025] Extract the target text of the target type from the retrieved text;
[0026] The text other than the target text in the searched text is segmented to obtain multiple word units;
[0027] The step of determining the output text corresponding to the prompt text based on the prompt text and the target phrase includes:
[0028] Based on the prompt text, the target phrase, and the target text, determine the output text corresponding to the prompt text.
[0029] Optionally, the step of segmenting the retrieved text into multiple word units includes:
[0030] The retrieved text is divided into multiple text blocks;
[0031] Each of the multiple text blocks is segmented into word units corresponding to the multiple text blocks;
[0032] The method of determining multiple initial word groups based on the order of the multiple word units in the search text includes:
[0033] Based on the word units corresponding to the multiple text blocks and the order of the multiple word units in their respective text blocks, the initial word groups corresponding to the multiple text blocks are determined respectively;
[0034] The step of selecting a target word group from the plurality of initial word groups based on the probability of the initial word groups includes:
[0035] Based on the probability of the initial word groups corresponding to the multiple text blocks, the target word groups corresponding to the multiple text blocks are selected from the initial word groups corresponding to the multiple text blocks.
[0036] Optionally, the initial phrases include first-type phrases and second-type phrases. The step of determining multiple initial phrases based on the order of the multiple word units in the search text includes:
[0037] Based on the order of the multiple word units in the search text, at least two adjacent word units among the multiple word units are merged to obtain the first type of word group;
[0038] Target words that meet preset conditions are determined from the word units and designated as the second type of word group.
[0039] Optionally, determining the output text corresponding to the prompt text based on the prompt text and the target phrase includes:
[0040] A prediction model is obtained by performing optimization operations on a large language model, wherein the optimization operations include at least one of model distillation, model quantization, model pruning, and operator fusion.
[0041] The prediction model determines the output text corresponding to the prompt text based on the prompt text and the target phrase.
[0042] Optionally, before determining the output text corresponding to the prompt text based on the prompt text and the target phrase, the method further includes:
[0043] If the target phrase contains single-sided brackets, then the brackets in the target phrase are corrected.
[0044] Optionally, determining multiple initial word groups based on the multiple word units includes:
[0045] Based on the plurality of word units and their probabilities, a plurality of initial word groups and their probabilities are determined by a composite function.
[0046] On the other hand, this application provides a data processing apparatus, the apparatus comprising:
[0047] The retrieval unit is used to retrieve the retrieval text corresponding to the prompt text based on the prompt text;
[0048] A word segmentation unit is used to segment the searched text into multiple word units, and the probabilities of the multiple word units are used to indicate the information entropy of the multiple word units.
[0049] A merging unit is used to determine multiple initial word groups based on the order of the multiple word units in the search text, wherein each initial word group includes at least one word unit, and the probability of the initial word group is determined by the probability of the included word units;
[0050] The selection unit is used to select a target word group from the plurality of initial word groups based on the probability of the initial word group;
[0051] The output text generation unit is used to determine the output text corresponding to the prompt text based on the prompt text and the target phrase.
[0052] Optionally, the merging unit includes:
[0053] A probability threshold determination unit is used to determine a probability threshold for the initial word group based on the target compression rate of the retrieved text and the probability of the initial word group before filtering the target word group from the plurality of initial word groups according to the probability of the initial word group.
[0054] The merging subunit is used to select the target word group from the plurality of initial word groups based on the probability threshold and the probability of the initial word group.
[0055] Optionally, the device further includes:
[0056] A compression ratio determination unit is configured to determine the target compression ratio of the search text based on the text length of the search text before determining the probability threshold of the initial word group based on the target compression ratio of the search text and the probability of the initial word group, wherein the target compression ratio and the text length are inversely correlated.
[0057] Optionally, the device further includes:
[0058] The first optimization unit is used to perform optimization operations on the initial model to obtain a compressed model. The optimization operations include at least one of model distillation, model quantization, model pruning, and operator fusion.
[0059] A probability calculation unit is used to calculate the probability of the plurality of word units using the compression model.
[0060] Optionally, the device further includes:
[0061] A text structure information determination unit is used to determine the text structure information corresponding to the prompt text;
[0062] A keyword determination unit is used to determine keywords corresponding to the text structure information from the plurality of word units;
[0063] The probability setting unit is used to set the probability of the keyword to a preset value.
[0064] Optionally, the word segmentation unit includes:
[0065] The text extraction unit is used to extract target text of the target type from the retrieved text;
[0066] The word segmentation subunit is used to segment other texts in the searched text, other than the target text, to obtain multiple word units;
[0067] The output text generation unit is specifically used for:
[0068] Based on the prompt text, the target phrase, and the target text, determine the output text corresponding to the prompt text.
[0069] Optionally, the word segmentation unit includes:
[0070] The segmentation unit is used to segment the retrieved text into multiple text blocks;
[0071] The word segmentation unit is used to segment the multiple text blocks into word units corresponding to the multiple text blocks respectively.
[0072] The merging unit is specifically used for:
[0073] Based on the word units corresponding to the multiple text blocks and the order of the multiple word units in their respective text blocks, the initial word groups corresponding to the multiple text blocks are determined respectively;
[0074] The selection unit is specifically used for:
[0075] Based on the probability of the initial word groups corresponding to the multiple text blocks, the target word groups corresponding to the multiple text blocks are selected from the initial word groups corresponding to the multiple text blocks.
[0076] Optionally, the initial phrases include first-type phrases and second-type phrases, and the merging unit includes:
[0077] The first word group determination unit is used to merge at least two adjacent word groups among the multiple word groups according to the order of the multiple word groups in the search text to obtain the first type of word group;
[0078] The second word group determination unit is used to determine target words that meet preset conditions from the word units, and to designate them as the second type of word groups.
[0079] Optionally, the output text generation unit includes:
[0080] The second optimization unit is used to perform optimization operations on the large language model to obtain a prediction model. The optimization operations include at least one of model distillation, model quantization, model pruning, and operator fusion.
[0081] The output text generation subunit is used to determine the output text corresponding to the prompt text based on the prompt text and the target phrase using the prediction model.
[0082] Optionally, the device further includes:
[0083] The bracket correction unit is used to correct the brackets in the target phrase if there are single-sided brackets in the target phrase before determining the output text corresponding to the prompt text based on the prompt text and the target phrase.
[0084] Optionally, the merging unit is specifically used for:
[0085] Based on the plurality of word units and their probabilities, a plurality of initial word groups and their probabilities are determined by a composite function.
[0086] On the other hand, this application provides a computer device, the device including a processor and a memory:
[0087] The memory is used to store computer programs and to transfer the computer programs to the processor;
[0088] The processor is configured to execute the data processing method described above according to instructions in the computer program.
[0089] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program for performing the data processing method described above.
[0090] On the other hand, embodiments of this application provide a computer program product including a computer program, which, when run on a computer device, causes the computer device to perform the data processing method.
[0091] As can be seen from the above technical solution, the retrieved text can be obtained based on the prompt text. The retrieved text is related to the prompt text and has more information. It can serve as knowledge text related to the domain of the prompt text, which helps improve the accuracy of text prediction. Segmenting the retrieved text yields multiple word units. The probabilities of these word units are used to indicate their information entropy. The higher the information entropy, the greater the information content of the word unit. Based on the order of these word units in the retrieved text, multiple initial word groups can be determined. Each initial word group includes at least one word unit. The probability of the initial word group is determined by the probabilities of the included word units. Therefore, the probability of the initial word group reflects the information entropy of the word units in the initial word group, and thus reflects the total information entropy of the initial word group. Based on the probabilities of the initial word groups, target word groups can be selected from them. That is, the total information entropy of the target word group meets certain selection criteria, and it has a large amount of information. It can serve as a summary of the retrieved text, reflecting the key information of the retrieved text. Compared to the retrieved text, the target word group has better redundant information and more key information, which is equivalent to compressing the retrieved text to obtain the target word group. Based on the prompt text and the target phrase, the output text corresponding to the prompt text can be determined. Since the target phrase contains the key information of the search text and has less data compared to the search text, the output text is related to both the prompt text and the search text. This realizes the domain expansion based on the search text during the output text generation process, ensuring high output accuracy. Moreover, the key information of the search text can be used without processing all the content of the search text. While ensuring high output accuracy, it also reduces the time consumption of text generation. Attached Figure Description
[0092] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0093] Figure 1 A schematic diagram illustrating an application scenario of a data processing method provided in an embodiment of this application;
[0094] Figure 2 A flowchart illustrating a data processing method provided in an embodiment of this application;
[0095] Figure 3 A schematic diagram illustrating the output text generation process provided in an embodiment of this application;
[0096] Figure 4 A schematic diagram illustrating another output text generation process provided in an embodiment of this application;
[0097] Figure 5 A schematic diagram of a compression process provided in an embodiment of this application;
[0098] Figure 6 This is a schematic diagram of another compression process provided in an embodiment of this application;
[0099] Figure 7 This is a schematic diagram illustrating a configuration provided in an embodiment of this application;
[0100] Figure 8 A schematic diagram illustrating the computation time provided in an embodiment of this application;
[0101] Figure 9 A schematic diagram illustrating an initial word group and its probability provided in an embodiment of this application;
[0102] Figure 10 This is a schematic diagram illustrating the data processing effect provided in an embodiment of this application;
[0103] Figure 11 A structural block diagram of a data processing apparatus provided in an embodiment of this application;
[0104] Figure 12 A structural diagram of a terminal device provided in an embodiment of this application;
[0105] Figure 13 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation
[0106] The embodiments of this application will now be described with reference to the accompanying drawings.
[0107] Currently, it's possible to provide users with relevant output text based on their prompts, enabling natural language interaction. However, the current output text generation process suffers from low accuracy and excessive time consumption.
[0108] To address the aforementioned technical problems, this application provides a data processing method and related apparatus. The output text is related to both the prompt text and the search text, enabling domain expansion based on the search text during the output text generation process. This ensures high output accuracy and allows the use of key information from the search text without processing all of its content. While maintaining high output accuracy, this also reduces the time required for text generation.
[0109] The data processing method provided in this application can be implemented using computer equipment, which can be a terminal device or a server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminal devices include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. Terminal devices and servers can be directly or indirectly connected via wired or wireless communication, and this application does not impose any limitations on this connection.
[0110] To facilitate understanding of the technical solutions provided in this application, the following section will introduce a data processing method provided in an embodiment of this application, in conjunction with a practical application scenario.
[0111] Figure 1 This illustration shows an application scenario of a data processing method provided in an embodiment of this application. The scenario includes a server 10 and a terminal device 20. The terminal device 20 has an application program installed for data processing. The server 10 and the terminal device 20 interact via a network. The server 10 or the terminal device 20 can function as a computer device, used to generate output text based on prompt text. The terminal device 20 interacts with the user, obtains the prompt text, and displays the output text. The following description uses the server 10 as an example of a computer device.
[0112] Server 10 can retrieve the corresponding search text based on the prompt text. The search text is related to the prompt text and has more information. It can be used as knowledge text related to the domain of the prompt text, which helps to improve the accuracy of text prediction.
[0113] Server 10 segments the search text to obtain multiple word units. The probabilities of these word units indicate their information entropy; higher information entropy indicates greater information content. Based on the order of these word units in the search text, multiple initial word groups can be determined. Each initial word group includes at least one word unit, and its probability is determined by the probabilities of its included word units. Therefore, the probability of an initial word group reflects the information entropy of its word units, and consequently, its total information entropy. Based on the probabilities of the initial word groups, target word groups can be selected. These target word groups, whose total information entropy meets certain selection criteria, possess significant information content and can serve as a summary of the search text, reflecting its key information. Compared to the search text, target word groups have better redundancy and more key information, essentially compressing the search text to obtain the target word groups.
[0114] Based on the prompt text and the target phrase, server 10 can determine the output text corresponding to the prompt text. Since the target phrase contains the key information of the search text and has less data than the search text, the output text is related to both the prompt text and the search text. This realizes the domain expansion based on the search text during the output text generation process, ensuring high output accuracy. Moreover, the key information of the search text can be used without processing all the content of the search text. While ensuring high output accuracy, it also reduces the time consumption of text generation.
[0115] Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of this application. In this embodiment, a server is used as the aforementioned computer device for description. The data processing method may include:
[0116] S101, obtain the search text corresponding to the prompt text based on the prompt text search.
[0117] In this embodiment, prompt text from the user can be obtained. This prompt text indicates the answer the user needs. Output text can be generated based on the prompt text, and a response to the user is provided through the output text to complete natural language interaction. Specifically, a large language model can be used to generate the output text based on the prompt text. Large language models, such as GPT-3 and LLaMA, are trained on massive amounts of data and use a large number of parameters to achieve language understanding and output. The number of parameters can be in the billions. These models can be applied to tasks such as answering questions, translating languages, and completing sentences, enabling them to be applied to natural language interaction in various fields, including chatbots and personal assistants.
[0118] In the process of generating output text based on prompt text, the generation of output text can be optimized through Retrieval-Augmented Generation (RAG). Specifically, refer to... Figure 3 The diagram illustrates a process for generating output text according to an embodiment of this application. First, a search text corresponding to the prompt text is obtained based on the prompt text. Then, the search text guides the generation of output text, thus providing a response to the prompt text. The search text is related to the prompt text and contains more information; it can serve as knowledge text related to the domain of the prompt text, adding more information to the generation of output text and helping to improve the accuracy of text prediction. The search text can come from a database, which may include text data from various domains or only text data of knowledge within a specific domain or organization. This text data is added to the database as original knowledge documents before the prompt text is obtained, so that the search text corresponding to the prompt text can be retrieved from the database.
[0119] In this way, in scenarios where output text is generated through a large language model, the large language model can reference other authoritative knowledge bases outside of the training data. Based on the already powerful functions of the large language model, RAG can extend it to access knowledge bases within a specific domain or organization without retraining the model. This improves the large language model cost-effectively, enabling it to maintain relevance, accuracy, and usability in various contexts.
[0120] S102, the retrieved text is segmented into multiple word units, and the probabilities of the multiple word units are used to indicate the information entropy of the multiple word units.
[0121] After obtaining the search text, output text can be generated based on the prompt text and the search text. However, directly processing the prompt text and search text to generate output text, such as using them as input to a large language model, can easily lead to an excessively large amount of input data for the model, resulting in a significant delay in outputting the output text. This is because search text typically contains a large amount of redundant or unnecessary information. Figure 3 As shown, phrases such as "As the saying goes...", "In addition...", and "It has been circulating in the A-share market for a long time..." can cause redundant or unnecessary information in the retrieved text to hinder the generation of accurate output text during the process of generating output text based on the retrieved text. Furthermore, the large amount of data in the retrieved text can slow down the reasoning process of the output text, resulting in a significant delay in the generation of the output text and a substantial increase in cost.
[0122] Therefore, in this embodiment of the application, redundancy in the search text can be removed by compressing the search text before generating the output text, thereby reducing the amount of data in the search text while retaining its information content. (Reference) Figure 4 The diagram illustrates another output text generation process provided in this embodiment of the application. Compressed text can be obtained by compressing the search text, which effectively removes redundant information, reduces the amount of information to be processed, and lowers latency and cost. The compression process of the search text can be referred to in S102-S104. The compression process in S102-S104 can be implemented by calling a text compression function or by using a text compression model. (Reference) Figure 5 The diagram shown is a schematic of a compression process provided in an embodiment of this application, wherein the retrieved text can be compressed to obtain compressed text. Specifically, the compressed text can be obtained by calling a text compression function.
[0123] Specifically, in the process of compressing search text, the search text can be segmented into multiple tokens. Probability prediction of these tokens determines their probability. The probabilities of each token indicate its information entropy. A higher information entropy indicates a greater amount of information in the token and its importance to the search text; conversely, a lower information entropy indicates a smaller amount of information and its lower importance. A token is the basic unit for processing and analyzing text, and can include words, punctuation marks, numbers, or specific strings of characters. For example, in the sentence "I'm happy.", "I'm," "happy," and "." can all be considered tokens.
[0124] In practice, target text of the target type can be extracted, but without segmentation or subsequent compression, only the text other than the target text in the search text is segmented into multiple word units, i.e., only the other text is compressed. The target type is a type whose key information depends on the overall structure; its key information is easily lost after segmentation and compression. Therefore, not compressing the target text of the target type preserves the complete information in the target text.
[0125] refer to Figure 6 The diagram illustrates another compression process provided in this application embodiment. The retrieved text is divided into target text and other text. After compression, the other text and the target text together constitute the compressed text. Target types include tables, Uniform Resource Locators (URLs), and HyperText Markup Language (HTML), while tables can be lightweight languages such as Markdown. The target text of the target type can be identified through regular expression matching (approximately 2ms), and then the other text outside the target file can be segmented to obtain multiple word units.
[0126] In practice, the retrieved text can be divided into multiple text chunks, and then each text chunk can be segmented to obtain the corresponding word units. (See reference...) Figure 5As shown, the compression operation for the retrieved text can be performed on a block-by-block basis. Multiple text blocks can be segmented in parallel to fully utilize computing resources and improve segmentation efficiency, or they can be segmented in batches (chunk-wise) to save memory and computing resources required for each processing step. Computing resources include, for example, Graphics Processing Units (GPUs). Alternatively, if only the target text in the retrieved text is segmented, these other texts can be divided into multiple blocks, and then segmented into word units for each block. This allows for batch or parallel processing of the other texts, essentially treating them as text to be compressed for subsequent processing, while the target text is not involved in the compression process.
[0127] In the process of segmenting the retrieved text, segments can be based on a preset number of characters, so that each text segment has approximately the same number of characters, which is beneficial for the balanced use of computing resources. However, this segmentation method can easily lead to a sentence or paragraph being divided into multiple text segments. Therefore, in this embodiment, after determining the initial segmentation position based on the preset number of characters, a segmentation stop token (chunk_end_tokens) before that segmentation position can be determined. The segmentation stop token is used to stop the segmentation in advance, and the position of the segmentation stop token is taken as the final segmentation position so that segmentation can be performed based on the final segmentation position.
[0128] refer to Figure 7 The diagram shown is a schematic of a setting provided in an embodiment of this application. The chunk end tokens can be specific punctuation marks, such as commas (,), periods (.), newlines (\n), or one or more combinations thereof. In this way, the text block formed by the text before the final chunk position has slightly fewer characters than the preset number, but it has complete sentences and paragraphs, which helps to generate coherent text blocks.
[0129] In practice, text blocks can be determined sequentially based on a preset character count, ensuring that the segmentation of each subsequent text block is based on the final segmentation position of the previous text block. This allows subsequent text blocks to be determined in the same way, ensuring that the character count of each text block is less than the preset character count. Alternatively, all initial segmentation positions can be determined at once based on the preset character count. After adjusting the initial segmentation positions to obtain the final segmentation positions, the number of text blocks before the final segmentation position will be less than the preset character count, while the number of text blocks after the final segmentation position may be greater than the preset character count. Within text blocks determined using the aforementioned methods, the character count of a single text block can be slightly greater or slightly less than the preset character count, resulting in small differences in character count between text blocks and not affecting the balanced utilization of computing resources.
[0130] After identifying multiple word units, their probabilities can be calculated. Specifically, the probabilities of multiple word units can be calculated using a compression model, or other methods. The compression model can be pre-trained and pre-optimized; for example, optimizing the initial model can yield a compressed model. Optimization reduces the model's weight, decreasing text compression time and optimizing output time. The compression model acts as an encoder, encoding word units to obtain their probabilities. The probability prediction time for determining the probabilities of word units can also be called encoding time.
[0131] Optimization methods can include at least one of model distillation, model quantization, model pruning, and operator fusion. Model distillation can transfer knowledge from a large, complex teacher model to a small, simple initial model. The initial model acts as a student model, learning from the teacher model's output to improve its performance, reducing computational cost and model size while maintaining high accuracy. Model quantization can convert floating-point parameters and calculations in the initial model to low-precision integer representations to obtain a compressed model, such as converting 32-bit floating-point numbers (FP32) to 8-bit integers (INT8), significantly reducing storage space and computational cost while increasing computational speed. Model pruning can remove parameters or connections from the initial model that have little impact on the output, resulting in a compressed model that reduces complexity and computational cost. Operator fusion can merge multiple adjacent operators in the initial model into a more efficient composite operator, such as fusing convolutional layers and activation function layers, reducing intermediate result storage and data transfer and improving computational efficiency.
[0132] As an example, the initial model includes a self-attention structure. By fusing operators within the self-attention structure of the initial model, a compressed model can be obtained; for example, by fusing convolutional layers and activation function layers, the time consumption of the probability calculation process can be optimized. (Reference) Figure 8The diagram illustrates the computation time consumption of an embodiment of this application. The processing of the retrieved text before the probability calculation of word units is called preprocessing, such as block processing, word segmentation, regular expression matching, and keyword determination. Taking a retrieved text data volume of 11k as an example, based on the first type of computing resources, the probability prediction time using the initial model is 84ms, of which the self-attention structure accounts for 60%, or 50ms. Using the optimized compressed model for probability prediction, the self-attention structure time is reduced, optimizing the overall probability prediction time by 30%, to approximately 55ms. Taking 24 repetitions of the self-attention calculation as an example, this is equivalent to saving 29s in 24 self-attention calculations, reducing each self-attention calculation from 2ms to 0.88ms. Based on the second type of computing resources, the probability prediction time using the initial model is 42ms. Using the compressed model for probability prediction, the overall probability prediction time can be optimized to 28ms.
[0133] During the compression of the retrieved text, certain keywords (force_token) can be forcibly retained. These keywords are word units that are important for the quality of the model's response. Keywords can be determined based on the text structure information corresponding to the prompt text. This text structure information indicates the required text structure for the prompt text. Therefore, the text structure information corresponding to the prompt text can be determined, and keywords corresponding to this text structure information can be identified from multiple word units. The probability of these keywords is set to a preset value, which can be a high probability value, such as 1, to ensure a higher probability of retention. This probability setting is equivalent to adding a key marker to the keywords, helping to ensure that paragraph structure and important information are not lost. Setting the probability of keywords can be performed before calculating the probability of word units; in this case, the probability of keywords will not be reset when calculating the probability of word units. Alternatively, setting the probability of keywords can be performed after calculating the probability of word units, in which case the probability of keywords can be modified to the preset value.
[0134] Generally, search text can include various structural information, such as textual structural information, which can be indicated by punctuation marks, and logical structural information, which can be indicated by object numbers. The textual structural information corresponding to the prompt text essentially indicates a certain aspect of the search text's structure. Keywords determined based on this are word units that indicate that aspect of structural information. Retaining these word units preserves that aspect of the structural information in the search text. In other words, the search text is compressed while retaining the necessary textual structure of the prompt text. This text compression is more targeted and accurate, thus making the determination of the output text more targeted and accurate.
[0135] Specifically, punctuation marks can include periods, line breaks, etc., and object numbering can include one or more of the following: title number, viewpoint number, article number, text number, etc. (Reference) Figure 7 As shown, the keyword (force_token) may include one or more of the following: period, newline character (\n), opinion number (e.g., opinion 1, opinion 2, opinion 3, etc.), article number (e.g., article 1, article 2, article 3, etc.), text number (e.g., text 1, text 2, text 3, etc.), title number (e.g., (I), (II), (III), etc.).
[0136] When the text structure information corresponding to the prompt text is used to indicate the structural information in the text form, the corresponding keywords can include at least one of the following: period, line break, etc. For example, if the prompt text corresponds to the summary extraction requirement, then during the compression of the search text, the structural information in the text form can be preserved, that is, the line break and the period before the line break are preserved. In this way, the search text is compressed only at the content level within paragraphs, while the structure between paragraphs is preserved.
[0137] When the text structure information corresponding to the prompt text is used to indicate logical structure information, the corresponding keywords can include at least one of the following: object number, or at least one of the following: period, line break, etc., added to the object number. For example, if the prompt text corresponds to the need for extracting viewpoints, then during the compression of the search text, the viewpoint-style structure information can be preserved. Keywords can include viewpoint numbers. In this way, the search text is compressed only for the content of each viewpoint separately, thus preserving the basic structure of the search text in terms of viewpoint presentation.
[0138] S103, based on the order of multiple word units in the retrieved text, determine multiple initial word groups according to the multiple word units, each initial word group includes at least one word unit, and the probability of the initial word group is determined by the probability of the included word units.
[0139] In this embodiment of the application, reference is made to Figure 5 As shown, based on the order of multiple word units in the retrieved text, multiple initial word groups can be determined according to the multiple word units. Each initial word group includes at least one word unit. The probability of the initial word group is determined by the probability of the included word units. Thus, the probability of the initial word group can reflect the information entropy of the word units in the initial word group, and thus reflect the total information entropy of the initial word group.
[0140] An initial phrase is the smallest lexical unit in natural language that has independent meaning and grammatical function. It is the basic linguistic unit used in natural language and can include one or more word units. Generally speaking, the granularity of an initial phrase is larger than a word unit but smaller than or equal to that of a sentence. For example, in the sentence "I'm happy.", "I'm" can be broken down into two words, "I" and "am," while "happy" is a single word, and the period (.") does not necessarily have to be a separate word. Of course, in "I'm happy.", "I amhappy" can also be considered a single word.
[0141] In this embodiment, the initial phrase may include a first type of phrase, which includes at least two word units. Based on the order of multiple word units in the search text, at least two adjacent word units can be merged to obtain the first type of phrase. The merging of at least two word units can be performed according to the semantics of the word units, giving the first type of phrase independent meaning and grammatical function. The merged first type of phrase is equivalent to a phrase extracted from the search text. In specific implementation, during the merging process, word units may not be reused, effectively cutting the search text to obtain the first type of phrase, with a larger granularity than the word unit itself. Alternatively, word units may be reused during the merging process; for example, a word unit can be merged with a preceding word unit into one phrase, and simultaneously merged with a subsequent word unit into another phrase. In this case, more initial phrases need to be removed to achieve the same compression rate.
[0142] The initial phrase can also include a second type of phrase, which consists of a single word unit. Target words meeting preset conditions can be identified from this word unit and included as the second type of phrase. Because the granularity of the second type of phrase is smaller than that of the first type, the overall granularity of the initial phrase can be controlled within a smaller range, preventing the loss of key information due to excessively large granularity. This is because when the granularity is too large, some information-rich word units, when merged into a sentence, are easily interfered with by other information-rich word units, thus reducing the probability of them belonging to the initial phrase and leading to the deletion of the original initial phrase.
[0143] Among them, the target word can be a keyword related to the prompt text. For example, the target word can be a word unit with a preset probability, which is conducive to the retention of keywords. The target word can also be a word unit with a relatively high probability, such as a word unit with a probability greater than 0.8. The target word can also be a word unit in the search text connected with a specific symbol, such as a comma or a pause mark. This is because of the difference between Chinese and English expressions. The processing method based on English is prone to underestimating the probability of the word group to which the word unit belongs. Therefore, adding word units can improve its retention probability.
[0144] refer to Figure 9 The diagram shown is a schematic diagram of an initial word group and its probability provided in an embodiment of this application. The horizontal axis represents each initial word group, and the vertical axis represents the probability of each initial word group. The initial word groups may include first-type word groups and second-type word groups, which will not be illustrated here.
[0145] In this embodiment of the application, when the searched text is processed in blocks, the process of determining the initial word group can be carried out separately according to the text blocks. That is, the initial word group corresponding to multiple text blocks can be determined separately according to the word units corresponding to multiple text blocks and the order of multiple word units in their respective text blocks. The determination of the initial word group corresponding to different text blocks can be carried out in batches or in parallel.
[0146] After determining multiple initial word groups, the probabilities of each initial word group can be determined based on the probabilities of the word units included in each initial word group. The probability of an initial word group can be the average or weighted average of the probabilities of the word units included in that initial word group, so that the probability of the initial word group can reflect the information entropy of the initial word group.
[0147] The two steps of determining the initial word group based on word units (merge_token_to_word) and determining the probability of the initial word group based on word unit probabilities (token_prob_to_word_prob) can be implemented separately by two functions, or they can be combined using a composite function (fuse). That is, multiple initial word groups and their probabilities can be determined using a composite function based on multiple word units and their probabilities. This achieves the fusion of operators determining the initial word groups and their probabilities in the post-processing after determining the word unit probabilities. (See reference...) Figure 8 As shown, this reduces the storage and data transfer of intermediate results, improving computational efficiency. For example, the calculation of the initial word groups and their probabilities accounts for approximately 21% of the entire compression process, taking about 35-40ms. After operator fusion, the Python version optimizes this to about 15-18ms, and the C++ version optimizes it to about 2ms, significantly reducing computation time.
[0148] S104, Based on the probability of the initial word groups, select the target word group from multiple initial word groups.
[0149] In this embodiment, a target phrase can be selected from multiple initial phrases based on the probability of the initial phrases. That is, the total information entropy of the target phrase meets certain screening conditions, and it has a large amount of information. It can be used as a summary of the search text and reflects the key information of the search text. Compared with the search text, the target phrase has better redundant information and more key information. It is equivalent to compressing the search text to obtain the target phrase. In this way, other phrases outside the target phrase are equivalent to being deleted as redundant information.
[0150] In this embodiment, target phrases can be determined based on a probability threshold. For example, initial phrases with a probability greater than or equal to the probability threshold can be used as target phrases. Specifically, the probability threshold of the initial phrases can be determined based on the target compression rate of the search text and the probability of the initial phrases. Then, target phrases are selected from multiple initial phrases based on the probability threshold and the probability of the initial phrases. The target compression rate of the search text is the target ratio of the amount of data in the search text after compression to the amount of data before retrieval. A higher target compression rate may lead to information loss, while a lower target compression rate has an insignificant compression effect. The process of selecting target phrases from the initial phrases is carried out with the ratio of the amount of data in the target phrases to the amount of data in the search text as the target compression rate.
[0151] In the process of calculating the probability threshold of the initial word group, the probabilities of the initial word group can be sorted to obtain a probability sequence. The target compression rate is used as a preset percentage, and the lowest probability of the first preset percentage in the probability sequence is determined as the probability threshold. Thus, the initial word group of the first preset percentage is determined as the target word group. This process is relatively time-efficient. However, since it does not consider the inconsistency of word units included in the initial word group, the actual ratio of the data volume of the target word group to the data volume of the retrieved text is usually close to the target compression rate, but may not be equal to the target compression rate.
[0152] Alternatively, in calculating the probability threshold of the initial word group, the probability threshold can be determined based on the target compression rate of the retrieved text, the probability of the initial word group, and the number of word units included in the initial word group. Specifically, the target compression rate can be used as a preset percentage of word units. The initial word groups are sorted to obtain a word group sequence. Based on the target compression rate and the number of word units included in the initial word groups, the top N initial word groups are determined. These top N initial word groups include a preset percentage of word units. The lowest probability corresponding to the top N initial word groups is used as the probability threshold, thereby determining the top N initial word groups as the target word groups. The ratio of the target word groups determined in this process to the actual data volume of the retrieved text is closer to the target compression rate.
[0153] In practice, search text can have a fixed target compression ratio, meaning different search texts share the same target compression ratio. Alternatively, since search texts of different lengths may contain varying amounts of redundant information, the target compression ratio can be determined based on the text length. The target compression ratio and text length are inversely correlated; the longer the text, the lower the target compression ratio, and vice versa. Dynamically adjusting the target compression ratio based on text length can effectively shorten longer search texts, ensuring response quality and accuracy while significantly reducing processing time. Specifically, when the search text length ranges from [1k, 4k), the target compression ratio can be set to 0.85; when the search text length ranges from [4k, 8k), the target compression ratio can be set to 0.8; when the search text length ranges from [8k, 16k], the target compression ratio can be set to 0.7, and so on.
[0154] After identifying the target phrase, if single-sided brackets exist within it, they can be corrected. For example, if a bracket is determined to be a necessary element, a corresponding opposite bracket can be added; conversely, if a bracket is determined to be a non-necessary element, it can be deleted. This addresses the issue of mismatched brackets in the input text affecting the quality of the generated output text. Identifying the target phrase from the initial phrase and correcting its brackets can serve as post-processing for retrieval compression.
[0155] When the search text is divided into multiple text blocks, the initial word groups corresponding to each text block can be filtered separately. That is, based on the probability of the initial word groups corresponding to each text block, the target word groups corresponding to each text block can be filtered out separately. The determination of the target word groups for multiple text blocks can be performed in parallel or in batches.
[0156] This method of compressing searched text by identifying target phrases is essentially an extractive summarization approach. It extracts target phrases from the searched text, ensuring high fluency and granularity (smaller than or equal to sentence granularity), thus introducing less redundant information. Furthermore, compared to generative summarization based on Natural Language Generation (NLG) technology, extractive summarization does not require autoregressive decoding, resulting in faster generation speeds. Therefore, combining extractive summarization with RAG scenarios leverages its capabilities to achieve high-precision text compression, ensuring high-quality responses based on prompts in RAG scenarios.
[0157] In summary, text compression quality can be ensured and the quality of the generated output text improved by using methods such as determining the target compression rate based on text length, not compressing target types, reducing phrase granularity by treating target words as second-category phrases, retaining keywords, and post-processing with single-sided brackets. Furthermore, data processing efficiency can be improved by processing multiple text blocks separately, optimizing the compression model, and fusing operators to reduce processing time, while minimizing the introduction of additional resource consumption and time. Thus, after optimizing both the time consumption and accuracy of the long text compression service, a balance between accuracy and performance gains is achieved.
[0158] In practice, when the search text length is 2k, the target compression ratio can be 0.85, and the compression time to determine the target phrase from the search text is approximately 30ms; when the search text length is 4.1k, the target compression ratio can be 0.8, and the compression time to determine the target phrase from the search text is approximately 54ms; when the search text length is 8k, the target compression ratio can be 0.7, and the compression time to determine the target phrase from the search text is approximately 100ms. (Reference) Figure 8 As shown, the compression time for the entire retrieved text is approximately 120ms. Compression time is an additional time introduced in the RAG scenario; however, the performance improvement brought by compression can usually cover this compression time, thus reducing the overall output text processing time.
[0159] S105, determine the output text corresponding to the prompt text based on the prompt text and the target phrase.
[0160] In this embodiment, the output text corresponding to the prompt text can be determined based on the prompt text and the target phrase. Since the target phrase contains the key information of the search text and has less data than the search text, the output text is related to both the prompt text and the search text, ensuring high output accuracy. Furthermore, the key information of the search text can be used without processing all the content of the search text, thus reducing the time consumption of text generation while ensuring high output accuracy.
[0161] The aforementioned steps of determining the search text (S101), compressing the search text (S102-S104), and generating the output text (S105) can be constructed through a full-process pipeline, enabling multiple steps to be executed automatically in sequence, balancing accuracy and performance gains.
[0162] When the target text of the target type is uncompressed, the target phrase does not contain the information of the target text. Therefore, the output text corresponding to the prompt text can be determined based on the prompt text, the target phrase, and the target text, so as to ensure that the generation of the output text fully considers the content of the search text.
[0163] Specifically, target words can be combined into compressed text, and then the output text corresponding to the prompt text can be determined based on the prompt text and the compressed text. More specifically, the compressed text and the prompt text can be merged into input text, and then the output text corresponding to the prompt text can be determined based on the input text. The compressed text can also be obtained by merging target phrases and target text.
[0164] The output text can be determined by a prediction model. This model can be based on a large language model, which determines the output text corresponding to the prompt text based on the prompt text and the target phrase. Alternatively, the prediction model can be optimized. This optimization can involve performing operations on the large language model, such as model distillation, model quantization, model pruning, and operator fusion, to reduce the amount of data and improve computational efficiency. The optimization process for the prediction model can be referenced from the optimization process of the aforementioned compressed model. The prompt text and target phrase constitute the input (prompt) data of the prediction model, guiding it to generate the corresponding output text.
[0165] Inference in large language models typically consists of two stages: prefilling and decoding. The prefilling stage processes the input data, converting it into a vector representation that the model can process, enabling it to "understand" the input. This generated vector representation can be stored in a key-value (KV) cache. The decoding stage processes the vector representation to generate the output text, progressively generating natural language text from the latent representation of the large language model. The decoding stage uses autoregressive decoding; each decoding stage generates one character, and after several decoding processes, the output text is generated. Therefore, the generation latency of the output text depends primarily on its length; the longer the output text, the more decoding steps are required.
[0166] In the inference process of a large language model, the decoding process generates each answer character by character, which takes a long time and generates many characters. Therefore, the number of decoding stages is very large, accounting for more than 90% of the entire inference process. In the prefill process, although the computation is large because all the input words need to be calculated at once, it is only a one-time process and accounts for less than 10% of the total inference time.
[0167] In large language model inference, four metrics are commonly used: Throughput, First Token Latency, Latency, and QPS (Requests per Second). These four performance metrics measure a system's service provisioning capability from four different perspectives. With the continuous improvement of large model training capabilities, online inference performance optimization has become increasingly critical. First Token Latency clearly reflects online inference performance and is one of the important goals of online inference performance optimization.
[0168] The most relevant factor to first-character latency is the length of the input data; the longer the input data, the higher the first-character latency. In RAG scenarios, long-context input is becoming the trend, resulting in a significant decrease in first-character latency performance. Based on this, through the aforementioned S102-S104, long-text compression of the retrieval text for RAG scenarios can be achieved, removing redundant information from the retrieval text and thus effectively reducing first-character latency.
[0169] refer to Figure 10 The diagram illustrates a data processing effect provided in an embodiment of this application. BERTScore is a metric for evaluating text generation quality. It utilizes a pre-trained BERT model to measure the semantic similarity between generated and reference text, providing a more effective method for evaluating text generation tasks. In actual testing, with a target compression rate of 0.7, the keyword being a newline character (\n), and chunk end tokens including periods and newlines, the quality of the output text generation was evaluated using BERTScore. The precision reached 0.84, the recall reached 0.81, and the F1 score reached 0.83, where the F1 score is the harmonic mean of precision and recall. This demonstrates that even with long text compression, high text generation accuracy, i.e., high response quality, can still be guaranteed.
[0170] Based on the data processing method provided in the embodiments of this application, the embodiments of this application also provide a data processing apparatus, see reference. Figure 11 The diagram shown is a structural block diagram of a data processing apparatus provided in an embodiment of this application. The data processing apparatus 1300 includes:
[0171] The retrieval unit 1301 is used to retrieve the retrieval text corresponding to the prompt text based on the prompt text.
[0172] The word segmentation unit 1302 is used to segment the searched text into multiple word units, and the probabilities of the multiple word units are used to indicate the information entropy of the multiple word units.
[0173] The merging unit 1303 is used to determine multiple initial word groups based on the order of the multiple word units in the search text, wherein each initial word group includes at least one word unit, and the probability of the initial word group is determined by the probability of the included word units.
[0174] Selection unit 1304 is used to select a target word group from the plurality of initial word groups based on the probability of the initial word group;
[0175] The output text generation unit 1305 is used to determine the output text corresponding to the prompt text based on the prompt text and the target phrase.
[0176] Optionally, the merging unit includes:
[0177] A probability threshold determination unit is used to determine a probability threshold for the initial word group based on the target compression rate of the retrieved text and the probability of the initial word group before filtering the target word group from the plurality of initial word groups according to the probability of the initial word group.
[0178] The merging subunit is used to select the target word group from the plurality of initial word groups based on the probability threshold and the probability of the initial word group.
[0179] Optionally, the device further includes:
[0180] A compression ratio determination unit is configured to determine the target compression ratio of the search text based on the text length of the search text before determining the probability threshold of the initial word group based on the target compression ratio of the search text and the probability of the initial word group, wherein the target compression ratio and the text length are inversely correlated.
[0181] Optionally, the device further includes:
[0182] The first optimization unit is used to perform optimization operations on the initial model to obtain a compressed model. The optimization operations include at least one of model distillation, model quantization, model pruning, and operator fusion.
[0183] A probability calculation unit is used to calculate the probability of the plurality of word units using the compression model.
[0184] Optionally, the device further includes:
[0185] A text structure information determination unit is used to determine the text structure information corresponding to the prompt text;
[0186] A keyword determination unit is used to determine keywords corresponding to the text structure information from the plurality of word units;
[0187] The probability setting unit is used to set the probability of the keyword to a preset value.
[0188] Optionally, the word segmentation unit includes:
[0189] The text extraction unit is used to extract target text of the target type from the retrieved text;
[0190] The word segmentation subunit is used to segment other texts in the searched text, other than the target text, to obtain multiple word units;
[0191] The output text generation unit is specifically used for:
[0192] Based on the prompt text, the target phrase, and the target text, determine the output text corresponding to the prompt text.
[0193] Optionally, the word segmentation unit includes:
[0194] The segmentation unit is used to segment the retrieved text into multiple text blocks;
[0195] The word segmentation unit is used to segment the multiple text blocks into word units corresponding to the multiple text blocks respectively.
[0196] The merging unit is specifically used for:
[0197] Based on the word units corresponding to the multiple text blocks and the order of the multiple word units in their respective text blocks, the initial word groups corresponding to the multiple text blocks are determined respectively;
[0198] The selection unit is specifically used for:
[0199] Based on the probability of the initial word groups corresponding to the multiple text blocks, the target word groups corresponding to the multiple text blocks are selected from the initial word groups corresponding to the multiple text blocks.
[0200] Optionally, the initial phrases include first-type phrases and second-type phrases, and the merging unit includes:
[0201] The first word group determination unit is used to merge at least two adjacent word groups among the multiple word groups according to the order of the multiple word groups in the search text to obtain the first type of word group;
[0202] The second word group determination unit is used to determine target words that meet preset conditions from the word units, and to designate them as the second type of word groups.
[0203] Optionally, the output text generation unit includes:
[0204] The second optimization unit is used to perform optimization operations on the large language model to obtain a prediction model. The optimization operations include at least one of model distillation, model quantization, model pruning, and operator fusion.
[0205] The output text generation subunit is used to determine the output text corresponding to the prompt text based on the prompt text and the target phrase using the prediction model.
[0206] Optionally, the device further includes:
[0207] The bracket correction unit is used to correct the brackets in the target phrase if there are single-sided brackets in the target phrase before determining the output text corresponding to the prompt text based on the prompt text and the target phrase.
[0208] Optionally, the merging unit is specifically used for:
[0209] Based on the plurality of word units and their probabilities, a plurality of initial word groups and their probabilities are determined by a composite function.
[0210] As can be seen from the above technical solution, the retrieved text can be obtained based on the prompt text. The retrieved text is related to the prompt text and has more information. It can serve as knowledge text related to the domain of the prompt text, which helps improve the accuracy of text prediction. Segmenting the retrieved text yields multiple word units. The probabilities of these word units are used to indicate their information entropy. The higher the information entropy, the greater the information content of the word unit. Based on the order of these word units in the retrieved text, multiple initial word groups can be determined. Each initial word group includes at least one word unit. The probability of the initial word group is determined by the probabilities of the included word units. Therefore, the probability of the initial word group reflects the information entropy of the word units in the initial word group, and thus reflects the total information entropy of the initial word group. Based on the probabilities of the initial word groups, target word groups can be selected from the multiple initial word groups. That is, the total information entropy of the target word group meets certain selection criteria, and it has a large amount of information. It can serve as a summary of the retrieved text, reflecting the key information of the retrieved text. Compared to the retrieved text, the target word group has better redundant information and more key information, which is equivalent to compressing the retrieved text to obtain the target word group. Based on the prompt text and the target phrase, the output text corresponding to the prompt text can be determined. Since the target phrase contains the key information of the search text and has less data compared to the search text, the output text is related to both the prompt text and the search text. This realizes the domain expansion based on the search text during the output text generation process, ensuring high output accuracy. Moreover, the key information of the search text can be used without processing all the content of the search text. While ensuring high output accuracy, it also reduces the time consumption of text generation.
[0211] This application also provides a computer device, which is the computer device described above, and may include a terminal device or a server. The aforementioned data processing device may be configured in the computer device. The computer device will now be described in conjunction with the accompanying drawings.
[0212] If the computer device is a terminal device, please refer to Figure 12 As shown, this application provides a terminal device, taking a mobile phone as an example:
[0213] Figure 12 This diagram illustrates a partial structural representation of a mobile phone related to the terminal device provided in this embodiment. (Reference) Figure 12 The mobile phone includes components such as a radio frequency (RF) circuit 1410, a memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a Wi-Fi module 1470, a processor 1480, and a power supply 1490. Those skilled in the art will understand that... Figure 12 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0214] The following is combined with Figure 12 A detailed introduction to each component of a mobile phone:
[0215] The RF circuit 1410 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1480; in addition, it transmits uplink data to the base station.
[0216] The memory 1420 can be used to store software programs and modules. The processor 1480 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 1420. The memory 1420 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 1420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0217] The input unit 1430 can be used to receive input numeric or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 1430 may include a touch panel 1431 and other input devices 1432.
[0218] The display unit 1440 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1440 may include a display panel 1441.
[0219] The mobile phone may also include at least one sensor 1450, such as a light sensor, a motion sensor, and other sensors.
[0220] Audio circuitry 1460, speaker 1461, and microphone 1462 provide an audio interface between the user and the mobile phone.
[0221] WiFi is a short-range wireless transmission technology. Through the WiFi module 1470, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access.
[0222] The processor 1480 is the control center of the mobile phone. It connects various parts of the mobile phone through various interfaces and lines. It performs various functions of the mobile phone and processes data by running or executing software programs and / or modules stored in the memory 1420 and calling data stored in the memory 1420.
[0223] The phone also includes a power supply 1490 (such as a battery) that powers the various components.
[0224] In this embodiment, the processor 1480 included in the terminal device also has the following functions:
[0225] The search text corresponding to the prompt text is obtained based on the prompt text search;
[0226] The retrieved text is segmented into multiple word units, and the probabilities of the multiple word units are used to indicate the information entropy of the multiple word units.
[0227] Based on the order of the multiple word units in the search text, multiple initial word groups are determined according to the multiple word units, each initial word group including at least one word unit, and the probability of the initial word group is determined by the probability of the included word units;
[0228] Based on the probability of the initial word groups, a target word group is selected from the plurality of initial word groups;
[0229] Based on the prompt text and the target phrase, determine the output text corresponding to the prompt text.
[0230] If the computer device is a server, this application embodiment also provides a server; please refer to [link to relevant documentation]. Figure 13 As shown, Figure 13 This is a structural diagram of a server 1500 provided in an embodiment of this application. The server 1500 can vary significantly due to different configurations or performance. It may include one or more processors 1522, such as a Central Processing Unit (CPU), a memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. The memory 1532 and storage media 1530 can be temporary or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the processor 1522 may be configured to communicate with the storage media 1530 and execute the series of instruction operations in the storage media 1530 on the server 1500.
[0231] Server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0232] The steps performed by the server in the above embodiments can be based on Figure 13 The server structure shown.
[0233] In addition, this application also provides a computer-readable storage medium for storing a computer program for executing the methods provided in the above embodiments.
[0234] This application also provides a computer program product including a computer program, which, when run on a computer device, causes the computer device to perform the method provided in the above embodiments.
[0235] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to computer programs. The aforementioned computer program can be stored in a computer-readable storage medium. When the computer program is executed, it performs the steps of the above method embodiments. The aforementioned computer-readable storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, or optical disk, etc., and other media that can store computer programs.
[0236] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0237] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Moreover, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data processing method, characterized in that, The method includes: The search text corresponding to the prompt text is obtained based on the prompt text search; The retrieved text is segmented into multiple word units, and the probabilities of the multiple word units are used to indicate the information entropy of the multiple word units. Based on the order of the multiple word units in the search text, multiple initial word groups are determined according to the multiple word units, each initial word group including at least one word unit, and the probability of the initial word group is determined by the probability of the included word units; Based on the probability of the initial word groups, a target word group is selected from the plurality of initial word groups; Based on the prompt text and the target phrase, determine the output text corresponding to the prompt text.
2. The method according to claim 1, characterized in that, The step of selecting a target word group from the plurality of initial word groups based on the probability of the initial word groups includes: Based on the target compression rate of the retrieved text and the probability of the initial word group, determine the probability threshold of the initial word group; Based on the probability threshold and the probability of the initial word group, the target word group is selected from the plurality of initial word groups.
3. The method according to claim 2, characterized in that, Before determining the probability threshold of the initial word group based on the target compression rate of the retrieved text and the probability of the initial word group, the method further includes: The target compression ratio of the retrieved text is determined based on the text length of the retrieved text, and the target compression ratio is inversely correlated with the text length.
4. The method according to claim 1, characterized in that, The method further includes: The initial model is optimized to obtain a compressed model. The optimization operation includes at least one of model distillation, model quantization, model pruning, and operator fusion. The probability of the multiple word units is calculated using the compression model.
5. The method according to claim 1, characterized in that, The method further includes: Determine the text structure information corresponding to the prompt text; Determine keywords corresponding to the text structure information from the plurality of word units; Set the probability of the keyword to a preset value.
6. The method according to any one of claims 1-5, characterized in that, The process of segmenting the retrieved text to obtain multiple word units includes: Extract the target text of the target type from the retrieved text; The text other than the target text in the searched text is segmented to obtain multiple word units; The step of determining the output text corresponding to the prompt text based on the prompt text and the target phrase includes: Based on the prompt text, the target phrase, and the target text, determine the output text corresponding to the prompt text.
7. The method according to claim 1, 4, or 5, characterized in that, The process of segmenting the retrieved text to obtain multiple word units includes: The retrieved text is divided into multiple text blocks; Each of the multiple text blocks is segmented into word units corresponding to the multiple text blocks; The method of determining multiple initial word groups based on the order of the multiple word units in the search text includes: Based on the word units corresponding to the multiple text blocks and the order of the multiple word units in their respective text blocks, the initial word groups corresponding to the multiple text blocks are determined respectively; The step of selecting a target word group from the plurality of initial word groups based on the probability of the initial word groups includes: Based on the probability of the initial word groups corresponding to the multiple text blocks, the target word groups corresponding to the multiple text blocks are selected from the initial word groups corresponding to the multiple text blocks.
8. The method according to any one of claims 1-5, characterized in that, The initial phrases include first-type phrases and second-type phrases. The determination of multiple initial phrases based on the order of the multiple word units in the search text includes: Based on the order of the multiple word units in the search text, at least two adjacent word units among the multiple word units are merged to obtain the first type of word group; Target words that meet preset conditions are determined from the word units and designated as the second type of word group.
9. The method according to any one of claims 1-5, characterized in that, The step of determining the output text corresponding to the prompt text based on the prompt text and the target phrase includes: A prediction model is obtained by performing optimization operations on a large language model, wherein the optimization operations include at least one of model distillation, model quantization, model pruning, and operator fusion. The prediction model determines the output text corresponding to the prompt text based on the prompt text and the target phrase.
10. The method according to any one of claims 1-5, characterized in that, Before determining the output text corresponding to the prompt text based on the prompt text and the target phrase, the method further includes: If the target phrase contains single-sided brackets, then the brackets in the target phrase are corrected.
11. The method according to any one of claims 1-5, characterized in that, The step of determining multiple initial word groups based on the multiple word units includes: Based on the plurality of word units and their probabilities, a plurality of initial word groups and their probabilities are determined by a composite function.
12. A data processing apparatus, characterized in that, The device includes: The retrieval unit is used to retrieve the retrieval text corresponding to the prompt text based on the prompt text; A word segmentation unit is used to segment the searched text into multiple word units, and the probabilities of the multiple word units are used to indicate the information entropy of the multiple word units. A merging unit is used to determine multiple initial word groups based on the order of the multiple word units in the search text, wherein each initial word group includes at least one word unit, and the probability of the initial word group is determined by the probability of the included word units; The selection unit is used to select a target word group from the plurality of initial word groups based on the probability of the initial word group; The output text generation unit is used to determine the output text corresponding to the prompt text based on the prompt text and the target phrase.
13. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute the data processing method according to any one of claims 1-11 according to instructions in the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the data processing method according to any one of claims 1-11.
15. A computer program product comprising a computer program, characterized in that, When it is run on a computer device, it causes the computer device to perform the data processing method according to any one of claims 1-11.