Text processing method and device, equipment and storage medium
By establishing the mapping relationship between the target string and the placeholder in the model and using the prefix tree mapping table in the streaming incremental output mode, the problem that the amount of text data input in the model affects the processing efficiency, achieving more efficient text processing.
Patent Information
- Application Number
- CN202510182579.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
The amount of text data input by the model affects the time-consuming process of the model, resulting in inefficient model processing.
By obtaining the search results corresponding to the problem text from the preset search library, identify the target string, and establish a mapping relationship between the target string and the placeholder. If the output mode of the target front end belongs to the streaming incremental output mode, build a prefix tree mapping table, replace the target string with placeholders, reduce the amount of model input data, and efficiently replace the placeholders in the model output through the prefix tree.
It effectively reduces the input data amount of the target text generation model, improves the model processing efficiency, and improves the processing efficiency in the streaming incremental output mode through an efficient placeholder replacement mechanism.
Smart Images

Figure CN120104720A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular to a text processing method, device, equipment and storage medium. Background Art
[0002] In the field of model application, the amount of text data input to the model affects the time consumption of the model processing process, that is, the character length of the model input affects the processing efficiency of the model.
[0003] Therefore, how to process the text input to the model to improve the processing efficiency of the model has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a text processing method, device, equipment and storage medium.
[0005] In a first aspect, an embodiment of the present disclosure provides a text processing method, the method comprising:
[0006] In response to a question text input by a user based on the target front end, a search result corresponding to the question text is obtained from a preset search library; wherein the search result includes at least one document fragment retrieved based on the question text;
[0007] Identifying at least one target string from the at least one document fragment according to a preset string recognition rule, and establishing a mapping relationship between the at least one target string and different placeholders;
[0008] If the output mode of the target front end belongs to the streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between the at least one target string and different placeholders, and the at least one target string in the at least one document fragment is replaced with the corresponding placeholder, so as to obtain a replaced document fragment corresponding to the at least one document fragment; wherein the prefix tree mapping table is used to store the mapping relationship between the at least one target string and different placeholders by using a prefix tree data structure, the node of the prefix tree is used to identify the placeholder, and the node carries the target string having a mapping relationship with the placeholder;
[0009] Inputting the replaced document fragment and the question text into a target text generation model, and outputting a text generation result after being processed by the target text generation model;
[0010] Based on the prefix tree mapping table, the placeholders in the text generation result are replaced with corresponding target strings to obtain a replaced text generation result; wherein the replaced text generation result corresponds to the answer content of the question text.
[0011] In an optional implementation manner, before establishing the mapping relationship between the at least one target character string and different placeholders, the method further includes:
[0012] A corresponding placeholder is generated for each of the at least one target character string; wherein the length of the placeholder is smaller than the length of the corresponding target character string.
[0013] In an optional implementation manner, after establishing the mapping relationship between the at least one target string and different placeholders, the method further includes:
[0014] Determine whether the output mode of the target front end belongs to the streaming incremental output mode;
[0015] If the output mode of the target front end does not belong to the streaming incremental output mode, constructing a character typical mapping table based on the mapping relationship between the at least one target character string and the placeholder; wherein the target front end is used to display the answer content corresponding to the question text;
[0016] Correspondingly, based on the mapping relationship between the at least one target string and different placeholders, replacing the placeholder in the text generation result with the corresponding target string to obtain the replaced text generation result includes:
[0017] Based on the character typical mapping table, the placeholders in the text generation result are replaced with corresponding target character strings to obtain a replaced text generation result.
[0018] In an optional implementation manner, replacing the placeholder in the text generation result with the corresponding target string based on the prefix tree mapping table to obtain the replaced text generation result includes:
[0019] If some characters of the placeholder are identified from the text generation result based on the prefix tree mapping table, the text generation result is intercepted until all characters of the placeholder are identified from the next text generation result, and then the placeholder is replaced with the corresponding target string to obtain the replaced text generation result.
[0020] In an optional implementation, the preset string recognition rule includes at least one regular expression, and the at least one regular expression is used to define a string feature.
[0021] In an optional implementation manner, before determining at least one target string from the at least one document segment according to a preset string recognition rule, the method further includes:
[0022] The at least one document segment is combined and processed according to a preset document format.
[0023] In an optional implementation manner, the step of inputting the replaced document fragment and the question text into a target text generation model and outputting a text generation result after being processed by the target text generation model further includes:
[0024] Generate input prompt words for a target text generation model based on the replaced document fragment, the question text, user role setting information, and model constraints;
[0025] Correspondingly, the replaced document fragment and the question text are input into a target text generation model, and a text generation result is output after being processed by the target text generation model, including:
[0026] The input prompt word is input into the target text generation model, and the text generation result is output after being processed by the target text generation model.
[0027] In a second aspect, the present disclosure provides a text processing device, the device comprising:
[0028] An acquisition module, configured to obtain, in response to a question text input by a user based on the target front end, a search result corresponding to the question text from a preset search library; wherein the search result includes at least one document fragment retrieved based on the question text;
[0029] an identification module, configured to identify at least one target string from the at least one document fragment according to a preset string identification rule, and establish a mapping relationship between the at least one target string and different placeholders;
[0030] A first construction module is used for constructing a prefix tree mapping table based on the mapping relationship between the at least one target string and different placeholders if the output mode of the target front end belongs to the streaming incremental output mode, and replacing the at least one target string in the at least one document fragment with the corresponding placeholders respectively to obtain a replaced document fragment corresponding to the at least one document fragment; wherein the prefix tree mapping table is used to store the mapping relationship between the at least one target string and different placeholders using a prefix tree data structure, the nodes of the prefix tree are used to identify the placeholders, and the nodes carry the target strings having a mapping relationship with the placeholders;
[0031] A model processing module, used for inputting the replaced document fragment and the question text into a target text generation model, and outputting a text generation result after being processed by the target text generation model;
[0032] A replacement module is used to replace the placeholders in the text generation result with corresponding target strings based on the prefix tree mapping table to obtain a replaced text generation result; wherein the replaced text generation result corresponds to the answer content of the question text.
[0033] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement a text processing method as provided in an embodiment of the present disclosure.
[0034] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the text processing method provided by the embodiment of the present disclosure.
[0035] In a fifth aspect, the present disclosure provides a computer program product, wherein the computer program product comprises a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor.
[0036] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:
[0037] In the text processing method provided by the embodiment of the present disclosure, first, in response to the question text input by the user based on the target front end, the search result corresponding to the question text is obtained from the preset search library. The search result includes at least one document fragment retrieved based on the question text, then, at least one target string is identified from the at least one document fragment according to the preset string recognition rule, and a mapping relationship between at least one target string and different placeholders is established. If the output mode of the target front end belongs to the streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between at least one target string and different placeholders, and at least one target string in at least one document fragment is replaced with a corresponding placeholder, so as to obtain a replaced document fragment corresponding to at least one document fragment, wherein the prefix tree mapping table is used to store the mapping relationship between at least one target string and different placeholders using a prefix tree data structure, the node of the prefix tree is used to identify the placeholder, and the node carries a target string having a mapping relationship with the placeholder, the replaced document fragment and the question text are input into the target text generation model, and the text generation result is output after being processed by the target text generation model, and the placeholder in the text generation result is replaced with the corresponding target string based on the prefix tree mapping table to obtain the replaced text generation result. Among them, the generated text result after replacement corresponds to the answer content of the question text.
[0038] After the disclosed embodiment retrieves the document fragment based on the question text input by the user and identifies the target string in the document fragment, a mapping relationship between the target string and the placeholder is established, and the target string in the document fragment is replaced with the placeholder, so as to reduce the input data volume of the target text generation model and improve the model processing efficiency. In addition, when it is determined that the output mode of the target front end is a streaming incremental output mode, the mapping relationship between the target string and the placeholder is stored in the form of a prefix tree mapping table, so as to efficiently complete the replacement of the placeholder for the text generation result output by the model based on the prefix tree mapping table, thereby improving the model processing efficiency as a whole. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0040] Figure 1 A flowchart of a text processing method provided by an embodiment of the present disclosure;
[0041] Figure 2 A schematic diagram of the structure of a prefix tree mapping table provided in an embodiment of the present disclosure;
[0042] Figure 3 A flowchart of another text processing method provided by an embodiment of the present disclosure;
[0043] Figure 4 A flowchart of another text processing method provided by an embodiment of the present disclosure;
[0044] Figure 5 A structural schematic diagram of a text processing device provided by an embodiment of the present disclosure;
[0045] Figure 6 A structural schematic diagram of a text processing device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0046] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0047] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0048] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0049] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0050] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0051] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0052] Retrieval-enhanced generation is a method that combines information retrieval technology with text generation technology. Its working principle is: first, documents related to the question text are retrieved from the pre-built knowledge base, and then the retrieved documents and the question text are input as input parameters to the text generation model to enhance the processing ability of the text generation model. However, since the retrieved documents in the pre-built knowledge base are usually long and contain a large number of characters, the text generation model needs to spend a lot of time to process these documents, resulting in low overall efficiency.
[0053] The disclosed embodiment provides a text processing method, specifically, in response to a question text input by a user based on a target front end, a search result corresponding to the question text is obtained from a preset search library. The search result includes at least one document fragment retrieved based on the question text, then, at least one target string is identified from at least one document fragment according to a preset string recognition rule, and a mapping relationship between at least one target string and different placeholders is established. If the output mode of the target front end belongs to a streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between at least one target string and different placeholders, and at least one target string in at least one document fragment is replaced with a corresponding placeholder, so as to obtain a replaced document fragment corresponding to at least one document fragment, wherein the prefix tree mapping table is used to store the mapping relationship between at least one target string and different placeholders using a prefix tree data structure, the node of the prefix tree is used to identify the placeholder, and the node carries a target string having a mapping relationship with the placeholder, the replaced document fragment and the question text are input into the target text generation model, and the text generation result is output after being processed by the target text generation model, and the placeholder in the text generation result is replaced with the corresponding target string based on the prefix tree mapping table to obtain the replaced text generation result. Among them, the generated text result after replacement corresponds to the answer content of the question text.
[0054] After the disclosed embodiment retrieves the document fragment based on the question text input by the user and identifies the target string in the document fragment, a mapping relationship between the target string and the placeholder is established, and the target string in the document fragment is replaced with the placeholder, so as to reduce the input data volume of the target text generation model and improve the model processing efficiency. In addition, when it is determined that the output mode of the target front end is a streaming incremental output mode, the mapping relationship between the target string and the placeholder is stored in the form of a prefix tree mapping table, so as to efficiently complete the replacement of the placeholder for the text generation result output by the model based on the prefix tree mapping table, thereby improving the model processing efficiency as a whole.
[0055] In order to facilitate understanding of the above embodiments, the present disclosure also provides a text processing method, which is applied to a client, and the client may include a target front end and a back end, the back end may be used to process business logic, and the target front end may be used to display an interactive interface. The method is described below in conjunction with a specific embodiment.
[0056] Figure 1 The flowchart of a text processing method provided by an embodiment of the present disclosure is shown in FIG. 1 , which can be executed by a text processing device, wherein the device can be implemented by software and / or hardware and can generally be integrated in an electronic device. Figure 1 As shown, the method includes:
[0057] S101: In response to a question text input by a user based on a target front end, a search result corresponding to the question text is obtained from a preset search library.
[0058] The search result includes at least one document fragment retrieved based on the question text.
[0059] The question text in the disclosed embodiments may be a textual expression input by a user when using a question-and-answer system, a retrieval system, etc. The question text requires corresponding answer content, that is, there is a corresponding relationship between the question text and the answer content, and the answer content is the answer given by the question system or the retrieval system to the question text.
[0060] The preset search library is a collection of information pre-set for the types of applications such as question-answering systems and search systems. The preset search library is used to store information in multiple fields. Taking the financial field search system as an example, the preset search library can be in the form of a database to store information related to the financial field. Taking the general question-answering system as an example, the preset search library can exist in the form of a knowledge base, which stores encyclopedia documents, academic papers and other materials.
[0061] The search result is the feedback result found from the preset search library for the question text input by the user. The search result may include a document fragment corresponding to the question text, and the document fragment may be content extracted from the preset search library.
[0062] In an optional implementation, since the document stored in the preset search library is long, at least one document fragment corresponding to the question text can be extracted from the document, and the document fragment constitutes the search result of the question text. The document fragment can be, for example, several natural paragraphs or several text lines in the document.
[0063] The target front end is used to receive the question text input by the user and display the answer content corresponding to the question text. Specifically, the target front end can be, for example, a web page, a front end interface, etc.
[0064] S102: Identify at least one target string from the at least one document segment according to a preset string identification rule, and establish a mapping relationship between the at least one target string and different placeholders.
[0065] The preset string recognition rule may be a rule pre-set for a target string, and the preset string recognition rule is used to determine the target string. Specifically, the preset string recognition rule may be used to recognize the target string in a document segment retrieved based on the question text.
[0066] The target string may be a string that does not require model processing, that is, the target string remains unchanged before and after model processing. For example, the target string may be a URL, a code block, or an image key.
[0067] In addition, if the model being processed is an external model (such as a model provided by a third party, a public model), in order to avoid information leakage, the target character string may be confidential content.
[0068] The target string can be identified by a preset string identification rule, and the preset string identification rule can be set in a variety of ways. In an optional implementation, the preset string identification rule can be a rule based on keyword matching, that is, by determining the keyword, a string in the document fragment whose similarity with the keyword is higher than a preset threshold is determined as the target string, and the preset threshold is a preset similarity value.
[0069] In another optional implementation, in order to identify a target string with a complex structure, the preset string recognition rule may further include at least one regular expression, and the regular expression is used to define a string feature. By defining the string feature, a target string in the document segment that meets the string feature may be identified.
[0070] In the embodiment of the present disclosure, after the target string is determined according to the preset string recognition rule, a corresponding placeholder is determined for the identified target string to establish a mapping relationship between the two, wherein the placeholder is a marking symbol used to replace the target string.
[0071] There are many ways to set a corresponding placeholder for the target string. In an optional implementation, a unified naming style can be used to set the placeholder for the target string, such as using uppercase letters plus brackets to set the placeholder.
[0072] In another optional implementation, in order to reduce the number of characters in the document fragment included in the search results, the character length of the placeholder can also be limited. Specifically, a corresponding placeholder is generated for the target string, and the length of the placeholder is less than the length of the corresponding target string. By limiting the length of the placeholder, the number of characters in the document fragment can be compressed after the placeholder is used to replace the target string in the document fragment, which can effectively reduce the processing burden of the subsequent target text generation model, thereby speeding up the speed of generating the answer content corresponding to the question text, and improving the model processing efficiency as a whole.
[0073] In an optional implementation, in order to facilitate determining the type of the target string according to the type of the placeholder so as to perform corresponding operations on the target string, different placeholders can also be set according to the type of the target string. For example, a placeholder starting with "IMG" is set for a string of image key type, and a placeholder starting with "URL" is set for a string of link type.
[0074] On the basis of the above embodiment, in order to facilitate subsequent expansion, the placeholder can also be set to a certain degree of extensibility. For example, the string of the image key type can be set to "IMG_%d", where "%d" is 1, 2, 3...n. For example, the string of the first image type identified in the document fragment can be set to "IMG_1", and the string of the second image type identified can be set to "IMG_2".
[0075] After determining how to set the placeholder, the target string is identified from the document fragment according to the preset string recognition rules, and a mapping relationship is established between the target string and its corresponding placeholder, that is, a one-to-one correspondence is established between the target string and its corresponding placeholder, that is, the correspondence between the two is recorded.
[0076] In the disclosed embodiment, the target string that the placeholder specifically replaces can be clearly identified by establishing a mapping relationship, so that the placeholder can be accurately restored to the original target string in a subsequent processing process.
[0077] In fact, on the basis of the above-mentioned embodiment, the identified repeated target string can also be replaced with the same placeholder based on the mapping relationship between the target string and the placeholder. Specifically, after identifying the target string from the document fragment according to the preset string recognition rule, the currently identified target string is compared with the previously identified target string based on the mapping relationship. If there is a target string that is the same as the currently identified target string in the previously identified target string, the placeholder corresponding to the target string replaces the currently identified target string. For example, if the target string M is currently identified, the target string is compared with the previously identified target string. If there is a target string N that is the same as the target string M in the previously identified target string, the placeholder corresponding to the target string N replaces the target string M, so as to achieve effective deduplication and reduce the number of placeholders.
[0078] S103: If the output mode of the target front end belongs to the streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between the at least one target string and different placeholders, and the at least one target string in the at least one document fragment is replaced with the corresponding placeholder to obtain a replaced document fragment corresponding to the at least one document fragment.
[0079] The prefix tree mapping table is used to store the mapping relationship between the at least one target string and different placeholders using a prefix tree data structure, the nodes of the prefix tree are used to identify the placeholders, and the nodes carry the target strings that have a mapping relationship with the placeholders.
[0080] The streaming incremental output mode may refer to an output mode in which the generated content is dynamically displayed to the user in real time as the processing progress of the question-answering system and the retrieval system progresses.
[0081] In an optional implementation, a method for determining whether the output mode of a target front end belongs to a streaming incremental output mode may be that when a question-answering system or a retrieval system receives a question text based on a user input of the target front end, the target front end may send identification information of whether the output mode of the target front end belongs to a streaming incremental output to the question-answering system or the retrieval system; if the identification information indicates that the output mode of the target front end belongs to a streaming incremental output, it may be determined that the output mode of the target front end belongs to a streaming incremental output mode.
[0082] In another optional implementation, the method for determining whether the output mode of the target front end belongs to the streaming incremental output mode can also be that, when determining whether the output mode of the target front end belongs to the streaming incremental output mode, a request is sent to the target front end to obtain whether the output mode of the target front end belongs to the streaming incremental output mode; if the identification information from the target front end indicates that the output mode of the target front end belongs to the streaming incremental output, it can be determined that the output mode of the target front end belongs to the streaming incremental output mode.
[0083] On the basis of the above embodiment, when the identification information returned by the target front end is received, the identification information may be recorded so as to be used for subsequently quickly determining the output mode of the target front end based on the identification information.
[0084] In an embodiment of the present disclosure, after determining that the output mode of the target front end belongs to the streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between at least one target string and different placeholders, that is, a prefix tree mapping table is constructed to store the mapping relationship between the target string and the placeholder.
[0085] The prefix tree may be a tree structure, in which a node is used to identify a placeholder, and the node also carries a target character string that has a mapping relationship with the placeholder identified by the node.
[0086] In an optional implementation, a prefix tree mapping table may be constructed by defining a prefix tree root node, a parent node, and a child node, and then filling the placeholder and the target string corresponding to the placeholder into the corresponding node according to the mapping relationship between the placeholder and the target string.
[0087] like Figure 2 The figure is a schematic diagram of the structure of a prefix tree mapping table provided by the embodiment of the present disclosure. First, a root node is set, corresponding nodes are set for the placeholders URL_1, URL_11, IMG_1 and IMG_11, and the target string corresponding to each placeholder is filled into the corresponding node.
[0088] In practical applications, in order to reduce the space occupied by the prefix tree mapping table, the embodiment of the present disclosure can also set a parent node for the same characters in multiple placeholders, and set child nodes under the corresponding parent node for different characters in the above multiple placeholders, so as to reduce the number of nodes and thus reduce the space occupied by the prefix tree mapping table.
[0089] In an optional implementation, after the target string is identified from the document fragment according to the preset string identification rule and a prefix tree mapping table is constructed, the prefix tree mapping table is traversed based on the target string to obtain a placeholder corresponding to the target string, and the target string in the document fragment is replaced with the corresponding placeholder to obtain a replaced document fragment. The target string does not exist in the replaced document fragment.
[0090] S104: Inputting the replaced document fragment and the question text into a target text generation model, and outputting a text generation result after being processed by the target text generation model.
[0091] The target text generation model is used for text generation. Specifically, the replaced document fragment and the question text input by the user are input into the target text generation model as input parameters. After being processed by the target text generation model, the text generation result is output. The text generation result is the output parameter corresponding to the replaced document fragment and the question text.
[0092] Among them, the text generation result can be specifically a token, and a token can refer to a basic unit generated by the target text generation model after a series of processing of the input parameters. The basic unit is used to constitute the answer content corresponding to the question text input by the user. Specifically, a token can be composed of one or more characters.
[0093] In an optional implementation, in order to improve the accuracy of the text generation results generated by the target text generation model, the user role information and model constraints may also be input into the target text generation model as input parameters of the target text generation model.
[0094] The user role information may be identity information set for the model. By setting the user role information, the target text generation model will adjust the angle of generating the text generation result based on the user role information. The model constraint condition may include restriction requirements set for the text generation result generated by the target text generation model. The model constraint condition may specifically be a length limit, language style requirements, etc. for the generated text generation result.
[0095] By inputting user role information and model constraints, the target text generation model can improve the accuracy of the generated text generation results and further enhance the user experience.
[0096] Specifically, based on the replaced document fragment, the question text input by the user, the user role setting information and the model constraints, an input prompt word for generating a target text generation model is constructed, and then the input prompt word is input into the target text generation model. The target text generation model analyzes and processes the input prompt word and outputs a text generation result.
[0097] S105: Based on the prefix tree mapping table, the placeholders in the text generation result are replaced with corresponding target strings to obtain a replaced text generation result.
[0098] The replaced text generation result corresponds to the answer content of the question text.
[0099] In an optional implementation, in the process of outputting text generation results in real time using the target text generation model, the text generation results are intercepted, and the placeholders in the text generation results are identified based on the prefix tree mapping table. Based on the identified placeholders, the node corresponding to the identified placeholder is determined from the prefix tree mapping table, and the target string carried by the node is obtained, and then the identified placeholder is restored to the target string to obtain the replaced text generation result, which is sent to the target front-end for rendering.
[0100] In another optional implementation, during the real-time output of text generation results by the target text generation model, if partial characters of a placeholder are identified from the text generation result based on a prefix tree mapping table, the text generation result can be intercepted until all characters of the placeholder are identified from the next text generation result, at which time the placeholder is replaced with the corresponding target string to obtain a replaced text generation result.
[0101] The replaced text generation result includes the target character string, and the replaced text generation result is used to form the answer content of the question text input by the user. Specifically, the replaced text generation result can be spliced in order from front to back to obtain the answer content.
[0102] Based on the above determination that the output mode of the target front end belongs to the streaming incremental output mode, by constructing a prefix tree mapping table for the mapping relationship between the target string and the placeholder to store the mapping relationship, the placeholder prefix (partial characters of the placeholder) contained in the text generation result can be timely identified during the restoration process, and the subsequent output text generation result can be intercepted, so as to complete the placeholder recognition process in the subsequent text generation result.
[0103] The text processing method provided by the embodiment of the present disclosure obtains the search result corresponding to the question text from the preset search library in response to the question text input by the user based on the target front end. The search result includes at least one document fragment retrieved based on the question text, and then, at least one target string is identified from the at least one document fragment according to the preset string recognition rule, and a mapping relationship between at least one target string and different placeholders is established. If the output mode of the target front end belongs to the streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between at least one target string and different placeholders, and at least one target string in at least one document fragment is replaced with a corresponding placeholder, so as to obtain a replaced document fragment corresponding to at least one document fragment, wherein the prefix tree mapping table is used to store the mapping relationship between at least one target string and different placeholders using a prefix tree data structure, the node of the prefix tree is used to identify the placeholder, and the node carries a target string having a mapping relationship with the placeholder, the replaced document fragment and the question text are input into the target text generation model, and the text generation result is output after being processed by the target text generation model, and the placeholder in the text generation result is replaced with the corresponding target string based on the prefix tree mapping table to obtain the replaced text generation result. Among them, the generated text result after replacement corresponds to the answer content of the question text.
[0104] After the disclosed embodiment retrieves the document fragment based on the question text input by the user and identifies the target string in the document fragment, a mapping relationship between the target string and the placeholder is established, and the target string in the document fragment is replaced with the placeholder, so as to reduce the input data volume of the target text generation model and improve the model processing efficiency. In addition, when it is determined that the output mode of the target front end is a streaming incremental output mode, the mapping relationship between the target string and the placeholder is stored in the form of a prefix tree mapping table, so as to efficiently complete the replacement of the placeholder for the text generation result output by the model based on the prefix tree mapping table, thereby improving the model processing efficiency as a whole.
[0105] In an optional implementation, when a question text input by a user based on a target front end is received, after obtaining the search results corresponding to the question text from a preset search library, the document fragments included in the search results can also be formatted to facilitate rapid determination of the target string based on preset string recognition rules.
[0106] Specifically, after obtaining the search result corresponding to the question text from the preset search library, if it is detected that the search result includes multiple document fragments, the multiple document fragments included in the search result can be combined and processed according to the preset document format to obtain a document object. The document object contains a target string, which can be used to identify the target string from the document object according to the preset string recognition rule.
[0107] In an optional implementation manner, the preset document format may be a format in which multiple document fragments are sequentially concatenated from front to back to combine and process the document fragments to obtain a document object.
[0108] In another optional implementation, the preset document format can also combine multiple document fragments in a knowledge grouping manner to obtain multiple document objects. Specifically, multiple document fragments included in the search results can be grouped according to knowledge relevance to obtain multiple document objects. After obtaining multiple document objects, subsequent text processing can be performed on the document objects.
[0109] In order to facilitate understanding of the content of the above embodiment, the present disclosure also provides another text processing method, referring to Figure 3 , which is a flow chart of another text processing method provided in an embodiment of the present disclosure.
[0110] S301: In response to a question text input by a user based on a target front end, a search result corresponding to the question text is obtained from a preset search library.
[0111] The search result includes at least one document fragment retrieved based on the question text.
[0112] S302: Determine at least one target string from the at least one document segment according to a preset string recognition rule, and establish a mapping relationship between the at least one target string and a placeholder.
[0113] The content of S301 - S302 may be understood by referring to the above embodiments.
[0114] S303: If the output mode of the target front end does not belong to the streaming incremental output mode, a character-typical mapping table is constructed based on the mapping relationship between the at least one target string and the placeholder, and the at least one target string in the at least one document fragment is replaced with the corresponding placeholder to obtain a replaced document fragment corresponding to the at least one document fragment.
[0115] The target front end is used to display the answer content corresponding to the question text. Specifically, the target front end can be, for example, a web page or a front end interface.
[0116] When it is determined that the output mode of the target front end does not belong to the streaming incremental output mode, the output mode of the target front end may be a full coverage output mode or other output modes.
[0117] In the disclosed embodiment, when it is determined that the output mode of the target front end does not belong to the streaming incremental output mode, a character-typical mapping table is constructed to store the mapping relationship between the target character string and the placeholder. The character-typical mapping table may refer to a table that stores the mapping relationship between the target character string and the placeholder in a dictionary data structure. The character-typical mapping table is easy to use and occupies less space.
[0118] In actual applications, since the mapping relationship between the target string and the placeholder may be recorded in multiple background processes, in order to avoid data inconsistency problems when the target string is replaced with the placeholder based on the mapping relationship, the embodiment of the present disclosure can construct a character-typical mapping table when the output mode of the target front-end does not belong to the streaming incremental output mode, so as to uniformly store the mapping relationship between the target string and the placeholder, thereby avoiding data inconsistency problems when the target string replaces the placeholder, and facilitating subsequent persistent use.
[0119] In an optional implementation, after the target string is identified from the document fragment according to the preset string recognition rules and a character-typical mapping table is constructed, the placeholder corresponding to the identified target string is determined based on the mapping relationship between the target string and the placeholder stored in the character-typical mapping table, and the identified target string is replaced with the corresponding placeholder to obtain a replaced document fragment.
[0120] S304: Inputting the replaced document fragment and the question text into a target text generation model, and outputting a text generation result after being processed by the target text generation model.
[0121] Specifically, the text generation result can be a token. A token can refer to a basic unit generated by the target text generation model after a series of processing of input parameters. The basic unit is used to constitute the answer content corresponding to the question text input by the user. Specifically, a token can be composed of one or more characters.
[0122] S305: replacing the placeholders in the text generation result with corresponding target character strings based on the character typical mapping table to obtain a replaced text generation result.
[0123] The replaced text generation result corresponds to the answer content of the question text.
[0124] In an optional implementation, during the process of the target text generation model outputting the text generation result in real time, when the output text generation result is pushed to the target front-end for rendering, if a placeholder consistent with the placeholder in the character-typical mapping table is identified from the current text generation result based on the character-typical mapping table, the identified placeholder is replaced with the corresponding target string based on the character-typical mapping table, and rendering is performed based on the previously output replaced text generation result and the restored target string to obtain the replaced text generation result. Among them, rendering based on the previously output replaced text generation result and the replaced target string can realize the function of overwriting the content displayed on the target front-end.
[0125] By constructing a character-typical mapping table based on the mapping relationship between the target string and the placeholder when it is determined that the output mode of the target front end does not belong to the streaming incremental output mode, it is possible to facilitate the replacement of the target string in subsequent document fragments and the restoration of the target string in the text generation result.
[0126] In order to facilitate understanding of the text processing method provided by the present disclosure, the present disclosure will take a retrieval system as an example to introduce the text processing method provided by the present disclosure. Figure 4 , which is another text processing flow diagram provided in an embodiment of the present disclosure.
[0127] In an optional implementation, the retrieval system receives a question text input by a user, and sends identification information of whether the front end belongs to a streaming incremental output mode to a back-end processing module. Based on the received question text, the retrieval system retrieves the retrieval results corresponding to the question text from the knowledge base, and the retrieval results include document fragment A, document fragment B, and document fragment C.
[0128] In an optional implementation, after obtaining document fragment A, document fragment B and document fragment C, document fragment A, document fragment B and document fragment C are combined and processed according to a preset document format to obtain a document object, so as to facilitate the subsequent identification of a target string from the document object based on preset string recognition rules.
[0129] In an optional implementation, after the document object is obtained, at least one target string is identified from the document object according to a preset string identification rule, and a mapping relationship between the at least one target string and a placeholder is established.
[0130] In an optional implementation, after the mapping relationship between the target string and the corresponding placeholder is established, it is determined whether the front end of the retrieval system belongs to the streaming incremental output mode. If it is determined based on the identification information that the front end belongs to the streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between the target string and different placeholders.
[0131] In another optional implementation, if it is determined based on the identification information that the front end does not belong to the streaming incremental output mode, a character typical mapping table is constructed based on the mapping relationship between the target character string and different placeholders.
[0132] After the prefix tree mapping table or the character typical mapping table is constructed, the target character string in the document object is replaced with the corresponding placeholder based on the prefix tree mapping table or the character typical mapping table to obtain a replaced document object.
[0133] In an optional implementation, based on the replaced document object, question text, user role setting information and model constraints, an input prompt word of the target text generation model is constructed. Then, the input prompt word is input into the text generation model, and the text generation result is output after being processed by the text generation model.
[0134] In an optional implementation, in the process of a text generation model outputting text generation results, if it is determined that the front end belongs to a streaming incremental output mode, the placeholder in the text generation result is identified based on the above-mentioned prefix tree mapping table. If partial characters of the placeholder are identified from the text generation result based on the prefix tree mapping table, the current text generation result is intercepted until all characters of the placeholder are identified from the next text generation result. After the complete placeholder is identified, the placeholder is restored to the corresponding target string based on the prefix tree mapping table to obtain a replaced text generation result, and the replaced text generation result is rendered and displayed on the front end.
[0135] In another optional implementation, in the process of the text generation model outputting the text generation result, if it is determined that the front end does not belong to the streaming incremental output mode, the text generation results are accumulated, the placeholders in the text generation results are identified based on the character typical mapping table, and the identified placeholders are restored to the target string, the previously displayed replaced text generation result is concatenated with the restored target string to obtain the current replaced text generation result, overwriting the previous replaced text generation result, and the front end is rendered and displayed based on the current replaced text generation result.
[0136] In order to implement the above embodiments, the present disclosure also proposes a text processing device. Figure 5 This is a schematic diagram of the structure of a text processing device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and can generally be integrated into an electronic device. Figure 5 As shown, the device comprises:
[0137] The acquisition module 501 is used to obtain the search results corresponding to the question text from the preset search library in response to the question text input by the user based on the target front end; wherein the search results include at least one document fragment retrieved based on the question text;
[0138] An identification module 502, configured to identify at least one target string from the at least one document fragment according to a preset string identification rule, and establish a mapping relationship between the at least one target string and different placeholders;
[0139] The first construction module 503 is used to construct a prefix tree mapping table based on the mapping relationship between the at least one target string and different placeholders if the output mode of the target front end belongs to the streaming incremental output mode, and replace the at least one target string in the at least one document fragment with the corresponding placeholders respectively to obtain a replaced document fragment corresponding to the at least one document fragment; wherein the prefix tree mapping table is used to store the mapping relationship between the at least one target string and different placeholders using a prefix tree data structure, the nodes of the prefix tree are used to identify the placeholders, and the nodes carry the target strings having a mapping relationship with the placeholders;
[0140] A model processing module 504, used for inputting the replaced document fragment and the question text into a target text generation model, and outputting a text generation result after being processed by the target text generation model;
[0141] The replacement module 505 is used to replace the placeholders in the text generation result with corresponding target strings based on the prefix tree mapping table to obtain a replaced text generation result; wherein the replaced text generation result corresponds to the answer content of the question text.
[0142] In an optional implementation, the device further includes:
[0143] The first generating module is used to generate corresponding placeholders for the at least one target character string respectively; wherein the length of the placeholder is smaller than the length of the corresponding target character string.
[0144] In an optional implementation, the device further includes:
[0145] A determination module, used to determine whether the output mode of the target front end belongs to a streaming incremental output mode;
[0146] A second construction module is used for constructing a character typical mapping table based on the mapping relationship between the at least one target character string and the placeholder when the output mode of the target front end does not belong to the streaming incremental output mode; wherein the target front end is used to display the answer content corresponding to the question text;
[0147] Accordingly, the replacement module is specifically used for:
[0148] Based on the character typical mapping table, the placeholders in the text generation result are replaced with corresponding target character strings to obtain a replaced text generation result.
[0149] In an optional implementation manner, the replacement module is specifically used to:
[0150] If some characters of the placeholder are identified from the text generation result based on the prefix tree mapping table, the text generation result is intercepted until all characters of the placeholder are identified from the next text generation result, and then the placeholder is replaced with the corresponding target string to obtain the replaced text generation result.
[0151] In an optional implementation, the preset string recognition rule includes at least one regular expression, and the at least one regular expression is used to define a string feature.
[0152] In an optional implementation, the device further includes:
[0153] The format processing module is used to combine and process the at least one document fragment according to a preset document format.
[0154] In an optional implementation, the device further includes:
[0155] A second generation module, configured to generate input prompt words of a target text generation model based on the replaced document fragment, the question text, user role setting information, and model constraints;
[0156] Accordingly, the model processing module is specifically used for:
[0157] The input prompt word is input into the target text generation model, and the text generation result is output after being processed by the target text generation model.
[0158] In the text processing device provided by the embodiment of the present disclosure, in response to the question text input by the user based on the target front end, the search result corresponding to the question text is obtained from the preset search library. The search result includes at least one document fragment retrieved based on the question text, and then, at least one target string is identified from the at least one document fragment according to the preset string recognition rule, and a mapping relationship between at least one target string and different placeholders is established. If the output mode of the target front end belongs to the streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between at least one target string and different placeholders, and at least one target string in at least one document fragment is replaced with a corresponding placeholder, and a replaced document fragment corresponding to at least one document fragment is obtained, wherein the prefix tree mapping table is used to store the mapping relationship between at least one target string and different placeholders using a prefix tree data structure, the node of the prefix tree is used to identify the placeholder, and the node carries a target string having a mapping relationship with the placeholder, the replaced document fragment and the question text are input into the target text generation model, and the text generation result is output after being processed by the target text generation model, and the placeholder in the text generation result is replaced with the corresponding target string based on the prefix tree mapping table to obtain the replaced text generation result. Among them, the generated text result after replacement corresponds to the answer content of the question text.
[0159] After the disclosed embodiment retrieves the document fragment based on the question text input by the user and identifies the target string in the document fragment, a mapping relationship between the target string and the placeholder is established, and the target string in the document fragment is replaced with the placeholder, so as to reduce the input data volume of the target text generation model and improve the model processing efficiency. In addition, when it is determined that the output mode of the target front end is a streaming incremental output mode, the mapping relationship between the target string and the placeholder is stored in the form of a prefix tree mapping table, so as to efficiently complete the replacement of the placeholder for the text generation result output by the model based on the prefix tree mapping table, thereby improving the model processing efficiency as a whole.
[0160] The text processing device provided in the embodiments of the present disclosure can execute the text processing method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0161] In addition to the above-mentioned method and apparatus, the embodiments of the present disclosure further provide a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the terminal device implements the text processing method described in the embodiments of the present disclosure.
[0162] The embodiments of the present disclosure further provide a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the text processing method described in the embodiments of the present disclosure is implemented.
[0163] In addition, the present disclosure also provides a text processing device, see Figure 6 As shown, it may include:
[0164] Processor 601, memory 602, input device 603 and output device 604. The number of processors 601 in the text processing device can be one or more. Figure 6 In some embodiments of the present disclosure, the processor 601, the memory 602, the input device 603 and the output device 604 may be connected via a bus or other means, wherein: Figure 6 The example of connecting through bus is taken in the following.
[0165] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing of the text processing device by running the software programs and modules stored in the memory 602. The memory 602 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc. In addition, the memory 602 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. The input device 603 can be used to receive input digital or character information, and generate signal input related to user settings and function control of the text processing device.
[0166] Specifically in this embodiment, the processor 601 will load the executable files corresponding to the processes of one or more applications into the memory 602 according to the following instructions, and the processor 601 will run the applications stored in the memory 602, thereby realizing the various functions of the above-mentioned text processing device.
[0167] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0168] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A text processing method, characterized in that: include: In response to a question text input by a user based on the target front end, a search result corresponding to the question text is obtained from a preset search library; wherein the search result includes at least one document fragment retrieved based on the question text; Identifying at least one target string from the at least one document fragment according to a preset string identification rule, and establishing a mapping relationship between the at least one target string and different placeholders; If the output mode of the target front end belongs to the streaming incremental output mode, a prefix tree mapping table is constructed based on the mapping relationship between the at least one target string and different placeholders, and the at least one target string in the at least one document fragment is replaced with the corresponding placeholder, so as to obtain a replaced document fragment corresponding to the at least one document fragment; wherein the prefix tree mapping table is used to store the mapping relationship between the at least one target string and different placeholders by using a prefix tree data structure, the node of the prefix tree is used to identify the placeholder, and the node carries the target string having a mapping relationship with the placeholder; Inputting the replaced document fragment and the question text into a target text generation model, and outputting a text generation result after being processed by the target text generation model; Based on the prefix tree mapping table, the placeholders in the text generation result are replaced with corresponding target strings to obtain a replaced text generation result; wherein the replaced text generation result corresponds to the answer content of the question text.
2. The method according to claim 1, characterized in that Before establishing the mapping relationship between the at least one target string and different placeholders, the method further includes: A corresponding placeholder is generated for each of the at least one target character string; wherein the length of the placeholder is smaller than the length of the corresponding target character string.
3. The method according to claim 1, characterized in that After establishing the mapping relationship between the at least one target string and different placeholders, the method further includes: Determine whether the output mode of the target front end belongs to the streaming incremental output mode; If the output mode of the target front end does not belong to the streaming incremental output mode, constructing a character typical mapping table based on the mapping relationship between the at least one target character string and the placeholder; wherein the target front end is used to display the answer content corresponding to the question text; Correspondingly, based on the mapping relationship between the at least one target string and different placeholders, replacing the placeholder in the text generation result with the corresponding target string to obtain the replaced text generation result includes: Based on the character typical mapping table, the placeholders in the text generation result are replaced with corresponding target character strings to obtain a replaced text generation result.
4. The method according to claim 1, characterized in that: The step of replacing the placeholder in the text generation result with the corresponding target string based on the prefix tree mapping table to obtain the replaced text generation result includes: If some characters of the placeholder are identified from the text generation result based on the prefix tree mapping table, the text generation result is intercepted until all characters of the placeholder are identified from the next text generation result, and then the placeholder is replaced with the corresponding target string to obtain the replaced text generation result.
5. The method according to claim 1, characterized in that The preset string recognition rule includes at least one regular expression, and the at least one regular expression is used to define a string feature.
6. The method according to claim 1, characterized in that Before determining at least one target string from the at least one document segment according to the preset string recognition rule, the method further includes: The at least one document segment is combined and processed according to a preset document format.
7. The method according to claim 1, characterized in that Before inputting the replaced document fragment and the question text into the target text generation model and outputting the text generation result after being processed by the target text generation model, the method further includes: Generate input prompt words for a target text generation model based on the replaced document fragment, the question text, user role setting information, and model constraints; Correspondingly, the replaced document fragment and the question text are input into a target text generation model, and a text generation result is output after being processed by the target text generation model, including: The input prompt word is input into the target text generation model, and the text generation result is output after being processed by the target text generation model.
8. A text processing device, characterized in that: The device comprises: An acquisition module, configured to obtain, in response to a question text input by a user based on the target front end, a search result corresponding to the question text from a preset search library; wherein the search result includes at least one document fragment retrieved based on the question text; an identification module, configured to identify at least one target string from the at least one document fragment according to a preset string identification rule, and establish a mapping relationship between the at least one target string and different placeholders; A first construction module is used for constructing a prefix tree mapping table based on the mapping relationship between the at least one target string and different placeholders if the output mode of the target front end belongs to the streaming incremental output mode, and replacing the at least one target string in the at least one document fragment with the corresponding placeholders respectively to obtain a replaced document fragment corresponding to the at least one document fragment; wherein the prefix tree mapping table is used to store the mapping relationship between the at least one target string and different placeholders using a prefix tree data structure, the nodes of the prefix tree are used to identify the placeholders, and the nodes carry the target strings having a mapping relationship with the placeholders; A model processing module, used for inputting the replaced document fragment and the question text into a target text generation model, and outputting a text generation result after being processed by the target text generation model; A replacement module is used to replace the placeholders in the text generation result with corresponding target strings based on the prefix tree mapping table to obtain a replaced text generation result; wherein the replaced text generation result corresponds to the answer content of the question text.
9. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the text processing method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the text processing method described in any one of claims 1 to 7.
11. A computer program product, characterized in that The computer program product comprises a computer program / instruction, and when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 7 is implemented.