Text processing method and apparatus
By refining the granularity of slot information to segment and bit operations, and combining it with a text processing model, the problem of incomplete slot information extraction in multi-turn dialogues in intelligent customer service systems has been solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2026-03-13
AI Technical Summary
Existing intelligent customer service systems are unable to effectively extract complete slot information through multiple rounds of dialogue when users express complex slot information, resulting in communication failures and a poor user experience.
The extraction of slot information is refined to two granularities: segment and position. Slot information is maintained through multi-turn dialogue, with segment/position operations performed each time to update until complete information is collected. A text processing model is then used to correct category information determination and generate text information.
It improves the ability to collect complex slot information in multi-turn dialogues, enhancing the communication efficiency and experience between users and intelligent voice systems.
Smart Images

Figure CN115147124B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of natural language processing technology, and particularly to text processing methods. One or more embodiments of this specification also relate to a text processing apparatus, a computing device, and a computer-readable storage medium. Background Technology
[0002] With the development of artificial intelligence (AI) technology, its applications are being incorporated into numerous fields. For example, in customer service systems, users initially interact with an AI chatbot. Only when the chatbot cannot resolve the user's needs is a human customer service representative provided. During the interaction with the chatbot, it collects the user's voice recordings, converts them into text, and extracts the necessary information from the text. However, when users express complex information, they may use multiple sentences, making it impossible for the chatbot to extract the information from the voice all at once. This leads to ineffective communication due to failed information extraction, resulting in a poor user experience. Therefore, extracting key information from complex dialogues is a pressing problem that needs to be solved. Summary of the Invention
[0003] In view of this, embodiments of this specification provide a text processing method. One or more embodiments of this specification also relate to a text processing apparatus, a computing device, a computer-readable storage medium, and a computer program, to address the technical deficiencies existing in the prior art.
[0004] According to a first aspect of the embodiments of this specification, a text processing method is provided, comprising:
[0005] Obtain the corrected text, the corresponding historical text, and the text to be processed;
[0006] The correction category information is determined based on the corrected text, the historical text, and the text to be processed.
[0007] Extract the correction text information from the correction text;
[0008] Target text information is generated based on the correction category information, the correction text information, and the text information to be processed.
[0009] According to a first aspect of the embodiments of this specification, a text processing method is provided, comprising:
[0010] Obtain the corrected text, the corresponding historical text, and the text to be processed;
[0011] The corrected text, the historical text, and the text to be processed are input into a text processing model, wherein the text processing model includes a determination unit, an extraction unit, and a generation unit. The determination unit determines correction category information based on the corrected text, the historical text, and the text to be processed. The extraction unit extracts the corrected text information from the corrected text. The generation unit generates target text information based on the correction category information, the corrected text information, and the text to be processed.
[0012] Obtain the target text information output by the text processing model.
[0013] According to a third aspect of the embodiments of this specification, a text processing apparatus is provided, comprising:
[0014] The acquisition module is configured to acquire the corrected text, the historical text corresponding to the corrected text, and the text information to be processed;
[0015] The determination module is configured to determine correction category information based on the corrected text, the historical text, and the text to be processed.
[0016] The extraction module is configured to extract the correction text information from the correction text;
[0017] The generation module is configured to generate target text information based on the correction category information, the correction text information, and the text information to be processed.
[0018] According to a fourth aspect of the embodiments of this specification, a text processing apparatus is provided, comprising:
[0019] The first acquisition module is configured to acquire the corrected text, the historical text corresponding to the corrected text, and the text information to be processed.
[0020] An input module is configured to input the corrected text, the historical text, and the text information to be processed into a text processing model. The text processing model includes a determining unit, an extraction unit, and a generating unit. The determining unit determines correction category information based on the corrected text, the historical text, and the text information to be processed. The extraction unit extracts the corrected text information from the corrected text. The generating unit generates target text information based on the correction category information, the corrected text information, and the text information to be processed.
[0021] The second acquisition module is configured to acquire the target text information output by the text processing model.
[0022] According to a fifth aspect of the embodiments of this specification, a computing device is provided, including a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor executes the computer instructions to implement the steps of the text processing method.
[0023] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions, which, when executed by a processor, implement the steps of the text processing method.
[0024] According to a seventh aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described text processing method.
[0025] The text processing method provided in this specification includes: obtaining corrected text, historical text corresponding to the corrected text, and text information to be processed; determining correction category information based on the corrected text, the historical text, and the text information to be processed; extracting corrected text information from the corrected text; and generating target text information based on the correction category information, the corrected text information, and the text information to be processed.
[0026] One embodiment of this specification implements the determination of corresponding correction category information based on the corrected text, historical text, and text information to be processed, so that there is corresponding correction category information under different correction category conditions. The target text information is generated based on the corresponding correction category information, corrected text information, and text information to be processed, which meets the need to collect complex slot information in multi-turn multi-talk scenarios in human-computer interaction scenarios. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating a text processing method provided in one embodiment of this specification;
[0028] Figure 2 This is a flowchart of another text processing method provided in one embodiment of this specification;
[0029] Figure 3 This is a schematic diagram of the architecture of a human-computer dialogue system provided in one embodiment of this specification;
[0030] Figure 4 This is a schematic diagram illustrating the framework of a text processing model provided by one embodiment of this specification;
[0031] Figure 5 This is a flowchart illustrating the processing procedure of a text processing method provided in one embodiment of this specification.
[0032] Figure 6This is a schematic diagram of the structure of a text processing device provided in one embodiment of this specification;
[0033] Figure 7 This is a schematic diagram of the structure of another text processing device provided in one embodiment of this specification;
[0034] Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0035] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0036] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to any or all possible combinations including one or more of the associated listed items.
[0037] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0038] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0039] Slot information: Slot information is the key information defined in the user request. In other words, it is the key information that the machine needs to collect in the human-computer interaction scenario. For example, when a user books a train ticket, the slot information collected by the machine can be the departure time, departure location, destination, etc.
[0040] Autoregressive model: A linear prediction model that uses a linear combination of random variables at several previous times to describe a random variable at a later time. This model is a sequential prediction and can only predict word by word in text generation.
[0041] Non-autoregressive models: Unlike autoregressive models, which require the generated words to predict the next word, non-autoregressive models break the sequential order of generation and can decode the entire target sentence, thus solving the sequential prediction performance problem caused by autoregression.
[0042] For intelligent customer service systems in voice interaction scenarios, the server side often needs to collect some key user information. For example, when a user buys a train ticket through an intelligent customer service system, the server side needs to collect key slot information such as the user's name, mobile phone number, and origin. Currently, traditional intelligent customer service systems extract slot information all at once based on the user's current statement, extracting the entire slot-level information and filling it into the corresponding slot. For example, if the user's current statement is "I want to buy a ticket to Beijing," then "Beijing" is extracted and filled into the "destination" slot. However, for some complex slot information, users often express this information through multiple statements. For example, when stating a mobile phone number, a user might say "188, then 1234, 5678." When the user communicates with the intelligent customer service system, the system cannot extract the complete mobile phone number at once; the final extraction result might be "5678." Therefore, the intelligent customer service system cannot fill the correct slot information into the corresponding slot, leading to communication failure and a poor user experience.
[0043] Based on this, this specification provides a text processing method that refines the extraction of slot information from the slot granularity to two granularities: segment and bit. The segment granularity involves dividing the slot into multiple segments for operation, while the bit granularity involves dividing the slot information into bits for operation. This allows for multi-turn dialogue communication to maintain slot information, updating the slot information with each segment / bit operation until the entire slot information is collected. This solves the problem that intelligent customer service systems or human-computer interaction systems cannot collect complex slot information expressed through multiple turns of dialogue. This specification also relates to a text processing device, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail in the following embodiments.
[0044] Figure 1 A flowchart of a text processing method according to an embodiment of this specification is shown, including steps 102 to 108.
[0045] Step 102: Obtain the corrected text, the corresponding historical text, and the text to be processed.
[0046] The corrected text can be understood as the text converted from the user's current statement. For example, if a user says, "My phone number is ***", the server can convert this into text, i.e., corrected text, which is "My phone number is ***". The historical text can be understood as the text converted from statements made in previous rounds before the corrected text. For example, if the user's interaction with the server is: "A: What is your phone number? B: My phone number is ***. A: Please confirm phone number ***. B: Confirm", then the corrected text is "Confirm". The historical text is the text converted from the dialogue content before the corrected text, typically 3-4 rounds prior. The pending text information can be understood as slot information extracted by the server based on the previous round of dialogue. For example, the pending text information could be "*** (phone number)".
[0047] In practical applications, when users interact with the intelligent customer service system via voice or other means through a terminal, the intelligent customer service system can run on the user's terminal or on a server that communicates with the terminal. Preferably, considering the terminal's operational capabilities and the user experience, in the embodiments of this specification, the intelligent customer service system runs on the server side.
[0048] The server receives the user's voice and converts it into text content. It then extracts key slot information from the text content. When the user describes the content of a slot through multiple rounds of voice, the server corrects the previously extracted text information based on the corrected text of the current round until complete and correct slot information is collected.
[0049] In one embodiment of this specification, a user purchases a train ticket through an intelligent customer service system. The dialogue between the user and the intelligent customer service system is as follows: "A: Please state your mobile phone number. B: My mobile phone number is 18812345678." At this time, the server side obtains the corrected text as "My mobile phone number is 18812345678" and the historical text as "Please state your mobile phone number." Since this is the first time the mobile phone number slot has been accessed, the text information to be processed is an 11-digit null value.
[0050] In another embodiment of this specification, a user purchases a train ticket through an intelligent customer service system. The dialogue between the user and the intelligent customer service system is as follows: "A: Please state your mobile phone number. B: My mobile phone number is 18812345678. A: Please confirm that your mobile phone number is 18812345678. B: No, it's 1239." The server side obtains the corrected text as "No, it's 1239," and the historical text as "Please state your mobile phone number. My mobile phone number is 18812345678. Please confirm that your mobile phone number is 18812345678." The text information to be processed is the mobile phone number slot information "18812345678" extracted from the previous rounds of dialogue.
[0051] To facilitate model processing of the acquired content, the corrected text, historical text, and text information to be processed can be combined into a text sequence for processing, thus simplifying subsequent model processing. Specifically, after acquiring the corrected text, the corresponding historical text, and the text information to be processed, the process further includes:
[0052] A sequence of text to be processed is generated based on the corrected text, the historical text, and the text information to be processed.
[0053] The text sequence to be processed can be understood as a text sequence generated by combining the corrected text, historical text, and text information to be processed.
[0054] In practical applications, the text sequence to be processed is a token sequence. When converting strings such as corrected text, historical text, and text information to be processed into token sequences, special words such as classification tokens and segmentation tokens are added to the sequence. After generating the text sequence to be processed, it can be input into a pre-trained model, which will then perform subsequent text processing and output the target text information.
[0055] In a specific embodiment of this specification, following the previous example, the corrected text, historical text, and text information to be processed are used to generate a text sequence to be processed. The text sequence to be processed is "[CLS] Please state your mobile phone number. [SEP] My mobile phone number is 18812345678. [SEP] Please confirm your mobile phone number is 18812345678. [SEP] Incorrect, it is 1239 [SEP] 18812345678. [SEP]", where [CLS] is the category token and [SEP] is the segmentation token.
[0056] Specifically, generating a text sequence to be processed based on the corrected text, the historical text, and the text information to be processed includes:
[0057] An initial text sequence is generated based on the corrected text, the historical text, and the text information to be processed;
[0058] The initial text sequence is preprocessed to obtain the text sequence to be processed.
[0059] The initial text sequence can be understood as a text sequence that has not undergone data preprocessing. The initial text sequence does not contain any special tokens. By preprocessing the initial text sequence, the text sequence to be processed can be obtained.
[0060] In practical applications, data preprocessing includes text segmentation, slot information processing, sequence padding, and special word padding. Text segmentation involves splitting the text sequence by character; if there are numeric strings, they are split by number. Slot information processing replaces the numeric strings in the server-side response with a specific token, i.e., it performs structured processing. Sequence padding extends the entire sequence to a specified length. For example, if the sequence corresponding to the text information to be processed is less than a specified number of characters, it is padded to the specified length; for example, if the text information to be processed is 188, its corresponding sequence is padded to 188[pad]. Special word padding adds a classification character [CLS] at the beginning of the text sequence and adds [SEP] between and at the end of sequences corresponding to different texts.
[0061] In a specific embodiment of this specification, following the previous example, the corrected text, historical text, and text information to be processed are combined into an initial text sequence of "Please state your mobile phone number. My mobile phone number is 18812345678. Please confirm your mobile phone number is 18812345678. No, it is 1239. 18812345678." The initial text sequence is preprocessed to obtain the text sequence to be processed: "[CLS]Please state your mobile phone number. [SEP]My mobile phone number is 18812345678. [SEP]Please confirm your mobile phone number is [UNK]. [SEP]No, it is 1239 [SEP]18812345678. [SEP]".
[0062] Step 104: Determine the correction category information based on the corrected text, the historical text, and the text information to be processed.
[0063] The correction category information can be understood as the operation category information for the text information to be processed. The correction category information is divided into five correction categories: Full Update, Full Clear, Append Content, Remain Unchanged, and Partial Update. Full Update can be understood as updating the entire text information to be processed based on the slot information extracted in the current round; Full Clear can be understood as clearing the entire text information to be processed; Append Content can be understood as adding the slot information extracted in the current round to the text information to be processed; Remain Unchanged can be understood as not updating the text information to be processed, keeping the current state unchanged; and Partial Update can be understood as partially correcting the text information to be processed. It is important to note that Full Update, Full Clear, Append Content, and Remain Unchanged are segment operation types, where a segment refers to a part of the slot information; Partial Update is a bit operation type, requiring bit-by-bit modification of the text information to be processed. The five preset correction categories can accommodate all possible modification scenarios.
[0064] In practical applications, after determining the correction category information based on the corrected text, historical text, and the text to be processed, the text to be processed can be corrected accordingly to obtain the correct key information. When the correction category information is a segment operation type, the corresponding operation is directly performed on the text to be processed, and the new text to be processed is output as the target text information; when the correction category information is a bit operation type, bit-by-bit text information is generated, which will generate new target text information.
[0065] In a specific embodiment of this specification, when the correction category information is predicted to be updated based on the corrected text, historical text, and text to be processed, the text to be processed is updated to all the key information extracted from the corrected text. If the text to be processed is "18812345678" and the key information extracted from the corrected text is "18812395678", then after performing all the update operations, the text to be processed is corrected to "18812395678".
[0066] In another specific embodiment of this specification, when the correction category information is predicted to be completely cleared based on the corrected text, historical text, and text information to be processed, the slot information in the text information to be processed is replaced with an empty value, that is, the text information to be processed is cleared. If the text information to be processed is "18812345789", then after performing the complete clearing operation, the text information to be processed is corrected to 11 empty characters.
[0067] In another specific embodiment of this specification, when the correction category information is predicted to be appended content based on the corrected text, historical text, and text information to be processed, key information extracted from the corrected text is added after the text information to be processed. If the text information to be processed is "1881234" and the key information extracted from the corrected text is "5678", then after performing the append operation, the text information to be processed is corrected to "18812345678".
[0068] In another specific embodiment of this specification, when the predicted correction category information based on the corrected text, historical text, and text information to be processed is to remain unchanged, the state of the current text information to be processed remains unchanged. If the text information to be processed is "18812345678", then after performing the unchanged operation, the result of the corrected text information to be processed is "18812345678".
[0069] In another specific embodiment of this specification, when the predicted correction category information is a partial update based on the corrected text, historical text, and text information to be processed, the text information that needs to be corrected in the text information to be processed is replaced with key information extracted from the corrected text. If the text information to be processed is "18812345678" and the key information extracted from the corrected text is "1239", then after performing the partial update operation, the text information to be processed is corrected to "18812395678".
[0070] Specifically, determining the correction category information based on the corrected text, the historical text, and the text to be processed includes:
[0071] The correction category weight corresponding to each correction category information in the correction category information set is calculated based on the text sequence to be processed, wherein the correction category information set includes at least two correction category information.
[0072] The target correction category information is determined from the correction category information set based on the correction category weight corresponding to each correction category information.
[0073] The correction category weight can be understood as the probability of executing the operation for that correction category information in this round. For example, given five correction category information types: a correction category weight of 0.2 for updating all, 0.2 for clearing all, 0.1 for adding content, 0.4 for keeping unchanged, and 0.1 for partial updating, the correction category to be executed in this operation can be determined based on the correction category weight corresponding to each correction category information type. The correction category information set can be understood as a pre-set information set including all types of correction categories. In the embodiments of this specification, the correction category information set includes five correction category information types. The correction category information to be executed in this operation is determined from the correction category information based on the correction category weight corresponding to each correction category information type.
[0074] In practical applications, the operation probability corresponding to each correction category can be calculated based on the sequence of text to be processed generated from the corrected text, historical text, and text to be processed information. The target correction category information is then selected based on the operation probability corresponding to each correction category information, and the target correction category information is the correction category that needs to be executed this time.
[0075] In a specific embodiment of this specification, following the example above, the correction category weights corresponding to the five correction category information in the correction category information set are calculated based on the text sequence to be processed. They are "All Update: 0.1, All Clear: 0.1, Append Content: 0.2, Remain Unchanged: 0, Partial Update: 0.6". If the target correction category information is determined to be "Partial Update", then the representation of all bits of the entire text information to be processed is extracted, and the content of each bit is predicted.
[0076] Specifically, the target correction category information is determined from the correction category information set based on the correction category weight corresponding to each correction category information, including:
[0077] The correction category information in the correction category information set is sorted in descending order according to the correction category weight corresponding to each correction category;
[0078] The target correction category information is determined based on the sorting results.
[0079] In practical applications, after determining the correction category weights corresponding to the five correction category information in the correction category information set, all correction category information can be sorted from high to low according to the correction category weights corresponding to each correction category information. The correction category information with the highest correction category weight is selected as the target correction category information according to the sorting result.
[0080] In a specific embodiment of this specification, there are five types of correction category information in the correction category information set. After calculating the correction category information corresponding to these five types of correction category information, these five types of correction category information are sorted from highest to lowest according to the correction category weight. The sorting result is "partial update: 0.6, append content: 0.2, full update: 0.1, full clear: 0.1, remain unchanged: 0". The correction category information ranked first is selected as the target correction category information.
[0081] Step 106: Extract the correction text information from the correction text.
[0082] After determining the correction category to be performed in this round, the corresponding correction category operation can be performed on the text information to be processed based on the correction text information in the correction text. The correction text information is key information extracted from the correction text. For example, if the correction text is "No, it's 1239", the extracted correction text information would be "1239".
[0083] In practical applications, when the correction category is set to "all cases," the correction text is a null character corresponding to the number of positions in the text to be processed. For example, if the correction text is "Incorrect, I'll repeat it," then the correction text is empty. The text to be processed is then updated with the correction text, thus clearing all text. When the correction category is set to "remain unchanged," the correction text can be the same as the text to be processed. For example, if the correction text is "Yes, correct," then the correction text can be the text to be processed, thus keeping the text to be processed unchanged.
[0084] Specifically, extracting the correction text information from the correction text includes:
[0085] The corrected text is annotated.
[0086] Extract the corrected text information from the corrected text based on the annotation results.
[0087] Among them, text annotation can be understood as classification annotation, which classifies each part of the corrected text and extracts the corrected text information.
[0088] In practical applications, text annotation for corrections can be performed using BIO (Block I / O) classification, where B represents the beginning of a slot, I represents the inside of a slot, and O represents non-slot information. Based on BIO classification annotation of the corrections, the corresponding slot information, i.e., the correction text information, can be extracted from the correction text information.
[0089] In a specific embodiment of this specification, following the example above, the corrected text "Incorrect, it is 1239" is labeled with BIO classification, and the corrected text information "1239" is extracted from the corrected text based on the labeling results.
[0090] Step 108: Generate target text information based on the correction category information, the correction text information, and the text information to be processed.
[0091] The target text information can be understood as the target text information after the text to be processed has been corrected according to the corrected text in this round. For example, when the corrected text is "No, it is 1239", the target text information generated after correcting the text to be processed "18812345678" is "18812395678".
[0092] In practical applications, after the target text information is generated, if the server continues to collect corrected text, the target text information will be used as the text information to be processed in the next round of corrected text.
[0093] In a specific embodiment of this specification, the target text information "18812395678" is generated based on the correction category information "partial update", the correction text information "1239", and the text information to be processed "18812345678".
[0094] In another specific embodiment of this specification, the target text information "18812395678" is generated based on the correction category information "All Update", the correction text information "18812395678", and the text information to be processed "18812345678".
[0095] Specifically, generating target text information based on the correction category information, the correction text information, and the text information to be processed includes:
[0096] Based on the correction category information and the correction text information, the target text information to be corrected is determined from the text information to be processed;
[0097] The target text information to be corrected is obtained by correcting the target text information in the text information to be processed based on the corrected text information.
[0098] The target text to be corrected can be understood as the text information that needs to be corrected in this round of text information to be processed. For example, when the correction category is partial update, the target text information to be corrected needs to be determined in the text information to be processed "18812345678" based on the correction text information "1239". If the target text information to be corrected is "1234", then the correction text information is used to correct the target text information to obtain the target text information.
[0099] In practical applications, when the correction category is "full update", the target text information to be corrected is the text information to be processed; when the correction category is "all cases", the target text information to be corrected is the text information to be processed; when the correction category is "add content", the empty character information in the text information to be processed is determined based on the corrected text information. For example, if the corrected text information is "1234" and the text information to be processed is "188H1H2H3H4H5H6H7H8", then the target text information to be processed is "H1H2H3H4"; when the correction category information is "unchanged", the target text information to be corrected can be the text information to be processed; it should be noted that when the correction category information is "partial update", the target text information to be corrected is a text information in the text information to be processed with the same number of characters as the corrected text information, and the representation of all bits of the entire text information to be processed is extracted. Based on the corrected text information and the target text information to be corrected, each bit of the text information to be processed is predicted to generate the target text information.
[0100] In a specific embodiment of this specification, following the example above, based on the correction category information "partial update" and the correction text information "1239", the target text information to be corrected in the text information to be processed "18812345678" is determined to be "1234". Based on the correction text information and the target text information to be corrected, the content of each bit of the text information to be processed is predicted, and the prediction result is "18812395678". Therefore, the target text information is "18812395678".
[0101] Specifically, determining the target text information to be corrected from the text information to be processed based on the correction category information and the correction text information includes:
[0102] Based on the correction category information and the correction text information, at least one piece of text information to be corrected is determined from the text information to be processed;
[0103] Calculate the correction similarity between the corrected text information and each piece of text information to be corrected;
[0104] The target text information to be corrected is determined based on the correction similarity corresponding to each text information to be corrected.
[0105] Among them, the correction similarity of the text information to be corrected can be understood as the similarity between the text information to be corrected and the corrected text information. The higher the correction similarity, the greater the probability of correction in this round. Based on the correction similarity corresponding to each text to be corrected, one with the highest correction probability in this round of text processing can be selected as the target text information to be corrected, and the subsequent correction operation can continue to be performed.
[0106] In practical applications, when the correction category information is updated completely, cleared completely, left unchanged, or appended, there is only one text information to be corrected. This eliminates the need to calculate the correction similarity, allowing direct identification of the target text information. However, when the correction category information is only partially updated, the text information to be processed includes multiple text information to be corrected. In this case, it is necessary to calculate the correction similarity for each text information to be corrected based on the corrected text information and each text information to be corrected, and then determine which text information is the target text information based on the correction similarity.
[0107] In a specific embodiment of this specification, when the correction category information is a partial update and the correction text information is "1239", the text information to be processed "18812345678" includes multiple text information to be corrected, namely "1881, 8812, 8123, 1234, 2345, 3456, and 5678". The correction similarity between the correction text information and each text information to be corrected is calculated. The result shows that the correction similarity of the text information to be corrected "1234" is the highest. Therefore, the text information to be corrected "1234" is determined to be the target text information to be corrected.
[0108] This specification provides a text processing method, including acquiring corrected text, corresponding historical text, and text information to be processed; determining correction category information based on the corrected text, historical text, and text information to be processed; extracting corrected text information from the corrected text; and generating target text information based on the correction category information, the corrected text information, and text information to be processed. This method refines the traditional slot-level state update of text information to be processed to a two-level operation granularity of segment and bit operations. It allows for multiple different correction scenarios based on different correction category information, accommodating all possible corrections and meeting the needs of collecting complex slot information in multi-turn presentations. This improves communication efficiency between users and intelligent voice systems, as well as the user experience.
[0109] This specification provides a text processing method, which is mainly used in the state update part to solve the problem that current human-computer dialogue systems cannot gradually improve the information of the same slot through multiple rounds of dialogue information. Figure 2 A flowchart of a text processing method according to an embodiment of this specification is shown, including steps 202 to 206.
[0110] Human-computer dialogue systems generally consist of four parts: dialogue understanding, state updating, dialogue strategy, and dialogue generation. The specific architecture is as follows: Figure 3 As shown, Figure 3This specification illustrates an embodiment of a human-computer dialogue system, which includes four processing units: a dialogue understanding unit, a state update unit, a dialogue strategy unit, and a dialogue generation unit. After a user sends a dialogue to the server, the server first needs to understand the dialogue, i.e., parse the user's expression, extract key information, and update the state. After updating the state, the server selects a dialogue strategy based on the degree of state update. For example, if state collection is complete, the next step can proceed; otherwise, the server continues to query the user. Different dialogue strategies correspond to different situations. Based on the selected dialogue strategy, the server determines the dialogue to be fed back to the user and sends it to the user.
[0111] Currently, state updates in human-computer interaction systems are based on the entire slot granularity. All slot information involved in a certain scenario is maintained uniformly. Each round of dialogue determines the slots that need to be updated and fills the extracted slot information into the corresponding slots to complete the modification. This modification generally involves filling, clearing, and replacing the entire slot. It can only be operated based on the information of one round of dialogue and cannot gradually improve the slots through multiple rounds of dialogue information. For example, it cannot achieve partial updates of slot information.
[0112] Step 202: Obtain the corrected text, the corresponding historical text, and the text to be processed.
[0113] Among them, the corrected text can be understood as the text converted from the user's expression in the current round, the historical text corresponding to the corrected text can be understood as the text converted from the user's expression in multiple rounds before the current round, and the text information to be processed can be understood as the slot information extracted based on the expression in previous rounds.
[0114] In practical applications, after obtaining the corrected text, historical text, and text to be processed information, all of them need to be input into the text processing model. Therefore, data preprocessing can be performed on the corrected text, historical text, and text to be processed information to generate a text sequence to be processed. Specifically, data preprocessing includes text segmentation, slot information processing, sequence filling, and special word filling. After data preprocessing, the generated text sequence to be processed is input into the text processing model.
[0115] In a specific embodiment of this specification, the corrected text is "Incorrect, it's 1239", and the historical text is "Please state your mobile phone number. My mobile phone number is 18812345678. Please confirm your mobile phone number is 18812345678". The text to be processed is "18812345678". Based on the corrected text, historical text, and text to be processed, data preprocessing is performed to generate a text sequence to be processed, which is then input into the text processing model.
[0116] Step 204: Input the corrected text, the historical text, and the text information to be processed into the text processing model, wherein the text processing model includes a determination unit, an extraction unit, and a generation unit. The determination unit determines the correction category information based on the corrected text, the historical text, and the text information to be processed. The extraction unit extracts the corrected text information from the corrected text. The generation unit generates target text information based on the correction category information, the corrected text information, and the text information to be processed.
[0117] In practical applications, the text processing model includes a determination unit, an extraction unit, and a generation unit. The target text information is output after the text sequence to be processed, which is generated based on the received corrected text, historical text, and text information to be processed, is processed by the determination unit, the extraction unit, and the generation unit.
[0118] Specifically, see Figure 4 , Figure 4 This specification illustrates a framework diagram of a text processing model according to an embodiment. The text sequence to be processed includes a text sequence corresponding to the corrected text, a text sequence corresponding to the historical text, and a sequence corresponding to the text information to be processed. After the text sequence to be processed is input into the text processing model, the determining unit in the model predicts the correction category information of the text information to be processed based on the text sequence; the extraction unit extracts the corrected text information from the corrected text based on the text sequence; and the generating unit corrects the text information to be processed based on the correction category information and the corrected text information, outputting the corresponding target text information as "18812395678". The text information to be processed is "18812345678", and the corrected text information is "39".
[0119] In a specific embodiment of this specification, following the example above, the text sequence to be processed is input into the text processing model. The determination unit in the text processing model determines the correction category information as "partial update" based on the text sequence to be processed, i.e., enters bit operation; the extraction unit extracts the correction text information "1239" from the text sequence to be processed; the generation unit uses non-autoregressive generation to extract the representation of all bits of the text information to be processed, performs bit content prediction based on the correction category information and the correction text information, and generates the sequence corresponding to the target text information, thereby obtaining the target text information.
[0120] In practical applications, before the text processing model is deployed to the system, data annotation and model training are required. In each round of human-computer dialogue to collect slot information, the extraction unit annotates the slot information corresponding to the corrected text, and the determination unit performs prediction training for the corrected category information. When the predicted corrected category information is partially updated, the target text information needs to be annotated. During model training, the extraction unit uses sequence labeling for positional classification training, with the operation type being task training for five corrected categories. Cross-entropy loss and the Adam optimization method can be used for model training and optimization.
[0121] Specifically, the text processing model can be trained through the following steps:
[0122] Acquire training correction text, correction text information, training history text corresponding to the training correction text, and training text information to be processed;
[0123] The training correction text, the training history text, and the training text to be processed are input into the text processing model to obtain the predicted text information output by the text processing model.
[0124] Calculate the model loss value based on the predicted text information and the corrected text information;
[0125] The model parameters of the text processing model are adjusted based on the model loss value, and the text processing model continues to be trained until the model training stops.
[0126] Among them, training correction text can be understood as the sample correction text used during model training, training history text can be understood as the sample history text used during model training, training text information to be processed can be understood as the sample text information to be processed used during model training, and correction text information can be understood as the standard target text information output by the model after this training correction.
[0127] In practical applications, the model is pre-trained using training samples (training corrected text, training history text corresponding to the training corrected text, and training text information to be processed) and labeled text (corrected text information). The model parameters are continuously optimized by calculating the model loss value. The model loss value is used to evaluate the degree to which the model's predicted value differs from the true value. In practical applications, the text processing model can be trained using a large number of training samples. By making predictions and calculating the model loss value each time, the model parameters can be continuously adjusted, thereby improving the model's output accuracy.
[0128] In a specific embodiment of this specification, training correction text, training history text, and training text to be processed are input into the text processing model to obtain the predicted text information output by the text processing model. The model loss value is calculated based on the predicted text information and the correction text information. The model parameters of the text processing model are adjusted based on the model loss value, and the next set of training correction text, training history text, and training text to be processed is input. This process continues until the model training stops.
[0129] The model training stopping conditions include:
[0130] The model loss value is less than a preset loss value threshold; and / or
[0131] The training rounds have reached the preset number of training rounds.
[0132] The preset loss threshold can be understood as the user-defined expected loss value. A loss value less than this threshold indicates that the current model has been successfully trained and meets the user's expectations. Training epochs refer to the number of times the model is trained using sample data; the preset training epochs are the number of times the model is trained using sample data, as set by the user. Once the preset number of training epochs has been reached, the model stops training.
[0133] In one specific embodiment provided in this specification, taking the stopping of training of a text processing model by the loss value being less than a preset loss threshold as an example, the preset loss threshold is 0.5. When the calculated loss value is less than 0.5, the text processing model is considered to have completed training.
[0134] In another specific embodiment provided in this specification, taking a preset number of training rounds to stop training information text processing as an example, the preset number of training rounds is 20 rounds. When the training rounds of the sample data reach 20 rounds, it is determined that the information text processing has been trained.
[0135] Step 206: Obtain the target text information output by the text processing model.
[0136] The target text information is the slot information obtained after the text has been corrected in this round of correction.
[0137] In a specific embodiment of this specification, following the example above, the target text information output by the text processing model is obtained, and the target text information is "18812395678".
[0138] This specification provides a text processing method, comprising: acquiring corrected text, historical text corresponding to the corrected text, and text information to be processed; inputting the corrected text, the historical text, and the text information to be processed into a text processing model, wherein the text processing model includes a determination unit, an extraction unit, and a generation unit; the determination unit determines correction category information based on the corrected text, the historical text, and the text information to be processed; the extraction unit extracts the corrected text information from the corrected text; the generation unit generates target text information based on the correction category information, the corrected text information, and the text information to be processed; and acquiring the target text information output by the text processing model. By using correction category information, the traditional slot-level state update is refined to a two-level operation granularity of segment operations and bit operations. Furthermore, a non-autoregressive generation method is used during bit operations to generate the content of all bits of the text information to be processed at once, avoiding the error accumulation caused by prediction errors in autoregressive generation methods. The method provided in this specification satisfies the need for slot information collection when users express the same slot information in multiple rounds, and also improves collection accuracy and efficiency.
[0139] The following is in conjunction with the appendix Figure 5 Taking the application of the text processing method provided in this specification in collecting license plate number slot information as an example, the text processing method will be further explained. Figure 5 The present specification shows a flowchart of a text processing method according to an embodiment of the present specification, with specific steps including steps 502 to 512.
[0140] Step 502: Obtain the corrected text, the historical text corresponding to the corrected text, and the text information to be processed.
[0141] In one embodiment of this specification, the corrected text "is Su D", the historical text "Please state your license plate number. Su D12345. Your license plate number is Su B12345", and the text information to be processed "Su B12345" are obtained.
[0142] Step 504: Generate a text sequence to be processed based on the corrected text, the historical text, and the text information to be processed.
[0143] In one embodiment of this specification, a text sequence to be processed is generated based on the corrected text, historical text, and text information to be processed. The text sequence to be processed is "[CLS] Please state your license plate number. Su D12345. Your license plate number is Su B12345. [SEP] It is Su D. [SEP] Su B12345 [SEP]".
[0144] Step 506: Calculate the correction category weight corresponding to each correction category information in the correction category information set based on the text sequence to be processed.
[0145] In one embodiment of this specification, the correction category weight corresponding to each correction category information in the correction category information set is calculated based on the text sequence to be processed. The correction category weights corresponding to each correction category information are as follows: all updates: 0.1, all cleared: 0.1, remain unchanged: 0.3, append content: 0.1, and partially updated: 0.4.
[0146] Step 508: Determine the target correction category information from the correction category information set according to the correction category weight corresponding to each correction category information.
[0147] In one embodiment of this specification, the target correction category information is determined to be partially updated based on the correction category weights corresponding to the five correction category information.
[0148] Step 510: Extract the correction text information from the correction text.
[0149] In one embodiment of this specification, the correction text information is extracted from the correction text, and the correction text information is "Su D".
[0150] Step 512: Generate target text information based on the correction category information, the correction text information, and the text information to be processed.
[0151] In one embodiment of this specification, the text information to be processed is corrected based on the correction category information and the correction text information. The target text to be corrected is determined to be "Su B". The representation of each bit of the text information to be processed "Su B12345" is extracted. The target text information is obtained as "Su D12345" by generating the text information bit by bit in a non-autoregressive manner according to the correction text information and the target text to be corrected "Su B".
[0152] This specification provides a text processing method, including acquiring corrected text, corresponding historical text, and text information to be processed; generating a text sequence to be processed based on the corrected text, historical text, and text information to be processed; calculating the correction category weight corresponding to each correction category information in a correction category information set based on the text sequence to be processed; determining the target correction category information in the correction category information set based on the correction category weight corresponding to each correction category information; extracting corrected text information from the corrected text; and generating target text information based on the correction category information, the corrected text information, and the text information to be processed. By using correction category information, the traditional slot-level state update is refined to a two-level operation granularity of segment operations and bit operations. Furthermore, a non-autoregressive generation method is used during bit operations to generate the content of all bits of the text information to be processed at once, avoiding the error accumulation caused by prediction errors in autoregressive generation methods. The method provided in this specification not only satisfies the need for slot information collection when the user expresses the same slot information in multiple rounds, but also improves the accuracy and efficiency of collection.
[0153] Corresponding to the above method embodiments, this specification also provides embodiments of a text processing device. Figure 6 A schematic diagram of the structure of a text processing apparatus according to an embodiment of this specification is shown. Figure 6 As shown, the device includes:
[0154] The acquisition module 602 is configured to acquire the corrected text, the historical text corresponding to the corrected text, and the text information to be processed;
[0155] The determination module 604 is configured to determine correction category information based on the corrected text, the historical text, and the text information to be processed.
[0156] Extraction module 606 is configured to extract correction text information from the correction text;
[0157] The generation module 608 is configured to generate target text information based on the correction category information, the correction text information, and the text information to be processed.
[0158] Optionally, the acquisition module 602 is further configured to:
[0159] A sequence of text to be processed is generated based on the corrected text, the historical text, and the text information to be processed.
[0160] Optionally, the acquisition module 602 is further configured to:
[0161] An initial text sequence is generated based on the corrected text, the historical text, and the text information to be processed;
[0162] The initial text sequence is preprocessed to obtain the text sequence to be processed.
[0163] Optionally, the determining module 604 is further configured to:
[0164] The correction category weight corresponding to each correction category information in the correction category information set is calculated based on the text sequence to be processed, wherein the correction category information set includes at least two correction category information.
[0165] The target correction category information is determined from the correction category information set based on the correction category weight corresponding to each correction category information.
[0166] Optionally, the determining module 604 is further configured to:
[0167] The correction category information in the correction category information set is sorted in descending order according to the correction category weight corresponding to each correction category;
[0168] The target correction category information is determined based on the sorting results.
[0169] Optionally, the extraction module 606 is further configured to:
[0170] The corrected text is annotated.
[0171] Extract the corrected text information from the corrected text based on the annotation results.
[0172] Optionally, the generation module 608 is further configured to:
[0173] Based on the correction category information and the correction text information, the target text information to be corrected is determined from the text information to be processed;
[0174] The target text information to be corrected is obtained by correcting the target text information in the text information to be processed based on the corrected text information.
[0175] Optionally, the generation module 608 is further configured to:
[0176] Based on the correction category information and the correction text information, at least one piece of text information to be corrected is determined from the text information to be processed;
[0177] Calculate the correction similarity between the corrected text information and each piece of text information to be corrected;
[0178] The target text information to be corrected is determined based on the correction similarity corresponding to each text information to be corrected.
[0179] The above is an illustrative scheme of the text processing apparatus of this embodiment. It should be noted that the technical solution of the text processing apparatus and the technical solution of the above-described text processing method belong to the same concept. For details not described in detail in the technical solution of the text processing apparatus, please refer to the description of the technical solution of the above-described text processing method.
[0180] Corresponding to the above method embodiments, this specification also provides embodiments of a text processing device. Figure 7 A schematic diagram of the structure of a text processing apparatus according to an embodiment of this specification is shown. Figure 7 As shown, the device includes:
[0181] The first acquisition module 702 is configured to acquire the corrected text, the historical text corresponding to the corrected text, and the text information to be processed.
[0182] Input module 704 is configured to input the corrected text, the historical text, and the text information to be processed into a text processing model, wherein the text processing model includes a determining unit, an extraction unit, and a generating unit. The determining unit determines correction category information based on the corrected text, the historical text, and the text information to be processed. The extraction unit extracts the corrected text information from the corrected text. The generating unit generates target text information based on the correction category information, the corrected text information, and the text information to be processed.
[0183] The second acquisition module 706 is configured to acquire the target text information output by the text processing model.
[0184] The above is an illustrative scheme of the text processing apparatus of this embodiment. It should be noted that the technical solution of the text processing apparatus and the technical solution of the above-described text processing method belong to the same concept. For details not described in detail in the technical solution of the text processing apparatus, please refer to the description of the technical solution of the above-described text processing method.
[0185] Figure 8 A structural block diagram of a computing device 800 according to an embodiment of this specification is shown. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.
[0186] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Wi-MAX interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0187] In one embodiment of this specification, the above-described components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0188] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 800 can also be a mobile or stationary server.
[0189] The processor 820 implements the text processing method when executing the computer instructions.
[0190] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described text processing method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described text processing method.
[0191] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the text processing method as described above.
[0192] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-described text processing method belong to the same concept, and all details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the above-described text processing method.
[0193] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described text processing method.
[0194] The above is an illustrative example of a computer program according to this embodiment. It should be noted that the technical solution of this computer program belongs to the same concept as the technical solution of the text processing method described above. Details not described in detail in the computer program's technical solution can be found in the description of the technical solution of the text processing method described above.
[0195] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0196] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0197] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0198] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0199] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A text processing method, comprising: Obtain the corrected text, the corresponding historical text, and the text to be processed; Correction category information is determined based on the corrected text, the historical text, and the text to be processed, wherein the correction category information includes segment operation type and bit operation type; Extract the correction text information from the correction text; Target text information is generated based on the correction category information, the correction text information, and the text information to be processed; The process of generating target text information based on the correction category information, the correction text information, and the text information to be processed includes: When the correction category information is a bit operation type, the representation of all bits of the text information to be processed is extracted, and bit prediction is performed on the text information to be processed according to the correction text information and the target text information to be corrected to generate target text information, wherein the target text information to be corrected is determined in the text information to be processed based on the correction category information and the correction text information.
2. The method as described in claim 1, after obtaining the corrected text, the historical text corresponding to the corrected text, and the text information to be processed, further includes: A sequence of text to be processed is generated based on the corrected text, the historical text, and the text information to be processed.
3. The method as described in claim 2, wherein generating a text sequence to be processed based on the corrected text, the historical text, and the text information to be processed comprises: An initial text sequence is generated based on the corrected text, the historical text, and the text information to be processed; The initial text sequence is preprocessed to obtain the text sequence to be processed.
4. The method as described in claim 2, wherein determining correction category information based on the corrected text, the historical text, and the text information to be processed includes: The correction category weight corresponding to each correction category information in the correction category information set is calculated based on the text sequence to be processed, wherein the correction category information set includes at least two correction category information. The target correction category information is determined from the correction category information set based on the correction category weight corresponding to each correction category information.
5. The method of claim 4, wherein the target correction category information is determined from the correction category information set according to the correction category weight corresponding to each correction category information, comprising: The correction category information in the correction category information set is sorted in descending order according to the correction category weight corresponding to each correction category; The target correction category information is determined based on the sorting results.
6. The method as described in claim 1, wherein extracting the correction text information from the correction text comprises: The corrected text is annotated. Extract the corrected text information from the corrected text based on the annotation results.
7. The method of claim 1, wherein generating target text information based on the correction category information, the correction text information, and the text information to be processed comprises: Based on the correction category information and the correction text information, the target text information to be corrected is determined from the text information to be processed; The target text information to be corrected is obtained by correcting the target text information in the text information to be processed based on the corrected text information.
8. The method of claim 7, wherein determining the target text information to be corrected in the text information to be processed based on the correction category information and the correction text information includes: Based on the correction category information and the correction text information, at least one piece of text information to be corrected is determined from the text information to be processed; Calculate the correction similarity between the corrected text information and each piece of text information to be corrected; The target text information to be corrected is determined based on the correction similarity corresponding to each text information to be corrected.
9. A text processing method, comprising: Obtain the corrected text, the corresponding historical text, and the text to be processed; The corrected text, the historical text, and the text to be processed are input into a text processing model, wherein the text processing model includes a determination unit, an extraction unit, and a generation unit. The determination unit determines correction category information based on the corrected text, the historical text, and the text to be processed. The extraction unit extracts the corrected text information from the corrected text. The generation unit generates target text information based on the correction category information, the corrected text information, and the text to be processed. The correction category information includes segment operation type and bit operation type. Obtain the target text information output by the text processing model; The generation unit is further configured to, when the correction category information is a bit operation type, extract the representation of all bits of the text information to be processed, perform bit prediction on the text information to be processed according to the correction text information and the target text information to be corrected, and generate target text information, wherein the target text information to be corrected is determined in the text information to be processed based on the correction category information and the correction text information.
10. The method of claim 9, wherein the text processing model is obtained by training through the following steps: Acquire training correction text, correction text information, training history text corresponding to the training correction text, and training text information to be processed; The training correction text, the training history text, and the training text to be processed are input into the text processing model to obtain the predicted text information output by the text processing model. Calculate the model loss value based on the predicted text information and the corrected text information; The model parameters of the text processing model are adjusted based on the model loss value, and the text processing model continues to be trained until the model training stops.
11. A computing device comprising a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor, when executing the computer instructions, performs the steps of the method according to any one of claims 1-8 or 9-10.
12. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-8 or 9-10.
Citation Information
Patent Citations
Punctuation processing method and device of speech recognition text
CN108564953A
Dialogue repairing method and device suitable for multiple rounds of dialogues, equipment and storage medium
CN113239152A