Text processing method and device and electronic equipment
By integrating the ASR post-processing process into the generative model and using character paths to constrain the output characters, the problem of mutual influence between modules in the existing technology is solved, and the readability and accuracy of the speech recognition results are improved.
Patent Information
- Application Number
- CN202510551055.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-09-19
AI Technical Summary
The post-processing model of existing speech recognition results consists of multiple independent modules. Adjustments between modules will affect each other, resulting in unstable output results. In addition, traditional algorithms have limited performance and high development and maintenance costs.
The ASR post-processing process is integrated into the generative model. The character path is determined based on the initial text, and the output characters of the generative model are constrained. The text processing generalization ability and character path of the generative model are used to match the output characters to avoid hallucination problems.
It improves the effects of punctuation, number conversion, dirty word filtering and text smoothing, improves the readability and accuracy of the target text, avoids mutual influence between modules, and reduces development and maintenance costs.
Smart Images

Figure CN120671666A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of large model technology, and in particular to a text processing method, device and electronic device. Background Art
[0002] In related technologies, post-processing of speech recognition results is implemented by multiple independent modules. Multiple modules independently optimize the speech recognition results in a certain order. Adjustment of any module will affect other modules, thereby affecting the final output result. Summary of the Invention
[0003] The present disclosure provides a text processing method, device, and electronic device to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, there is provided a text processing method, comprising:
[0005] Determining prompt information based on the initial text corresponding to the target voice;
[0006] Inputting the prompt information into a generative model, using character paths as constraints, and generating output characters corresponding to the initial text based on the prompt information;
[0007] The type of the output characters includes at least one of a text type, a number type, a punctuation type, and a space character type, and the output characters constitute the target text corresponding to the target speech;
[0008] The character path includes at least one character in the initial text.
[0009] In the above solution, the step of determining the prompt information based on the initial text corresponding to the target speech includes:
[0010] Determining prompt information based on the initial text corresponding to the target speech and the task identifier;
[0011] Different task identifiers correspond to different character paths.
[0012] In the above solution, the method further includes:
[0013] The character path corresponding to the task identifier is determined according to the initial text.
[0014] In the above solution, determining the character path corresponding to the task identifier according to the initial text includes at least one of the following:
[0015] In response to the task identifier representing a punctuation task, determining the character path based on a concatenation of all characters in the initial text;
[0016] In response to the task identifier representing a digital conversion task, determining the character path based on a concatenation of characters other than the digital intention character in the initial text;
[0017] In response to the task identifier representing a full-type task, determining the character path based on a concatenation of characters other than the numeric intention character in the initial text;
[0018] In response to the task identifier representing a full-type task, the character path is determined based on a concatenation of characters other than numeric intention characters and modal particles in the initial text.
[0019] In the above solution, the prompt information is input into the generative model, and the character path is used as a constraint to generate the output characters corresponding to the initial text based on the prompt information, including:
[0020] The generative model infers the initial text based on the task identifier in the prompt information, and constrains the inference result based on the character path to generate output characters that match the character path; the target text composed of the output characters is the initial text corrected by the generative model.
[0021] In the above solution, the training process of the generative model includes:
[0022] Determining first-type task data based on a task identifier corresponding to the punctuation task and input data;
[0023] Inputting the first type of task data into the generative model to obtain first prediction data;
[0024] Adjust the weight of the generative model based on the first prediction data and first identification data corresponding to the input data.
[0025] In the above solution, the training process of the generative model includes:
[0026] Determining second-type task data based on a task identifier and input data corresponding to the digital conversion task;
[0027] Inputting the second type of task data into the generative model to obtain second prediction data;
[0028] Based on the second prediction data and second identification data corresponding to the input data, the weight of the generative model is adjusted.
[0029] In the above solution, the training process of the generative model includes:
[0030] Determining third-type task data based on task identifiers and input data corresponding to all-type tasks; wherein the all-type tasks include at least two of a punctuation task, a number conversion task, and a modal particle removal task;
[0031] Inputting the third type of task data into the generative model to obtain third prediction data;
[0032] Based on the third prediction data and third identification data corresponding to the input data, the weight of the generative model is adjusted.
[0033] According to a second aspect of the present disclosure, a text processing device is provided, the device comprising:
[0034] a determining unit, configured to determine prompt information based on an initial text corresponding to the target speech;
[0035] an inference unit, configured to input the prompt information into a generative model, and generate output characters corresponding to the initial text based on the prompt information, using character paths as constraints;
[0036] The type of the output characters includes at least one of a text type, a number type, a punctuation type, and a space character type, and the output characters constitute the target text corresponding to the target speech;
[0037] The character path includes at least one character in the initial text.
[0038] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0039] at least one processor; and
[0040] a memory communicatively connected to the at least one processor; wherein,
[0041] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.
[0042] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0044] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0045] Figure 1 A schematic diagram of an ASR post-processing model in related art is shown;
[0046] Figure 2 A first optional flow chart of the text processing method provided by an embodiment of the present disclosure is shown;
[0047] Figure 3 A second optional flow chart of the text processing method provided by an embodiment of the present disclosure is shown;
[0048] Figure 4 A third optional flow chart of the text processing method provided by the embodiment of the present disclosure is shown;
[0049] Figure 5 A fifth optional flow chart of the text processing method provided by the embodiment of the present disclosure is shown;
[0050] Figure 6 A fifth optional flow chart of the text processing method provided by the embodiment of the present disclosure is shown;
[0051] Figure 7 A processing schematic diagram provided by an embodiment of the present disclosure is shown;
[0052] Figure 8 An optional structural diagram of a text processing device provided by an embodiment of the present disclosure is shown;
[0053] Figure 9 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0054] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.
[0055] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0056] In the following description, the terms "first\second" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.
[0057] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by those skilled in the art in the art of this disclosure. The terms used in this disclosure are only for the purpose of describing the embodiments of this disclosure and are not intended to limit this disclosure.
[0058] It should be understood that in the various embodiments of the present disclosure, the size of the serial number of each implementation process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.
[0059] In related technologies, the basic task of Automatic Speech Recognition (ASR) requires the conversion from speech to text. The output recognition result text is usually the normalized text of the actual transcription of the speech. The reason for outputting the normalized text is that the text used in the training of the speech recognition model has mostly been normalized. However, the recognition result text has not been denormalized, and there will be problems such as no punctuation, no Arabic numeral conversion, and no text smoothing such as removing redundant modal particles, which affects the user experience. Therefore, in order to improve the readability of the recognition result text, the recognition result text will be denormalized through the post-processing model.
[0060] Figure 1 A schematic diagram of an ASR post-processing model in related art is shown.
[0061] In related technologies, such as Figure 1 As shown, the ASR post-processing model consists of multiple independent modules: the punctuation module, the Arabic numeral conversion module, the dirty word filtering algorithm / rule module, and the text smoothing module. Each independent module can be a combination of a model and rules to complete different tasks. For example, the punctuation module is typically a Long Short-Term Memory (LSTM) model, and the Arabic numeral conversion module is usually a combination of an algorithm and regular expressions.
[0062] However, in the existing ASR post-processing model, each module is optimized independently, and there is a data processing sequence between each module. Modifications to any module may affect at least one other module. In addition, traditional algorithm models and algorithm performance are limited, and the rule base is getting larger and larger, and the development and maintenance costs are also increasing.
[0063] For example, the output of the ASR model is "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995." After being processed by the post-processing model, the resulting text is "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995." Or, the output of the ASR model is "I explained my origins in detail." After being processed by the post-processing model, the resulting text is "I explained my origins in detail."
[0064] In response to the defects existing in the related art, the embodiments of the present disclosure provide a text processing method, which integrates the ASR post-processing process into the generative model. At the same time, in order to avoid the hallucination problem of the generative model in the task, the character path is determined based on the initial text output by the ASR model, and the output characters of the generative model are constrained to ensure that the intention or characters of the processing results of the generative model match the initial text.
[0065] Figure 2 A first optional flow chart of the text processing method provided by an embodiment of the present disclosure is shown, and will be explained according to each part.
[0066] Step S201: determining prompt information based on the initial text corresponding to the target speech.
[0067] In some embodiments, the initial text is text obtained after the target speech is recognized by the ASR model; the prompt information includes information (prompt) input to the generative model, which is used to instruct the generative model to generate relevant content or complete a corresponding task. The prompt information includes the initial text.
[0068] In some optional embodiments, a carrier that implements the text processing method (hereinafter referred to as the carrier) can add corresponding identifiers to the initial text based on the functions that need to be implemented or the tasks that need to be completed by the generative model to obtain the prompt information.
[0069] The carrier can be a computer program, electronic circuit, database, mobile application, electronic device, cloud computing platform, distributed system, artificial intelligence framework, mathematical model, automation tool and microcontroller, etc., which can implement software or hardware of algorithm and method process.
[0070] Step S202: input the prompt information into a generative model.
[0071] In some embodiments, the carrier inputs the prompt information into the generative model, so that the generative model uses the character path as a constraint and generates output characters corresponding to the initial text based on the prompt information. The output characters include at least one of a text type, a number type, a punctuation type, and a whitespace character type, and the output characters constitute the target text corresponding to the target speech.
[0072] In some embodiments, the character path is determined based on the initial text and may include at least one character in the initial text to constrain the output characters generated by the generative model so that the output characters include at least at least one character in the initial text, or so that the output characters match the initial text; the matching may include at least one of the output character being the same as at least one character in the initial text, the intention corresponding to the target text being the same as the intention corresponding to the initial text, or the order of multiple output characters in the target text being the same as the order of corresponding characters in the initial text.
[0073] In this way, through the text processing method provided by the embodiment of the present disclosure, the initial text output by the ASR model is processed based on the generalization ability of text processing of the generative model, thereby improving the effects of punctuation, number conversion, dirty word filtering and text smoothing; at the same time, the output characters of the generative model are constrained based on the character path, so that the output characters match the initial text, avoiding the hallucination problem and improving the readability and accuracy of the target text.
[0074] Figure 3 A second optional flow chart of the text processing method provided by the embodiment of the present disclosure is shown, and will be explained according to each step.
[0075] Step S301: determining prompt information based on the initial text corresponding to the target speech and the task identifier.
[0076] In some embodiments, the task identifier may include at least one of a punctuation task identifier, a digital conversion task identifier, and a full-type task identifier; the full-type task includes at least two of a punctuation task, a digital conversion task, and a modal particle removal task.
[0077] In some embodiments, the carrier may determine the prompt information by adding a task identifier to the initial text. Alternatively, the carrier may add at least one task identifier to the initial text. If the carrier adds two task identifiers to the initial text, the tasks corresponding to the two task identifiers may or may not overlap. For example, the carrier may add a punctuation task identifier and a digital conversion task identifier to the initial text to obtain the prompt information. Alternatively, the carrier may add a punctuation task identifier and a full-type task identifier to the initial text to obtain the prompt information.
[0078] Step S302: input the prompt information into a generative model.
[0079] In some embodiments, the carrier inputs the prompt information into the generative model, so that the generative model uses the character path as a constraint and generates output characters corresponding to the initial text based on the prompt information.
[0080] In specific implementation, in response to the task identifier representing a punctuation task, the type of the output character includes at least a text type and a punctuation type; in response to the task identifier representing a digital conversion task, the type of the output character includes at least a text type and a digital type; in response to the task identifier representing an all-type task, the type of the output character includes at least two of a text type, a digital type, a punctuation type and a blank character type.
[0081] The punctuation task involves adding punctuation to the initial text to produce the target text, so the output characters include punctuation characters. The number conversion task involves converting characters representing numerical intent in the initial text into Arabic numerals to produce the target text, such as converting year characters into Arabic numerals. Whitespace characters correspond to modal particles in the initial text, meaning that modal particles in the initial text are converted into whitespace characters after being processed by the generative model.
[0082] In some embodiments, the character path is determined based on the initial text and may include at least one character in the initial text to constrain the output characters generated by the generative model so that the output characters include at least at least one character in the initial text, or so that the output characters match the initial text; the matching may include at least one of the output character being the same as at least one character in the initial text, the intention corresponding to the target text being the same as the intention corresponding to the initial text, or the order of multiple output characters in the target text being the same as the order of corresponding characters in the initial text.
[0083] In this way, through the text processing method provided by the embodiment of the present disclosure, the initial text output by the ASR model is processed based on the generalization ability of text processing of the generative model, thereby improving the effects of punctuation, number conversion, dirty word filtering and text smoothing; at the same time, the output characters of the generative model are constrained based on the character path, so that the output characters match the initial text, avoiding the hallucination problem and improving the readability and accuracy of the target text.
[0084] Figure 4 A third optional flow chart of the text processing method provided by an embodiment of the present disclosure is shown, and will be explained according to each step.
[0085] Step S401: Determine a prompt message based on the initial text corresponding to the target voice and the task identifier.
[0086] The specific step flow of step S401 is the same as that of step S301, and will not be repeated here.
[0087] Step S402: Determine the character path corresponding to the task identifier according to the initial text.
[0088] In some embodiments, different task identifiers correspond to different character paths; the carrier determines multiple characters constituting the character path from the initial text based on the task identifier, and concatenates the multiple characters to obtain the character path. Specifically, the initial text output by the ASR model has the ability to constrain the target voice. When it is not desired that the target text output by the generative model is significantly different from the content of the initial text or the target voice, determining the character path based on the initial text can achieve the constraint on the output characters of the generative model.
[0089] Specifically, when the task identifier represents a punctuation task, the character path is determined based on the concatenation of all characters in the initial text; that is, the characters included in the character path correspond one-to-one to the characters included in the initial text, and the order of the multiple characters in the character path is the same as the order of the characters in the initial text.
[0090] Specifically, when the task identifier represents a number conversion task, the character constraint is determined based on the concatenation of the characters other than the number-intention characters in the initial text; the number-intention characters include characters (or text characters) with the intention of numbers, such as "one, two, three, four, five, six, seven, eight, nine, ten", or "one, two, three, four, five, six, seven, eight, nine, ten", etc. The characters in the character path correspond one-to-one to the characters other than the number-intention characters in the initial text, and the order of the multiple characters in the character path is the same as the order of the characters in the initial text.
[0091] Specifically, when the task identifier represents an all-type task, the character path is determined based on the concatenation of the characters other than the number-intention characters in the initial text; that is, the character path is the same as the character path corresponding to the number conversion task.
[0092] Alternatively, specifically, when the task identifier represents an all-type task, the character path is determined based on the concatenation of the characters other than the number-intention characters and the modal particles in the initial text. The modal particles include function words used to express emotions, attitudes or tones, usually located at the beginning, end or pause of a sentence, without承担 actual lexical meaning, but can affect the tone and emotional color of the sentence.
[0093] Step S403: input the prompt information into the generative model.
[0094] In some embodiments, the carrier inputs the prompt information into the generative model so that the generative model infers the initial text based on the task identifier in the prompt information, and constrains the inference result based on the character path to generate output characters that match the character path; the target text composed of the output characters is the initial text corrected by the generative model.
[0095] Among them, the output characters that match the character path include the output characters that match (or are the same as) the characters in the character path, the order of multiple output characters is the same as the order of corresponding characters in the character path, and the intentions of multiple output characters are the same as the intentions of corresponding characters in the character path.
[0096] In some embodiments, the target speech is input into the ASR model, and an initial text is obtained through recognition. The initial text is a conversion of the target speech, which includes modal particles, characters intended to be numbers, and does not include punctuation marks; the generative model is used to correct the initial text to obtain the target text; the generative model corrects the initial text by implementing a task corresponding to the task identifier based on the task identifier. If the prompt information includes a task identifier for a punctuation task, the generative model adds punctuation marks at the corresponding position in the initial text, and constrains the output characters based on the character path so that the output characters match the characters in the character path; if If the prompt information includes a task identifier for a digital conversion task, the generative model converts the digital intended characters in the initial text into Arabic numerals, and constrains the non-digital intended characters in the initial text based on the character path, so that the output characters match the characters in the character path; if the prompt information includes a task identifier for a full-type task, the generative model removes the modal particles in the initial text, adds punctuation marks at the corresponding positions in the initial text, and converts the digital intended characters in the initial text into at least two Arabic numerals, and constrains the non-character type characters in the initial text based on the character path, so that the output characters match the characters in the character path.
[0097] In this way, through the text processing method provided by the embodiment of the present disclosure, the initial text output by the ASR model is processed based on the text processing generalization capability of the generative model and the constraints of the character path, thereby improving the effects of punctuation, number conversion, dirty word filtering and text smoothing; at the same time, the output characters of the generative model are constrained based on the character path, so that the output characters match the initial text, avoiding the hallucination problem and improving the readability and accuracy of the target text.
[0098] Figure 5A fifth optional flow chart of the text processing method provided by an embodiment of the present disclosure is shown, and will be explained according to each step.
[0099] Step S501: punctuation task training.
[0100] In some embodiments, the carrier determines first type of task data based on the task identifier and input data corresponding to the punctuation task; inputs the first type of task data into the generative model to obtain first prediction data; and adjusts the weight of the generative model based on the first prediction data and the first identification data corresponding to the input data.
[0101] In a specific implementation, different tasks correspond to different task identifiers. The input data may be training data, and the first prediction data may be data after the generative model adds punctuation to the input data; the first identification data is data after the expected punctuation is added to the expected position of the input data.
[0102] In some embodiments, the carrier can determine a loss function based on the first prediction data and the first identification data, adjust the weight of the generative model based on the loss function, and obtain prediction data again based on the adjusted generative model until the loss function is less than a threshold or the number of training times reaches an upper limit.
[0103] In other embodiments, the carrier may determine first prompt information based on the first prediction data and the first identification data, input the first prompt information into a generative model, and instruct the generative model to generate new prediction data based on the first prompt information. The first prompt information includes position and type prompt information of punctuation in the first prediction data, thereby instructing the generative model to add punctuation to the input data based on the position and type indication information.
[0104] For example, the task identifier corresponding to the punctuation task can be "<|punc-prediction|>". Assuming that the input data obtained based on the ASR model is "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995", the first type of task data can be "<|startofturn|>user\nAh, she was born on August 18, 1986, and her brother was born on March 1, 1995<|endofturn|><|punc-prediction|><|startofturn|>model\nAh, she was born on August 18, 1986, and her brother was born on March 1, 1995.<|endofturn|>"; among them, "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995." is the first identifier data.
[0105] Step S502: digital conversion task training.
[0106] In some embodiments, the carrier determines second type task data based on the task identifier and input data corresponding to the digital conversion task; inputs the second type task data into the generative model to obtain second prediction data; and adjusts the weight of the generative model based on the second prediction data and the second identification data corresponding to the input data.
[0107] Specifically, the second prediction data may be the data after the generative model converts the digital intention characters in the input data into Arabic numerals; the second identification data may be the data after the digital intention characters after the input data are converted into the expected Arabic numerals.
[0108] In some embodiments, the carrier can determine a loss function based on the second prediction data and the second identification data, adjust the weight of the generative model based on the loss function, and obtain prediction data again based on the adjusted generative model until the loss function is less than a threshold or the number of training times reaches an upper limit.
[0109] In other embodiments, the carrier may determine second prompt information based on the second prediction data and the second identification data, input the second prompt information into the generative model, and instruct the generative model to generate new prediction data based on the second prompt information. The second prompt information includes prompt information corresponding to the converted Arabic numerals in the second prediction data, such as incorrect position, incorrect form, or incorrect content, to instruct the generative model to convert the numeric intended characters in the input data into more accurate Arabic numerals based on the prompt information corresponding to the converted Arabic numerals.
[0110] For example, the task identifier corresponding to the digital conversion task can be "<|arabic-num-norm|>". Assuming that the input data obtained based on the ASR model is "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995", the second type of task data can be "<|startofturn|>user\nAh, she was born on August 18, 1986, and her brother was born on March 1, 1995<|endofturn|><|arabic-num-norm|><|startofturn|>model\nAh, she was born on August 18, 1986, and her brother was born on March 1, 1995.<|endofturn|>"; among them, "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995." is the second identification data.
[0111] Since there are also digital intent characters in some words, such as "one by one", "kill two birds with one stone", "all directions", etc., if the digital intent characters in them are converted into Arabic numerals, it will instead affect the accuracy of the target text. Therefore, during the training process, the prediction data of the generative model needs to be corrected through the loss function or the second hint information to prevent the generative model from converting the digital intent characters in idioms into Arabic numerals. In this case, the prediction data corresponding to the digital intent characters is the characters themselves.
[0112] In some embodiments, the carrier can perform punctuation task training and digital conversion task training in the order shown in this disclosure. After executing step S501, the generative model has the ability to complete the punctuation task. On this basis, during the process of performing the digital conversion task, the generative model will perform the punctuation task on the input data corresponding to the digital conversion task, that is, add punctuation to the input data.
[0113] Step S503, full-type task training.
[0114] In some embodiments, the carrier determines the third type of task data based on the task identifier and input data corresponding to the full-type task; inputs the third type of task data into the generative model to obtain the third prediction data; and adjusts the weights of the generative model based on the third prediction data and the third identifier data corresponding to the input data.
[0115] Specifically, the third prediction data can be data of at least two of adding punctuation to the generative model, converting the digital intent characters in the input data into Arabic numerals, and removing modal particles; the third identifier data is the data after adding the expected punctuation at the expected position and converting the digital intent characters after the input data into the expected Arabic numerals.
[0116] In some embodiments, the carrier can determine the loss function based on the third prediction data and the third identifier data, adjust the weights of the generative model based on the loss function, and obtain the prediction data again based on the adjusted generative model until the loss function is less than the threshold or the number of training times reaches the upper limit.
[0117] In other embodiments, the carrier may determine third prompt information based on the third prediction data and the third identification data, input the third prompt information into the generative model, and instruct the generative model to generate new prediction data based on the third prompt information. The third prompt information includes at least two of the following: position prompt information and type prompt information of punctuation marks in the third prediction data, prompt information corresponding to the converted Arabic numerals, and prompt information related to modal particles; thereby instructing the generative model to convert the numeric characters in the input data into more accurate Arabic numerals based on the prompt information corresponding to the converted Arabic numerals.
[0118] For example, the task identifier corresponding to the full-type task can be "<|full-norm|>". Assuming that the input data obtained based on the ASR model is "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995", the third-type task data can be "<|startofturn|>user\nAh, she was born on August 18, 1986, and her brother was born on March 1, 1995<|endofturn|><|full-norm|><|startofturn|>model\nShe was born on August 18, 1986, and her brother was born on March 1, 1995.<|endofturn|>"; among them, "She was born on August 18, 1986, and her brother was born on March 1, 1995." is the third identifier data.
[0119] In some optional embodiments, the carrier can train a new, untrained generative model so that the trained generative model can achieve one of the above-mentioned punctuation tasks, digital conversion tasks and all-type tasks; or, the carrier can use a pre-trained generative model as an initial model (base model) to train the initial model for one of the punctuation tasks, digital conversion tasks and all-type tasks, so that the trained initial model (i.e., the generative model) can adapt to one of the punctuation tasks, digital conversion tasks and all-type tasks.
[0120] In some embodiments, step S501, step S502 and step S503 can be independent training tasks, that is, they are executed according to the functions that the generative model needs to implement (or the tasks that the generative model needs to complete), that is, only at least one of step S501, step S502 and step S503 is executed; training can also be performed in a certain order, that is, punctuation task training is performed first (step S501), and then digital conversion task training is performed (step S502), and finally, after punctuation task training and digital conversion task training, all types of task training is performed (step S503).
[0121] In this way, through the text processing method provided by the embodiment of the present disclosure, based on the generalization ability of text processing of the generative model, the generative model is trained on punctuation tasks, number conversion tasks and all-type tasks respectively, so that the three tasks of adding punctuation, number conversion and removing modal particles are implemented by the generative model, thereby improving the post-processing performance of the ASR model and avoiding the problem of mutual influence between multiple models in related technologies.
[0122] Figure 6 FIG. 5 shows a fifth optional flow chart of the text processing method provided by the embodiment of the present disclosure. Figure 7 The processing diagram provided by the embodiment of the present disclosure is shown and will be described according to each part.
[0123] Step S601: determining prompt information based on the initial text corresponding to the target speech and the task identifier.
[0124] In some embodiments, the carrier determines a corresponding task identifier based on the function that the generative model needs to implement, and determines prompt information based on the task identifier and the initial text.
[0125] Taking the initial text as "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995" and the task identifier corresponding to the full-type task as an example, the corresponding prompt information is "<|startofturn|>user\nAh, she was born on August 18, 1986, and her brother was born on March 1, 1995<|endofturn|><|full-norm|><|startofturn|>model\n".
[0126] Step S602: determining a character path based on the initial text and the task identifier.
[0127] In some embodiments, the carrier determines, based on the task identifier, that at least one character in the initial text constitutes the character path.
[0128] For example, if the task identifier is the task identifier corresponding to all types of tasks, then the character string obtained by linking other character strings other than modal particles and digital intention characters in the initial text is determined to be the character path. In the embodiment disclosed in the present disclosure, the character path is "She was born on the year-month-day and her brother was born on the year-month-day".
[0129] Correspondingly, if the task identifier is the identifier of a punctuation task, the character path is the same as the initial text, that is, "Ah, she was born on August 18, 1986, and her brother was born on March 1, 1995"; if the task identifier is the identifier of a digital conversion task, the character path is the character string connected with other character strings in the initial text except for the digital intention characters, that is, "Ah, she was born on year, month, and day, and her brother was born on year, month, and day".
[0130] Step S603: Input the prompt information and the character path into the generative model to obtain the target text.
[0131] In some embodiments, the generative model is a model trained based on Steps S501 to S503.
[0132] In some embodiments, the carrier inputs the prompt information into the generative model, causing the generative model to process the initial text in the prompt information based on the task identifier in the prompt information, and to constrain the output characters based on the character path, such that the output characters match the characters in the character path.
[0133] Specifically, if the generative model determines that the task corresponding to the task identifier is a full-type task, it removes the filler words "ah ah" from the initial text, or converts the filler words in the initial text to blank characters; it converts the digital intent characters in the initial text to Arabic numerals and adds punctuation in the initial text. In the above process, only the changes to the filler words and digital intent characters are involved, and other characters remain unchanged. However, to avoid the problem of hallucinations, the generative model uses the character path as a constraint to ensure that the other output characters (i.e., the output characters) match the characters in the character path.
[0134] For example, the generative model converts "ah ah" to blank characters and outputs them, and based on the initial text and the character path, determines that the next output character is "she", and then based on the character path, determines that the subsequent three output characters are "was born in"; the character after "was born in" is a digital type character, and the result of the output data conversion task is output, i.e., "86", and based on the initial text and the character path, determines that the next output character is "year", and outputs sequentially. After the output character is "day", according to the punctuation task result, "," is output, and the above similar operations are repeated until "." is output. All the output characters constitute the target text "She was born on August 18, 1986, and her younger brother was born on March 1, 1995.". Constraining the output characters based on the character path can avoid the hallucination problem of the generative model, such as outputting "he" instead of "she".
[0135] In this way, through the text processing method provided by the embodiment of the present invention, the advantage of the generative model being "more intelligent" than the traditional model and the generative model's ability to generalize text processing are utilized, and the generative model is used to post-process the speech recognition results of the ASR model (i.e., the initial text), and the punctuation task, number conversion task, and modal particle removal task are concentrated in the generative model. The generative model implements the functions corresponding to the above three tasks, realizes text regularization and optimization processing of the initial text, and improves the readability of the target text; in addition, the output characters are constrained based on the character path, or the search strategy when the generative model outputs is restricted based on the character path, so as to avoid the "hallucination" problem of the generative model, ensure that the final target text and the intention and content of the initial text remain unchanged, and ensure the accuracy of the target text.
[0136] Figure 8 An optional structural diagram of a text processing device provided by an embodiment of the present disclosure is shown, and will be described according to each part.
[0137] In some embodiments, the text processing apparatus 700 includes a determination unit 701 and an inference unit 702 .
[0138] The determining unit 701 is configured to determine prompt information based on the initial text corresponding to the target speech;
[0139] The inference unit 702 is configured to input the prompt information into a generative model, and generate output characters corresponding to the initial text based on the prompt information using character paths as constraints;
[0140] The type of the output characters includes at least one of a text type, a number type, a punctuation type, and a space character type, and the output characters constitute the target text corresponding to the target speech;
[0141] The character path includes at least one character in the initial text.
[0142] The determining unit 701 is specifically configured to determine prompt information based on the initial text corresponding to the target speech and the task identifier.
[0143] The determining unit 701 is further configured to determine the character path corresponding to the task identifier according to the initial text; different task identifiers correspond to different character paths.
[0144] The determining unit 701 is specifically configured to do at least one of the following:
[0145] In response to the task identifier representing a punctuation task, determining the character path based on a concatenation of all characters in the initial text;
[0146] In response to the task identifier representing a digital conversion task, determining the character path based on a concatenation of characters other than digital characters in the initial text;
[0147] In response to the task identifier representing a full-type task, determining the character path based on a concatenation of characters other than digits in the initial text;
[0148] In response to the task identifier representing a full-type task, the character path is determined based on a concatenation of characters other than numerals and modal particles in the initial text.
[0149] The inference unit 702 is specifically used to infer the initial text based on the task identifier in the prompt information, and constrain the inference result based on the character path to generate output characters that match the character path; the target text composed of the output characters is the initial text corrected by the generative model.
[0150] In some embodiments, the text processing apparatus 700 may further include a training unit 703, wherein the training unit 703 is configured to determine first-type task data based on a task identifier corresponding to the punctuation task and input data;
[0151] Inputting the first type of task data into the generative model to obtain first prediction data;
[0152] Adjust the weight of the generative model based on the first prediction data and first identification data corresponding to the input data.
[0153] The training unit 703 is further configured to determine second-type task data based on the task identifier and input data corresponding to the digital conversion task;
[0154] Inputting the second type of task data into the generative model to obtain second prediction data;
[0155] Based on the second prediction data and second identification data corresponding to the input data, the weight of the generative model is adjusted.
[0156] The training unit 703 is further configured to determine third-type task data based on task identifiers and input data corresponding to all-type tasks; the all-type tasks include at least two of the punctuation task, the number conversion task, and the modal particle removal task;
[0157] Inputting the third type of task data into the generative model to obtain third prediction data;
[0158] Based on the third prediction data and third identification data corresponding to the input data, the weight of the generative model is adjusted.
[0159] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0160] Figure 9 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0161] like Figure 9 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0162] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0163] The computing unit 801 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the text processing method. For example, in some embodiments, the text processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the text processing method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the text processing method by any other suitable means (e.g., via firmware).
[0164] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0165] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0166] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0167] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0168] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0169] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0170] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0171] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.
[0172] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A text processing method, comprising: Determining prompt information based on the initial text corresponding to the target voice; Inputting the prompt information into a generative model, using character paths as constraints, and generating output characters corresponding to the initial text based on the prompt information; The type of the output characters includes at least one of a text type, a number type, a punctuation type, and a space character type, and the output characters constitute the target text corresponding to the target speech; The character path includes at least one character in the initial text.
2. The method according to claim 1, wherein determining the prompt information based on the initial text corresponding to the target speech comprises: The prompt information is determined based on the initial text corresponding to the target speech and the task identifier.
3. The method according to claim 2, further comprising: Determining the character path corresponding to the task identifier according to the initial text; Different task identifiers correspond to different character paths.
4. The method according to claim 3, wherein determining the character path corresponding to the task identifier according to the initial text comprises at least one of the following: In response to the task identifier representing a punctuation task, determining the character path based on a concatenation of all characters in the initial text; In response to the task identifier representing a digital conversion task, determining the character path based on a concatenation of characters other than the digital intention character in the initial text; In response to the task identifier representing a full-type task, determining the character path based on a concatenation of characters other than the numeric intention character in the initial text; In response to the task identifier representing a full-type task, the character path is determined based on a concatenation of characters other than numeric intention characters and modal particles in the initial text.
5. The method according to claim 2, wherein the prompt information is input into a generative model, and the character path is used as a constraint to generate output characters corresponding to the initial text based on the prompt information, comprising: The generative model infers the initial text based on the task identifier in the prompt information, and constrains the inference result based on the character path to generate output characters that match the character path; the target text composed of the output characters is the initial text corrected by the generative model.
6. The method according to claim 1, wherein the training process of the generative model comprises: Determining first-type task data based on a task identifier corresponding to the punctuation task and input data; Inputting the first type of task data into the generative model to obtain first prediction data; Adjust the weight of the generative model based on the first prediction data and first identification data corresponding to the input data.
7. The method according to claim 1, wherein the training process of the generative model comprises: Determining second-type task data based on a task identifier and input data corresponding to the digital conversion task; Inputting the second type of task data into the generative model to obtain second prediction data; Based on the second prediction data and second identification data corresponding to the input data, the weight of the generative model is adjusted.
8. The method according to any one of claims 1, 6 or 7, wherein the training process of the generative model comprises: Determining third-type task data based on task identifiers and input data corresponding to all types of tasks; The full range of tasks includes at least two of the punctuation task, the number conversion task, and the modal particle removal task; Inputting the third type of task data into the generative model to obtain third prediction data; Based on the third prediction data and third identification data corresponding to the input data, the weight of the generative model is adjusted.
9. A text processing device, comprising: a determining unit, configured to determine prompt information based on an initial text corresponding to the target speech; an inference unit, configured to input the prompt information into a generative model, and generate output characters corresponding to the initial text based on the prompt information, using character paths as constraints; The type of the output characters includes at least one of a text type, a number type, a punctuation type, and a space character type, and the output characters constitute the target text corresponding to the target speech; The character path includes at least one character in the initial text.
10. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 8.