Voice editing method and device, storage medium and electronic device
By using replacement modules, insert modules and comprehensive modules in the intelligent speech recognition system to process user editing commands and statements, the problem of inefficient speech text editing in the prior art is solved, and a more efficient speech editing process is achieved.
Patent Information
- Application Number
- CN202110873669.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-07-30
AI Technical Summary
When editing voice text, the existing intelligent voice recognition system has a long process and a lot of information that needs to be input, resulting in inefficient editing.
By obtaining the edit commands entered by the user and the statements to be edited, it is determined whether the edit command is a descriptive command. If not, it is used as a target statement, input the pre-trained replacement module and insertion module for processing, output candidate replacement statements and insertion statements, and the target candidate statements are selected by the comprehensive module, and the user feedbacks the selection instructions and replaces them.
There is no need for the user to specify the modified text location, just enter a small amount of information to achieve text editing, shorten the voice editing process, improve editing efficiency, and provide users with better services.
Smart Images

Figure CN113591441B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech editing method and device, a storage medium and an electronic device. Background Art
[0002] With the maturity and development of voice technology, the application fields of voice technology are becoming more and more extensive. At present, most smart terminal devices and smart car devices are integrated with intelligent voice recognition systems. Intelligent voice recognition systems can convert the collected user voice into the text content required by the user, thereby providing more convenient services for user communication.
[0003] When the current intelligent speech recognition system converts speech into speech text, there may be errors in text conversion or the speech text needs to be optimized. When editing speech text, users usually need to manually select or repeat the text to be edited, and then enter the modified content to complete the text editing. The current method of editing text through speech has a long process and requires a lot of information to be entered, resulting in low editing efficiency. Summary of the invention
[0004] In view of this, the present invention provides a voice editing method and device, a storage medium and an electronic device. The present invention can predict the text content that the user needs to change based on the target sentence input by the user, and provide the user with a sentence that is more in line with the modification intention. The text editing is realized without specifying the location of the modified text. The entire editing process only requires the user to input a small amount of information, which effectively shortens the editing process of text using voice and effectively improves the efficiency of voice editing.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0006] A voice editing method, comprising:
[0007] Obtaining an editing command input by a user and a sentence to be edited, wherein the sentence to be edited is a sentence selected by the user in the text to be edited, the text to be edited is a text obtained by converting the converted speech input by the user into text, and the editing command is a text obtained by converting the command speech input by the user based on the sentence to be edited into text;
[0008] Determining whether the editing command is a descriptive command;
[0009] If the editing command is not a descriptive command, determining the editing command as a target sentence;
[0010] Inputting the target sentence and the sentence to be edited into a pre-trained replacement module and an insertion module;
[0011] Triggering the insertion module to process the target sentence and the sentence to be edited, and outputting M candidate insertion sentences, where M is a positive integer;
[0012] Triggering the replacement module to process the target sentence and the sentence to be edited, and outputting N candidate replacement sentences and empty phrase prediction probabilities, where N is a positive integer;
[0013] Inputting the empty phrase prediction probability, each of the candidate replacement sentences and each of the candidate insertion sentences into a preset integration module;
[0014] Triggering the integration module to determine a target candidate sentence from each of the candidate replacement sentences and each of the candidate insertion sentences, and displaying each of the target candidate sentences to the user;
[0015] receiving a selection instruction fed back by the user based on each of the target candidate sentences, and determining whether the selection instruction includes a sentence identifier;
[0016] If the selection instruction includes a sentence identifier, the sentence to be edited is replaced by the target candidate sentence corresponding to the sentence identifier in the selection instruction.
[0017] The above method may, optionally, further include:
[0018] When receiving a voice conversion instruction sent by a user, collecting the converted voice of the user, and calling a preset voice conversion module to convert the converted voice into text;
[0019] The text is input into a preset spoken language removal module, so that the spoken language removal module marks the spoken words in the text and obtains a marking sequence corresponding to the text, removes the spoken words in the text based on the marking sequence, and displays the text after the spoken words are removed as a text to be edited to the user.
[0020] In the above method, optionally, the step of determining whether the editing command is a descriptive command includes:
[0021] Matching the editing command with each preset regular expression;
[0022] Determine whether there is a regular expression corresponding to the editing command;
[0023] If there is a regular expression corresponding to the editing command, determining that the editing command is a descriptive command;
[0024] If there is no regular expression corresponding to the editing command, it is determined that the editing command is not a descriptive command.
[0025] In the above method, optionally, triggering the insertion module to process the target sentence and the sentence to be edited and outputting M candidate insertion sentences includes:
[0026] The insertion module performs word segmentation processing on the sentence to be edited to obtain at least two insertion positions corresponding to the sentence to be edited;
[0027] For each insertion position of the sentence to be edited, insert the target sentence into the insertion position to obtain a first candidate sentence corresponding to the insertion position;
[0028] Inputting each of the first candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a first candidate score for each of the first candidate sentences;
[0029] First candidate sentences are selected in descending order of first candidate scores until the number of selected first candidate sentences is M, and each selected first candidate sentence is determined as a candidate insertion sentence.
[0030] In the above method, optionally, the triggering of the replacement module to process the target sentence and the sentence to be edited, and outputting N candidate replacement sentences and empty phrase prediction probabilities, includes:
[0031] The replacement module processes the target sentence and the sentence to be edited based on a neural network model to obtain vectors corresponding to the target sentence and the sentence to be edited, and processes the vectors based on a preset vocabulary limitation strategy to construct a search tree corresponding to the sentence to be edited, wherein the search tree includes a plurality of child nodes, and the words in each of the child nodes are composed of characters in the sentence to be edited;
[0032] Based on a preset beam search strategy, searching each subnode in the search tree to generate a plurality of erroneous short sentences and empty phrase prediction probabilities corresponding to the sentence to be edited;
[0033] Determine the replacement probability of each erroneous short sentence, and select the erroneous short sentences in descending order of replacement probability, until the number of selected erroneous phrases is consistent with the preset number of short sentences, and determine each selected erroneous short sentence as a target erroneous short sentence;
[0034] For each of the target erroneous sentences, determining the content corresponding to the target erroneous sentence in the sentence to be edited, and replacing the content corresponding to the target erroneous sentence in the sentence to be edited with the target sentence, thereby obtaining a replacement example sentence corresponding to the target erroneous sentence;
[0035] When the sentence to be edited and the target sentence satisfy any one of the preset supplementary rules, generating at least one supplementary example sentence based on the sentence to be edited and the target sentence;
[0036] Determine each of the replacement example sentences and each of the supplementary example sentences as a second candidate sentence, and input each of the second candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a second candidate score for each of the second candidate sentences;
[0037] Second candidate sentences are selected in descending order of the second candidate scores until the number of selected second candidate sentences reaches N, and each of the selected second candidate sentences is determined as a candidate replacement sentence.
[0038] In the above method, optionally, the step of triggering the synthesis module to determine a target candidate statement from each of the candidate replacement statements and each of the candidate insertion statements includes:
[0039] The synthesis module determines a sentence score for each of the candidate replacement sentences and each of the candidate insertion sentences;
[0040] Determine a sentence score with the largest value among the sentence scores of the candidate replacement sentences, and determine the sentence score with the largest value as the first sentence score;
[0041] Determine the sentence score with the smallest value among the sentence scores of the candidate insertion sentences, and determine the sentence score with the smallest value as the second sentence score;
[0042] Determining whether the second sentence score is greater than the first sentence score;
[0043] If the second sentence score is greater than the first sentence score, determining the first number of replacement sentences and the first number of insertion sentences based on a preset first selection rule, selecting candidate replacement sentences in descending order of the sentence scores of the candidate replacement sentences until the number of the selected candidate replacement sentences is equal to the first number of replacement sentences, and selecting candidate insertion sentences in descending order of the sentence scores of the candidate insertion sentences until the number of the selected candidate insertion sentences is equal to the first number of insertion sentences, and determining the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences;
[0044] If the second sentence score is not greater than the first sentence score, determining whether the empty phrase prediction probability is within a preset first interval;
[0045] If the empty phrase prediction probability is within the first interval, selecting candidate replacement sentences according to the sentence scores of the candidate replacement sentences from high to low until the number of the selected candidate replacement sentences is equal to the first number of replacement sentences, and selecting candidate insertion sentences according to the sentence scores of the candidate insertion sentences from high to low until the number of the selected candidate insertion sentences is equal to the first number of insertion sentences, and determining the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences;
[0046] If the empty phrase prediction probability is not within the first interval, determining whether the empty phrase prediction probability is within a preset second interval;
[0047] If the empty phrase prediction probability is within the second interval, determining a second number of replacement sentences and a second number of insertion sentences based on a preset second candidate rule, selecting candidate replacement sentences in descending order of sentence scores of the candidate replacement sentences until the number of selected candidate replacement sentences equals the second number of replacement sentences, and selecting candidate insertion sentences in descending order of sentence scores of the candidate insertion sentences until the number of selected candidate insertion sentences equals the second number of insertion sentences, and determining the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences;
[0048] If the empty phrase prediction probability is not within the second interval, a third number of replacement sentences and a third number of insertion sentences are determined based on a preset third candidate rule, and candidate replacement sentences are selected in descending order of their sentence scores until the number of selected candidate replacement sentences is equal to the third number of replacement sentences, and candidate insertion sentences are selected in descending order of their sentence scores until the number of selected candidate insertion sentences is equal to the third number of insertion sentences, and the selected candidate insertion sentences and the selected candidate replacement sentences are both determined as target candidate sentences.
[0049] The above method may optionally further include:
[0050] If it is determined that the editing command is a descriptive command, the description type of the descriptive command is determined, and an editing operation corresponding to the description type is performed on the sentence to be edited.
[0051] A voice editing device, comprising:
[0052] an acquisition unit, configured to acquire an editing command input by a user and a sentence to be edited, wherein the sentence to be edited is a sentence selected by the user in the text to be edited, the text to be edited is a text obtained by converting the converted speech input by the user into text, and the editing command is a text obtained by converting the command speech input by the user based on the sentence to be edited into text;
[0053] A judging unit, used for judging whether the editing command is a descriptive command;
[0054] a determining unit, configured to determine the editing command as a target sentence if the editing command is not a descriptive command;
[0055] A first input unit, used for inputting the target sentence and the sentence to be edited into a pre-trained replacement module and an insertion module;
[0056] A first triggering unit is used to trigger the insertion module to process the target sentence and the sentence to be edited, and output M candidate insertion sentences, where M is a positive integer;
[0057] A second triggering unit is used to trigger the replacement module to process the target sentence and the sentence to be edited, and output N candidate replacement sentences and empty phrase prediction probabilities, where N is a positive integer;
[0058] A second input unit, used to input the empty phrase prediction probability, each of the candidate replacement sentences and each of the candidate insertion sentences into a preset integration module;
[0059] A display unit, used to trigger the integration module to determine a target candidate sentence from each of the candidate replacement sentences and each of the candidate insertion sentences, and to display each of the target candidate sentences to the user;
[0060] A receiving unit, configured to receive a selection instruction fed back by the user based on each of the target candidate sentences, and determine whether the selection instruction includes a sentence identifier;
[0061] A replacement unit is used to replace the sentence to be edited with a target candidate sentence corresponding to the sentence identifier in the selection instruction if the selection instruction includes a sentence identifier.
[0062] A collection unit, configured to collect the converted voice of the user upon receiving a voice conversion instruction sent by the user, and call a preset voice conversion module to convert the converted voice into text;
[0063] A removal unit is used to input the text into a preset spoken language removal module, so that the spoken language removal module marks the spoken words in the text and obtains a marking sequence corresponding to the text, removes the spoken words in the text based on the marking sequence, and displays the text after the spoken words are removed as a text to be edited to the user.
[0064] The above device may optionally further include:
[0065] A collection unit, configured to collect the converted voice of the user upon receiving a voice conversion instruction sent by the user, and call a preset voice conversion module to convert the converted voice into text;
[0066] A removal unit is used to input the text into a preset spoken language removal module, so that the spoken language removal module marks the spoken words in the text and obtains a marking sequence corresponding to the text, removes the spoken words in the text based on the marking sequence, and displays the text after the spoken words are removed as a text to be edited to the user.
[0067] In the above device, optionally, the judging unit includes:
[0068] A matching subunit, used for matching the editing command with various preset regular expressions;
[0069] A first judging subunit, used for judging whether there is a regular expression corresponding to the editing command;
[0070] a first determining subunit, configured to determine that the editing command is a descriptive command if a regular expression corresponding to the editing command exists;
[0071] The second determining subunit is configured to determine that the editing command is not a descriptive command if there is no regular expression corresponding to the editing command.
[0072] In the above device, optionally, the first trigger unit includes:
[0073] An obtaining subunit is used for the insertion module to perform word segmentation processing on the sentence to be edited to obtain at least two insertion positions corresponding to the sentence to be edited;
[0074] An inserting subunit, for inserting the target sentence into each insertion position of the sentence to be edited, to obtain a first candidate sentence corresponding to the insertion position;
[0075] an output subunit, configured to input each of the first candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a first candidate score for each of the first candidate sentences;
[0076] The first selection subunit is used to select first candidate sentences in descending order of first candidate scores until the number of selected first candidate sentences is M, and determine each selected first candidate sentence as a candidate insertion sentence.
[0077] In the above device, optionally, the second triggering unit includes:
[0078] A construction subunit is used for the replacement module to process the target sentence and the sentence to be edited based on the neural network model to obtain vectors corresponding to the target sentence and the sentence to be edited, and to process the vectors based on a preset vocabulary limitation strategy to construct a search tree corresponding to the sentence to be edited, wherein the search tree contains a plurality of child nodes, and the words in each of the child nodes are composed of characters in the sentence to be edited;
[0079] A first generating subunit, configured to search each subnode in the search tree based on a preset beam search strategy to generate a plurality of erroneous short sentences and empty phrase prediction probabilities corresponding to the sentence to be edited;
[0080] The third determination subunit is used to determine the replacement probability of each erroneous short sentence, and select the erroneous short sentences in the order of the replacement probability from high to low, until the number of the selected erroneous phrases is consistent with the preset number of short sentences, and each selected erroneous short sentence is determined as a target erroneous short sentence;
[0081] a replacement subunit, configured to determine, for each target erroneous sentence, a content corresponding to the target erroneous sentence in the sentence to be edited, and replace the content corresponding to the target erroneous sentence in the sentence to be edited with the target sentence, thereby obtaining a replacement example sentence corresponding to the target erroneous sentence;
[0082] A second generating subunit, configured to generate at least one supplementary example sentence based on the sentence to be edited and the target sentence when the sentence to be edited and the target sentence satisfy any one of the preset supplementary rules;
[0083] a fourth determination subunit, configured to determine each of the replacement example sentences and each of the supplementary example sentences as a second candidate sentence, and input each of the second candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a second candidate score for each of the second candidate sentences;
[0084] The second selection subunit is used to select second candidate sentences in descending order of the second candidate scores until the number of selected second candidate sentences is N, and determine each selected second candidate sentence as a candidate replacement sentence.
[0085] In the above device, optionally, the display unit comprises:
[0086] A fifth determination subunit, used by the integration module to determine a sentence score of each of the candidate replacement sentences and each of the candidate insertion sentences;
[0087] A sixth determining subunit, configured to determine a sentence score with the largest value among the sentence scores of the candidate replacement sentences, and determine the sentence score with the largest value as the first sentence score;
[0088] a seventh determination subunit, configured to determine a sentence score with the smallest value among the sentence scores of the candidate insertion sentences, and determine the sentence score with the smallest value as the second sentence score;
[0089] A second judging subunit, configured to judge whether the second sentence score is greater than the first sentence score;
[0090] a third selection subunit, configured to determine a first number of replacement statements and a first number of insertion statements based on a preset first selection rule if the second statement score is greater than the first statement score, and select candidate replacement statements in descending order of the statement scores of the candidate replacement statements until the number of the selected candidate replacement statements is equal to the first number of replacement statements, and select candidate insertion statements in descending order of the statement scores of the candidate insertion statements until the number of the selected candidate insertion statements is equal to the first number of insertion statements, and determine the selected candidate insertion statements and the selected candidate replacement statements as target candidate statements;
[0091] an eighth determination subunit, configured to determine whether the empty phrase prediction probability is within a preset first interval if the second sentence score is not greater than the first sentence score;
[0092] a fourth selection subunit, configured to select candidate replacement sentences in descending order of sentence scores of the candidate replacement sentences, until the number of the selected candidate replacement sentences is equal to the first number of replacement sentences, and select candidate insertion sentences in descending order of sentence scores of the candidate insertion sentences, until the number of the selected candidate insertion sentences is equal to the first number of insertion sentences, if the empty phrase prediction probability is within the first interval, and determine the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences;
[0093] a ninth determining subunit, configured to determine whether the empty phrase prediction probability is located in a preset second interval if the empty phrase prediction probability is not located in the first interval;
[0094] a fifth selection subunit, configured to determine, if the empty phrase prediction probability is within the second interval, a second number of replacement sentences and a second number of insertion sentences based on a preset second candidate rule, and select candidate replacement sentences in descending order of sentence scores of the candidate replacement sentences until the number of selected candidate replacement sentences equals the second number of replacement sentences, and select candidate insertion sentences in descending order of sentence scores of the candidate insertion sentences until the number of selected candidate insertion sentences equals the second number of insertion sentences, and determine the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences;
[0095] The sixth selection subunit is used to determine the third number of replacement sentences and the third number of insertion sentences based on a preset third candidate rule if the empty phrase prediction probability is not within the second interval, and select candidate replacement sentences in descending order of their sentence scores until the number of selected candidate replacement sentences is equal to the third number of replacement sentences, and select candidate insertion sentences in descending order of their sentence scores until the number of selected candidate insertion sentences is equal to the third number of insertion sentences, and determine the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences.
[0096] The above device may optionally further include:
[0097] The execution unit is used to determine the description type of the descriptive command if it is determined that the editing command is a descriptive command, and perform an editing operation corresponding to the description type on the sentence to be edited.
[0098] A storage medium includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the above-mentioned voice editing method.
[0099] An electronic device includes a memory and one or more instructions, wherein the one or more instructions are stored in the memory and are configured to be executed by one or more processors as the above-mentioned voice editing method.
[0100] Compared with the prior art, the present invention has the following advantages:
[0101] The present invention provides a speech editing method and device, a storage medium and an electronic device. The method comprises: obtaining an editing command input by a user and a sentence to be edited, when the editing command is not a descriptive command, determining that the editing command is a target sentence, inputting the target sentence and the sentence to be edited into a replacement module and an insertion module, and inputting N candidate replacement sentences and empty phrase prediction probabilities output by the replacement module and M candidate insertion sentences output by the insertion module into a comprehensive module, so that the comprehensive module selects a target candidate sentence and displays each target candidate sentence to the user, and when a selection instruction fed back by the user is received and a sentence identifier is included in the selection instruction, the sentence to be edited is replaced by the target candidate sentence corresponding to the sentence identifier, thereby eliminating the need for the user to specify the position of a text to be modified, and only requiring the correct text to be input during the editing process to predict the text content that the user wants to modify, thereby generating a text that better meets the user's modification intention. The present invention effectively shortens the speech editing process, and only requiring the input of a small amount of information during the speech editing process, thereby making the speech editing simpler, thereby improving the efficiency of the speech editing, and providing users with better quality services. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0103] Figure 1 A method flow chart of a voice editing method provided by an embodiment of the present invention;
[0104] Figure 2 Another method flow chart of a voice editing method provided by an embodiment of the present invention;
[0105] Figure 3 A method flow chart of another method of a voice editing method provided by an embodiment of the present invention;
[0106] Figure 4 A scene example diagram of a voice editing method provided by an embodiment of the present invention;
[0107] Figure 5 A schematic diagram of the structure of a voice editing device provided by an embodiment of the present invention;
[0108] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0109] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0110] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0111] In order to provide users with a simpler editing method, the present invention provides a voice editing method, which allows users to edit text through voice without having to specify the location of erroneous text through voice or manual selection, and the editing process is simple. While improving the efficiency of voice editing, it also provides great convenience to users, provides users with better services, and improves the efficiency of editing.
[0112] The embodiment of the present invention provides a voice editing method, which can be applied in an intelligent voice editing system. The intelligent voice editing system can be constructed by an intelligent computer device. The execution subject of the present invention is a server or processor of the intelligent voice editing system. One of the flow charts of the method provided by the present invention is as follows: Figure 1 The specific instructions are as follows:
[0113] S101: Obtain an editing command input by a user and a statement to be edited.
[0114] The sentence to be edited in the present invention is the sentence selected by the user in the text to be edited, the text to be edited is the text after the conversion voice input by the user is converted into text, and the editing command is the text after the command voice input by the user based on the sentence to be edited is converted into text, wherein the conversion voice is the voice that the user needs to convert into text content.
[0115] There are many ways to obtain the editing commands input by the user and the sentences to be edited. One of the ways is to display the sentences in the text to the user one by one. When the command voice input by the user is received, the sentence currently displayed to the user is used as the sentence to be edited, and the command voice is converted into text to obtain the editing command, wherein the text converted into the command voice input by the user can be a phrase or a sentence.
[0116] S102, determining whether the editing command is a descriptive command; if the editing command is not a descriptive command, executing S103; if the editing command is a descriptive command, executing S112.
[0117] The specific process of determining whether an editing command is a descriptive command is as follows:
[0118] Matching the editing command with each preset regular expression;
[0119] Determine whether there is a regular expression corresponding to the editing command;
[0120] If there is a regular expression corresponding to the editing command, determining that the editing command is a descriptive command;
[0121] If there is no regular expression corresponding to the editing command, it is determined that the editing command is not a descriptive command.
[0122] Different regular expressions correspond to different syntaxes, and different syntaxes correspond to different types of descriptive commands. Descriptive commands may specifically include insert commands, replace commands, and delete commands, and different types of commands correspond to different regular expressions. For example, the syntax corresponding to the insert command may be: "Insert B in front of A", "Add D after C", and other similar sentences can all be the syntax corresponding to the insert command; the syntax corresponding to the replace command may be: "Replace F with E", "Change Y to W", and other similar sentences can all be used as the syntax corresponding to the replace command; the syntax of the delete command may be: "Delete Q", "Remove T", and other similar sentences can all be used as the syntax corresponding to the delete command.
[0123] S103: Determine the editing command as a target sentence.
[0124] When the editing command is not a descriptive command, the editing command can be determined as the content that the user needs to edit, and the editing command can be determined as the target sentence, that is, the user needs to replace the content in the sentence to be edited with the target sentence or needs to insert the target sentence into the sentence to be edited; preferably, the target sentence in this scheme is a phrase, and the target sentence is the correct text entered by the user.
[0125] S104: Input the target sentence and the sentence to be edited into a pre-trained replacement module and insertion module.
[0126] The replacement module and the insertion module in the present invention are both pre-trained models, wherein the replacement module can be a module constructed using text generation models such as GPT-2, Seq2Seq based on RNN, or masked language models such as BERT model, and the insertion module can be a module constructed using text generation models such as GPT-2, Seq2Seq based on RNN, or masked language models such as BERT model; preferably, the replacement module and the insertion module in the present invention are both modules constructed using the GTP-2 model.
[0127] The training of the insertion module and the replacement module is explained. When training the insertion module and the replacement module, a data set is first constructed for the insertion module and the replacement module. The data set is used to train the insertion module and the replacement module. The process of constructing the data set includes: generating replacement samples and insertion samples. Common replacement errors in life include homophony, synonymy, loss, repetition, etc., and replacement samples can be generated from these directions.
[0128] The specific process of generating replacement samples is as follows: collect a large amount of text from online forum posts, first divide the sentences according to punctuation, then remove sentences containing English, special characters, sentences that are too long, too short, and truncated, and finally perform word segmentation and part-of-speech tagging on the sentences; for each sentence obtained, select a random number of phrases with random lengths as error intervals, replace the correct phrases in the interval with the error phrases, and obtain the error original text as the input of the replacement sample. Select the correct phrase in one of the intervals as the target phrase, which is also used as the input of the replacement sample. The error phrase in the interval is used as the output of the replacement sample. The method of generating an incorrect phrase based on a correct phrase includes: (1) pronoun: parsing the correct phrase into pinyin, randomly replacing certain vowels and consonants in the pinyin, or randomly adding or deleting the pinyin of a certain word, and then generating text based on the pinyin; (2) synonym: randomly selecting several words in the correct phrase as the words to be modified, finding synonyms with high cosine similarity with the word embedding of the word to be modified in the vector space based on word embedding, and replacing the word to be modified with the synonym; (3) missing: randomly selecting characters or words from the correct phrase and deleting them; (4) repetition: concatenating the incorrect phrase generated according to (1), (2), (3) with the correct phrase, or repeating the correct phrase twice to obtain the incorrect phrase.
[0129] The specific process of generating insertion samples is as follows: collect a large amount of text from online forum posts, first divide the sentences according to punctuation, then remove sentences containing English, special characters, sentences that are too long, too short, and truncated, and finally segment the sentences and tag the parts of speech; for each sentence obtained, select a random number of phrases with random lengths from the sentence as the error interval, delete the correct phrases in the interval to obtain the error original text as the input of the insertion sample. Select the correct phrase in one of the intervals as the target phrase, which is also used as the input of the insertion sample, and replace the empty phrase (i.e. the sentence end symbol " <eos>”) as the output of the inserted sample.
[0130] The training process of the replacement module includes: using the text generation model, pre-training on the Chinese Wikipedia corpus, the training task is to predict the current word based on the previous text, and stop training when the perplexity is lower than the threshold. The model is fine-tuned on the constructed dataset, the training task is the generation task, the wrong original text and the correct phrase are used as input, and the corresponding wrong phrase (including empty phrase) is generated. The fine-tuning training process includes 5-10 rounds of fixed training, and the model with the best performance on the validation set is selected as the final model.
[0131] The process of training the insertion module includes: using the text generation model, pre-training on the Chinese Wikipedia corpus, the training task is to predict the current word based on the previous context, and stop training when the perplexity is lower than the threshold. Fine-tuning is performed on the constructed dataset, and the training task is the same as the pre-training task. The fine-tuning training process includes 5-10 rounds of fixed training, and the model with the best performance on the validation set is selected as the final model.
[0132] S105 , triggering the insertion module to process the target statement and the statement to be edited, and outputting M candidate insertion statements, where M is a positive integer.
[0133] After receiving the target sentence and the sentence to be edited, the insertion module processes the target sentence and the sentence to be edited, thereby outputting M candidate insertion sentences, where M can be set according to actual needs. The specific process is as follows: Figure 2 As shown, Figure 2 The steps in the figure are all the processes executed by the insertion module, and the specific descriptions are as follows:
[0134] S201: Perform word segmentation processing on the sentence to be edited to obtain at least two insertion positions corresponding to the sentence to be edited.
[0135] The insertion module performs word segmentation processing on the sentence to be edited so as to divide the sentence to be edited into at least one word segment, and then determines the number of insertion positions according to the number of word segments, wherein the number of insertion positions is one more than the number of word segments. Specifically, when the number of word segments of the sentence to be edited is 1, the number of insertion positions is 2, specifically in front of and behind the word segment; specifically, there is an insertion position before and after each word segment of the sentence to be edited, and there is an insertion position between two word segments.
[0136] S202: For each insertion position of the sentence to be edited, insert the target sentence into the insertion position to obtain a first candidate sentence corresponding to the insertion position.
[0137] After placing the target sentence in the insertion position, the first candidate sentence corresponding to the insertion position can be obtained. For example, if the sentence to be edited is "Today is very good" and the target sentence is "the weather", the insertion position of the sentence to be edited can be determined as "-Today is -very-good-", where "-" represents the insertion position. After inserting the target sentence in each insertion position, the first candidate sentences obtained are: 1. The weather is very good today; 2. The weather is very good today; 3. The weather is very good today; 4. The weather is very good today.
[0138] S203: Input each of the first candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a first candidate score for each of the first candidate sentences.
[0139] Continuing with the description in S203, the sentence scoring model is used to calculate the first candidate score of each first candidate sentence. Assume that the first candidate score of the first candidate sentence numbered 1 is 20, the first candidate score of the first candidate sentence numbered 2 is 90, the first candidate score of the first candidate sentence numbered 3 is 19, and the first candidate score of the first candidate sentence numbered 4 is 25.
[0140] The sentence scoring model calculates the first candidate score of the first candidate sentence according to a preset scoring formula, wherein the scoring formula is specifically:
[0141]
[0142] Wherein, language_model_score(s) represents the first candidate score of the first candidate sentence; s represents the first candidate sentence; l represents the total number of words in the sentence; w i represents the i-th word in the sentence; p(w1) represents the probability that the first word is w1; p(w i |w1...w i-1 ) means that the words from the 1st word to the i-1th word are w1w2...w i-1 In the case of i probability.
[0143] It should be noted that the specific form of the above scoring formula after expansion is:
[0144] language_model_score(s)=log(p(w1)p(w2|w1)p(w3|w1w2)...p(w l |w1...w l-1 )) / l;
[0145] The first candidate score of each first candidate sentence is calculated by using the above scoring formula.
[0146] S204 , selecting first candidate sentences in descending order of first candidate scores until the number of selected first candidate sentences is M, and determining each selected first candidate sentence as a candidate insertion sentence.
[0147] Continuing with the description in S203, the first candidate sentences are arranged in descending order according to the first candidate scores, so that the queue obtained is: the first candidate sentence numbered 2 is ranked first, the first candidate sentence numbered 4 is ranked second, the first candidate sentence numbered 1 is ranked third, and the first candidate sentence numbered 3 is ranked fourth; when M is 2, the candidate insertion sentences selected are: the first candidate sentence numbered 2: "The weather is very good today", and the first candidate sentence numbered 4: "The weather is very good today"; preferably, when M is 5, the first candidate sentences numbered 1, 2, 3 and 4 are all selected as candidate insertion sentences, and the missing candidate insertion sentences can be filled with empty sentences, that is, there is an empty sentence among the 5 candidate insertion sentences at this time.
[0148] In the method provided by the embodiment of the present invention, by calculating the first candidate score of each first candidate sentence and determining the candidate replacement sentence in the first candidate sentence according to the first candidate score, a sentence that is more in line with the sentence context and emotion can be provided to the user, so that the obtained candidate replacement sentence can better meet the needs of the user, thereby improving the accuracy and efficiency of voice editing.
[0149] S106: triggering the replacement module to process the target sentence and the sentence to be edited, and outputting N candidate replacement sentences and empty phrase prediction probabilities, where N is a positive integer.
[0150] After receiving the target statement and the statement to be edited, the replacement module performs the following operations:
[0151] The replacement module processes the target sentence and the sentence to be edited based on a neural network model to obtain vectors corresponding to the target sentence and the sentence to be edited, and processes the vectors based on a preset vocabulary limitation strategy to construct a search tree corresponding to the sentence to be edited, wherein the search tree contains multiple child nodes, and the words in each of the child nodes are composed of words in the sentence to be edited.
[0152] The neural network model of the present invention can be a text generation model such as GPT-2, Seq2Seq based on RNN, or a BERT model. The present invention generates each sub-node in the search tree based on a vocabulary restriction strategy, so that the words in each sub-node are composed of the text in the sentence to be edited.
[0153] Based on the preset beam search strategy, each child node in the search tree is searched to generate a plurality of erroneous short sentences and empty phrase prediction probabilities corresponding to the sentence to be edited; the beam search strategy in the present invention limits the size of the beam, which can be represented by n, wherein the value of n is associated with the best training result of the replacement module during training. When searching each child node in the search tree based on the beam search strategy, the probability of each child node on the path can be multiplied, and then normalized according to the depth of the child node, the probability of each erroneous short sentence being replaced can be obtained, and the normalization process is: (log(p1p2...p l )) / l, where p1 represents the probability corresponding to each child node; l represents the number of child nodes contained in the wrong sentence. Furthermore, the prediction probability of the empty phrase is the probability that the candidate wrong sentence with the highest probability of being replaced is an empty phrase, and the empty phrase is the sentence end symbol" <eos>”.
[0154] Determine the replacement probability of each erroneous short sentence, and select erroneous short sentences in descending order of replacement probability, until the number of selected erroneous phrases is consistent with the preset number of short sentences, and determine each selected erroneous short sentence as a target erroneous short sentence. The erroneous short sentence in the present invention can be a sentence or a phrase.
[0155] For each target erroneous sentence, the content corresponding to the target erroneous sentence is determined in the sentence to be edited, and the target sentence is used to replace the content corresponding to the target erroneous sentence in the sentence to be edited, thereby obtaining a replacement example sentence corresponding to the target erroneous sentence.
[0156] When the sentence to be edited and the target sentence satisfy any one of the preset supplementary rules, at least one supplementary example sentence is generated based on the sentence to be edited and the target sentence; wherein the supplementary example sentence is a sentence generated according to the supplementary rules satisfied by the sentence to be edited and the target sentence; there are multiple supplementary rules here, specifically pronunciation supplementary rules, context alignment supplementary rules, etc.; the pronunciation supplementary rule is specifically: if the phonetic similarity between the target sentence and a phrase in the sentence to be edited is lower than a threshold, then it is determined that the sentence to be edited and the target sentence satisfy the pronunciation supplementary rules in each supplementary rule, and the phrase is replaced with the target sentence, and the replaced sentence is used as the supplementary example sentence; wherein the phonetic similarity can be calculated based on the high-dimensional encoding of vowels and consonants; the context alignment supplementary rule is specifically: if the head and tail of the target sentence are consistent with the head and tail of a phrase in the sentence to be edited, then it is determined that the sentence to be edited and the target sentence satisfy the context alignment supplementary rules in each supplementary rule, and the phrase is replaced with the target sentence, and the replaced sentence is used as the supplementary example sentence. It should be further explained that the target error sentence and the sentence to be edited may satisfy multiple supplementary rules at the same time. In the present invention, there is also a situation that the target error sentence and the sentence to be edited do not satisfy any supplementary rule. When this happens, there is no need to generate a supplementary example sentence.
[0157] Each of the replacement example sentences and each of the supplementary example sentences are determined as second candidate sentences, and each of the second candidate sentences is input into a preset sentence scoring model, so that the sentence scoring model outputs a second candidate score for each of the second candidate sentences; for the description of the sentence scoring model, please refer to Figure 2 The relevant instructions in will not be repeated here.
[0158] Select the second candidate sentences in descending order of the second candidate scores until the number of selected second candidate sentences reaches N, and determine each selected second candidate sentence as a candidate replacement sentence. For instructions on selecting the second candidate sentence, refer to Figure 2 The description about selecting the first candidate sentence will not be repeated here.
[0159] S107: input the empty phrase prediction probability, each of the candidate replacement sentences, and each of the candidate insertion sentences into a preset integration module.
[0160] S108: triggering the integration module to determine a target candidate sentence from each of the candidate replacement sentences and each of the candidate insertion sentences, and presenting each of the target candidate sentences to the user.
[0161] When each target candidate sentence is displayed to the user, each target candidate sentence may be displayed to the user, or each target candidate sentence may be displayed to the user one by one.
[0162] The comprehensive module determines at least one target candidate sentence, and the method of displaying the determined target candidate sentence to the user can be specifically as follows: assigning a number to each target candidate sentence, sorting the target candidate sentences according to the number, so as to obtain a candidate sentence list, and displaying the candidate sentence list to the user, and also playing the voice corresponding to each target candidate sentence in the candidate sentence list to the user one by one in the order of the number in the form of voice.
[0163] The process of the synthesis module determining the target candidate sentence is as follows: Figure 3 As shown, Figure 3 These are all contents executed by the comprehensive module, and the specific descriptions are as follows:
[0164] S301: Determine a statement score for each of the candidate replacement statements and each of the candidate insertion statements.
[0165] The sentence scores of the candidate replacement sentence and the candidate insertion sentence are the scores calculated using the sentence scoring model in S203.
[0166] S302. Determine the statement score with the largest value among the statement scores of the candidate replacement statements, and determine the statement score with the largest value as the first statement score; and determine the statement score with the smallest value among the statement scores of the candidate insertion statements, and determine the statement score with the smallest value as the second statement score.
[0167] S303, determine whether the second sentence score is greater than the first sentence score; if the second sentence score is greater than the first sentence score, execute S304; if the second sentence score is not greater than the first sentence score, execute S305.
[0168] S304. Determine the first number of replacement statements and the first number of insertion statements based on a preset first selection rule, and select candidate replacement statements in descending order of the statement scores of the candidate replacement statements until the number of selected candidate replacement statements is equal to the first number of replacement statements, and select candidate insertion statements in descending order of the statement scores of the candidate insertion statements until the number of selected candidate insertion statements is equal to the first number of insertion statements, and determine the selected candidate insertion statements and the selected candidate replacement statements as target candidate statements.
[0169] The first number of replacement statements is the number of statements selected from the candidate replacement statements; the first number of insertion statements is the number of statements selected from the candidate insertion statements.
[0170] The first selection rule sets specific values of the first number of replacement statements and the first number of insertion statements. Preferably, the first number of replacement statements in the present invention is 1, and the first number of insertion statements is 3. The first number of replacement statements and the second number of replacement statements can be set according to actual needs. The first number of insertion statements in the present invention is greater than the first number of replacement statements.
[0171] S305, determining whether the empty phrase prediction probability is within a preset first interval; if the empty phrase prediction probability is within the first interval, executing S306; if the empty phrase prediction probability is not within the first interval, executing S307.
[0172] The first interval in the present invention is a half-open and half-closed interval, such as (0.98, 1]. Optionally, the value range of the empty phrase prediction probability in the present invention is between 0 and 1.
[0173] S306. Select candidate replacement statements in descending order of their statement scores until the number of selected candidate replacement statements equals the first number of replacement statements, and select candidate insertion statements in descending order of their statement scores until the number of selected candidate insertion statements equals the first number of insertion statements, and determine the selected candidate insertion statements and the selected candidate replacement statements as target candidate statements.
[0174] Regarding the first number of replacement statements and the second number of insertion statements, reference may be made to the description of S304 , which will not be described in detail here.
[0175] S307, determining whether the empty phrase prediction probability is within a preset second interval; if the empty phrase prediction probability is within the second interval, executing S308; if the empty phrase prediction probability is not within the second interval, executing S309.
[0176] The second interval is a half-open and half-closed interval, and the second interval may specifically be [0.5, 0.98).
[0177] S308. Determine the second number of replacement statements and the second number of insertion statements based on a preset second candidate rule, and select candidate replacement statements in descending order of the statement scores of the candidate replacement statements until the number of selected candidate replacement statements is equal to the second number of replacement statements, and select candidate insertion statements in descending order of the statement scores of the candidate insertion statements until the number of selected candidate insertion statements is equal to the second number of insertion statements, and determine the selected candidate insertion statements and the selected candidate replacement statements as target candidate statements.
[0178] The second number of replacement statements is the number of statements selected from the candidate replacement statements; the second number of insertion statements is the number of statements selected from the candidate insertion statements. The second selection rule sets specific values of the second number of replacement statements and the second number of insertion statements. Preferably, the second number of replacement statements in the present invention is 2, and the second number of insertion statements is 2; the second number of replacement statements and the second number of insertion statements can be set according to actual needs, and the second number of insertion statements in the present invention is the same as the second number of replacement statements.
[0179] S309. Determine the third number of replacement statements and the third number of insertion statements based on a preset third candidate rule, and select candidate replacement statements in descending order of the statement scores of the candidate replacement statements until the number of selected candidate replacement statements is equal to the third number of replacement statements, and select candidate insertion statements in descending order of the statement scores of the candidate insertion statements until the number of selected candidate insertion statements is equal to the third number of insertion statements, and determine the selected candidate insertion statements and the selected candidate replacement statements as target candidate statements.
[0180] When the empty phrase prediction probability is not in the second interval, it can be determined that the empty phrase prediction probability is in the third interval, wherein the third interval is a half-open and half-closed interval, and the third interval can specifically be [0, 0.5).
[0181] The third number of replacement statements is the number of statements selected from the candidate replacement statements; the third number of insertion statements is the number of statements selected from the candidate insertion statements. The third selection rule sets specific values of the third number of replacement statements and the third number of insertion statements. Preferably, the third number of replacement statements in the present invention is 3, and the third number of insertion statements is 1; the third number of replacement statements and the third number of insertion statements can be set according to actual needs, and the third number of insertion statements in the present invention is the same as the third number of replacement statements.
[0182] In the method provided by an embodiment of the present invention, a target candidate sentence is selected based on the empty phrase prediction probability, the sentence score of each candidate insertion sentence, and the sentence score of each candidate insertion sentence, wherein the number of candidate insertion sentences included in the target candidate sentence and the number of candidate insertion sentences are selected according to different situations. The present invention provides a variety of selection rules to be applicable to a variety of different scenarios, thereby improving the applicability of the present invention and providing users with more appropriate target candidate sentences.
[0183] S109, receiving the selection instruction fed back by the user based on each of the target candidate sentences, and determining whether the selection instruction contains a sentence identifier; if the selection instruction contains a sentence identifier, executing S110; if the selection instruction does not contain a sentence identifier, executing S111.
[0184] The sentence identifier can be the sentence number of the target candidate sentence or a confirmation identifier, wherein the sentence number is a unique identifier of the target candidate sentence. Exemplarily, if the selection instruction is the first sentence, the selection instruction includes a sentence identifier, which is a sentence number, indicating that the user selects the first target candidate sentence. If a selection instruction sent by the user is received when the third target candidate sentence is displayed to the user, and the selection instruction is confirmation, the selection instruction includes a sentence identifier, which is a confirmation indication, i.e., it indicates that the user selects the third target candidate sentence.
[0185] S110: Replace the sentence to be edited with the target candidate sentence corresponding to the sentence identifier in the selection instruction.
[0186] The target candidate sentence corresponding to the sentence identifier contained in the selection instruction replaces the sentence to be marked, thereby editing the text to be edited by voice.
[0187] S111. Obtain an operation identifier in the selection instruction, and execute an operation corresponding to the operation identifier.
[0188] The operation identifier in the present invention includes but is not limited to a cancel identifier, a return identifier or a send identifier, etc. Different operation identifiers correspond to different operations. For example, the cancel identifier corresponds to a cancel operation, that is, canceling the editing and no longer editing the sentence to be edited; the return identifier corresponds to a return operation, that is, no longer editing the sentence to be edited and returning to the previous operation, such as re-displaying the sentence before the sentence to be edited; the send identifier corresponds to a send operation, that is, no longer editing the sentence to be edited and sending the current text to be edited.
[0189] S112: Determine the description type of the descriptive command, and perform an editing operation corresponding to the description type on the sentence to be edited.
[0190] The description type of the descriptive command includes but is not limited to an insert type, a delete type or a replace type. The editing operation corresponding to the insert type is an insert operation. For example, if the specific content of the descriptive command is to insert A before B, then the position of B will be determined in the statement to be edited, and A will be inserted in front of B; the editing operation corresponding to the delete type is a delete operation. For example, if the specific content of the descriptive command is to delete C, then C will be determined in the statement to be edited and deleted from the statement to be edited; the editing operation corresponding to the replace type is a replace operation. For example, if the specific content of the descriptive command is to replace D with F, then D will be determined in the statement to be edited, and F will replace D.
[0191] In the method provided by the embodiment of the present invention, the editing command input by the user and the sentence to be edited are obtained. When the editing command is not a descriptive command, the editing command is determined to be a target sentence, the target sentence and the sentence to be edited are input into the replacement module and the insertion module, and the N candidate replacement sentences and the empty phrase prediction probability output by the replacement module and the M candidate insertion sentences output by the insertion module are all input into the comprehensive module, so that the comprehensive module selects the target candidate sentence, and each target candidate sentence is displayed to the user, and the selection instruction fed back by the user is received, and when the sentence identifier is included in the selection instruction, the target candidate sentence corresponding to the sentence identifier is replaced with the sentence to be edited. The present invention can predict the text content that the user wants to modify through the target sentence input by the user, without the user specifying the specific location of the text to be modified, and provides the user with a sentence that is more semantically consistent and more contextually coherent, so as to edit the text. The method provided by the present invention can edit the text without inputting a descriptive command, and shortens the process of traditional voice editing, thereby improving the efficiency of voice editing, and providing users with better quality services.
[0192] In the present invention, before obtaining the editing command input by the user and the sentence to be edited, the following contents are also included:
[0193] When receiving a voice conversion instruction sent by a user, collecting the converted voice of the user, and calling a preset voice conversion module to convert the converted voice into text;
[0194] The text is input into a preset spoken language removal module, so that the spoken language removal module marks the spoken words in the text and obtains a marking sequence corresponding to the text, removes the spoken words in the text based on the marking sequence, and displays the text after the spoken words are removed as a text to be edited to the user.
[0195] It should be noted that when the user needs to convert speech into text, a speech conversion instruction is sent to the intelligent speech editing system to enable the intelligent speech editing system to start working. After the intelligent speech editing system starts working, the speech acquisition module in the intelligent speech system starts to acquire the conversion speech of the user, where the conversion speech is the speech that the user needs to convert into text. Preferably, the speech acquisition module is set in the intelligent terminal.
[0196] In the present invention, the spoken words are the spoken words inadvertently inserted by the user that have no influence on the sentence, including but not limited to "um", "ah", the adverb "just", the conjunction "then", and the pronoun "this", etc. When the text after removing the spoken words is displayed to the user as the text to be edited, it can be displayed to the user sentence by sentence. During the process of displaying to the user, when the edited speech input by the user is acquired, the sentence currently displayed to the user is used as the sentence to be edited, and an editing command to convert the edited speech into text is generated.
[0197] In the method provided by the embodiment of the present invention, the spoken words in the speech input by the user are removed by using the spoken word removal module, thereby improving the accuracy of the data, reducing the data processing amount of the replacement module and the insertion module, and automatically removing the spoken words without the need for the user to operate, providing a more convenient operation method for the user.
[0198] The spoken language removal module in the present invention is constructed using models that can be used for annotation, such as BERT model, RNN model, LSTM model, CRF model, etc., and the spoken language removal module is a pre-trained module. Preferably, the spoken language removal module in the present invention is constructed using the BERT model, and the process of training the spoken language removal module constructed using the BERT model is described, specifically: first, BERT is pre-trained on Chinese Wikipedia data, and the training tasks are masked language model tasks and predicting the next sentence tasks, and the training is stopped when the perplexity is lower than the threshold. Then, a linear layer is added to the top of the last hidden layer of BERT, and then fine-tuned on the pre-constructed data set, and the training task is a sequence annotation task. When the accuracy and recall on the verification set are stable, the training is stopped, and the training of the spoken language removal module is completed at this time. The pre-constructed data set here is a spoken word database, which records the common insertion positions of different spoken words, including the end of a sentence, before a noun phrase, and randomly. When constructing a spoken word database, a large amount of text can be collected from online forum posts. First, sentence segmentation is performed based on punctuation, and then sentences containing English and special characters, sentences that are too long, too short, and truncated are removed. Finally, the sentences are segmented and part-of-speech tagged. According to the part-of-speech tag sequence, regular expressions are used to find the noun phrases in the sentence. For each noun phrase and each common insertion position where they appear, some spoken word tagging samples are constructed. Specifically, the spoken word is inserted into a specific position of the sentence (the end of the sentence, before a noun phrase, randomly, etc.) as the input of the sample, and a tag sequence of the same length as the input is generated as the output. The spoken word position is marked as 1, and the non-spoken word position is marked as 0. A random number of spoken words will be inserted into a sentence. When inserting a spoken word, there is a certain probability that a comma will be inserted before and after the spoken word. This comma is marked as 1 in the tag sequence as the output.
[0199] The method provided in the embodiment of the present invention can be applied in a variety of scenarios. The following is a specific example to illustrate the application of the present invention in a practical scenario.
[0200] Scenario Example 1:
[0201] Take the voice editor on the mobile phone as an example, in which the user uses the wired headset to perform editing operations, refer to Figure 4 , is an application scenario instance diagram, and the specific description is as follows:
[0202] 401. When the input button is clicked, or the middle button on the headphone cable is long pressed, the system transcribes the voice input into text and displays it in the edit box.
[0203] 402. After completing the input, the system selects the first sentence that has not been edited as the current sentence, and then sends the paragraph input this time to the spoken word removal model, and replaces the returned result into the edit box.
[0204] 403. When the reading button is clicked, or the middle button of the headphone cable is short-pressed, the system starts reading from the current sentence. If reading has already started, it will be paused. The system will automatically select the sentence currently being read.
[0205] 404. When the current / next button is clicked, or the plus / minus button on the headphone cable is pressed briefly, the system selects and reads the previous / next sentence. When the plus button on the headphone cable is pressed for a long time, the system starts reading from the beginning and selects the first sentence. When the edit box is clicked, the sentence at the corresponding position is selected, and the system will highlight and read the sentence. When the area where the sentence is located is selected again, the cursor will be placed at the specific location clicked.
[0206] 405. When the edit button is clicked, or the headphone line minus sign is long pressed, the system will pause reading and transcribe the voice input editing command into text, and send it together with the current sentence to the descriptive command processing module on the server side. The system receives the return result of the descriptive command processing module or the comprehensive module, replaces the original sentence with the best result on the interface, and reads the result aloud. The system enters the result selection mode, and the button style and the function of the headphone line button change.
[0207] 406. When the current / next button is clicked, or the plus / minus sign on the headphone line is short pressed, the system switches to the previous / next result, displays and reads the result.
[0208] 407. When the OK button is clicked, or the headphone cable button is short pressed, the system selects the result. When exiting the result selection mode, the button style and the headphone cable button function return to the original state.
[0209] 408. When the cancel button is clicked, or the middle button of the headphone cable is long pressed, the system terminates the current editing, changes the current sentence back to the original sentence, exits the result selection mode, and the button style and the function of the headphone cable button return to their original state.
[0210] Scenario Example 2:
[0211] Taking the in-vehicle voice editor as an example, the in-vehicle voice editor can interact using pure voice. The systems described in this embodiment are all intelligent voice editing systems, which are specifically described as follows:
[0212] 501. The system recognizes the user's voice. When the user says a command word related to voice input, such as "input", "send a message to xx", etc., the voice input process is started; the system recognizes the user's voice and can use voice collection, that is, collect the user's voice, and use the voice conversion module to convert the collected voice into corresponding text.
[0213] 502. The system receives and recognizes the paragraph spoken by the user, sends it to the spoken word removal model, and then receives the returned result. The system reads the paragraph aloud from the beginning.
[0214] 503. When a human voice or a specific user's voice is recognized, the playback is paused. During this stage, all the content spoken by the user is regarded as an editing command. The command spoken by the user is recognized and sent to the descriptive command processing module of the server together with the current sentence. The system receives the return result of the descriptive command processing module or the comprehensive module, and reads the editing results in sequence (reading the number first when reading each result).
[0215] 504. When the user says "OK", the result currently being read is selected. When the user says a specific number, the result with the corresponding number is selected. When the user says "Cancel", the current editing is canceled.
[0216] 505. Continue reading from the current sentence and jump to 503 after recognizing the human voice.
[0217] 506. After reading to the end of the paragraph, if the user says "OK", the editing of this paragraph ends. After that, all the contents are regarded as input contents. Jump to 502 and repeat the above steps until the user says a command to terminate the input, such as "Send".
[0218] Scenario Example 3:
[0219] Taking VR scene as an example, VR scene is combined with visual interface. Users can determine the modification range of text by head movement or eye movement. At the same time, intelligent editing technology is used to infer the specific position and range of modification, providing an efficient text editing method for VR devices. The specific process is as follows:
[0220] 601. After receiving the voice input paragraph, the system will send it to the spoken word removal model, and then receive the returned results and display them in the text editing box of the VR device.
[0221] 602. After receiving the user's voice, if the user gazes at a blank area, the user intends to input, and jump to 601; if the user gazes at the text that has been input, the intention is to edit, and the sentence and editing command near the gaze position are sent to the descriptive command processing module on the server side. The system receives the return result of the descriptive command processing module or the comprehensive module, and displays the alternative editing results through the candidate box. Here, the eye acquisition module can be used to collect the user's eye movements, and then determine whether the user intends to input or edit.
[0222] 603. When the user is heard to say a specific number, the result is selected. When the user is heard to say "cancel", the current editing is canceled.
[0223] When applied to VR scenes, in order to reduce errors, the line spacing of sentences can be appropriately increased, or each sentence can be displayed on a different page.
[0224] and Figure 1 Corresponding to the method shown in the figure, the present invention also provides a voice editing device to support Figure 1 The method shown in the figure is applied in practice. The device can be set in an intelligent voice editing system. The device can be composed of a computer terminal or an intelligent device. The structural diagram of the device is shown in FIG. Figure 5 The specific instructions are as follows:
[0225] The acquisition unit 701 is used to acquire an editing command input by a user and a sentence to be edited, wherein the sentence to be edited is a sentence selected by the user in the text to be edited, the text to be edited is a text obtained by converting the converted speech input by the user into text, and the editing command is a text obtained by converting the command speech input by the user based on the sentence to be edited into text;
[0226] A judging unit 702, configured to judge whether the editing command is a descriptive command;
[0227] A determination unit 703, configured to determine the editing command as a target sentence if the editing command is not a descriptive command;
[0228] A first input unit 704, used to input the target sentence and the sentence to be edited into a pre-trained replacement module and an insertion module;
[0229] A first triggering unit 705 is used to trigger the insertion module to process the target sentence and the sentence to be edited, and output M candidate insertion sentences, where M is a positive integer;
[0230] A second triggering unit 706 is used to trigger the replacement module to process the target sentence and the sentence to be edited, and output N candidate replacement sentences and empty phrase prediction probabilities, where N is a positive integer;
[0231] A second input unit 707, used to input the empty phrase prediction probability, each of the candidate replacement sentences and each of the candidate insertion sentences into a preset integration module;
[0232] A display unit 708, used to trigger the integration module to determine a target candidate sentence from each of the candidate replacement sentences and each of the candidate insertion sentences, and to display each of the target candidate sentences to the user;
[0233] A receiving unit 709 is configured to receive a selection instruction fed back by the user based on each of the target candidate sentences, and determine whether the selection instruction includes a sentence identifier;
[0234] The replacement unit 710 is configured to replace the sentence to be edited with the target candidate sentence corresponding to the sentence identifier in the selection instruction if the selection instruction includes a sentence identifier.
[0235] In the device provided by the embodiment of the present invention, the editing command input by the user and the sentence to be edited are obtained. When the editing command is not a descriptive command, the editing command is determined to be a target sentence, the target sentence and the sentence to be edited are input into the replacement module and the insertion module, and the N candidate replacement sentences and the empty phrase prediction probability output by the replacement module and the M candidate insertion sentences output by the insertion module are all input into the comprehensive module, so that the comprehensive module selects the target candidate sentence and displays each target candidate sentence to the user. When a selection instruction fed back by the user is received and a sentence identifier is included in the selection instruction, the sentence to be edited is replaced by the target candidate sentence corresponding to the sentence identifier. Thus, the user does not need to specify the location of the erroneous text, and the user can edit the text by only inputting a small amount of information through voice, which effectively shortens the text editing process using voice and improves the efficiency of voice editing; and provides a more convenient editing method for the user, which can predict the text content that the user wants to modify according to the target sentence input by the user, and provide the user with more suitable examples, thereby providing the user with better service and improving the efficiency of editing.
[0236] In the device provided in the embodiment of the present invention, the device may also be configured as:
[0237] A collection unit, configured to collect the converted voice of the user upon receiving a voice conversion instruction sent by the user, and call a preset voice conversion module to convert the converted voice into text;
[0238] A removal unit is used to input the text into a preset spoken language removal module, so that the spoken language removal module marks the spoken words in the text and obtains a marking sequence corresponding to the text, removes the spoken words in the text based on the marking sequence, and displays the text after the spoken words are removed as a text to be edited to the user.
[0239] In the device provided by the embodiment of the present invention, the determination unit 702 of the device may be configured as follows:
[0240] A matching subunit, used for matching the editing command with various preset regular expressions;
[0241] A first judging subunit, used for judging whether there is a regular expression corresponding to the editing command;
[0242] a first determining subunit, configured to determine that the editing command is a descriptive command if a regular expression corresponding to the editing command exists;
[0243] The second determining subunit is configured to determine that the editing command is not a descriptive command if there is no regular expression corresponding to the editing command.
[0244] In the device provided by the embodiment of the present invention, the first trigger unit 705 of the device may be configured as follows:
[0245] An obtaining subunit is used for the insertion module to perform word segmentation processing on the sentence to be edited to obtain at least two insertion positions corresponding to the sentence to be edited;
[0246] An inserting subunit, for inserting the target sentence into each insertion position of the sentence to be edited, to obtain a first candidate sentence corresponding to the insertion position;
[0247] an output subunit, configured to input each of the first candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a first candidate score for each of the first candidate sentences;
[0248] The first selection subunit is used to select first candidate sentences in descending order of first candidate scores until the number of selected first candidate sentences is M, and determine each selected first candidate sentence as a candidate insertion sentence.
[0249] In the device provided by the embodiment of the present invention, the second trigger unit 706 of the device may be configured as follows:
[0250] A construction subunit is used for the replacement module to process the target sentence and the sentence to be edited based on the neural network model to obtain vectors corresponding to the target sentence and the sentence to be edited, and to process the vectors based on a preset vocabulary limitation strategy to construct a search tree corresponding to the sentence to be edited, wherein the search tree contains a plurality of child nodes, and the words in each of the child nodes are composed of characters in the sentence to be edited;
[0251] A first generating subunit, configured to search each subnode in the search tree based on a preset beam search strategy to generate a plurality of erroneous short sentences and empty phrase prediction probabilities corresponding to the sentence to be edited;
[0252] The third determination subunit is used to determine the replacement probability of each erroneous short sentence, and select the erroneous short sentences in the order of the replacement probability from high to low, until the number of the selected erroneous phrases is consistent with the preset number of short sentences, and each selected erroneous short sentence is determined as a target erroneous short sentence;
[0253] a replacement subunit, configured to determine, for each target erroneous sentence, a content corresponding to the target erroneous sentence in the sentence to be edited, and replace the content corresponding to the target erroneous sentence in the sentence to be edited with the target sentence, thereby obtaining a replacement example sentence corresponding to the target erroneous sentence;
[0254] A second generating subunit is used to generate at least one supplementary example sentence based on the sentence to be edited and the target sentence when the sentence to be edited and the target sentence satisfy any one of the preset supplementary rules;
[0255] a fourth determination subunit, configured to determine each of the replacement example sentences and each of the supplementary example sentences as a first candidate sentence, and input each of the second candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a second candidate score for each of the second candidate sentences;
[0256] The second selection subunit is used to select second candidate sentences in descending order of the second candidate scores until the number of selected second candidate sentences is N, and determine each selected second candidate sentence as a candidate replacement sentence.
[0257] In the device provided in the embodiment of the present invention, the display unit 708 of the device may be configured as follows:
[0258] A fifth determination subunit, used by the integration module to determine a sentence score of each of the candidate replacement sentences and each of the candidate insertion sentences;
[0259] A sixth determining subunit, configured to determine a sentence score with the largest value among the sentence scores of the candidate replacement sentences, and determine the sentence score with the largest value as the first sentence score;
[0260] a seventh determination subunit, configured to determine a sentence score with the smallest value among the sentence scores of the candidate insertion sentences, and determine the sentence score with the smallest value as the second sentence score;
[0261] A second judging subunit, configured to judge whether the second sentence score is greater than the first sentence score;
[0262] a third selection subunit, configured to determine a first number of replacement statements and a first number of insertion statements based on a preset first selection rule if the second statement score is greater than the first statement score, and select candidate replacement statements in descending order of the statement scores of the candidate replacement statements until the number of the selected candidate replacement statements is equal to the first number of replacement statements, and select candidate insertion statements in descending order of the statement scores of the candidate insertion statements until the number of the selected candidate insertion statements is equal to the first number of insertion statements, and determine the selected candidate insertion statements and the selected candidate replacement statements as target candidate statements;
[0263] an eighth determination subunit, configured to determine whether the empty phrase prediction probability is within a preset first interval if the second sentence score is not greater than the first sentence score;
[0264] a fourth selection subunit, configured to select candidate replacement sentences in descending order of sentence scores of the candidate replacement sentences, until the number of the selected candidate replacement sentences is equal to the first number of replacement sentences, and select candidate insertion sentences in descending order of sentence scores of the candidate insertion sentences, until the number of the selected candidate insertion sentences is equal to the first number of insertion sentences, if the empty phrase prediction probability is within the first interval, and determine the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences;
[0265] a ninth determining subunit, configured to determine whether the empty phrase prediction probability is located in a preset second interval if the empty phrase prediction probability is not located in the first interval;
[0266] a fifth selection subunit, configured to determine, if the empty phrase prediction probability is within the second interval, a second number of replacement sentences and a second number of insertion sentences based on a preset second candidate rule, and select candidate replacement sentences in descending order of sentence scores of the candidate replacement sentences until the number of selected candidate replacement sentences equals the second number of replacement sentences, and select candidate insertion sentences in descending order of sentence scores of the candidate insertion sentences until the number of selected candidate insertion sentences equals the second number of insertion sentences, and determine the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences;
[0267] The sixth selection subunit is used to determine the third number of replacement sentences and the third number of insertion sentences based on a preset third candidate rule if the empty phrase prediction probability is not within the second interval, and select candidate replacement sentences in descending order of their sentence scores until the number of selected candidate replacement sentences is equal to the third number of replacement sentences, and select candidate insertion sentences in descending order of their sentence scores until the number of selected candidate insertion sentences is equal to the third number of insertion sentences, and determine the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences.
[0268] In the device provided in the embodiment of the present invention, the device may also be configured as:
[0269] The execution unit is used to determine the description type of the descriptive command if it is determined that the editing command is a descriptive command, and perform an editing operation corresponding to the description type on the sentence to be edited.
[0270] An embodiment of the present invention further provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the above-mentioned voice editing method.
[0271] The embodiment of the present invention further provides an electronic device, the structural diagram of which is shown in FIG. Figure 6 As shown, it specifically includes a memory 801 and one or more instructions 802, wherein the one or more instructions 802 are stored in the memory 801 and are configured to be executed by one or more processors 803 to perform the following operations:
[0272] Obtaining an editing command input by a user and a sentence to be edited, wherein the sentence to be edited is a sentence selected by the user in the text to be edited, the text to be edited is a text obtained by converting the converted speech input by the user into text, and the editing command is a text obtained by converting the command speech input by the user based on the sentence to be edited into text;
[0273] Determining whether the editing command is a descriptive command;
[0274] If the editing command is not a descriptive command, determining the editing command as a target sentence;
[0275] Inputting the target sentence and the sentence to be edited into a pre-trained replacement module and an insertion module;
[0276] Triggering the insertion module to process the target sentence and the sentence to be edited, and outputting M candidate insertion sentences, where M is a positive integer;
[0277] Triggering the replacement module to process the target sentence and the sentence to be edited, and outputting N candidate replacement sentences and empty phrase prediction probabilities, where N is a positive integer;
[0278] Inputting the empty phrase prediction probability, each of the candidate replacement sentences and each of the candidate insertion sentences into a preset integration module;
[0279] Triggering the integration module to determine a target candidate sentence from each of the candidate replacement sentences and each of the candidate insertion sentences, and displaying each of the target candidate sentences to the user;
[0280] receiving a selection instruction fed back by the user based on each of the target candidate sentences, and determining whether the selection instruction includes a sentence identifier;
[0281] If the selection instruction includes a sentence identifier, the sentence to be edited is replaced by the target candidate sentence corresponding to the sentence identifier in the selection instruction.
[0282] The specific implementation processes and derivative methods of the above-mentioned embodiments are all within the protection scope of the present invention.
[0283] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0284] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0285] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.< / eos> < / eos>
Claims
1. A voice editing method, characterized in that: include: Acquire an editing command input by a user and a sentence to be edited, wherein the sentence to be edited is a sentence selected by the user in the text to be edited, the text to be edited is a text obtained by converting the converted speech input by the user into text, and the editing command is a text obtained by converting the command speech input by the user based on the sentence to be edited into text; Determining whether the editing command is a descriptive command, wherein the descriptive command includes an insert command, a replace command, and a delete command; If the editing command is not a descriptive command, determining the editing command as a target sentence; Inputting the target sentence and the sentence to be edited into a pre-trained replacement module and an insertion module; Triggering the insertion module to process the target sentence and the sentence to be edited, and outputting M candidate insertion sentences, where M is a positive integer; Triggering the replacement module to process the target sentence and the sentence to be edited, and outputting N candidate replacement sentences and empty phrase prediction probabilities, where N is a positive integer; Inputting the empty phrase prediction probability, each of the candidate replacement sentences and each of the candidate insertion sentences into a preset integration module; Triggering the integration module to determine a target candidate sentence from each of the candidate replacement sentences and each of the candidate insertion sentences, and displaying each of the target candidate sentences to the user; receiving a selection instruction fed back by the user based on each of the target candidate sentences, and determining whether the selection instruction includes a sentence identifier; If the selection instruction includes a sentence identifier, replacing the sentence to be edited with the target candidate sentence corresponding to the sentence identifier in the selection instruction; The determining whether the editing command is a descriptive command includes: Matching the editing command with each preset regular expression; Determine whether there is a regular expression corresponding to the editing command; If there is a regular expression corresponding to the editing command, determining that the editing command is a descriptive command; If there is no regular expression corresponding to the editing command, determining that the editing command is not a descriptive command; The triggering of the replacement module to process the target sentence and the sentence to be edited, and outputting N candidate replacement sentences and empty phrase prediction probabilities, includes: The replacement module processes the target sentence and the sentence to be edited based on a neural network model to obtain vectors corresponding to the target sentence and the sentence to be edited, and processes the vectors based on a preset vocabulary limitation strategy to construct a search tree corresponding to the sentence to be edited, wherein the search tree includes a plurality of child nodes, and the words in each of the child nodes are composed of characters in the sentence to be edited; Based on a preset beam search strategy, searching each subnode in the search tree to generate a plurality of erroneous short sentences and empty phrase prediction probabilities corresponding to the sentence to be edited; Determine the replacement probability of each erroneous short sentence, and select the erroneous short sentences in descending order of replacement probability, until the number of selected erroneous phrases is consistent with the preset number of short sentences, and determine each selected erroneous short sentence as a target erroneous short sentence; For each of the target erroneous sentences, determining the content corresponding to the target erroneous sentence in the sentence to be edited, and replacing the content corresponding to the target erroneous sentence in the sentence to be edited with the target sentence, thereby obtaining a replacement example sentence corresponding to the target erroneous sentence; When the sentence to be edited and the target sentence satisfy any one of the preset supplementary rules, generating at least one supplementary example sentence based on the sentence to be edited and the target sentence; Determine each of the replacement example sentences and each of the supplementary example sentences as a second candidate sentence, and input each of the second candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a second candidate score for each of the second candidate sentences; Second candidate sentences are selected in descending order of the second candidate scores until the number of selected second candidate sentences reaches N, and each of the selected second candidate sentences is determined as a candidate replacement sentence.
2. The method according to claim 1, characterized in that Before obtaining the editing command input by the user and the statement to be edited, it also includes: When receiving a voice conversion instruction sent by a user, collecting the converted voice of the user, and calling a preset voice conversion module to convert the converted voice into text; The text is input into a preset spoken language removal module, so that the spoken language removal module marks the spoken words in the text and obtains a marking sequence corresponding to the text, removes the spoken words in the text based on the marking sequence, and displays the text after the spoken words are removed as a text to be edited to the user.
3. The method according to claim 1, characterized in that The triggering of the insertion module processes the target sentence and the sentence to be edited, and outputs M candidate insertion sentences, including: The insertion module performs word segmentation processing on the sentence to be edited to obtain at least two insertion positions corresponding to the sentence to be edited; For each insertion position of the sentence to be edited, insert the target sentence into the insertion position to obtain a first candidate sentence corresponding to the insertion position; Inputting each of the first candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a first candidate score for each of the first candidate sentences; First candidate sentences are selected in descending order of first candidate scores until the number of selected first candidate sentences is M, and each selected first candidate sentence is determined as a candidate insertion sentence.
4. The method according to claim 1, characterized in that The triggering of the synthesis module to determine a target candidate statement from each of the candidate replacement statements and each of the candidate insertion statements includes: The synthesis module determines a sentence score for each of the candidate replacement sentences and each of the candidate insertion sentences; Determine a sentence score with the largest value among the sentence scores of the candidate replacement sentences, and determine the sentence score with the largest value as the first sentence score; Determine the sentence score with the smallest value among the sentence scores of the candidate insertion sentences, and determine the sentence score with the smallest value as the second sentence score; Determining whether the second sentence score is greater than the first sentence score; If the second sentence score is greater than the first sentence score, determining the first number of replacement sentences and the first number of insertion sentences based on a preset first selection rule, selecting candidate replacement sentences in descending order of the sentence scores of the candidate replacement sentences until the number of the selected candidate replacement sentences is equal to the first number of replacement sentences, and selecting candidate insertion sentences in descending order of the sentence scores of the candidate insertion sentences until the number of the selected candidate insertion sentences is equal to the first number of insertion sentences, and determining the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences; If the second sentence score is not greater than the first sentence score, determining whether the empty phrase prediction probability is within a preset first interval; If the empty phrase prediction probability is within the first interval, selecting candidate replacement sentences according to the sentence scores of the candidate replacement sentences from high to low until the number of the selected candidate replacement sentences is equal to the first number of replacement sentences, and selecting candidate insertion sentences according to the sentence scores of the candidate insertion sentences from high to low until the number of the selected candidate insertion sentences is equal to the first number of insertion sentences, and determining the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences; If the empty phrase prediction probability is not within the first interval, determining whether the empty phrase prediction probability is within a preset second interval; If the empty phrase prediction probability is within the second interval, determining a second number of replacement sentences and a second number of insertion sentences based on a preset second candidate rule, selecting candidate replacement sentences in descending order of sentence scores of the candidate replacement sentences until the number of selected candidate replacement sentences equals the second number of replacement sentences, and selecting candidate insertion sentences in descending order of sentence scores of the candidate insertion sentences until the number of selected candidate insertion sentences equals the second number of insertion sentences, and determining the selected candidate insertion sentences and the selected candidate replacement sentences as target candidate sentences; If the empty phrase prediction probability is not within the second interval, a third number of replacement sentences and a third number of insertion sentences are determined based on a preset third candidate rule, and candidate replacement sentences are selected in descending order of their sentence scores until the number of selected candidate replacement sentences is equal to the third number of replacement sentences, and candidate insertion sentences are selected in descending order of their sentence scores until the number of selected candidate insertion sentences is equal to the third number of insertion sentences, and the selected candidate insertion sentences and the selected candidate replacement sentences are both determined as target candidate sentences.
5. The method according to claim 1, characterized in that Also includes: If it is determined that the editing command is a descriptive command, the description type of the descriptive command is determined, and an editing operation corresponding to the description type is performed on the sentence to be edited.
6. A voice editing device, characterized in that: include: an acquisition unit, configured to acquire an editing command input by a user and a sentence to be edited, wherein the sentence to be edited is a sentence selected by the user in the text to be edited, the text to be edited is a text obtained by converting the converted speech input by the user into text, and the editing command is a text obtained by converting the command speech input by the user based on the sentence to be edited into text; A judging unit, used to judge whether the editing command is a descriptive command, wherein the descriptive command includes an insert command, a replace command and a delete command; a determining unit, configured to determine the editing command as a target sentence if the editing command is not a descriptive command; A first input unit, used for inputting the target sentence and the sentence to be edited into a pre-trained replacement module and an insertion module; A first triggering unit is used to trigger the insertion module to process the target sentence and the sentence to be edited, and output M candidate insertion sentences, where M is a positive integer; A second triggering unit is used to trigger the replacement module to process the target sentence and the sentence to be edited, and output N candidate replacement sentences and empty phrase prediction probabilities, where N is a positive integer; A second input unit, used to input the empty phrase prediction probability, each of the candidate replacement sentences and each of the candidate insertion sentences into a preset integration module; A display unit, used to trigger the integration module to determine a target candidate sentence from each of the candidate replacement sentences and each of the candidate insertion sentences, and to display each of the target candidate sentences to the user; A receiving unit, configured to receive a selection instruction fed back by the user based on each of the target candidate sentences, and determine whether the selection instruction includes a sentence identifier; a replacement unit, configured to replace the sentence to be edited with a target candidate sentence corresponding to the sentence identifier in the selection instruction if the selection instruction includes a sentence identifier; The judging unit may be configured as follows: A matching subunit, used for matching the editing command with various preset regular expressions; A first judging subunit, used for judging whether there is a regular expression corresponding to the editing command; a first determining subunit, configured to determine that the editing command is a descriptive command if a regular expression corresponding to the editing command exists; a second determining subunit, configured to determine that the editing command is not a descriptive command if there is no regular expression corresponding to the editing command; The second trigger unit can be configured as: A construction subunit is used for the replacement module to process the target sentence and the sentence to be edited based on the neural network model to obtain vectors corresponding to the target sentence and the sentence to be edited, and to process the vectors based on a preset vocabulary limitation strategy to construct a search tree corresponding to the sentence to be edited, wherein the search tree contains a plurality of child nodes, and the words in each of the child nodes are composed of characters in the sentence to be edited; A first generating subunit, configured to search each subnode in the search tree based on a preset beam search strategy to generate a plurality of erroneous short sentences and empty phrase prediction probabilities corresponding to the sentence to be edited; The third determination subunit is used to determine the replacement probability of each erroneous short sentence, and select the erroneous short sentences in the order of the replacement probability from high to low, until the number of the selected erroneous phrases is consistent with the preset number of short sentences, and each selected erroneous short sentence is determined as a target erroneous short sentence; a replacement subunit, configured to determine, for each target erroneous sentence, a content corresponding to the target erroneous sentence in the sentence to be edited, and replace the content corresponding to the target erroneous sentence in the sentence to be edited with the target sentence, thereby obtaining a replacement example sentence corresponding to the target erroneous sentence; A second generating subunit is used to generate at least one supplementary example sentence based on the sentence to be edited and the target sentence when the sentence to be edited and the target sentence satisfy any one of the preset supplementary rules; a fourth determination subunit, configured to determine each of the replacement example sentences and each of the supplementary example sentences as a second candidate sentence, and input each of the second candidate sentences into a preset sentence scoring model, so that the sentence scoring model outputs a second candidate score for each of the second candidate sentences; The second selection subunit is used to select second candidate sentences in descending order of the second candidate scores until the number of selected second candidate sentences is N, and determine each selected second candidate sentence as a candidate replacement sentence.
7. A storage medium, characterized in that: The storage medium includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the voice editing method according to any one of claims 1 to 5.
8. An electronic device, characterized in that: It comprises a memory and one or more instructions, wherein the one or more instructions are stored in the memory and are configured to be executed by one or more processors to perform the speech editing method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Text editing method, device and system and terminal equipment
CN107861932A
Text error correction method and device, computer equipment and storage medium
CN111859921A