A method and device for acquiring target text based on speech recognition, and a storage medium

By obtaining and updating the recognition text of voice clips in the outgoing call robot, the problems of incomplete voice recognition and hysteresis of outgoing call robots are solved, and fast and accurate user voice text acquisition is achieved, improving user intention analysis and interactive experience.

CN115050370BActive Publication Date: 2025-08-26BEIJING WATERDROP TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210703643.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-08-26
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

When existing outbound call robots are voice recognition, improper setting of duration threshold results in incomplete user voice recognition or hysteresis, affecting the accuracy of user intention analysis and user experience.

Method used

By obtaining the recognition text of the first voice clip and judging its integrity, if it is incomplete, the recognition text of the second voice clip will be obtained within the preset time interval and updated until the recognition text is complete. The target text is finally obtained using the preset time interval and integrity constraints.

Benefits of technology

On the premise of ensuring timely response from the outgoing call robot, it quickly and accurately obtains the complete recognition text of user voice, improving the accuracy of user intention analysis and user experience of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115050370B_ABST
    Figure CN115050370B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for acquiring a target text based on speech recognition, a storage medium, and a computer device. The method comprises: acquiring a first recognized text corresponding to a first speech segment, using the first recognized text as a text to be judged, and judging whether the text to be judged is a complete text; when the text to be judged is not a complete text, judging whether a second recognized text corresponding to a second speech segment is acquired within a preset time interval; when the result is negative, using the text to be judged as the target text; when the result is positive, updating the text to be judged based on the second recognized text, and returning to the step of judging whether the text to be judged is a complete text, until the text to be judged is a complete text, or when the result is negative, the target text is obtained. The present application can quickly and accurately obtain the complete recognized text of the user's speech while ensuring that the outbound call robot responds in a timely manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to a method and device for acquiring target text based on speech recognition, a storage medium, and a computer device. Background Art

[0002] With the continuous development of speech recognition and artificial intelligence technologies, outbound call robots have emerged. They use speech recognition technology to convert user speech into text and, using artificial intelligence, automatically respond to the user's speech based on the recognized text. This significantly reduces the workload of customer service personnel.

[0003] The integrity of an outbound call robot's speech recognition directly impacts the accuracy of its analysis of user intent, which in turn affects the accuracy of the outbound call robot's responses. In the prior art, outbound call robots typically set a duration threshold for speech recognition. When the user pauses speaking for longer than the duration threshold, the robot assumes the user has finished speaking. The robot then performs speech recognition on the user's speech, generating the corresponding text. The robot then matches the recognized text to the outbound call robot's response. However, with this approach, if the duration threshold is too low, the user's "gasping" behavior may cause speech recognition to begin before the user has finished speaking, resulting in an inability to fully interpret the user's intent. If the duration threshold is too long, the robot's responses may be delayed. Consequently, existing outbound call robots often experience low user satisfaction and a poor user experience when interacting with users.

[0004] Therefore, how to quickly and accurately obtain the complete recognition text of the user's voice, improve the accuracy of parsing the user's intention, and thus enhance the user experience during the human-computer interaction process has become a technical problem that needs to be solved urgently in this field. Summary of the Invention

[0005] In view of this, the present application provides a target text acquisition method and device based on speech recognition, a storage medium, and a computer device, which can quickly and accurately obtain the complete recognition text of the user's voice while ensuring that the outbound call robot responds in a timely manner, thereby improving the accuracy of the user's intention analysis and effectively improving the user experience during the human-computer interaction process.

[0006] According to one aspect of the present application, a method for acquiring target text based on speech recognition is provided, comprising:

[0007] Obtaining a first recognized text corresponding to the first voice segment, using the first recognized text as a text to be determined, and determining whether the text to be determined is a complete text;

[0008] When the text to be determined is not a complete text, determining whether a second recognized text corresponding to the second voice segment is obtained within a preset time interval;

[0009] When the result is no, the text to be judged is used as the target text;

[0010] When the result is yes, the text to be judged is updated based on the second recognition text, and returns to the step of judging whether the text to be judged is a complete text, until the text to be judged is a complete text, or ends when the result is no, to obtain the target text.

[0011] According to another aspect of the present application, a target text acquisition device based on speech recognition is provided, comprising:

[0012] a text acquisition module, configured to acquire a first recognized text corresponding to a first voice segment, use the first recognized text as a text to be determined, and determine whether the text to be determined is a complete text;

[0013] a determination module, configured to determine whether a second recognized text corresponding to a second voice segment is obtained within a preset time interval when the text to be determined is not a complete text;

[0014] A target text determination module, configured to use the text to be determined as the target text when the result is negative;

[0015] The return module is used to update the text to be judged based on the second recognition text when the result is yes, and return to the step of judging whether the text to be judged is a complete text, until the text to be judged is a complete text, or end when the result is no, to obtain the target text.

[0016] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the above-mentioned target text acquisition method based on speech recognition is implemented.

[0017] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the above-mentioned method for acquiring target text based on speech recognition when executing the program.

[0018] By means of the above technical solution, the present application provides a method and apparatus for acquiring a target text based on speech recognition, a storage medium, and a computer device. First, a first speech segment and a first recognized text corresponding to the first speech segment can be acquired. Next, the first recognized text can be used as the text to be determined, and a determination can be made as to whether the text to be determined is a complete text. If, after determination, it is found that the text to be determined is not a complete text, a determination can be made as to whether a second recognized text corresponding to a second speech segment has been acquired within a preset time interval. If the second recognized text corresponding to the second speech segment has not been acquired within the preset time interval, the text to be determined can be directly used as the target text. If the second recognized text corresponding to the second speech segment has been acquired within the preset time interval, the text to be determined can be updated based on the second recognized text, so that the updated text to be determined includes the first recognized text and the second recognized text. Thereafter, a determination can be made again as to whether the updated text to be determined is a complete text. The above process is repeated until the updated text to be determined is a complete text, or until the second recognized text has not been received within the preset time interval. After this, the target text can be obtained based on the updated text to be determined. In the embodiment of the present application, after each time the recognized text is obtained, it is determined whether the recognized text is a complete text. If it is not a complete text, it is determined whether a new recognized text is received within a preset time interval. The target text is finally obtained by jointly constraining the integrity and the preset time interval. Under the premise of ensuring that the outbound call robot responds in a timely manner, the complete recognized text of the user's voice can be obtained quickly and accurately, thereby improving the accuracy of the analysis of the user's intention and effectively improving the user experience in the human-computer interaction process.

[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0021] Figure 1 A flow chart of a method for acquiring target text based on speech recognition provided in an embodiment of the present application is shown;

[0022] Figure 2 A flow chart of another method for acquiring target text based on speech recognition provided in an embodiment of the present application is shown;

[0023] Figure 3A flow chart of another method for acquiring target text based on speech recognition provided in an embodiment of the present application is shown;

[0024] Figure 4 A schematic structural diagram of a target text acquisition device based on speech recognition provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0025] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0026] In this embodiment, a method for acquiring target text based on speech recognition is provided. Figure 1 As shown, the method includes:

[0027] Step 101: obtaining a first recognized text corresponding to a first voice segment, using the first recognized text as a text to be determined, and determining whether the text to be determined is a complete text;

[0028] The target text acquisition method based on speech recognition provided in the embodiment of the present application can be applied to the outbound robot scenario, and can ensure that even in the case of "heavy breathing" when the user is speaking, the complete text corresponding to the user's voice can still be obtained. In the embodiment of the present application, the outbound robot can set a preset recognition time before performing text recognition on the user's voice. When the user stops speaking for longer than the preset recognition time, the user's voice can be recognized to obtain the corresponding recognition text. First, the outbound robot can use speech recognition technology to recognize the user's voice. When the user stops speaking for the first time for longer than the preset recognition time, a first voice segment and a first recognized text corresponding to the first voice segment can be obtained. Then, the first recognized text can be used as the text to be judged, and a preset method can be used to determine whether the text to be judged is a complete text. Here, the complete text refers to the text corresponding to the complete words spoken by the user. For example, when the first recognized text is "I want", after the user says "I want", the subsequent pause time is greater than the preset recognition time, but "I want" does not express the user's complete intention. In fact, the user's words are not finished at this time, so the first recognized text is not a complete text; when the first recognized text is "I want to know what the specific content of this business includes", the subsequent pause time is greater than the preset recognition time. At this time, the user has expressed the complete intention, so the first recognized text is a complete text.

[0029] Step 102: When the text to be determined is not a complete text, determining whether a second recognized text corresponding to a second voice segment is obtained within a preset time interval;

[0030] Step 103: when the result is no, the text to be judged is used as the target text;

[0031] In this embodiment, if it is determined that the text to be determined is a complete text, then the text to be determined is directly output as the target text. If it is determined that the text to be determined is not a complete text, the accuracy of the outbound call robot matching the response text based on the text to be determined (that is, the first recognition text) is low. Therefore, it can be determined whether the second recognition text corresponding to the second voice segment is obtained within the preset time interval. Here, the preset time interval can be obtained based on the human-computer conversation log. For example, the time when the user did not express the complete intention each time, but continued to express it later, can be found in the human-computer conversation log. The average value of the pause time can be calculated and used as the preset time interval. In addition, the maximum value of the pause time can also be used as the preset time interval. The purpose of setting the preset time interval in the embodiment of the present application is to ensure that the text to be determined is not a complete text, and when the user does not speak for a long time later, the timeliness of the outbound call robot's response can be improved. Although the embodiment of the present application has a high accuracy in determining whether the text to be determined is a complete text, in order to avoid the outbound call robot from falling into infinite waiting when the user stops speaking even before finishing speaking for various reasons, when the second recognized text corresponding to the second voice segment is not obtained within the preset time interval, the text to be determined can be directly used as the target text.

[0032] Step 104, when the result is yes, update the text to be judged based on the second recognized text, and return to the step of judging whether the text to be judged is a complete text, until the text to be judged is a complete text, or end when the result is no, to obtain the target text.

[0033] In this embodiment, when a second recognized text corresponding to a second voice segment is obtained within a preset time interval, the text to be determined can be updated based on the second recognized text so that the updated text to be determined includes the first recognized text and the second recognized text. Thereafter, it can be determined whether the updated text to be determined is a complete text. If it is a complete text, the updated text to be determined is output as the target text. If it is not a complete text, it is again determined whether a second recognized text corresponding to a new second voice segment is obtained within the preset time interval. Here, the first voice segment can be the first voice segment of the user recognized, and all user voice segments after the first voice segment can be referred to as second voice segments. If a second recognized text is obtained again within the preset time interval, the second recognized text can be used to continue updating the previously updated text to be determined, obtaining an updated text to be determined (i.e., the first recognized text + two second recognized texts) again. The above process is repeated until the updated text to be determined is satisfied that it is a complete text, or when no second recognized text is received within the preset time interval. After this, the target text can be obtained based on the updated text to be determined. Subsequent outbound call robots can directly match response texts based on the target text, which can greatly improve the accuracy of response text matching.

[0034] By applying the technical solution of this embodiment, first, a first voice segment and a first recognized text corresponding to the first voice segment can be obtained. Next, the first recognized text can be used as the text to be judged, and it can be determined whether the text to be judged is a complete text. If it is determined that the text to be judged is not a complete text, it can be determined whether a second recognized text corresponding to the second voice segment has been obtained within a preset time interval. If the second recognized text corresponding to the second voice segment has not been obtained within the preset time interval, the text to be judged can be directly used as the target text. When the second recognized text corresponding to the second voice segment has been obtained within the preset time interval, the text to be judged can be updated based on the second recognized text so that the updated text to be judged includes the first recognized text and the second recognized text. Thereafter, it can be determined again whether the updated text to be judged is a complete text... The above process is repeated until the updated text to be judged is a complete text, or until the second recognized text is not received within the preset time interval. After that, the target text can be obtained based on the updated text to be judged. In the embodiment of the present application, after each time the recognized text is obtained, it is determined whether the recognized text is a complete text. If it is not a complete text, it is determined whether a new recognized text is received within a preset time interval. The target text is finally obtained by jointly constraining the integrity and the preset time interval. Under the premise of ensuring that the outbound call robot responds in a timely manner, the complete recognized text of the user's voice can be obtained quickly and accurately, thereby improving the accuracy of the analysis of the user's intention and effectively improving the user experience in the human-computer interaction process.

[0035] Furthermore, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, another method for obtaining target text based on speech recognition is provided, such as Figure 2 As shown, the method includes:

[0036] Step 201: Obtain a first recognized text corresponding to a first voice segment, and use the first recognized text as a text to be determined;

[0037] In this embodiment, first, a first speech segment and a first recognized text corresponding to the first speech segment may be obtained. Then, the first recognized text may be used as the text to be determined.

[0038] Step 202: Based on the preset matching text, determine whether the text to be determined is an incomplete text; if the result is yes, output the determination result; if the result is no, input the text to be determined into a text integrity recognition model, and determine whether the text to be determined is a complete text based on the model output result of the text integrity recognition model;

[0039] In this embodiment, the text to be judged can first be judged based on the preset matching text to determine whether the text to be judged is an incomplete text. In other words, the role of the preset matching text is to identify whether the text to be judged is an incomplete text. Specifically, after the judgment based on the preset matching text, two output results can be obtained. When the output result is "yes", it indicates that the text to be judged is an incomplete text. When the output result is "no", it indicates that the preset matching text has not determined that the text to be judged is an incomplete text. The preset matching text can specifically be a set of commonly used words, vocabulary, or phrases in incomplete texts. The text to be judged is matched with these preset words, vocabulary, or phrases. If the match is successful, the output result is "yes", indicating that the text to be judged is incomplete. If the match is not successful, the output result is "no", indicating that the preset matching text has not determined that the text to be judged is an incomplete text. Therefore, the text to be judged can be further input into a text integrity recognition model, which can continue to judge the integrity of the text to be judged and determine whether the text to be judged is a complete text. The embodiment of the present application sets two judgment steps, and when the judgment result of the first step is an incomplete text, the judgment result is directly returned, thereby ensuring the accuracy and effectiveness of the system's response.

[0040] Step 203: When the text to be determined is not a complete text, determining whether a second recognized text corresponding to the second voice segment is obtained within a preset time interval;

[0041] In this embodiment, if it is found through determination that the text to be determined is not a complete text, it may be determined whether a second recognized text corresponding to the second voice segment is obtained within a preset time interval.

[0042] Step 204: when the result is negative, the text to be judged is used as the target text;

[0043] Step 205, when the result is yes, update the text to be judged based on the second recognized text, and return to the step of judging whether the text to be judged is a complete text, until the text to be judged is a complete text, or end when the result is no, to obtain the target text.

[0044] In this embodiment, when the second recognition text corresponding to the second voice segment is not obtained within the preset time interval, the text to be judged can be directly used as the target text. When the second recognition text corresponding to the second voice segment is obtained within the preset time interval, the text to be judged can be updated according to the second recognition text, so that the updated text to be judged contains the first recognition text and the second recognition text. Afterwards, it can be judged again whether the updated text to be judged is a complete text... Repeat the above process until it is satisfied that the updated text to be judged is a complete text, or it ends when the second recognition text is not received within the preset time interval, and then the target text can be obtained according to the updated text to be judged after the end.

[0045] In an embodiment of the present application, optionally, the "based on the preset matching text, determining whether the text to be judged is an incomplete text" in step 202 includes: based on the precise matching mode, determining whether the text to be judged contains the first preset matching text; and / or, based on the local matching mode, determining whether the end of the sentence of the text to be judged contains the second preset matching text; and / or, based on the regular matching mode, determining whether the text to be judged contains the third preset matching text; when there is an included result, judging that the text to be judged is an incomplete text.

[0046] In this embodiment, when making a judgment based on a preset matching text, it can be implemented by at least one of the following three modes, which can specifically be an exact matching mode, a partial matching mode, and a regular matching mode.

[0047] The precise matching mode can predetermine a first preset matching text. The first preset matching text may include single words, vocabulary, and sentences with obvious non-ending features that are frequently used in human-computer conversation logs, such as "then you," "then I," "that," "consider," "oh," "I me," "this," "then," and "wait a moment." If the first preset matching text exists in the text to be judged, it can be determined that the text to be judged is incomplete.

[0048] The local matching mode can predetermine a second preset matching text. This second preset matching text may include words, phrases, and sentences with non-ending features such as transitions frequently used by historical users in human-computer conversation logs. When the second preset matching text appears at the end of the text to be judged, it can be considered that the text to be judged is incomplete. For example, if the user's voice says "I heard, but", and the user pauses for longer than the preset recognition time after saying "but", the outbound call robot will end the previous user voice recognition and obtain the text to be judged "I heard, but". At this time, the word corresponding to the end of the text to be judged is "but", which clearly indicates that the user did not fully express their intention. If this text to be judged is used to directly match the response text, the accuracy of the response text can be imagined. If "but" is used as the second preset matching text, when "but" is matched at the end of the recognized text, the recognized text is considered incomplete. Subsequently, the complete target text is gradually obtained, which can greatly improve the matching accuracy of the response text and enhance the user interaction experience.

[0049] The regular matching mode can predetermine a third preset matching text, which can be some preset regular expressions, such as "Oh this | listen to me". If the text to be judged successfully matches the third preset matching text, the text to be judged can also be considered as an incomplete text.

[0050] If the judgment is based on the preset matching text, and only one of the exact matching mode, partial matching mode and regular matching mode is used to implement it, then as long as the match is successful, the text to be judged can be judged as incomplete text; if at least two modes are included for implementation, then as long as one of the modes is successfully matched, the text to be judged can be judged as incomplete text.

[0051] In an embodiment of the present application, optionally, before the step 202 of "determining whether the text to be determined is an incomplete text based on a preset matching text", the method further includes: obtaining a human-computer conversation log, determining a historical voice record corresponding to a historical user from the human-computer conversation log, and adding an attribute tag to the historical voice record, wherein the attribute tag includes a complete tag and an incomplete tag; identifying a first voice record with a pause time greater than a preset pause time from the historical voice record with a complete tag, determining a target position where the pause time is greater than the preset pause time from the first voice record, determining a first text based on the target position, and determining the first preset matching text based on the first text; and / or, identifying a second text corresponding to the end of each voice from the historical voice record with an incomplete tag, and using the second text whose occurrence number is greater than a first number threshold as the second preset matching text; and / or, determining a regular expression whose occurrence number is greater than a second number threshold from the historical voice record with an incomplete tag, and determining the third preset matching text based on the regular expression.

[0052] In this embodiment, when determining the preset matching text, the human-computer conversation log can first be obtained from the database. The human-computer conversation log can include the conversation content in the form of voice between the outbound call robot and the user. Then, the historical voice records corresponding to the historical users can be found from the human-computer conversation log. These historical voice records can include complete voice records, that is, voice records corresponding to the complete words spoken by the user. In addition, they can also include incomplete voice records, that is, voice records corresponding to incomplete words spoken by the user. Therefore, attribute tags can be added to these historical voice records respectively. Here, the attribute tags can include complete tags and incomplete tags. Specifically, a complete tag is added to the historical voice records that are complete voice records, and an incomplete tag is added to the historical voice records that are incomplete voice records.

[0053] After adding the attribute tags, first, you can find the historical voice records with pause times from the historical voice records with complete tags, and determine the first voice record with a pause time greater than the preset pause time from the historical voice records with pause times. Then, determine the target position from each first voice record. Here, the target position is the position in each first voice record where the pause time is greater than the preset pause time. For example, the conversation content corresponding to the first voice record is "Oh, (pause 2 seconds) I don't have time now." Then the position where there is a pause of 2 seconds is the target position. Afterwards, the first text can be determined based on the target position. The first text can specifically be the last word, vocabulary, sentence, etc. before the target position. For example, in the above example, "Oh" can be the first text. Finally, the first preset matching text can be determined based on the first text. Specifically, the first text that appears more times can be used as the first preset matching text, and the first text that appears less times can be discarded.

[0054] Furthermore, after adding attribute tags, the second text corresponding to the end of each speech in the historical voice recordings with incomplete tags can be found. The second text whose occurrence count exceeds the first threshold can then be used as the second preset matching text. The second text can be a single word, vocabulary, or sentence. For example, if the historical voice recording with incomplete tags is "I heard, but," then "but" can be the second text.

[0055] In addition, regular expressions with a number of occurrences greater than a second threshold value can be found from these historical voice records with incomplete labels, and then the third preset matching text can be further determined based on these regular expressions.

[0056] In an embodiment of the present application, optionally, after the "adding attribute labels to the historical voice records", the method further includes: converting the historical voice records into corresponding historical texts, performing word segmentation processing on the historical texts, and determining the word vector combination corresponding to each of the historical texts based on the word segmentation processing results; inputting the word vector combination into the initial recognition model respectively, and determining the model loss value of the initial recognition model based on the output prediction results and the attribute labels corresponding to the historical texts; when the model loss value is greater than the preset loss threshold, adjusting the model parameters of the initial recognition model, and returning to the step of inputting the word vector combination into the initial recognition model respectively, until the model loss value is less than or equal to the preset loss threshold, and obtaining the text integrity recognition model.

[0057] In this embodiment, after adding attribute labels to each historical voice record, the historical voice record can also be converted into corresponding historical texts, and then the initial recognition model can be trained using these historical texts to obtain a text integrity recognition model. First, each historical text can be segmented. Specifically, each historical text can be segmented according to vocabulary or according to individual characters. Then, the word vector corresponding to each segmented word can be determined from the preset word vector table based on the segmentation result. When segmenting words based on vocabulary, the one-dimensional vector corresponding to each character in each vocabulary can be determined from the preset word vector table, and then the one-dimensional vector corresponding to each character in the vocabulary can be spliced ​​to obtain the word vector corresponding to the vocabulary; when segmenting words based on individual characters, the one-dimensional vector corresponding to each character can be determined from the preset word vector table, that is, the word vector corresponding to the segmented word. After determining the word vector corresponding to each segmented word, the word vector corresponding to each segmented word can be combined to finally obtain the word vector combination corresponding to each historical text. After obtaining the word vector combination corresponding to each historical text, the word vector combination corresponding to each historical text can be input into the initial recognition model to obtain a prediction result, which can be a probability value. Then, based on the prediction result and the attribute label corresponding to the historical text, the model loss value corresponding to the initial recognition model can be calculated. Here, the complete label in the attribute label can be assigned a value of 1, and the incomplete label in the attribute label can be assigned a value of 0. After the assignment, it is convenient to calculate the model loss value. When the model loss value is greater than the preset loss threshold, it means that the model accuracy of the initial recognition model has not reached the expected goal, and the model parameters of the initial recognition model need to be adjusted. After adjusting the parameters, the word vector combination is input into the initial recognition model after adjusting the parameters again, and the model loss value is calculated again until the model loss value is less than or equal to the preset loss threshold. The initial recognition model with a model loss value less than or equal to the preset loss threshold is used as the text integrity recognition model.

[0058] In an embodiment of the present application, optionally, the "updating the text to be judged based on the second recognized text" in step 205 includes: splicing the first recognized text and the second recognized text according to the time when the texts were generated, and updating the text to be judged based on the splicing result.

[0059] In this embodiment, the specific process of updating the text to be determined using the second recognition text can be as follows: After obtaining the second recognition text, the first recognition text and the second recognition text can be spliced ​​together, and the text to be determined can be updated based on the splicing result. During the splicing, the first recognition text obtained can be placed in front, and the second recognition text obtained can be placed in the back, that is, the splicing is performed according to the time when the recognition texts were generated. For example, if the first recognition text obtained first is "I want to consult" and the second recognition text obtained later is "How to handle business A", then after obtaining the second recognition text, the first recognition text and the second recognition text can be spliced ​​together to obtain the updated text to be determined "I want to consult how to handle business A". Similarly, if the updated text to be determined is still not a complete text, a new second recognition text can be obtained again. In this case, the second recognition text can be spliced ​​after the first recognition text and the previous second recognition text. For example, if the first recognition text is "I want", the first second recognition text is "consult", and the second second recognition text is "detailed information about business A", then the splicing result is "I want to consult detailed information about business A". In this embodiment of the present application, the recognition text obtained after the first recognition text can be referred to as the second recognition text.

[0060] In an embodiment of the present application, optionally, before "determining whether the second recognition text corresponding to the second voice segment is obtained within the preset time interval" in step 203, the method further includes: starting a preset timer and setting the timing of the preset timer to the preset time interval; after "when the result is no" in step 204, the method further includes: clearing the timing of the preset timer.

[0061] In this embodiment, before determining whether the second recognized text corresponding to the second voice segment is obtained within a preset time interval, a preset timer can be started, and the preset timer can be used to determine the preset time interval, wherein the timing of the preset timer can be set to the preset time interval. In addition, if the second recognized text is not obtained within the preset time interval, the timing of the preset timer can be reset to zero and wait for the next start. The embodiment of the present application can simply and conveniently determine the preset time interval by using the preset timer.

[0062] Furthermore, the present application provides another method for acquiring target text based on speech recognition, such as Figure 3 As shown, the method includes:

[0063] First, set the speech recognition cumulative result variable C, obtain the speech recognition result from the speech recognition system, accumulate it to the speech recognition cumulative result variable C, and then input the speech recognition cumulative result into the text classification system. If it is determined that the text corresponding to the current speech recognition cumulative result is incomplete, start the timer T and wait for the output of subsequent speech recognition results. If a new speech recognition result is obtained from the speech recognition system before the timer T expires, continue to accumulate the speech recognition result into the speech recognition cumulative result variable C and then input it into the text classification system. If it is determined that the text expression corresponding to the speech recognition cumulative result is complete, clear the timer T, output the final speech recognition cumulative result, and clear the speech recognition cumulative result variable C. If no new speech recognition result is obtained from the speech recognition system when the timer T expires, clear the timer, output the final speech recognition cumulative result, and clear the speech recognition cumulative result variable C.

[0064] For example: the voice recognition system outputs the result "I have time", but in fact the user just pauses, and there is a subsequent "I don't have time" in the real voice. When the voice recognition system obtains "I have time" for the first time, it stores it in the voice recognition cumulative result variable C, and makes a completeness judgment on the "I have time" in the current voice recognition cumulative result variable C. The text classification system believes that this sentence is incomplete, so it starts the timer T, and continues to receive the subsequent user voice recognition results "I don't have time", and then splices "I have no time" and the previous "I have time" into "I don't have time for the time being", inputs it into the text classification system, and judges whether the text expression is complete. If the result is complete, the final voice recognition cumulative result is output and sent to the outbound robot, so that the outbound robot makes a corresponding response. The embodiment of the present application can quickly and accurately obtain the complete recognition text of the user's voice under the premise of ensuring the timely response of the outbound robot, thereby improving the accuracy of the analysis of the user's intention and effectively improving the user experience during the human-computer interaction process.

[0065] Further, as Figure 1 The specific implementation of the method, the embodiment of the present application provides a target text acquisition device based on speech recognition, such as Figure 4 As shown, the device includes:

[0066] a text acquisition module, configured to acquire a first recognized text corresponding to a first voice segment, use the first recognized text as a text to be determined, and determine whether the text to be determined is a complete text;

[0067] a determination module, configured to determine whether a second recognized text corresponding to a second voice segment is obtained within a preset time interval when the text to be determined is not a complete text;

[0068] A target text determination module, configured to use the text to be determined as the target text when the result is negative;

[0069] The return module is used to update the text to be judged based on the second recognition text when the result is yes, and return to the step of judging whether the text to be judged is a complete text, until the text to be judged is a complete text, or end when the result is no, to obtain the target text.

[0070] Optionally, the text acquisition module includes:

[0071] A judging unit, configured to judge whether the text to be judged is an incomplete text based on a preset matching text;

[0072] A result output unit, configured to output the judgment result when the result is yes;

[0073] The text input unit is used to input the text to be judged into a text integrity recognition model when the result is no, and determine whether the text to be judged is a complete text based on the model output result of the text integrity recognition model.

[0074] Optionally, the judgment unit is used to: determine whether the text to be judged contains a first preset matching text based on a precise matching mode; and / or, determine whether the end of the sentence of the text to be judged contains a second preset matching text based on a local matching mode; and / or, determine whether the text to be judged contains a third preset matching text based on a regular matching mode; when there is an inclusion result, judge that the text to be judged is an incomplete text.

[0075] Optionally, the device further comprises:

[0076] a tag adding module configured to obtain a human-computer conversation log, determine a historical voice recording corresponding to a historical user from the human-computer conversation log, and add attribute tags to the historical voice recordings before determining whether the text to be determined is an incomplete text based on the preset matching text;

[0077] a first recognition module configured to identify, from historical voice recordings with complete labels, a first voice recording having a pause time that is greater than a preset pause time, determine a target position where the pause time is greater than the preset pause time from the first voice recording, determine a first text based on the target position, and determine the first preset matching text based on the first text; and / or,

[0078] A second recognition module is configured to identify, from historical speech recordings with incomplete labels, a second text corresponding to the end of each speech, and use the second text whose number of occurrences is greater than a first count threshold as the second preset matching text; and / or,

[0079] The third recognition module is used to determine a regular expression whose number of occurrences is greater than a second number threshold from the historical voice records with incomplete tags, and determine the third preset matching text based on the regular expression.

[0080] Optionally, the device further comprises:

[0081] A word segmentation module is used to convert the historical voice record into a corresponding historical text after adding attribute tags to the historical voice record, perform word segmentation processing on the historical text, and determine a word vector combination corresponding to each historical text based on the word segmentation processing result;

[0082] A loss value determination module is used to input the word vector combination into the initial recognition model respectively, and determine the model loss value of the initial recognition model based on the output prediction result and the attribute label corresponding to the historical text;

[0083] A parameter adjustment module is used to adjust the model parameters of the initial recognition model when the model loss value is greater than a preset loss threshold, and return to the step of inputting the word vector combination into the initial recognition model respectively until the model loss value is less than or equal to the preset loss threshold, thereby obtaining the text integrity recognition model.

[0084] Optionally, the return module is configured to: concatenate the first recognized text and the second recognized text according to text generation time, and update the text to be determined based on the concatenation result.

[0085] Optionally, the device further comprises:

[0086] A starting module, configured to start a preset timer and set the timing of the preset timer to the preset time interval before determining whether the second recognized text corresponding to the second voice segment is obtained within the preset time interval;

[0087] Accordingly, the device further comprises:

[0088] The reset module is used to reset the timing of the preset timer when the result is no.

[0089] It should be noted that for other corresponding descriptions of the functional units involved in the target text acquisition device based on speech recognition provided in the embodiment of the present application, please refer to Figures 1 to 3The corresponding description in the method will not be repeated here.

[0090] Based on the above Figures 1 to 3 The method shown in FIG. 1 is a method for performing the above-mentioned operation. Accordingly, the embodiment of the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned operation is performed. Figures 1 to 3 The target text acquisition method based on speech recognition is shown.

[0091] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each implementation scenario of the present application.

[0092] Based on the above Figures 1 to 3 The method shown, and Figure 4 In order to achieve the above-mentioned purpose, the embodiment of the present application further provides a computer device, which can be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figures 1 to 3 The target text acquisition method based on speech recognition is shown.

[0093] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a Wi-Fi module, etc. The user interface may include a display, an input unit such as a keyboard, etc., and the optional user interface may also include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a Wi-Fi interface), etc.

[0094] Those skilled in the art will understand that the computer device structure provided in this embodiment does not constitute a limitation on the computer device, and may include more or fewer components, or a combination of certain components, or different component arrangements.

[0095] The storage medium may also include an operating system and a network communication module. An operating system is a program that manages and stores the hardware and software resources of a computer device, supporting the execution of information processing programs and other software and / or programs. The network communication module facilitates communication between components within the storage medium, as well as with other hardware and software within the physical device.

[0096] From the above description of the embodiments, those skilled in the art will clearly understand that the present application can be implemented using software plus the necessary general-purpose hardware platform, or it can be implemented using hardware. First, a first voice segment and a first recognized text corresponding to the first voice segment can be obtained. Next, the first recognized text can be used as the text to be determined, and a determination can be made as to whether the text to be determined is complete. If the determination reveals that the text to be determined is not complete, a determination can be made as to whether a second recognized text corresponding to a second voice segment has been obtained within a preset time interval. If the second recognized text corresponding to the second voice segment has not been obtained within the preset time interval, the text to be determined can be directly used as the target text. If the second recognized text corresponding to the second voice segment has been obtained within the preset time interval, the text to be determined can be updated based on the second recognized text, so that the updated text to be determined includes both the first recognized text and the second recognized text. Thereafter, a determination can be made again as to whether the updated text to be determined is complete. The above process is repeated until the updated text to be determined is complete, or until the second recognized text has not been received within the preset time interval. After this, the target text can be obtained based on the updated text to be determined. In the embodiment of the present application, after each time the recognized text is obtained, it is determined whether the recognized text is a complete text. If it is not a complete text, it is determined whether a new recognized text is received within a preset time interval. The target text is finally obtained by jointly constraining the integrity and the preset time interval. Under the premise of ensuring that the outbound call robot responds in a timely manner, the complete recognized text of the user's voice can be obtained quickly and accurately, thereby improving the accuracy of the analysis of the user's intention and effectively improving the user experience in the human-computer interaction process.

[0097] Those skilled in the art will understand that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily required to implement the present application. Those skilled in the art will understand that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the implementation scenario description, or can be changed accordingly and located in one or more devices different from the implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0098] The serial numbers of the above application are for descriptive purposes only and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure only discloses several specific implementation scenarios of the present application, but the present application is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present application.

Claims

1. A method for acquiring target text based on speech recognition, characterized in that: include: Obtaining a first recognized text corresponding to the first voice segment, using the first recognized text as a text to be determined, and determining whether the text to be determined is a complete text; When the text to be determined is not a complete text, determining whether a second recognized text corresponding to the second voice segment is obtained within a preset time interval; When the result is no, the text to be judged is used as the target text; When the result is yes, the text to be determined is updated based on the second recognized text, and the process returns to the step of determining whether the text to be determined is a complete text, until the text to be determined is a complete text, or the process ends when the result is no, thereby obtaining the target text; The determining whether the text to be determined is a complete text includes: Based on a preset matching text, determining whether the text to be determined is an incomplete text through at least one of an exact matching mode, a partial matching mode, and a regular matching mode, wherein the preset matching text includes words, vocabulary, and speech that are commonly used in incomplete texts; When the result is yes, output the judgment result; When the result is no, the text to be judged is input into a text integrity recognition model, and whether the text to be judged is a complete text is determined based on a model output result of the text integrity recognition model.

2. The method according to claim 1, characterized in that The determining whether the text to be determined is an incomplete text based on the preset matching text includes: Based on the precise matching mode, determining whether the text to be determined contains a first preset matching text; and / or, Based on the local matching pattern, determining whether the end of the sentence of the text to be judged contains a second preset matching text; and / or, Based on the regular matching pattern, determining whether the text to be judged contains a third preset matching text; When there is an included result, it is determined that the text to be judged is an incomplete text.

3. The method according to claim 2, characterized in that Before determining whether the text to be determined is an incomplete text based on the preset matching text, the method further includes: Obtaining a human-computer conversation log, determining a historical voice record corresponding to a historical user from the human-computer conversation log, and adding attribute tags to the historical voice record, wherein the attribute tags include complete tags and incomplete tags; Identifying, from historical voice recordings with complete labels, a first voice recording having a pause time that is greater than a preset pause time, determining, from the first voice recording, a target position where the pause time is greater than the preset pause time, determining a first text based on the target position, and determining the first preset matching text based on the first text; and / or, Identifying the second text corresponding to the end of each speech in the historical speech recordings with incomplete labels, and using the second text whose number of occurrences is greater than the first number threshold as the second preset matching text; and / or, A regular expression whose occurrence times is greater than a second threshold value is determined from the historical voice records with incomplete tags, and the third preset matching text is determined based on the regular expression.

4. The method according to claim 3, characterized in that After adding attribute tags to the historical voice records, the method further includes: Converting the historical voice records into corresponding historical texts, performing word segmentation processing on the historical texts, and determining a word vector combination corresponding to each of the historical texts based on the word segmentation processing results; Inputting the word vector combinations into an initial recognition model respectively, and determining a model loss value of the initial recognition model based on the output prediction results and the attribute labels corresponding to the historical texts; When the model loss value is greater than the preset loss threshold, the model parameters of the initial recognition model are adjusted, and the process returns to the step of inputting the word vector combination into the initial recognition model respectively, until the model loss value is less than or equal to the preset loss threshold, thereby obtaining the text integrity recognition model.

5. The method according to claim 1, characterized in that The updating of the to-be-determined text based on the second recognized text includes: The first recognition text and the second recognition text are spliced ​​according to the text generation time, and the text to be judged is updated based on the splicing result.

6. The method according to claim 1, characterized in that Before determining whether a second recognized text corresponding to the second voice segment is acquired within a preset time interval, the method further includes: Starting a preset timer and setting the timing of the preset timer to the preset time interval; Accordingly, when the result is negative, the method further includes: The timing of the preset timer is cleared.

7. A target text acquisition device based on speech recognition, characterized in that: include: a text acquisition module, configured to acquire a first recognized text corresponding to a first voice segment, use the first recognized text as a text to be determined, and determine whether the text to be determined is a complete text; a determination module, configured to determine whether a second recognized text corresponding to a second voice segment is obtained within a preset time interval when the text to be determined is not a complete text; A target text determination module, configured to use the text to be determined as the target text when the result is negative; a return module configured to, when the result is yes, update the text to be determined based on the second recognized text, and return to the step of determining whether the text to be determined is a complete text, until the text to be determined is a complete text, or terminate when the result is no, to obtain the target text; The text acquisition module includes: a judgment unit, configured to judge whether the text to be judged is an incomplete text based on a preset matching text by using at least one of an exact matching mode, a partial matching mode, and a regular matching mode, wherein the preset matching text includes single words, vocabulary, and speech terms commonly used in incomplete texts; A result output unit, configured to output the judgment result when the result is yes; The text input unit is used to input the text to be judged into a text integrity recognition model when the result is no, and determine whether the text to be judged is a complete text based on the model output result of the text integrity recognition model.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Voice conversation processing method and system

    CN112995419A

  • Real-time speech recognition method and device, equipment and medium

    CN113516994A