A method, device, equipment and storage medium for switching speech rights
By collecting user voice and video data in the smart device and using the pre-trained decision-making network model to make the right to speak, the problem of misswitching the right to speak in the smart device is solved, and the fluency and user experience of human-computer dialogue are improved.
Patent Information
- Application Number
- CN202211073708.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-02
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-09-02
AI Technical Summary
The voice acquisition device in the prior art smart device may mistakenly collect the user's input voice, resulting in the wrong switching of voice rights, reducing the fluency of human-computer dialogue and user experience.
When the user inputs a voice during the playback of the target response voice, the target response voice is paused, and the user's target voice data and target video data are collected, and the user's target voice data are inputted into the target decision network model obtained in advance to make a switching decision on the voice, and the voice power is switched based on the decision result.
It improves the accuracy of voice switching, enhances the fluency of human-computer dialogue, and improves user experience.
Smart Images

Figure CN115440220B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to computer technology, and in particular to a method, device, equipment and storage medium for switching speech rights. Background Art
[0002] With the rapid development of computer technology, intelligent devices can be used for human-computer dialogue, thus saving labor costs. In the process of human-computer dialogue, users often interrupt the intelligent device to speak, so it is necessary to switch the right of speech in time so that the user can obtain the right of speech to express himself, thereby improving the user experience.
[0003] Currently, smart devices usually detect whether there is user input voice during the process of playing the answer voice. If there is user input voice, the speaking right is switched to the user.
[0004] However, in the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0005] The voice collection device in the smart device may mistakenly collect the user's input voice, for example, the user's voice is collected even though the user did not actually speak; or, the user actually spoke but his real intention was not to interrupt the voice playback operation of the smart device, for example, what the user said was not about interruption, or the user was just communicating with other people, etc., which may lead to the erroneous switching of the right to speak, thereby reducing the fluency of the human-computer dialogue and the user experience. Summary of the invention
[0006] The embodiments of the present invention provide a method, device, equipment and storage medium for switching the right to speak, so as to improve the accuracy of switching the right to speak, thereby improving the fluency of human-computer dialogue and enhancing the user experience.
[0007] In a first aspect, an embodiment of the present invention provides a method for switching a speaking right, including:
[0008] If a user input voice is detected during the playback of the target response voice, the playback of the target response voice is paused, and the user's target voice data and target video data are collected;
[0009] Inputting the target voice data and the target video data into a pre-trained target decision network model to make a decision on the switching of the speaking right;
[0010] Based on the output of the target decision network model, a target decision result is determined, and the current discourse right is switched based on the target decision result.
[0011] In a second aspect, an embodiment of the present invention further provides a speaking right switching device, including:
[0012] A data collection module, for pausing the playing of the target answering voice and collecting the user's target voice data and target video data if a user input voice is detected during the playing of the target answering voice;
[0013] A switching decision module, used for inputting the target voice data and the target video data into a pre-trained target decision network model to make a switching decision on the speaking right;
[0014] The speech right switching module is used to determine the target decision result based on the output of the target decision network model, and switch the current speech right based on the target decision result.
[0015] In a third aspect, an embodiment of the present invention further provides an electronic device, the electronic device comprising:
[0016] one or more processors;
[0017] A memory for storing one or more programs;
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for switching the right to speak as provided in any embodiment of the present invention.
[0019] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for switching speech rights as provided in any embodiment of the present invention.
[0020] One embodiment of the above invention has the following advantages or beneficial effects:
[0021] When the user input voice is detected during the playback of the target response voice, the playback of the target response voice is paused, and the user's target voice data and target video data are collected, and the target voice data and target video data are input into the pre-trained target decision network model to make a decision on switching the right to speak, and the target decision result is determined based on the output of the target decision network model. Therefore, by utilizing the target decision network model, based on the user's target voice data and target video data, it is possible to more accurately determine whether the user actually intends to interrupt the voice playback, so that based on the target decision result, the current right to speak can be switched more accurately, thereby improving the accuracy of switching the right to speak, thereby improving the fluency of human-computer dialogue, and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0023] Figure 1 is a flow chart of a method for switching speaking rights provided by an embodiment of the present invention;
[0024] Figure 2 is a structural example diagram of a target decision network model involved in an embodiment of the present invention;
[0025] Figure 3 is an example diagram of a human-computer dialogue involved in an embodiment of the present invention;
[0026] Figure 4 is a flow chart of another method for switching speaking rights provided by an embodiment of the present invention;
[0027] Figure 5 is a structural example diagram of another target decision network model involved in an embodiment of the present invention;
[0028] Figure 6 It is a structural schematic diagram of a speech right switching device provided by an embodiment of the present invention;
[0029] Figure 7 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0031] Figure 1 This is a flow chart of a method for switching the right to speak provided by an embodiment of the present invention. This embodiment is applicable to the case where the right to speak is switched during the process of playing a response voice on a smart device. The method can be executed by a device for switching the right to speak, which can be implemented by software and / or hardware and integrated into an electronic device. Figure 1 As shown, the method specifically comprises the following steps:
[0032] S110: If a user input voice is detected during the playback of the target response voice, the playback of the target response voice is paused, and the user's target voice data and target video data are collected.
[0033] Among them, the target response voice may refer to the response voice currently being broadcast by the smart device. The target response voice may be determined based on the inquiry voice input by the user. Alternatively, during the first interaction, the active greeting voice of the smart device may also be used as the target response voice, so as to detect whether there is a user interruption during the playback of the active greeting voice of the first interaction. The target voice data may refer to the audio data input by the user. The target video data may refer to the data of the collected user image, such as the user's full body image data or the user's facial image data.
[0034] Specifically, during each human-computer interaction, the smart device can collect the current inquiry voice input by the user based on the voice collection device, and after the current inquiry voice collection is completed, the target response voice that matches the collected current inquiry voice can be determined; or, during the first interaction, the active greeting voice of the smart device can also be used as the target response voice, and the target response voice can be played through the playback device. The smart device can respond to the start playback instruction of the target response voice and start real-time detection of whether there is user input voice during the playback of the target response voice, so as to determine whether the user speaks during the playback of the target response voice. If the user input voice is detected, the playback operation of the target response voice can be immediately paused, and the user's target voice data can be collected by the voice collection device and the user's target video data can be collected by the camera.
[0035] Exemplarily, this embodiment can collect the user's target voice data and target video data based on a preset collection method. For example, the current collection time period can be determined based on the preset total collection time, and the user's target audio data and target video data in the current collection time period can be collected. The preset total collection time period can also be divided based on the preset number of decisions to obtain the preset collection time period for each collection, and the current collection time period for the current collection can be determined based on the preset collection time period, and the user's target audio data and target video data in the current collection time period can be collected. For example, the preset total collection time period is 10 seconds, and the preset number of times threshold is 5 times, then the preset collection time period for each collection is 2 seconds, so that the current time can be used as the start collection time, and the preset total collection time period can be used as the time interval to determine the current collection time period, and the user's target audio data and target video data in the current collection time period, that is, within 10 seconds, can be collected, so that a unified switching decision can be made later. Alternatively, you can also use the current moment as the start collection moment, preset the collection duration as the time interval to determine the current collection time period, and collect the user's target audio data and target video data within the current collection time period, that is, within 2 seconds, so that 5 switching decisions can be made later.
[0036] It should be noted that before using the technical solutions disclosed in each embodiment of the present disclosure, the type, scope of use, and usage scenarios of the personal information involved in the present disclosure should be informed to the user in an appropriate manner in accordance with relevant laws and regulations, and the user's authorization should be obtained. For example, the user voice data and user video data collected in this embodiment are collected after informing the user in advance and obtaining the user's authorization.
[0037] S120, inputting the target voice data and the target video data into a pre-trained target decision network model to make a decision on switching the right to speak.
[0038] Among them, the target decision network model can be a classification network model that is pre-trained based on sample data and is used to decide whether to switch the right to speak.
[0039] Specifically, the collected target voice data and target video data can be input into the target decision network model, so that the target decision network model can combine the target voice data and the target video data to comprehensively determine whether the user's current speaking behavior is a behavior that really wants to interrupt the voice broadcast operation of the smart device, thereby improving the accuracy of switching the right to speak.
[0040] For example, Figure 2 A structural example diagram of a target decision network model is given. Figure 2 As shown, the target decision network model may include: a speech feature extraction submodel, a video feature extraction submodel and a decision submodel. Among them, the speech feature extraction submodel can be any network architecture for extracting speech feature information. For example, the speech feature extraction submodel can be but not limited to a Wav2Vec network model. The video feature extraction submodel can be any network architecture for extracting video feature information. For example, the video feature extraction submodel can be but not limited to a Video2Vec network model. The decision submodel can be a network model that performs decision classification based on target speech feature information and target video feature information.
[0041] Exemplarily, S120 may include: inputting the target voice data into the voice feature extraction submodel to perform voice feature extraction, and obtaining the extracted target voice feature information; inputting the target video data into the video feature extraction submodel to perform video feature extraction, and obtaining the extracted target video feature information; inputting the target voice feature information and the target video feature information into the decision submodel to perform a speech right switching decision, and obtain the target decision result.
[0042] For example, Figure 2As shown, the decision sub-model may include: an information splicing layer and a fully connected layer. Accordingly, inputting the target voice feature information and the target video feature information into the decision sub-model to make a decision on the switching of the discourse right and obtain the target decision result may include: inputting the target voice feature information and the target video feature information into the information splicing layer to splice the feature information and obtain the spliced target feature information; inputting the target feature information into the fully connected layer to perform decision classification and obtain the target decision result.
[0043] Specifically, the information concatenation layer can concatenate the input target speech feature information and target video feature information in terms of vector dimension. For example, the feature vector dimension corresponding to the target speech feature information is (1, 1024), and the feature vector dimension corresponding to the target video feature information is (1, 1024). Then, the feature vector dimension corresponding to the concatenated target feature information is 1024 + 1024 = 2048. The fully connected layer can map the target feature information into a two-dimensional vector [x1, x2], where x1 represents the probability of switching the right to speak, that is, the probability that the user really wants to interrupt, and x2 represents the probability of not switching the right to speak, that is, the probability that the user really does not want to interrupt, and the sum of the two is equal to 1. The fully connected layer can output the predicted probability x1 of switching the right to speak, or it can determine whether x1 is greater than a preset threshold. If so, it directly outputs the right to speak switching result, otherwise it outputs the right to speak retention result. By using the information concatenation layer and the fully connected layer, the decision sub-model can more conveniently and accurately determine the target decision result, thereby improving the efficiency and accuracy of the right to speak switching.
[0044] S130: Determine a target decision result based on the output of the target decision network model, and switch the current discourse right based on the target decision result.
[0045] The target decision result may include: a speaking right switching result or a speaking right reservation result. The current speaking right may refer to the end that currently has the speaking right. In this embodiment, the speaking right switching is performed during the voice broadcasting process of the smart device, so that the current speaking right belongs to the smart device, not the user.
[0046] Specifically, the target decision network model can directly output the final target decision result, or it can also output the predicted probability of switching the right to speak. When outputting the predicted probability, it is necessary to detect whether the output predicted probability is greater than a preset threshold (such as 0.5). If so, the target decision result is determined to be the right to speak switching result, otherwise the target decision result is determined to be the right to speak retention result. Based on the obtained target decision result, it can be accurately determined whether the current right to speak needs to be switched, thereby improving the accuracy of the right to speak switch.
[0047] Exemplarily, "switching the current speaking right based on the target decision result" in S130 may include: if the target decision result is a speaking right switching result, switching the current speaking right to the user; if the target decision result is a speaking right retention result, continuing to play the target response voice.
[0048] Specifically, when the target decision result is a right-to-speak switching result, it indicates that the user really wants to interrupt the broadcast operation of the smart device. At this time, the current right-to-speak can be switched to the user, so that the user obtains the current right-to-speak, and the smart device stops playing and waits for the user to express himself. By switching the current right-to-speak in time, the user's resistance can be avoided, thereby improving the user experience. When the target decision result is a right-to-speak reservation result, it indicates that the user does not really want to interrupt the broadcast operation of the smart device, for example, because of a data collection error or the user is just communicating with others. At this time, the target response voice can continue to be played to improve the fluency of the human-computer dialogue.
[0049] Exemplarily, continuing to play the target response voice may include: if the current number of decisions is less than a preset number threshold, then based on the preset collection time, updating the current collection time period, and returning to execute the operation of collecting the user's target voice data and target video data in S110 based on the updated current collection time period; if the current number of decisions is equal to the preset number threshold, then continuing to play the target response voice.
[0050] Among them, the preset number threshold may be a preset total number of allowed switching decisions. The preset collection duration may refer to the time length of user data that needs to be collected for each decision. The preset collection duration may be determined based on the preset total collection duration and the preset number threshold. The preset total collection duration may be a preset total time length allowed for collecting user data. It should be noted that the more user voice data and user videos are collected, the more accurate the decision result output by the target decision network model, but the longer the switching time is, so that the right to speak cannot be switched more timely.
[0051] Specifically, this embodiment can collect user data multiple times based on a shorter preset collection time, and make the current switching decision based on all the target voice data and target video data currently collected, so that the target decision result can be obtained in time as the result of the speech right switching, and then the speech right can be switched more timely, without waiting until the preset total collection time to switch the speech right, so the switching efficiency is improved while ensuring the switching accuracy, and the user experience is improved. If the current number of decisions is equal to the preset number threshold, it means that the target decision result made each time within the preset total collection time is a speech right retention result. At this time, the target response voice can continue to be played to avoid waiting indefinitely, thereby improving the fluency of human-computer dialogue and improving the user experience.
[0052] The technical solution of this embodiment, when detecting the user input voice during the playback of the target response voice, pauses the playback of the target response voice, collects the user's target voice data and target video data, and inputs the target voice data and target video data into a pre-trained target decision network model to make a decision on switching the right to speak, and determines the target decision result based on the output of the target decision network model. By utilizing the target decision network model, based on the user's target voice data and target video data, it is possible to more accurately determine whether the user actually intends to interrupt the voice playback, so that based on the target decision result, the current right to speak can be switched more accurately, thereby improving the accuracy of switching the right to speak, thereby improving the fluency of human-computer dialogue, and enhancing the user experience.
[0053] Based on the above technical solution, Figure 2 For the target decision network model in, the training process of the target decision network model may include: obtaining overlapping interaction sample data and label decision results corresponding to the overlapping interaction sample data; inputting the overlapping interaction sample data into the preset decision network model to make a switching decision on the right to speak, and obtaining the output decision result based on the output of the preset decision network model; determining the training error based on the output decision result and the label decision result, and back-propagating the training error to the preset decision network model, and adjusting the network parameters in the preset decision network model; when the preset convergence conditions are met, determining that the training of the preset decision network model is completed, and obtaining the target decision network model.
[0054] The overlapping interaction sample data may include: sample voice data and sample video data of the user during the overlapping dialogue interaction. Figure 3 An example diagram of a human-computer dialogue is given. This embodiment can extract the original dialogue log and determine, based on the dialogue timestamp between the user and the smart device in the dialogue log, whether there is an overlapping conversation when the user speaks during the process of the smart device playing the answering voice, for example, Figure 3 The conversations in the dotted box are overlapping conversations. By comparing the timestamps in the overlapping conversations, the time period where the intersection occurs is determined, and the user voice data and user video data in the time period are used as sample voice data and sample video data, respectively. For example, the sample voice data and sample video data are Figure 3 The user's voice data and video data within 14s to 16s in the sample voice data can be manually labeled based on the semantic information in the sample voice data and the user's action expression in the sample video data to determine whether the user really wants to interrupt the user, thereby obtaining an accurate label decision result.
[0055] Specifically, the training error can be determined based on the training function according to the output decision results and label decision results of the preset decision network model, and the training error can be back-propagated to the preset decision network model to adjust the network parameters in the preset decision network model until the preset convergence conditions are met. For example, when the number of iterations reaches the preset number or the training error converges, it is determined that the training of the preset decision network model is completed. At this time, the preset decision network model that has completed the training can be used as the target decision network model. By using overlapping interaction sample data and corresponding label decision results for model training, the accuracy of the target decision network model switching decision can be guaranteed, thereby ensuring the accuracy of the speech right switching.
[0056] Figure 4 This is a flow chart of another method for switching the right to speak provided by an embodiment of the present invention. Based on the above embodiments, this embodiment optimizes the step of "inputting the target voice data and the target video data into the pre-trained target decision network model to make a decision on switching the right to speak". The explanations of the terms that are the same or corresponding to the above embodiments are not repeated here.
[0057] See also Figure 4 Another method for switching the right to speak provided in this embodiment specifically includes the following steps:
[0058] S410: If a user input voice is detected during the playback of the target response voice, the playback of the target response voice is paused, and the user's target voice data and target video data are collected.
[0059] S420: Determine the target response text corresponding to the target response voice.
[0060] Specifically, this embodiment can convert the target response speech based on a speech-to-text tool to obtain the converted target response text. This embodiment can also determine the target response text corresponding to the target response speech based on the pre-stored response texts and the mapping relationship between the response text and the response speech.
[0061] S430, inputting the target response text, target voice data and target video data into the pre-trained target decision network model to make a decision on the switching of the speaking right.
[0062] Specifically, the target response text, the collected target voice data and the target video data can be input into the target decision network model, so that the target decision network model can combine the target response text, the target voice data and the target video data at the same time to more accurately determine whether the user's current speaking behavior is a behavior that really wants to interrupt the voice broadcast operation of the smart device, thereby further improving the accuracy of the switching of the right to speak.
[0063] For example, Figure 5 A structural example diagram of a target decision network model is given. Figure 5 As shown, the target decision network model may include: a text feature extraction submodel, a speech feature extraction submodel, a video feature extraction submodel and a decision submodel. Among them, the text feature extraction submodel can be any network architecture for extracting text feature information. For example, the text feature extraction submodel can be but not limited to a BERT (Bidirectional Encoder Representation from Transformers) pre-trained language representation model. The speech feature extraction submodel can be any network architecture for extracting speech feature information. For example, the speech feature extraction submodel can be but not limited to a Wav2Vec network model. The video feature extraction submodel can be any network architecture for extracting video feature information. For example, the video feature extraction submodel can be but not limited to a Video2Vec network model. The decision submodel can be a network model for making decisions and classifications based on target text feature information, target speech feature information and target video feature information.
[0064] Exemplarily, S430 may include: inputting the target response text into a text feature extraction submodel to perform text feature extraction, and obtaining extracted target text feature information; inputting the target voice data into a voice feature extraction submodel to perform voice feature extraction, and obtaining extracted target voice feature information; inputting the target video data into a video feature extraction submodel to perform video feature extraction, and obtaining extracted target video feature information; inputting the target text feature information, target voice feature information, and target video feature information into a decision submodel to perform a speech right switching decision, and obtaining a target decision result.
[0065] For example, Figure 5 As shown, the decision sub-model may include: an information splicing layer and a fully connected layer. Accordingly, inputting the target text feature information, the target voice feature information, and the target video feature information into the decision sub-model to make a decision on the switching of the discourse right and obtain the target decision result may include: inputting the target text feature information, the target voice feature information, and the target video feature information into the information splicing layer to splice the feature information and obtain the spliced target feature information; inputting the target feature information into the fully connected layer to perform decision classification and obtain the target decision result.
[0066] Specifically, the information concatenation layer can concatenate the input target text feature information, target voice feature information, and target video feature information in terms of vector dimension. For example, if the feature vector dimension corresponding to the target text feature information is (1,768), the feature vector dimension corresponding to the target voice feature information is (1,1024), and the feature vector dimension corresponding to the target video feature information is (1,1024), then the feature vector dimension corresponding to the concatenated target feature information is 768+1024+1024=2816. The fully connected layer can map the target feature information into a two-dimensional vector [x1,x2], where x1 represents the probability of switching the right to speak, that is, the probability that the user really wants to interrupt, and x2 represents the probability of not switching the right to speak, that is, the probability that the user really does not want to interrupt, and the sum of the two is equal to 1. The fully connected layer can output the predicted probability x1 of switching the right to speak, or can determine whether x1 is greater than a preset threshold. If so, the right to speak switching result is directly output, otherwise the right to speak retention result is output. By utilizing the information splicing layer and the fully connected layer, the decision-making sub-model can determine the target decision result more conveniently and accurately, thereby improving the efficiency and accuracy of switching the right to speak.
[0067] S440: Determine a target decision result based on the output of the target decision network model, and switch the current discourse right based on the target decision result.
[0068] The technical solution of this embodiment, by inputting the target response text, target voice data and target video data into the target decision network model, enables the target decision network model to make more accurate switching decisions on the right to speak based on the target response text, target voice data and target video data, thereby further improving the accuracy of switching the right to speak.
[0069] Based on the above technical solution, Figure 5 For the target decision network model in, the training process of the target decision network model may include: obtaining overlapping interaction sample data and label decision results corresponding to the overlapping interaction sample data; inputting the overlapping interaction sample data into the preset decision network model to make a switching decision on the right to speak, and obtaining the output decision result based on the output of the preset decision network model; determining the training error based on the output decision result and the label decision result, and back-propagating the training error to the preset decision network model, and adjusting the network parameters in the preset decision network model; when the preset convergence conditions are met, determining that the training of the preset decision network model is completed, and obtaining the target decision network model.
[0070] The overlapping interaction sample data may include: sample response texts during overlapping interactions, as well as sample voice data and sample video data of the user. This embodiment may extract the original conversation logs, and may determine, based on the conversation timestamps between the user and the smart device in the conversation logs, that there is an overlapping conversation when the user speaks during the process of the smart device playing the response voice. The response text corresponding to the response voice in the overlapping conversation is used as the sample response text, for example, Figure 3 The answer text corresponding to the answer voice of the smart device within 12s to 16s. By comparing the timestamps in the overlapping conversations, the time period of the intersection is determined, and the user voice data and user video data in the time period are used as sample voice data and sample video data, respectively. For example, the sample voice data and sample video data are Figure 3 The user's voice data and video data within 14s to 16s in the sample response text and sample voice data and the user's action expression and other information in the sample video data can be manually labeled to determine whether the user really wants to interrupt the user, thereby obtaining an accurate label decision result.
[0071] Specifically, based on the training function, the training error can be determined according to the output decision results and label decision results of the preset decision network model, and the training error can be back-propagated to the preset decision network model, and the network parameters in the preset decision network model can be adjusted until the preset convergence conditions are met. For example, when the number of iterations reaches the preset number or the training error converges, it is determined that the training of the preset decision network model is completed. At this time, the preset decision network model that has completed the training can be used as the target decision network model, so that the accuracy of the switching decision of the target decision network model can be guaranteed through model training, and the accuracy of the switching of the right to speak can be further guaranteed.
[0072] The following is an embodiment of a speech right switching device provided in an embodiment of the present invention. The device and the speech right switching methods of the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the speech right switching device, please refer to the embodiment of the above-mentioned speech right switching method.
[0073] Figure 6 This is a schematic diagram of the structure of a speech right switching device provided by an embodiment of the present invention. This embodiment is applicable to the case where the speech right is switched during the process of playing the answering voice by the smart device. Figure 6 As shown, the device specifically includes: a data collection module 610, a switching decision module 620 and a speech right switching module 630.
[0074] Among them, the data acquisition module 610 is used to pause the playback of the target response voice and collect the user's target voice data and target video data if the user input voice is detected during the playback of the target response voice; the switching decision module 620 is used to input the target voice data and the target video data into the pre-trained target decision network model to make a switching decision on the right to speak; the right to speak switching module 630 is used to determine the target decision result based on the output of the target decision network model, and switch the current right to speak based on the target decision result.
[0075] The technical solution of this embodiment, when detecting the user input voice during the playback of the target response voice, pauses the playback of the target response voice, collects the user's target voice data and target video data, and inputs the target voice data and target video data into a pre-trained target decision network model to make a decision on switching the right to speak, and determines the target decision result based on the output of the target decision network model. By utilizing the target decision network model, based on the user's target voice data and target video data, it is possible to more accurately determine whether the user actually intends to interrupt the voice playback, so that based on the target decision result, the current right to speak can be switched more accurately, thereby improving the accuracy of switching the right to speak, thereby improving the fluency of human-computer dialogue, and enhancing the user experience.
[0076] Optionally, the switching decision module 620 includes:
[0077] A target response text determination unit, used to determine a target response text corresponding to the target response voice;
[0078] The switching decision unit is used to input the target response text, the target voice data and the target video data into a pre-trained target decision network model to make a switching decision on the right to speak.
[0079] Optionally, the target decision network model includes: a text feature extraction sub-model, a speech feature extraction sub-model, a video feature extraction sub-model and a decision sub-model;
[0080] The switching decision unit is specifically used to: input the target response text into the text feature extraction submodel to perform text feature extraction, and obtain the extracted target text feature information; input the target voice data into the voice feature extraction submodel to perform voice feature extraction, and obtain the extracted target voice feature information; input the target video data into the video feature extraction submodel to perform video feature extraction, and obtain the extracted target video feature information; input the target text feature information, the target voice feature information and the target video feature information into the decision submodel to perform a switching decision on the right to speak, and obtain the target decision result.
[0081] Optionally, the decision sub-model includes: an information concatenation layer and a fully connected layer;
[0082] The switching decision unit is also specifically used to: input the target text feature information, the target voice feature information and the target video feature information into the information splicing layer to splice the feature information to obtain the spliced target feature information; input the target feature information into the fully connected layer to perform decision classification to obtain the target decision result.
[0083] Optionally, the device further includes: a target decision network model training module, which is used to:
[0084] Obtain overlapping interaction sample data and label decision results corresponding to the overlapping interaction sample data, wherein the overlapping interaction sample data includes: sample response text during overlapping dialogue interaction and sample voice data and sample video data of the user; input the overlapping interaction sample data into a preset decision network model to make a decision on switching the right to speak, and obtain an output decision result based on the output of the preset decision network model; determine a training error based on the output decision result and the label decision result, and back-propagate the training error to the preset decision network model to adjust the network parameters in the preset decision network model; when a preset convergence condition is met, determine that the training of the preset decision network model is completed, and obtain a target decision network model.
[0085] Optionally, the speaking right switching module 630 includes:
[0086] A speaking right switching unit, configured to switch the current speaking right to the user if the target decision result is a speaking right switching result;
[0087] The speech right reservation unit is used to continue playing the target response voice if the target decision result is a speech right reservation result.
[0088] Optionally, the speech right reservation unit is specifically used to:
[0089] If the current number of decisions is less than the preset number threshold, the current collection time period is updated based on the preset collection duration, and the operation of collecting the user's target voice data and target video data is returned based on the updated current collection time period; if the current number of decisions is equal to the preset number threshold, the target response voice continues to be played.
[0090] The speech right switching device provided in the embodiment of the present invention can execute the speech right switching method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the speech right switching method.
[0091] It is worth noting that in the embodiment of the above-mentioned speech right switching device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0092] Figure 7 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 7 A block diagram of an exemplary electronic device 12 suitable for use in implementing embodiments of the present invention is shown. Figure 7 The electronic device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0093] like Figure 7 As shown, the electronic device 12 is in the form of a general purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).
[0094] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0095] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0096] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 7 not shown, usually called a "hard drive"). Although Figure 7Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The system memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.
[0097] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0098] The electronic device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), may communicate with one or more devices that enable a user to interact with the electronic device 12, and / or may communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the electronic device 12 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the electronic device 12 via a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0099] The processing unit 16 executes various functional applications and data processing by running the program stored in the system memory 28, for example, implementing a method for switching the right to speak provided by the embodiment of the present invention, the method comprising:
[0100] If a user input voice is detected during the playback of the target response voice, the playback of the target response voice is paused, and the user's target voice data and target video data are collected;
[0101] Inputting the target voice data and the target video data into a pre-trained target decision network model to make a decision on the switching of the speaking right;
[0102] Based on the output of the target decision network model, a target decision result is determined, and the current discourse right is switched based on the target decision result.
[0103] Of course, those skilled in the art can understand that the processor can also implement the technical solution of the speech right switching method provided by any embodiment of the present invention.
[0104] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method for switching the right to speak provided in any embodiment of the present invention are implemented. The method includes:
[0105] If a user input voice is detected during the playback of the target response voice, the playback of the target response voice is paused, and the user's target voice data and target video data are collected;
[0106] Inputting the target voice data and the target video data into a pre-trained target decision network model to make a decision on the switching of the speaking right;
[0107] Based on the output of the target decision network model, a target decision result is determined, and the current discourse right is switched based on the target decision result.
[0108] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0109] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0110] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0111] Computer program code for performing the operations of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0112] It should be understood by those skilled in the art that the modules or steps of the present invention described above can be implemented by a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, optionally, they can be implemented by a program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0113] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for switching speaking rights, characterized in that: include: If a user input voice is detected during the playback of the target response voice, the playback of the target response voice is paused, and the user's target voice data and target video data are collected; Determine a target response text corresponding to the target response voice; Inputting the target response text, the target voice data and the target video data into a pre-trained target decision network model to make a decision on the switching of the speaking right; Determine a target decision result based on an output of the target decision network model, and switch the current speaking right based on the target decision result; The switching of the current speaking right based on the target decision result includes: If the target decision result is a speaking right switching result, the current speaking right is switched to the user to stop playing the target answering voice and wait for the user to speak.
2. The method according to claim 1, characterized in that The target decision network model includes: a text feature extraction sub-model, a speech feature extraction sub-model, a video feature extraction sub-model and a decision sub-model; The step of inputting the target response text, the target voice data and the target video data into a pre-trained target decision network model to make a decision on the switching of the speaking right includes: Inputting the target response text into the text feature extraction sub-model to extract text features, and obtaining extracted target text feature information; Inputting the target speech data into the speech feature extraction submodel to extract speech features, and obtaining extracted target speech feature information; Inputting the target video data into the video feature extraction sub-model to extract video features, and obtaining extracted target video feature information; The target text feature information, the target voice feature information and the target video feature information are input into the decision sub-model to make a decision on the switching of the speaking right, and obtain a target decision result.
3. The method according to claim 2, characterized in that The decision sub-model includes: an information splicing layer and a fully connected layer; The step of inputting the target text feature information, the target voice feature information and the target video feature information into the decision sub-model to make a decision on the switching of the speaking right and obtain a target decision result includes: Inputting the target text feature information, the target voice feature information and the target video feature information into the information splicing layer to splice the feature information, and obtaining the spliced target feature information; The target feature information is input into the fully connected layer for decision classification to obtain a target decision result.
4. The method according to claim 1, characterized in that The training process of the target decision network model includes: Obtaining overlapping interaction sample data and label decision results corresponding to the overlapping interaction sample data, wherein the overlapping interaction sample data includes: sample response text during the overlapping interaction of the dialogue and sample voice data and sample video data of the user; Inputting the overlapping interaction sample data into a preset decision network model to make a decision on the switching of the speaking right, and obtaining an output decision result based on the output of the preset decision network model; Determine a training error based on the output decision result and the label decision result, and back-propagate the training error to the preset decision network model to adjust network parameters in the preset decision network model; When the preset convergence condition is met, it is determined that the training of the preset decision network model is completed, and the target decision network model is obtained.
5. The method according to any one of claims 1 to 4, characterized in that: The switching of the current speaking right based on the target decision result includes: If the target decision result is a speech right reservation result, the target response voice continues to be played.
6. The method according to claim 5, characterized in that The continuing to play the target answer voice includes: If the current number of decisions is less than a preset number threshold, the current collection time period is updated based on the preset collection duration, and the operation of collecting the user's target voice data and target video data is returned based on the updated current collection time period; If the current number of decisions is equal to the preset number threshold, the target response voice continues to be played.
7. A speech right switching device, characterized in that: include: A data collection module, for pausing the playing of the target answering voice and collecting the user's target voice data and target video data if a user input voice is detected during the playing of the target answering voice; The switching decision module includes: a target response text determination unit, which is used to determine the target response text corresponding to the target response voice; A switching decision unit, used for inputting the target response text, the target voice data and the target video data into a pre-trained target decision network model to make a switching decision on the speaking right; A speech right switching module, used to determine a target decision result based on the output of the target decision network model, and switch the current speech right based on the target decision result; The speaking right switching module includes: a speaking right switching unit, which is used to switch the current speaking right to the user if the target decision result is a speaking right switching result, so as to stop playing the target answering voice and wait for the user to express himself.
8. An electronic device, characterized in that: The electronic device comprises: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the speech right switching method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for switching the right to speak as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Multimodal interaction method and system for storytelling robot
CN109359177A