An intelligent voice interaction method, device, electronic device and storage medium

By fine-grained segmentation and tone acceptance processing of user voice streams, the problem of long response time of existing intelligent voice interaction solutions is solved, and faster response and better user experience is achieved.

CN115132192BActive Publication Date: 2025-05-30ALIBABA INNOVATION PRIVATE LIMITED
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110309632.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-23
Publication Date
2025-05-30
Estimated Expiration
2041-03-23

AI Technical Summary

Technical Problem

The existing intelligent voice interaction solution has a long response time, which leads to an increase in the time for users to wait for reply, affecting the user experience.

Method used

Early response is achieved by dividing the user's voice stream into multiple voice segments in a finer granularity, and playing the tone to inherit the voice file before performing the overall response.

Benefits of technology

Reduces the time for users to wait for reply and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115132192B_ABST
    Figure CN115132192B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an intelligent voice interaction method, device, electronic device and storage medium. According to the scheme provided in the embodiments of the present application, the voice stream sent by the user is obtained, and the voice stream is divided into multiple first voice segments according to the first silent interval duration, and the voice stream is divided into at least one second voice segment according to the second silent interval duration, wherein the first silent interval duration is less than the second silent interval duration; the tone continuation voice file corresponding to the first voice segment is obtained; before responding to the second voice segment, the tone continuation voice file corresponding to the first voice segment is played. The scheme in the present application divides the user's voice stream into finer granularity, and responds in advance by tone continuation before making an overall response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to an intelligent voice interaction method, device, electronic device and storage medium. Background Art

[0002] Using a conversational robot to interact with users through voice is a common method in the industry. This type of interaction is usually a one-time long connection process, and users are more sensitive to the robot's response delay. However, the current common intelligent voice interaction solutions in the industry often have a response time of more than 2 seconds due to the long link, and users often cannot get an immediate response.

[0003] Based on this, a faster voice interaction solution is needed to improve the user experience. Summary of the invention

[0004] In view of this, an embodiment of the present application provides an intelligent voice interaction solution to at least partially solve the above-mentioned problems.

[0005] According to a first aspect of an embodiment of the present application, an intelligent voice interaction method is provided, the method comprising:

[0006] Acquire a voice stream sent by a user, and divide the voice stream into a plurality of first voice segments according to a first silence interval duration;

[0007] Dividing the voice stream into at least one second voice segment according to a second silence interval duration, wherein the first silence interval duration is shorter than the second silence interval duration;

[0008] Obtaining a voice file corresponding to the first voice segment; and

[0009] Before responding to the second voice segment, play the tone-continuing voice file corresponding to the first voice segment.

[0010] According to a second aspect of an embodiment of the present application, an intelligent voice interaction device is provided, the device comprising:

[0011] A voice acquisition module acquires the voice stream sent by the user;

[0012] A first segmentation module, which obtains a voice stream sent by a user and segments the voice stream into a plurality of first voice segments according to a first silence interval duration;

[0013] A second segmentation module is configured to segment the voice stream into at least one second voice segment according to a second silence interval duration, wherein the first silence interval duration is shorter than the second silence interval duration;

[0014] A file acquisition module that acquires the tone-connecting voice file corresponding to the first voice segment;

[0015] An interaction module that plays the tone-connecting voice file corresponding to the first voice segment before responding to the second voice segment.

[0016] According to the third aspect of the embodiments of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the intelligent voice interaction method described in the first aspect.

[0017] According to the fourth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the intelligent voice interaction method described in the first aspect.

[0018] According to the solution provided by the embodiments of the present application, a voice stream sent by a user is acquired, and the voice stream is segmented into multiple first voice segments according to a first silence interval duration, and at least one second voice segment is segmented according to a second silence interval duration, where the first silence interval duration is less than the second silence interval duration; the tone-connecting voice file corresponding to the first voice segment is acquired; before responding to the second voice segment, the tone-connecting voice file corresponding to the first voice segment is played. By segmenting the user's voice stream with a finer granularity and making an early response in a tone-connecting manner before making an overall response, the waiting time of the user for a reply is reduced, and the user experience is improved. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0020] Figure 1 It is a schematic diagram of the process involved in a current voice conversation;

[0021] Figure 2 It is a schematic diagram of the process of an intelligent voice interaction method provided by the embodiments of the present application;

[0022] Figure 3 It is a schematic diagram of a method for dividing voice segments provided by the embodiments of the present application;

[0023] Figure 4 A schematic flowchart of a tone connection link involved in an embodiment of the present application;

[0024] Figure 5 A schematic logic diagram of a multimodal detection algorithm provided by an embodiment of the present application;

[0025] Figure 6 A schematic structural diagram of an intelligent voice interaction device provided by an embodiment of the present application;

[0026] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0027] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.

[0028] The following further illustrates the specific implementation of the embodiments of the present application with reference to the accompanying drawings of the embodiments of the present application.

[0029] Currently, voice dialogue robots are all based on text dialogue robots, and text and voice modality conversion is achieved by adding Automatic Speech Recognition (ASR) and Text-to-speech (TTS) to achieve the ability of voice dialogue, as Figure 1 shown, Figure 1 is a schematic flowchart of the process involved in current voice dialogue. In this way, the interaction link is long, and the response time is often relatively long, usually reaching more than 2S. Based on this, the present application provides a faster voice interaction solution to improve the user experience.

[0030] As Figure 2 shown, Figure 2 is a schematic flowchart of an intelligent voice interaction method provided by an embodiment of the present application, including:

[0031] S201, obtaining the voice stream sent by the user.

[0032] S203, dividing the voice stream into multiple first voice segments according to the first silent interval duration.

[0033] When a user interacts with an intelligent robot, the speech stream usually naturally carries silent intervals of different durations due to pauses between utterances. Such pauses usually indicate semantic differences in phrases or sentences expressed by the user's language. In traditional solutions, the speech stream is segmented based on such pauses to achieve overall recognition and response to the problems the user wants to solve.

[0034] The silent interval means that the volume of the sound at any time point within the duration of the interval of the speech stream does not exceed a certain value, and it does not need to be completely silent.

[0035] In this application, the specific way of dividing the speech stream based on the silent interval duration can be that if the silent interval duration between two time points exceeds a preset value, it is divided into different speech segments; conversely, if the silent interval duration between two time points does not exceed the preset value, it is divided into the same speech segment.

[0036] For example, the user's speech stream is "Excuse me, how to update Taobao app", and its corresponding speech stream is as Figure 3 shown, Figure 3 which is a schematic diagram of a way to divide speech segments provided by an embodiment of this application. The silent interval duration between the characters "oh" and "tao" is 300ms, the interval duration between "oh" and "tao" is "500ms", and the silent interval duration between "app" and "zen" is 250ms.

[0037] Based on this, the first silent interval duration can be preset to 200ms, so that the speech stream can be divided into multiple first speech segments such as "Excuse me", "xia oh", "Taobao app", and "how to update" according to the first silent interval duration.

[0038] S205, divide the speech stream into at least one second speech segment according to the second silent interval duration, where the first silent interval duration is less than the second silent interval duration.

[0039] Continuing the previous example, the speech stream can be divided into a second speech segment "Excuse me, how to update Taobao app" based on the second silent interval duration. As Figure 3 shown in Figure 3 the dotted line intervals in are the multiple first speech segments obtained by dividing the speech stream, and the silent segments corresponding to the first silent interval duration are between the multiple first speech segments; the solid line interval is a second speech segment obtained by using another division method (i.e., based on the second silent interval).

[0040] In practical applications, the first silent interval duration and the second silent interval duration can be preset based on empirical statistics according to actual needs. For example, usually, the second silent interval duration is 800 ms, while the first silent interval duration is 400 ms or a smaller 200 ms.

[0041] In the embodiments of the present application, the second silent interval duration is usually used for the division of a whole sentence. The obtained second speech segment is usually a complete sentence to achieve a complete understanding of the user and give a corresponding complete answer. That is, in some scenarios, if the user only says one sentence, there may be only one obtained second speech segment. If the user says multiple sentences, there may be multiple corresponding second speech segments.

[0042] The first silent interval duration is used to perform a finer-grained division of the speech stream. The obtained first speech segment is usually a semantic unit (usually corresponding to a phrase or an incomplete sentence) so as to make a response based on a finer semantic unit. Obviously, since the first silent interval duration is less than the second silent interval duration, the duration of the obtained second speech segment is longer than that of the first speech segment.

[0043] S207, obtain the tone connection speech file corresponding to the first speech segment.

[0044] Specifically, the semantics of the first speech segment can be determined first, and then the segment type of the first speech segment can be determined according to the semantics, where the segment type includes an in-sentence segment or an end-of-sentence segment; then, the tone connection speech file corresponding to the segment type can be obtained.

[0045] Since the tone connection speech file is not a solution to the user's problem (in fact, the speech file corresponding to the second speech segment is usually a solution to the user's problem), the tone connection speech file does not need to be too long and can be implemented using a short file.

[0046] That is, the tone connection speech file can be a speech file not exceeding a preset duration (for example, the playback duration does not exceed 2 s), or the number of characters included in the tone connection speech file does not exceed a preset number (for example, does not exceed 5 characters), etc.

[0047] When the type of the first speech segment is an in-sentence segment, some short whispered words can be used to respond to the user. For example, "um" or "oh" etc. can be used to indicate listening to the user's expression and drive the user to continue expressing. Also, a special tone can be included in the tone connection speech file corresponding to the in-sentence segment, and interact with the user in a statement manner.

[0048] When the first voice segment is a sentence-ending segment, the modal particles such as "thinking" or "responding" can be used to make the user feel that the expression is understood and respected. In order to better interact with the user in the sentence-ending segment, it is necessary to correctly understand the user's intention corresponding to a sentence-ending segment. The user's intention is classified, so that the tone-continuing voice file corresponding to the user's user intention type can be obtained.

[0049] Specifically, after dividing the speech stream and obtaining a plurality of first speech segments, wherein the plurality of first speech segments obtained by division may include a plurality of sentence-ending segments, the speech stream between a sentence-ending segment and a previous sentence-ending segment (i.e., another sentence-ending segment that is earlier in the event sequence and has the shortest time interval with the sentence-ending segment) can be determined as the complete sentence corresponding to the sentence-ending segment. If a sentence-ending segment does not have a previous sentence-ending segment, the speech stream between the initial time and the sentence-ending segment can be determined as the complete sentence corresponding to the sentence-ending segment.

[0050] Furthermore, the semantics of the complete sentence can be determined, and the type of user intent represented by the complete sentence can be determined based on the semantics of the complete sentence, and the tone continuation voice file corresponding to the user intent type can be obtained. The correspondence between the user intent type and the tone continuation voice file can be preset, and the user intent type can include such as greeting, confirmation, denial or issuing instructions, as well as other types. As shown in Table 1, Table 1 is a schematic table of user intent types and continuation words contained in the tone continuation voice file provided in an embodiment of the present application.

[0051]

[0052]

[0053] In addition, it should be noted that when obtaining the tone-continuation voice file, some potentially risky continuation words can be avoided based on the user intent type. For example, if the user intent type is "instruction", then the use of risky words with clear meanings such as "good" and "OK" can be avoided to interact, so as to avoid the risk of verbally promising to solve the user's problem but not actually being able to solve it.

[0054] Furthermore, in one embodiment, the acquired tone-continuation voice file may be an audio file directly acquired; or the text corresponding to the first voice segment may be first acquired, and then the tone-continuation audio corresponding to the text may be synthesized, and the tone-continuation audio may be determined as the corresponding tone-continuation voice file.

[0055] In one embodiment, the tone type in the tone connection voice file corresponding to the obtained first voice segment can also correspond to the user intention type, thereby further improving the user experience.

[0056] Specifically, it can be that the tone connection voice file of the same conjunction itself has multiple different types of tones; or, it can also be that the tone type corresponding to the user intention type is adopted when synthesizing the voice file based on the conjunction, so as to obtain a tone connection voice file containing a certain tone corresponding to the user intention type.

[0057] For example, for user intention types such as "greeting", "thanking", "confirming", etc., a positive tone (for example, a happy tone) can be used to express to the user. For the tone connection types of "denying", "instructing", and "others", a flat or declarative tone can be used for reply.

[0058] S209, play the tone connection voice file corresponding to the first voice segment before replying to the second voice segment.

[0059] As described above, the division of the first voice segment is for tone connection and does not involve the actual user question. The question involved by the specific user is the answer voice segment obtained by processing the second voice segment in the conventional manner. The specific processing process can be referred to as Figure 1 the processing flow shown.

[0060] In other words, the reply to the first voice segment (i.e., the tone connection link) and the reply to the second voice segment (i.e., the traditional voice Q&A link) are actually two different and independent links that do not affect each other. As Figure 4 shown, Figure 4 is a schematic flow chart of a tone connection link involved in an embodiment of the present application. In this schematic diagram, the tone connection link does not depend on the reply link of the traditional voice robot, can run in parallel with the traditional link, and all the instant replies performed by the tone connection link occur before the traditional link (i.e., replying to the second voice segment) answers the user, and will not affect the traditional reply link.

[0061] According to the solution provided by the embodiments of the present application, a voice stream sent by a user is obtained, and the voice stream is segmented into multiple first voice segments according to a first silent interval duration, and at least one second voice segment is segmented according to a second silent interval duration, where the first silent interval duration is less than the second silent interval duration; a tone connection voice file corresponding to the first voice segment is obtained; and the tone connection voice file corresponding to the first voice segment is played before answering the second voice segment. By segmenting the user's voice stream with a finer granularity and making an early response in a tone connection manner before making an overall answer, the waiting time of the user for a reply is reduced, and the user experience is improved.

[0062] In one embodiment, since there are usually multiple sentence segments obtained by division, if a reply is made for each sentence segment, it is easy to cause user fatigue. Therefore, it is necessary to correspondingly control the reply frequency of the first voice segment of the sentence segment type to further improve the user experience.

[0063] Specifically, a corresponding decision algorithm can be used to determine whether the first voice segment of the sentence segment type meets a preset condition, and a sentence connection is made only when the condition is met to realize a reply to the user, thereby controlling the frequency of the sentence connection.

[0064] For example, the historical reply frequency or historical reply ratio for answering the first voice segment in the voice stream is determined. When the historical reply frequency or historical reply ratio does not exceed a preset value, it is determined that the voice segment meets the preset condition.

[0065] Suppose a total of 30 first voice segments are obtained by dividing a voice stream, and 3 replies have been made before. Then it can be known that the historical reply frequency is 3, or the historical reply ratio is 10%. Therefore, if the preset value of the historical reply frequency is set to 5 or the preset value of the historical reply ratio is set to 0.2, a reply can be made to the sentence segment.

[0066] For another example, a random number corresponding to the first voice segment can also be given. When the random number does not exceed a preset range, it is determined that the first voice segment meets the preset condition. By actually controlling the range of the random number, the probability that a first voice segment may be replied is also controlled. For example, assume that the random number generator can give all random numbers from 1 to 100 with equal probability. Then the preset range of the random numbers that can be replied can be set to [70, 100], thereby actually controlling the probability that the first voice segment may be replied to be about 30%, avoiding a too high frequency of in-sentence replies.

[0067] As described above, in the embodiments of the present application, different response methods need to be adopted for in-sentence fragments and end-of-sentence fragments. Therefore, in one embodiment, an end-of-sentence detection algorithm can also be used to specifically detect in-sentence fragments and end-of-sentence fragments. Specifically, the semantic features of the text corresponding to the first speech fragment can be determined, and the audio features of the first speech fragment can be extracted; the first speech fragment can be classified according to the semantic features and the audio features, and the fragment type of the first speech fragment can be determined according to the classification result.

[0068] For example, the semantic features can be classified by a pre-trained text semantic classifier to obtain a semantic classification result; and the audio features can be classified by a pre-trained audio classifier to obtain an audio feature classification result. For example, a pre-trained Bert text pre-training structure can be used to classify the text of the speech fragment, and a bidirectional long short-term memory network (Long Short-Term Memory, LSTM) can be used to obtain the audio representation of the Micro-turn. And finally, the Attention mechanism in deep learning is used to fuse the semantic features and the audio features to perform binary classification on the first speech fragment.

[0069] As Figure 5 shown, Figure 5 FIG. 10 is a logical schematic diagram of a multi-modal detection algorithm provided by an embodiment of the present application for detecting the fragment type of the first speech fragment. In this algorithm, a multi-task training form is adopted, that is, classifiers of different modalities (that is, including a classifier on semantic features and a classifier on audio features) are trained simultaneously, and then their prediction results are fused to obtain a multi-modal prediction result.

[0070] Further, in the input audio data, in order to avoid excessive data, a complete first speech fragment or the last few seconds (such as 5 s or 3 s) of the first speech fragment can be adopted, and at the same time, the corresponding silent fragment (that is, the silent fragment corresponding to the last and adjacent first silent interval duration in the time series) is merged to obtain a speech sample to be classified as the input.

[0071] The silent segment is added to capture audio information ignored by ASR such as breathing sounds, thereby improving the classification effect. Next, feature extraction is performed on the original audio: the speech sample to be recognized is equally segmented to obtain a plurality of equally long speech sub-samples. For example, the speech sample to be recognized is framed with a window of 40 ms in length to obtain a speech sub-sample with a frame length of 40 ms, and then audio features are extracted from each frame of the speech sub-sample.

[0072] Specifically, we can select MFCC

[10] , logfbank, F0, Intensity and other features widely used in speech recognition tasks, as well as their delta and delta-delta values ​​to represent the dynamic change information of the features, and we can also use audio processing libraries such as librosa and praat to extract features. Then, we input the extracted audio features into the bidirectional LSTM to obtain the audio representation of the first speech segment, and then integrate the text prediction model to obtain the prediction results through the fully connected layer and Softmax.

[0073] Through the text pre-training structure and multimodal feature fusion algorithm, the sentence-end segment level expressions in the voice stream are identified to determine whether the user has completed the expression. The multiple first voice segments obtained by division can be more accurately classified into sentence-end segments and sentence-mid segments, thereby achieving more accurate subsequent responses and further improving the user experience.

[0074] The intelligent voice interaction method of this embodiment can be executed by any appropriate electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PCs, etc.

[0075] In a second aspect of the embodiments of the present application, an intelligent voice interaction device is also provided. Figure 6 As shown, Figure 6 A schematic diagram of the structure of an intelligent voice interaction device provided in an embodiment of the present application includes:

[0076] The voice acquisition module 609 acquires the voice stream sent by the user;

[0077] The first segmentation module 601 obtains a voice stream sent by a user, and segments the voice stream into a plurality of first voice segments according to a first silence interval duration;

[0078] A second segmentation module 603 is configured to segment the voice stream into at least one second voice segment according to a second silence interval duration, wherein the first silence interval duration is shorter than the second silence interval duration;

[0079] An acquisition module 605 is used to acquire a tone-continuing voice file corresponding to the first voice segment;

[0080] The file interaction module 607 plays the tone-continuing voice file corresponding to the first voice segment before responding to the second voice segment.

[0081] Optionally, the acquisition module 605 determines the semantics of the first voice segment, determines the segment type of the first voice segment according to the semantics, wherein the segment type includes a mid-sentence segment or a sentence-end segment; and obtains a tone-continuing voice file corresponding to the segment type.

[0082] Optionally, the obtaining module 605 obtains a tone-connecting speech file corresponding to the segment type and not exceeding a preset duration, or obtains a tone-connecting speech file corresponding to the segment type and having a character count not exceeding a preset number.

[0083] Optionally, when the segment type is an end-of-sentence segment, the obtaining module 605 determines the complete sentence corresponding to the end-of-sentence segment from the speech stream; determines the user intention type represented by the complete sentence; and obtains a tone-connecting speech file corresponding to the user intention type.

[0084] Optionally, the obtaining module 605 determines the previous end-of-sentence segment in the speech stream, and determines the speech stream between the previous end-of-sentence segment and the end-of-sentence segment as the complete sentence corresponding to the end-of-sentence segment.

[0085] Optionally, the obtaining module 605 determines the semantics of the complete sentence, and determines the user intention type represented by the complete sentence according to the semantics of the complete sentence.

[0086] Optionally, in the device, the user intention type includes at least one of greeting, confirmation, denial, or instruction.

[0087] Optionally, the obtaining module 605 determines the semantic features of the text corresponding to the first speech segment, and extracts the audio features of the first speech segment; classifies the first speech segment according to the semantic features and the audio features, and determines the segment type of the first speech segment according to the classification result.

[0088] Optionally, the obtaining module 605 classifies the semantic features according to a pre-trained text semantic classifier to obtain a semantic classification result; and classifies the audio features according to a pre-trained audio classifier to obtain an audio feature classification result; fuses the semantic classification result and the audio feature classification result to determine the segment type of the first speech segment.

[0089] Optionally, the obtaining module 605 combines the first speech segment and the corresponding silent segment to obtain a speech sample to be classified; equally divides the speech sample to obtain a plurality of equally long speech sub-samples, and extracts the audio features of the first speech segment from the divided speech sub-samples.

[0090] Optionally, when the segment type is an in - sentence segment, the obtaining module 605 determines whether the first voice segment meets a preset condition. When the first voice segment meets the preset condition, a tone - connecting voice file corresponding to the segment type is obtained.

[0091] Optionally, the obtaining module 605 determines the historical response frequency or historical response ratio for responding to the first voice segment in the voice stream. When the historical response frequency or historical response ratio does not exceed a preset value, it is determined that the voice segment meets the preset condition; or, the random number corresponding to the first voice segment is determined. When the random number does not exceed a preset range, it is determined that the first voice segment meets the preset condition.

[0092] Optionally, the obtaining module 605 obtains the text corresponding to the first voice segment, synthesizes a tone - connecting audio corresponding to the text, and determines the tone - connecting audio as the corresponding tone - connecting voice file.

[0093] The intelligent voice interaction device in this embodiment is used to implement the corresponding intelligent voice interaction processing method in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here. In addition, the function implementation of each module in the intelligent voice interaction device in this embodiment can refer to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated here either.

[0094] Refer to Figure 7 , which shows a schematic structural diagram of an electronic device according to an embodiment of the present application. The specific implementation of the electronic device is not limited in the specific embodiments of the present application.

[0095] As Figure 7 shown, the electronic device may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.

[0096] Wherein:

[0097] The processor 502, the communication interface 504, and the memory 506 communicate with each other through the communication bus 508.

[0098] The communication interface 504 is used to communicate with other electronic devices or servers.

[0099] The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the foregoing intelligent voice interaction method embodiments.

[0100] Specifically, the program 510 may include program code, and the program code includes computer operation instructions.

[0101] The processor 502 may be a central processing unit (CPU), or a specific application integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type, such as one or more CPUs; or may be of different types, such as one or more CPUs and one or more ASICs.

[0102] A memory 506 is used to store a program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0103] The program 510 is specifically configured to cause the processor 502 to perform the following operations:

[0104] Obtain a voice stream sent by a user;

[0105] Segment the voice stream into a plurality of first voice segments according to a first silence interval duration; and,

[0106] Segment the voice stream into at least one second voice segment according to a second silence interval duration, where the first silence interval duration is less than the second silence interval duration;

[0107] Obtain a tone-connecting voice file corresponding to the first voice segment; and

[0108] Before answering the second voice segment, play the tone-connecting voice file corresponding to the first voice segment.

[0109] For the specific implementation of each step in the program 510, reference may be made to the corresponding steps and units in the above-mentioned embodiments of the intelligent voice interaction method, which will not be elaborated herein. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated herein.

[0110] Through the electronic device of this embodiment, obtain the voice stream sent by the user, divide the voice stream into multiple first voice segments according to the first silent interval duration, and divide the voice stream into at least one second voice segment according to the second silent interval duration, where the first silent interval duration is less than the second silent interval duration; obtain the tone connection voice file corresponding to the first voice segment; play the tone connection voice file corresponding to the first voice segment before answering the second voice segment. By dividing the user's voice stream into finer granularity and making an early response in a tone connection manner before making an overall response, the waiting time of the user for a reply is reduced, and the user experience is improved.

[0111] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0112] The method according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or an FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the intelligent voice interaction method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the intelligent voice interaction method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the intelligent voice interaction method shown herein.

[0113] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0114] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application shall be defined by the claims.

Claims

1. An intelligent voice interaction method, include: Get the voice stream sent by the user; Dividing the voice stream into a plurality of first voice segments according to a first silence interval duration; Dividing the voice stream into at least one second voice segment according to a second silence interval duration, wherein the first silence interval duration is shorter than the second silence interval duration; Obtaining a voice file corresponding to the first voice segment; and Before responding to the second voice segment, playing the tone-continuing voice file corresponding to the first voice segment; Among them, obtaining the tone continuation voice file corresponding to the first voice segment includes: determining the semantics of the first voice segment, determining the segment type of the first voice segment according to the semantics, wherein the segment type includes a mid-sentence segment or a sentence-end segment; obtaining the tone continuation voice file corresponding to the segment type.

2. The method according to claim 1, in, Obtaining a voice file corresponding to the segment type, including: Acquire a tone-continuing voice file corresponding to the segment type and having a duration not exceeding a preset time, or acquire a tone-continuing voice file corresponding to the segment type and having a number of characters not exceeding a preset number.

3. The method according to claim 1, in, When the segment type is a sentence-end segment, obtaining a tone-continuing voice file corresponding to the segment type includes: Determining a complete sentence corresponding to the sentence end segment from the speech stream; Determining the type of user intent represented by the complete sentence; Obtain a voice file corresponding to the user's intention type.

4. The method according to claim 3, in, Determining a complete sentence corresponding to the sentence end segment from the speech stream includes: A previous sentence-end segment in the speech stream is determined, and the speech stream between the previous sentence-end segment and the sentence-end segment is determined as a complete sentence corresponding to the sentence-end segment.

5. The method according to claim 3, in, Determine the type of user intent represented by the complete sentence, including: The semantics of the complete sentence is determined, and the type of user intent represented by the complete sentence is determined based on the semantics of the complete sentence.

6. The method according to claim 5, in, The user intention type includes at least one of greeting, confirming, denying or issuing an instruction.

7. The method according to claim 1, in, Determining the segment type of the first speech segment according to the semantics includes: Determining semantic features of the text corresponding to the first speech segment, and extracting audio features of the first speech segment; The first speech segment is classified according to the semantic feature and the audio feature, and the segment type of the first speech segment is determined according to the classification result.

8. The method according to claim 7, classifying the first speech segment according to the semantic feature and the audio feature, and determining the segment type of the first speech segment according to the classification result, include: Classifying the semantic features according to a pre-trained text semantic classifier to obtain a semantic classification result; Further, classify the audio features according to a pre-trained audio classifier to obtain an audio feature classification result; Fuse the semantic classification result and the audio feature classification result to determine the segment type of the first speech segment.

9. The method according to claim 7, wherein, extracting the audio features of the first speech segment includes: merging the first speech segment and its corresponding silent segment to obtain a speech sample to be classified; equally segment the speech sample to obtain a plurality of equally long speech sub-samples, and extract the audio features of the first speech segment from the segmented speech sub-samples.

10. The method according to claim 1, wherein, when the segment type is an in-sentence segment, obtain a tone-connecting speech file corresponding to the segment type, and the method includes: determine whether the first speech segment meets a preset condition, and when the first speech segment meets the preset condition, obtain a tone-connecting speech file corresponding to the segment type.

11. The method according to claim 10, determining whether the first speech segment meets a preset condition, includes: determine the historical response frequency or historical response ratio of the response to the first speech segment in the speech stream, and when the historical response frequency or historical response ratio does not exceed a preset value, determine that the speech segment meets the preset condition; or, determine the random number corresponding to the first speech segment, and when the random number does not exceed a preset range, determine that the first speech segment meets the preset condition.

12. The method according to claim 1, obtaining the tone-connecting speech file corresponding to the first speech segment, includes: obtain the text corresponding to the first speech segment, synthesize a tone-connecting audio corresponding to the text, and determine the tone-connecting audio as the corresponding tone-connecting speech file.

13. An intelligent voice interaction device, including: a voice acquisition module, which acquires a speech stream sent by a user; a first segmentation module, which acquires the speech stream sent by the user and segments the speech stream into a plurality of first speech segments according to a first silent interval duration; a second segmentation module, which segments the speech stream into at least one second speech segment according to a second silent interval duration, wherein the first silent interval duration is less than the second silent interval duration; a file acquisition module, which acquires the tone-connecting speech file corresponding to the first speech segment, including: determining the semantics of the first speech segment, determining the segment type of the first speech segment according to the semantics, wherein the segment type includes an in-sentence segment or an end-of-sentence segment; obtaining the tone-connecting speech file corresponding to the segment type; an interaction module, which plays the tone-connecting speech file corresponding to the first speech segment before responding to the second speech segment.

14. An electronic device, including: a processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the intelligent voice interaction method described in any one of claims 1-12.

15. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the intelligent voice interaction method described in any one of claims 1-12.

Citation Information

Patent Citations

  • Voice recognition terminal and voice recognition method using computer terminal

    JP2014191030A