Method and apparatus for detecting speech, device, medium, and program product

By combining audio and text information and utilizing speech activity detection and machine learning models to dynamically adjust the silence duration, the problem of delay or premature truncation in speech detection is solved, enabling more accurate and faster determination of the end position of speech and improving the user experience.

WO2026012246A1PCT designated stage Publication Date: 2026-01-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/106409
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-11
Filing Date
2025-07-01
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing speech detection technologies are prone to delays or premature truncation when determining the end of speech, which affects the user experience.

Method used

By combining audio and text information, and utilizing speech activity detection, speech recognition models, and machine learning models, the duration of silence is dynamically adjusted, and the end position of speech is determined based on user characteristics and semantic features.

Benefits of technology

It improves the accuracy and speed of speech detection, reduces waiting time, avoids semantic incompleteness and premature truncation, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025106409_15012026_PF_FP_ABST
    Figure CN2025106409_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method and apparatus for detecting speech, a device, a medium, and a program product. The method comprises acquiring an audio for a user's query, wherein the audio uses speech corresponding to the query as a start part. The method further comprises: determining a text for the query on the basis of the speech. The method further comprises on the basis of the speech and the text, determining an end position of the speech corresponding to the query in the audio.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices, media, and procedures for detecting speech.

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202410931166.9, filed on July 11, 2024, entitled “Method, Apparatus, Device, Medium and Procedure for Detecting Speech”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The embodiments of this disclosure generally relate to the field of audio processing, and more specifically to methods, apparatus, devices, media, and program products for detecting speech. Background Technology

[0004] Machine learning is becoming increasingly important in people's daily lives and is gradually becoming an indispensable tool. More and more jobs are starting to use machine learning models. For example, word processing, image processing, and audio processing are all beginning to utilize machine learning models. The development of machine learning technology has greatly improved the efficiency of multimodal data processing.

[0005] With the rapid development of machine learning technology, the processing of various multimodal data has become faster and more accurate. For example, in audio processing, pre-deployed audio processing models can assist users in tasks such as automatic pitch adjustment and audio segment extraction. Furthermore, to meet the evolving needs of speech recognition technology, machine learning models are increasingly being used in speech detection. Summary of the Invention

[0006] Embodiments of this disclosure provide a method, apparatus, device, medium, and program product for detecting speech.

[0007] According to a first aspect of this disclosure, a method for detecting speech is provided. The method includes acquiring audio of a user's query, the audio beginning with the speech of the query. The method also includes determining text corresponding to the query based on the speech. Furthermore, the method includes determining the end position of the speech in the audio corresponding to the query based on the audio and the text.

[0008] In a second aspect of this disclosure, an apparatus for detecting speech is provided. The apparatus includes an audio acquisition module configured to acquire audio for a user's query, the audio beginning with the speech of the query; a text determination module configured to determine text for the query based on the speech; and an end position determination module configured to determine the end position of the speech in the audio corresponding to the query, based on the audio and the text.

[0009] In a third aspect of this disclosure, an electronic device is provided, including at least one processor; and a storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to the first aspect of this disclosure.

[0010] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0011] In a fifth aspect of this disclosure, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0012] It should be understood that the content described in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0014] Figure 1 illustrates a schematic diagram of an example environment in which the devices and / or methods of the embodiments of this disclosure may be implemented;

[0015] Figure 2 illustrates a schematic diagram of an example method for detecting speech according to an embodiment of the present disclosure;

[0016] Figure 3 illustrates a schematic diagram of one embodiment of speech detection according to an embodiment of the present disclosure;

[0017] Figure 4 illustrates a schematic diagram of another embodiment for detecting speech according to an embodiment of the present disclosure;

[0018] Figure 5 illustrates a schematic diagram of yet another embodiment of speech detection according to an embodiment of the present disclosure;

[0019] Figure 6 illustrates a schematic block diagram of an apparatus for detecting speech according to an embodiment of the present disclosure;

[0020] Figure 7 illustrates a schematic block diagram of an example device suitable for implementing embodiments of the present disclosure.

[0021] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and provisions. Upon receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media performing the operations of this disclosed technical solution, based on the prompt message.

[0023] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0024] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0025] There are still many problems to be solved in the process of speech detection. For example, taking the interaction between a user and a computing device or an application installed on the device (such as a voice assistant) as an example, the user can issue a wake-up command to the computing device or application and then start a conversation. The user can issue various query commands to the computing device or application according to their own needs, such as asking questions or engaging in conversation. The computing device or application can then pick up the user's conversation until the user finishes speaking. Then, the computing device or application will turn off the sound pickup after a predetermined time. For example, after waking up the computing device or application, the user might ask, "What will the weather be like tomorrow?" The computing device or application will then pick up the user's speech and turn off the sound pickup after the user has been silent for a period of time, and then continue to process the relevant commands in the user's speech.

[0026] In traditional solutions, when users interact with computing devices or their installed applications via voice, the computing device or backend server relies on Voice Activity Detection (VAD) technology to identify whether a user is speaking in the environment. These operations are typically used in the preprocessing for Automatic Speech Recognition (ASR). When someone is detected speaking in the current environment, after a predetermined time (e.g., 500 milliseconds), the computing device or application activates its microphone to receive the user's speech. After the user finishes speaking, VAD detects a period of silence, such as 1.5 seconds, then the computing device or application deactivates its microphone and sends the previously captured audio to the server for subsequent tasks such as recognition. However, the duration of silence in this traditional solution is generally set to a fixed value, which can lead to two scenarios: increased latency and premature termination. One scenario is that the user has finished speaking but still has to wait for a fixed duration before termination, resulting in a hard delay. If the termination duration is long, the overall latency increases, impacting the user experience. Another scenario is that some users speak slowly and may pause to think. If the preset pause time has been exceeded, the device assumes the user has finished speaking. However, the user has not actually finished speaking, resulting in premature pause and a poor user experience.

[0027] To address at least the aforementioned and other potential problems, embodiments of this disclosure propose a method for detecting speech. In this method, audio of a user's query is first acquired at a computing device. This audio begins with the user's query speech. The computing device then processes the speech to determine the text corresponding to the user's query. Finally, the computing device uses the determined text and the audio of the user's query to determine the end position of the speech in the audio corresponding to the user's query. This method, by combining the audio of the user's query and the corresponding text during speech detection, allows for faster and more accurate determination of the end position of the query speech, thereby reducing waiting time during the interaction process, avoiding semantic incompleteness caused by premature termination, and improving the user experience.

[0028] Embodiments of the present disclosure will now be described in further detail with reference to the accompanying drawings. Figure 1 illustrates an example environment in which the devices and / or methods of the embodiments of the present disclosure may be implemented. In environment 100, computing device 106 first acquires audio 108 of a user's query 102. Audio 108 begins with the voice 104 of the user's query. Then, computing device 106 uses the voice 104 to determine text 110 of the user's query 102. After determining the text 110, computing device 106 further uses audio 108 to determine the end position 112 of the voice 104 in audio 108 corresponding to the user's query 102.

[0029] Examples of computing device 106 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.

[0030] As shown in Figure 1, computing device 106 can acquire audio 108 corresponding to user query 102. In one example, the audio 108 corresponding to user query 102 is received directly by computing device 106 from another computing device. In another example, the audio 108 corresponding to user query 102 is obtained directly by computing device 106, such as audio information including speech received through a microphone. Additionally, computing device 106 can acquire a voice dataset corresponding to the user and then determine the user's characteristics based on the voice dataset. As mentioned above, user authorization is required before obtaining the user's voice data.

[0031] In this scenario, computing device 106 can further utilize the identified user characteristics to determine the end position 112 of the query speech. For example, this can be achieved by preprocessing the speech dataset and extracting acoustic features (such as Mel-frequency cepstral coefficients), and performing feature selection or optimization, to further determine the end position 112 of the audio 108 based on the user characteristics. For instance, the duration of the maximum permissible silence period can be extended based on the user characteristics.

[0032] The computing device 106 can also utilize the voice 104 to further determine the text 110 of the user's query 102. In some embodiments, the user's query 102 may contain not only the user's own voice information but also voice information of other objects in the background. The computing device 106 accepts all voice information, and then, according to a predetermined function, can filter out the voice information of other objects, retaining only the user's voice information. The above examples are for illustrative purposes only and are intended to limit the scope of this disclosure.

[0033] Finally, the computing device 106 will further determine the end position 112 of the voice 104 in the audio 108 corresponding to the user's query 102 based on the determined text 110 and the audio 108 for the user's query 102.

[0034] In some embodiments, audio 108 is determined by a VAD (Voice over Analytical Audio), for example, by determining the start point of the speech using a VAD, and then the speech after the start point is transmitted as audio 108 to a computing device, whereby the computing device 106 determines the end point 112 of the speech. Additionally, text 110 is generated by converting audio 104 into text using the functionality of an automatic speech recognition model.

[0035] The above process describes determining the end position of the speech corresponding to the user's query using computing device 106. Additionally, determining the end position of the speech corresponding to the user's query can be performed by an application running on computing device 106.

[0036] This method combines the audio of the user's query with the corresponding text during speech detection, enabling a faster and more accurate determination of the end point of the query's speech. This reduces waiting time during the interaction process, avoids semantic incompleteness caused by premature pausing, and improves the user experience.

[0037] The foregoing description, with reference to FIG1, illustrates an example environment in which the devices and / or methods of the embodiments of the present disclosure may be implemented. The following description, with reference to FIG2, illustrates an example method 200 for detecting speech according to an embodiment of the present disclosure. This method can be applied to the example environment in FIG1 or any suitable environment, and can be executed by computing device 106 or any suitable computing device. Additionally, the method can also be executed by an application running on computing device 106.

[0038] As shown in Figure 2, in example method 200, at box 202, computing device 106 acquires audio 108 for a user's query 102, where audio 108 begins with the speech of the query. In one example, the audio processed by this method is audio transmitted by a speech activity detection model. When the speech activity detection model detects user speech input, it uses audio beginning with the speech input in this method to determine the end position of the user's query. For example, the speech activity detection model determines that the user has entered speech corresponding to the user's query after detecting speech such as 500 milliseconds. Therefore, the speech activity detection model determines the start point of the speech, and then computing device 106 uses the audio after that start point as audio 108 to be processed. In another example, computing device 106 receives audio from another device with the determined speech as the start part.

[0039] In some embodiments, when the computing device 106 is a terminal device, the user can turn on the microphone on the terminal device to receive user voice information. For example, the user's query could be "Please recommend a good book that you think is worth reading," and the user can then speak the above content. The microphone of the terminal device can then receive audio including the user's query. In some embodiments, when the computing device 106 is a server, it can receive audio 108 from the user's terminal device, which includes voice information corresponding to the user's query 102. The above examples are merely for describing this disclosure and are not intended to specifically limit this disclosure.

[0040] At box 204, computing device 106 determines text 110 for user query 102 based on speech 104. To more accurately determine the end position 112, computing device 106 further utilizes the text 110 corresponding to the speech in audio 108. This text 110 also corresponds to user query 102. In some embodiments, computing device 106 determines the features of each audio frame in audio 108 and then uses these features to determine the corresponding text. For example, computing device 106 uses a speech recognition model to convert audio 108 into text 110. During speech activity detection, each audio frame in audio 108 can be labeled to identify whether it is a speech frame or a non-speech frame, such as a silence frame. For example, if an audio frame is a speech frame, it can be labeled as 1. If the audio frame is a non-speech frame, it can be labeled as 0. Therefore, the speech recognition model can determine the text corresponding to speech by processing only the audio frames labeled as 1. In some embodiments, a pre-established mapping relationship between speech frames and text exists. Then, the computing device 106 determines the text corresponding to the speech frame based on the mapping relationship. The above example is only for describing this disclosure and is not intended to limit the specific scope of this disclosure.

[0041] In some embodiments, the user's query 102 may include not only the user's own voice 104, but also voice information of other objects in the background. The computing device 106 can filter out the voice information of other objects and convert only the user's voice into corresponding text. For example, the user's voice features can be determined based on the user's voice, thereby distinguishing it from the voice features of other objects. Then, the user's voice features are used to obtain the voice information corresponding to that user. The above examples are only used to describe this disclosure and are intended to limit this disclosure.

[0042] At box 206, computing device 106 determines the end position 112 of the speech 104 in audio 108 corresponding to the user's query 102 based on audio 108 and text 110. After obtaining audio 108 and text 110 corresponding to the speech 104 in audio 108, computing device 106 uses this information to determine the end position 112 of the speech 104 corresponding to the user's query 102.

[0043] In some embodiments, computing device 106 can determine audio features for audio based on the audio. For example, computing device 106 can divide the audio into multiple audio frames. For example, during speech activity detection, computing device 106 divides audio 108 into multiple audio frames. Then, computing device 106 processes the multiple audio frames to determine audio features for each of the multiple audio frames. Computing device 106 can determine the audio features for each audio frame using methods such as Fourier transform.

[0044] Furthermore, the computing device 106 can also utilize the text to further determine semantic features for the text. For example, it can process the features of the text to determine its semantic features. In one example, the computing device 106 determines the semantic features of the text using a machine learning model. In another example, the computing device 106 determines the semantic features of the text using a bag-of-words model. The above examples are merely for describing this disclosure and are not intended to limit it. After determining the audio features and semantic features, the computing device 106 can use the determined audio features and semantic features to determine the end position of the speech in the audio corresponding to the user's query. In one example, the computing device 106 processes the audio features and semantic features according to a pre-trained machine learning model to determine the end position of the speech in the audio corresponding to the user's query. In another example, the computing device 106 obtains a pre-established mapping relationship between audio features, semantic features, and the end position of the speech. Then, the computing device 106 determines the end position of the speech corresponding to the user's query based on this mapping relationship. The above examples are merely for describing this disclosure and are not intended to limit it.

[0045] In some embodiments, when determining the end position of the speech in the audio corresponding to the query, the computing device first predicts the end position of the speech in the audio corresponding to the query using audio features and semantic features. If the end position can be predicted using the audio features and semantic features, the audio before the end position is determined as the speech segment used to fully express the user's query. Additionally, this speech segment is sent to subsequent processes for further processing, such as performing related operations corresponding to the user's query. If the computing device 106 cannot predict the end position of the speech in the audio corresponding to the query using the audio features and semantic features, it is necessary to further determine whether the duration of the silent audio segment in the audio reaches a predetermined duration, where the predetermined duration is the maximum allowed silent period when determining the end position. For example, the duration of audio frames marked as silent frames is counted. If the end position is not predicted within the predetermined duration, the end position can be determined at the predetermined duration.

[0046] In some embodiments, the computing device 106 can also dynamically adjust the predetermined duration based on whether the semantics of the text obtained from the audio are complete. For example, when determining the end position of the speech corresponding to the query in the audio, the computing device 106 can use semantic features to determine whether the semantics of the user's query 102 are complete. If the semantics are complete, the predetermined duration can remain unchanged. If the semantics of the query are incomplete, the computing device can adjust the length of the predetermined duration, for example, by increasing the duration. In one example, if the user is thinking about a question while speaking, the semantics of the text obtained from a continuous speech segment may not be complete, and the user may continue describing after a period of silence. In this case, appropriately adjusting the predetermined duration for silence based on the semantic incompleteness allows the user's speech to be truncated prematurely, improving the user experience.

[0047] In some embodiments, when determining the end position of the speech corresponding to the query in the audio, the predetermined duration can also be adjusted based on other information, such as user characteristics. In this case, the computing device first acquires the user characteristics corresponding to the user. Then, the computing device 106 adjusts the length of the predetermined duration based on the acquired user characteristics. For example, each speaker has a different speaking speed; some speak very quickly, some speak very slowly, and sometimes there may be longer pauses. Therefore, adjusting the predetermined duration or the maximum allowed silence period when determining the end position based on user characteristics can also avoid prematurely truncating the user's speech, thereby improving the user experience.

[0048] In some embodiments, when determining the end position of the speech in the audio corresponding to the query using audio and text, the computing device can process the audio and text using a pre-trained speech truncation model to provide the end position of the speech in the audio corresponding to the query. In some embodiments, the speech truncation model is pre-stored in the computing device or received from another computing device.

[0049] In some embodiments, the speech truncation model is trained in computing device 106. In this process, computing device 106 acquires the original audio of a sample query and the sample text corresponding to the sample query. The sample text includes an identifier indicating the end of the sample query, which has the end position of the sample speech in the sample audio corresponding to the sample query. For example, the sample text might be "Please recommend a good book that you think is worth reading," and the sample speech might be the speech of the above text spoken by a user. During training, an identifier "eos" can be added to the sample text, and this identifier has a frame index corresponding to the audio frame of "book" as the sample position. The computing device then uses the sample audio and sample text to train the speech truncation model. For example, the speech truncation model can be adjusted based on the position predicted by the speech truncation model and the original position. The above examples are merely for describing this disclosure and are not intended to limit the specific scope of this disclosure.

[0050] This method combines the audio of the user's query with the text of the audio during speech detection, enabling the rapid and accurate determination of the end position of the query's speech. Furthermore, by incorporating the semantic features of the query and the user's characteristics, the determined end position becomes more precise, reducing the error rate of the speech truncation model and improving the user experience.

[0051] The above description, with reference to FIG2, illustrates an example method 200 for detecting speech according to an embodiment of the present disclosure. The following description, with reference to FIG3, illustrates an embodiment of speech detection according to an embodiment of the present disclosure.

[0052] In Example 300 shown in Figure 3, a sentence 302 spoken by the user during interaction with the computing device 106, without pause or with a short pause, is given: "Guess the last four digits of my girlfriend's ID card." Before speaking, the user has activated the voice interaction function of the computing device or the application installed on it using a specified voice command or a specified button on the computing device, and the computing device or the application installed on it has also enabled the sound pickup function.

[0053] When the user speaks this sentence, there is no long pause (e.g., 100 milliseconds) between each word. After receiving the sentence spoken by the user, the computing device predicts the end position 304. After predicting the end position 304, the computing device turns off the sound pickup function.

[0054] In some embodiments, when a user speaks, there may be a pause and a period of silence. When the computing device predicts the end position after the user finishes speaking, it does not know whether the user has finished speaking. In this case, the computing device needs to wait for a period of time before it can predict the end position.

[0055] For example, after a user says the sentence 306, "Please recommend a good book that you think is worth reading," a silent audio segment 308 occurs. The duration of this silent audio segment (e.g., 600 milliseconds) does not exceed a predetermined duration, such as 1.5 seconds. If, after 600 milliseconds, the user is detected saying again, "I want to choose a reason to describe why this book is worth reading" 310, the calculation device predicts position 312 as the end position. In this example, if the end position cannot be predicted within the predetermined duration of the silent frame, the end position can be determined based on the silence truncation duration given by speech activity detection.

[0056] The above description, with reference to FIG3, is a schematic diagram of one embodiment of speech detection according to embodiments of the present disclosure. The following description, with reference to FIG4, is a schematic diagram of another embodiment of speech detection according to embodiments of the present disclosure.

[0057] In Example 400 shown in Figure 4, a paused sentence 402 is given when the user interacts with the computing device: "Please recommend a good book that you think is worth reading." Before speaking, the user has activated the voice interaction function of the computing device or the application installed on it using a specified voice command or a specified button on the computing device, and the computing device or the application installed on it has also turned on the sound pickup function.

[0058] It can be seen that the semantics of statement 402 are complete, and the computing device will predict that the end of statement 402 is the end position 406 after a waiting time of less than a predetermined duration (e.g., 1.5 seconds) of silent audio segment 504.

[0059] In some embodiments, the computing device may further determine whether statement 408 is semantically complete based on the semantic features of statement 408. If it is not complete, the predetermined duration of the silent audio segment is adjusted appropriately, for example, to 2.5 seconds, and then the end position 412 is predicted after the silent audio segment 410 of the adjusted duration has elapsed.

[0060] The above description, with reference to FIG4, is a schematic diagram of another embodiment of speech detection according to embodiments of the present disclosure. The following description, with reference to FIG5, is a schematic diagram of yet another embodiment of speech detection according to embodiments of the present disclosure.

[0061] In Example 500 shown in Figure 5, a paused sentence 504 is given when the user interacts with the computing device, which reads: "Your object." Before speaking, the user has activated the voice interaction function of the computing device or the application installed on it using a specified voice command or a specified button on the computing device, and the computing device or the application installed on it has also enabled the sound pickup function.

[0062] The computing device can further predict the end position of a sentence by incorporating user features 502 into a speech truncation model. For example, if the user features indicate that the user's speech rate is slow, the device can automatically increase the predetermined duration of the silent audio segment. After the silent audio segment 506, when the computing device detects the sentence 508 "What's your name again?", it further combines it with sentence 504 to determine whether the semantics are complete. If the semantics are complete, the end position 510 is predicted at the end of the sentence.

[0063] Additionally, if the computing device detects that the semantics are still incomplete, the predetermined duration of the silent audio segment is further adjusted.

[0064] The above description, with reference to FIG. 5, illustrates a schematic diagram of yet another embodiment of a speech detection device according to an embodiment of the present disclosure. The following description, with reference to FIG. 6, illustrates a schematic block diagram of a speech detection device 600 according to an embodiment of the present disclosure.

[0065] As shown in Figure 6, the device 600 includes an audio acquisition module 610 configured to acquire audio for a user's query, the audio starting with the query's speech; a text determination module 620 configured to determine text for the query based on the speech; and an end position determination module 630 configured to determine the end position of the speech in the audio corresponding to the query based on the audio and the text.

[0066] In some embodiments, the end position determination module 630 further includes: an audio feature determination module configured to determine audio features for the audio based on the audio; a semantic feature determination module configured to determine semantic features for the query based on the text; and an end position determination module configured to determine the end position of the speech in the audio corresponding to the query based on the audio features and the semantic features.

[0067] In some embodiments, the audio feature determination module further includes: an audio segmentation module configured to determine multiple audio frames of the audio by segmenting the audio; and an audio feature determination module configured to determine audio features for each of the multiple audio frames based on the multiple audio frames.

[0068] In some embodiments, the end position determination module further includes: an end position prediction module configured to predict the end position of the speech in the audio corresponding to the query based on audio features and semantic features; and an end position determination module configured to determine the end position at a predetermined duration if the end position has not been predicted when the duration of a silent audio segment in the audio reaches a predetermined duration, wherein the predetermined duration is the duration of the maximum allowed silent period when the end position is determined.

[0069] In some embodiments, the end position determination module further includes: a semantic integrity determination module, configured to determine whether the semantics of the query are complete based on semantic features; and a predetermined duration length adjustment module, configured to adjust the length of the predetermined duration in response to semantic incompleteness of the query.

[0070] In some embodiments, the end position determination module further includes: a user feature acquisition module configured to acquire user features corresponding to the user; and a predetermined market length adjustment module configured to adjust the length of the predetermined duration based on the user features.

[0071] In some embodiments, the end position determination module 630 further includes a speech truncation model application module, configured to determine the end position of the speech in the audio corresponding to the query by applying the audio and text to a speech truncation model.

[0072] In some embodiments, the speech truncation model application module further includes: a speech truncation model training module, configured to acquire a school-based audio file containing a sample query and sample text corresponding to the sample query, wherein the sample text includes an identifier indicating the end of the sample query, and an identifier indicating the end position of the sample speech in the sample audio for the sample query; and to train a speech truncation model based on the sample audio and sample text.

[0073] In some embodiments, audio is determined by speech activity detection.

[0074] Figure 7 shows a schematic block diagram of an example device 700 that can be used to implement embodiments of the present disclosure. The computing device 106 in Figure 1 can be implemented using device 700. As shown, device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. Various programs and data required for the operation of device 700 may also be stored in RAM 703. CPU 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0075] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0076] The various processes and handling described above, such as method 200, can be executed by processing unit 701. For example, in some embodiments, method 200 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of method 200 described above can be performed.

[0077] This disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0078] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0079] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0080] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0081] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0082] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0083] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0084] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0085] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for detecting speech, comprising: Obtain the audio of the user's query, wherein the audio begins with the speech of the query; Based on the voice, determine the text for the query; as well as Based on the audio and the text, determine the end position of the speech in the audio that corresponds to the query.

2. The method according to claim 1, wherein determining the end position of the speech in the audio corresponding to the query based on the audio and the text comprises: Based on the audio, determine the audio features specific to the audio. Based on the text, determine the semantic features for the query; as well as Based on the audio features and the semantic features, the end position of the speech in the audio corresponding to the query is determined.

3. The method of claim 2, wherein determining the audio features based on the audio includes: By dividing the audio, multiple audio frames of the audio are determined; as well as Based on the plurality of audio frames, determine the audio features for each of the plurality of audio frames.

4. The method according to claim 2, wherein determining the end position of the speech in the audio corresponding to the query based on the audio features and the semantic features includes: Based on the audio features and the semantic features, the end position of the speech in the audio corresponding to the query is predicted; as well as In response to the fact that the end position has not been predicted when the duration of the silent audio segment in the audio reaches a predetermined duration, the end position is determined at the predetermined duration, wherein the predetermined duration is the maximum allowed duration of the silent period when the end position is determined.

5. The method according to claim 4, wherein determining the end position of the speech in the audio corresponding to the query based on the audio features and the semantic features further includes: Based on the semantic features, determine whether the semantics of the query are complete; as well as In response to the semantic incompleteness of the query, the length of the predetermined duration is adjusted.

6. The method according to claim 4, wherein determining the end position of the speech corresponding to the query in the audio based on the audio features and the semantic features further includes: Obtain the user characteristics corresponding to the user; as well as Based on the user characteristics, the length of the predetermined duration is adjusted.

7. The method of claim 1, wherein determining the end position of the speech in the audio corresponding to the query based on the audio and the text comprises: By applying the audio and the text to a speech truncation model, the end position of the speech in the audio corresponding to the query is determined.

8. The method according to claim 7, wherein training the speech truncation model comprises: Obtain the school-based audio containing the sample query and the sample text corresponding to the sample query, wherein the sample text includes an identifier indicating the end of the sample query, the identifier having the end position of the sample speech in the sample audio for the sample query; and The speech truncation model is trained based on the sample audio and the sample text.

9. The method of claim 1, wherein the audio is determined by speech activity detection.

10. An apparatus for detecting speech, comprising: The audio acquisition module is configured to acquire audio in response to a user's query, wherein the audio begins with the speech of the query. The text determination module is configured to determine the text for the query based on the speech. as well as The end position determination module is configured to determine the end position of the speech in the audio that corresponds to the query, based on the audio and the text.

11. An electronic device, comprising: At least one processor; as well as A storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, the computer program implementing the method according to any one of claims 1-9 when executed by a processor.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Man-machine conversation detection method and device

    CN108257616A

  • Method and device for processing voice request

    CN110400576A

  • Voice endpoint detection method and device and electronic equipment

    CN111583912A

  • Voice endpoint detection method and device, equipment and storage medium

    CN114155839A

  • Voice detection method and device and storage medium

    CN115240716A