Voice activity detection method, device, computer-readable storage medium and equipment
By segmenting and keyframe analysis of audio data, and identifying and processing voice fragments containing user voice, the problem of low overall recognition efficiency of voice in the intelligent customer service system is solved, and more efficient computing resource utilization and semantic recognition efficiency are achieved.
Patent Information
- Application Number
- CN202111231199.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-10-22
AI Technical Summary
In the intelligent customer service system, voice calls include user voice and silent audio bands, which are low overall recognition efficiency, resulting in low computing resource utilization and low recognition efficiency.
By segmenting the audio data, voice fragments containing user voice are determined based on the keyframes in each audio clip, and these voice fragments are semantically recognized to improve recognition efficiency.
It improves the utilization rate of computing resources and semantic recognition efficiency, reduces labor costs, and realizes the intelligence and automation of telephone customer service.
Smart Images

Figure CN113990304B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular to a voice activity detection method, a voice activity detection device, a computer-readable storage medium, and an electronic device. Background Art
[0002] With the development of intelligent customer service technology, customer service is no longer solely dependent on manual work, but can automatically recognize the user's voice and match a suitable reply based on the analysis of the user's voice. Specifically, it is necessary to first obtain the call voice in the process of providing customer service, and recognize the call voice to determine the semantics, and then match the relevant replies based on the semantics. However, in the call voice, there are audio segments containing the user's voice and silent audio segments that do not contain the user's voice. If the call voice is recognized as a whole, it is easy to cause the problem of low recognition efficiency.
[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0004] The purpose of the present application is to provide a voice activity detection method, a voice activity detection device, a computer-readable storage medium and an electronic device, which can segment audio data and determine the voice segments containing user voice in the audio data based on the key frames in each audio segment, and then perform semantic recognition on the voice segments containing user voice, thereby improving the utilization of computing resources and the efficiency of semantic recognition.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0006] According to one aspect of the present application, a voice activity detection method is provided, comprising:
[0007] When receiving audio data, dividing the audio data into a plurality of audio segments;
[0008] Determining a key frame of each audio clip in the plurality of audio clips according to a preset key frame rule;
[0009] Determining a voice segment containing a user's voice from a plurality of audio segments according to key frames of each audio segment;
[0010] Perform semantic recognition on the speech fragments and match the corresponding speech replies based on the semantic recognition results.
[0011] In an exemplary embodiment of the present application, before dividing the audio data into a plurality of audio segments, the method further includes:
[0012] detecting specific audio frames containing noise in the audio data;
[0013] The specific audio frame is subjected to denoising processing based on at least one preset audio signal to obtain denoised audio data.
[0014] In an exemplary embodiment of the present application, the preset key frame rule is used to limit the number of key frames and key frame positions in each audio segment.
[0015] In an exemplary embodiment of the present application, determining a voice segment containing a user's voice from a plurality of audio segments according to key frames of each audio segment includes:
[0016] The key frames of each audio clip are segmented at the byte level according to a preset number of bytes to obtain a set of byte groups corresponding to each key frame; wherein each byte group in the set of byte groups contains a plurality of bytes, and each byte group corresponds to the same number of bytes;
[0017] A voice segment containing the user's voice is determined from the multiple audio segments according to the byte group set corresponding to each key frame.
[0018] In an exemplary embodiment of the present application, determining a voice segment containing a user voice from a plurality of audio segments according to a set of byte groups corresponding to each key frame includes:
[0019] Reorganize each byte group set at the frame level within the set to obtain a reorganized frame corresponding to each audio clip;
[0020] Performing speech detection on the reconstructed frames corresponding to each audio segment to obtain speech detection results corresponding to each audio segment; wherein the speech detection results are used to characterize the relationship between the target audio segment and the user's speech, and the target audio segment is any audio segment among the audio segments;
[0021] The voice segment containing the user's voice is determined according to the voice detection results corresponding to each audio segment.
[0022] In an exemplary embodiment of the present application, the voice segment includes a first type of voice segment and a second type of voice segment, and the voice segment containing the user voice is determined according to the voice detection results corresponding to each audio segment, including:
[0023] Determine a previous audio segment that is temporally continuous with the target audio segment;
[0024] If the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they contain the user's speech, it is determined that there is no speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's speech;
[0025] If the speech detection result of the previous audio segment indicates that it does not contain the user's voice, and the speech detection result of the target audio segment indicates that it contains the user's voice, it is determined that there is a speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's voice.
[0026] In an exemplary embodiment of the present application, semantic recognition is performed on a speech segment, including:
[0027] If there are temporally continuous speech segments, the temporally continuous speech segments are merged into the speech to be recognized;
[0028] The speech to be recognized is converted into text information, and semantic recognition is performed on the text information to obtain the semantic recognition result.
[0029] According to one aspect of the present application, a voice activity detection device is provided, comprising:
[0030] An audio segmentation unit, configured to segment the audio data into a plurality of audio segments when receiving the audio data;
[0031] A key frame determining unit, configured to determine a key frame of each audio segment in the plurality of audio segments according to a preset key frame rule;
[0032] A voice segment determining unit, used to determine a voice segment containing a user's voice from a plurality of audio segments according to key frames of each audio segment;
[0033] The semantic recognition unit is used to perform semantic recognition on the voice fragments and match the corresponding voice reply based on the semantic recognition results.
[0034] In an exemplary embodiment of the present application, the above-mentioned device further includes:
[0035] The denoising unit is used to detect a specific audio frame containing noise in the audio data before the audio segmentation unit segments the audio data into a plurality of audio segments; and to perform denoising on the specific audio frame based on at least one preset audio signal to obtain denoised audio data.
[0036] In an exemplary embodiment of the present application, the preset key frame rule is used to limit the number of key frames and key frame positions in each audio segment.
[0037] In an exemplary embodiment of the present application, the voice segment determining unit determines a voice segment containing a user voice from a plurality of audio segments according to key frames of each audio segment, including:
[0038] The key frames of each audio clip are segmented at the byte level according to a preset number of bytes to obtain a set of byte groups corresponding to each key frame; wherein each byte group in the set of byte groups contains a plurality of bytes, and each byte group corresponds to the same number of bytes;
[0039] A voice segment containing the user's voice is determined from the multiple audio segments according to the byte group set corresponding to each key frame.
[0040] In an exemplary embodiment of the present application, the voice segment determining unit determines a voice segment containing the user voice from multiple audio segments according to the byte group set corresponding to each key frame, including:
[0041] Reorganize each byte group set at the frame level within the set to obtain a reorganized frame corresponding to each audio clip;
[0042] Performing speech detection on the reconstructed frames corresponding to each audio segment to obtain speech detection results corresponding to each audio segment; wherein the speech detection results are used to characterize the relationship between the target audio segment and the user's speech, and the target audio segment is any audio segment among the audio segments;
[0043] The voice segment containing the user's voice is determined according to the voice detection results corresponding to each audio segment.
[0044] In an exemplary embodiment of the present application, the voice segment includes a first type of voice segment and a second type of voice segment, and the voice segment determination unit determines the voice segment containing the user voice according to the voice detection results corresponding to each audio segment, including:
[0045] Determine a previous audio segment that is temporally continuous with the target audio segment;
[0046] If the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they contain the user's speech, it is determined that there is no speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's speech;
[0047] If the speech detection result of the previous audio segment indicates that it does not contain the user's voice, and the speech detection result of the target audio segment indicates that it contains the user's voice, it is determined that there is a speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's voice.
[0048] In an exemplary embodiment of the present application, the semantic recognition unit performs semantic recognition on the speech segment including:
[0049] When there are temporally continuous speech segments, merging the temporally continuous speech segments into speech to be recognized;
[0050] The speech to be recognized is converted into text information, and semantic recognition is performed on the text information to obtain the semantic recognition result.
[0051] According to one aspect of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above methods by executing the executable instructions.
[0052] According to one aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, any one of the above methods is implemented.
[0053] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above-mentioned various optional implementations.
[0054] The exemplary embodiments of the present application may have some or all of the following beneficial effects:
[0055] In the voice activity detection method provided in an example implementation of the present application, when audio data is received, the audio data can be divided into multiple audio segments; the key frames of each audio segment in the multiple audio segments are determined according to preset key frame rules; the voice segment containing the user's voice is determined from the multiple audio segments according to the key frames of each audio segment; the voice segment is semantically recognized, and the corresponding voice reply is matched based on the semantic recognition result. According to the above scheme description, on the one hand, the present application can divide the audio data, and determine the voice segment containing the user's voice in the audio data based on the key frames in each audio segment, and then semantically recognize the voice segment containing the user's voice, thereby improving the utilization of computing resources and the efficiency of semantic recognition. On the other hand, the present application can reduce labor costs and realize the intelligence and automation of telephone customer service.
[0056] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0058] Figure 1 A schematic diagram showing an exemplary system architecture of a voice activity detection method and a voice activity detection device to which the embodiments of the present application can be applied;
[0059] Figure 2 A schematic diagram showing the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown;
[0060] Figure 3 The following schematically shows a flow chart of a voice activity detection method according to an embodiment of the present application;
[0061] Figure 4 A schematic diagram of a waveform of audio data according to an embodiment of the present application is schematically shown;
[0062] Figure 5 The following schematically shows a flow chart of a voice activity detection method according to an embodiment of the present application;
[0063] Figure 6 The structure diagram of a voice activity detection system according to an embodiment of the present application is schematically shown;
[0064] Figure 7 The structural block diagram of a voice activity detection device according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0065] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present application will be more comprehensive and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical solutions of the present application may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present application.
[0066] In addition, the accompanying drawings are only schematic illustrations of the present application and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0067] Figure 1 A schematic diagram of a system architecture of an exemplary application environment in which a voice activity detection method and a voice activity detection device according to an embodiment of the present application can be applied is shown.
[0068] like Figure 1 As shown, the system architecture 100 may include one or more of terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with display screens, including but not limited to desktop computers, portable computers, smart phones, tablet computers, etc. It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.
[0069] The voice activity detection method provided in the embodiment of the present application is generally executed by the server 105, and accordingly, the voice activity detection device is generally arranged in the server 105. However, it is easy for those skilled in the art to understand that the voice activity detection method provided in the embodiment of the present application can also be executed by the terminal device 101, 102 or 103, and accordingly, the voice activity detection device can also be arranged in the terminal device 101, 102 or 103, and this is not particularly limited in the present exemplary embodiment. For example, in an exemplary embodiment, the server 105 can divide the audio data into multiple audio segments when receiving audio data; determine the key frames of each audio segment in the multiple audio segments according to a preset key frame rule; determine the voice segment containing the user's voice from the multiple audio segments according to the key frames of each audio segment; perform semantic recognition on the voice segment, and match the corresponding voice reply based on the semantic recognition result.
[0070] Figure 2 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.
[0071] It should be noted that Figure 2 The computer system 200 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0072] like Figure 2 As shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 202 or a program loaded from a storage part 208 into a random access memory (RAM) 203. Various programs and data required for system operation are also stored in the RAM 203. The CPU 201, the ROM 202, and the RAM 203 are connected to each other via a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.
[0073] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as needed. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as needed, so that a computer program read therefrom is installed into the storage section 208 as needed.
[0074] In particular, according to an embodiment of the present application, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 209, and / or installed from a removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, various functions defined in the method and apparatus of the present application are executed.
[0075] This exemplary embodiment provides a voice activity detection method. The voice activity detection method can be applied to the server 105, or to one or more of the terminal devices 101, 102, 103, which is not particularly limited in this exemplary embodiment. Figure 3 As shown, the voice activity detection method may include the following steps S310 to S340.
[0076] Step S310: When audio data is received, the audio data is divided into a plurality of audio segments.
[0077] Step S320: determining a key frame of each audio segment in the plurality of audio segments according to a preset key frame rule.
[0078] Step S330: determining a voice segment including the user's voice from the multiple audio segments according to the key frames of each audio segment.
[0079] Step S340: Perform semantic recognition on the voice segment and match the corresponding voice reply based on the semantic recognition result.
[0080] Implementation Figure 3 The method shown can segment the audio data and determine the voice segments containing the user's voice in the audio data based on the key frames in each audio segment, and then perform semantic recognition on the voice segments containing the user's voice, thereby improving the utilization of computing resources and the efficiency of semantic recognition. In addition, it can also reduce labor costs and realize the intelligence and automation of telephone customer service.
[0081] Next, the above steps of this exemplary embodiment are described in more detail.
[0082] In step S310, when audio data is received, the audio data is divided into a plurality of audio segments.
[0083] Among them, the audio data may be speech data containing one or more segments of user voices, and the audio segments obtained by segmentation may be segments of equal length or segments of unequal length. The present application does not limit the number of audio segments. In addition, the audio segment may contain N1 audio frames, where N1 (e.g., 2) is a positive integer; the length of each audio frame may be represented by Tms, where T (e.g., 10, 20, 30) is a positive integer; each audio frame may include N2 sampling points, where N2 (e.g., 80, 160, 240) is a positive integer; each audio frame contains N3 bytes, where N3 (e.g., 160, 320, 640) is a positive integer. For example, the audio segment in the present application contains 2 audio frames, each audio frame is 10ms long, each audio frame contains 80 sampling points, and each audio frame contains 160 bytes.
[0084] Specifically, the multiple audio clips obtained by segmentation can be referred to Figure 4 , Figure 4 The waveform diagram of audio data according to an embodiment of the present application is schematically shown. Figure 4As shown, the horizontal axis of the waveform diagram can represent time, and the vertical axis can represent power. By dividing the audio data into multiple audio segments, determining the key frame of each audio segment in the multiple audio segments according to a preset key frame rule, and determining the voice segment containing the user's voice from the multiple audio segments according to the key frame of each audio segment, the timestamp indication of the voice segment corresponding to the waveform diagram can be as follows:
[0085]
[0086]
[0087] Before dividing the audio data into a plurality of audio segments in step S310, the method further includes: detecting a specific audio frame containing noise in the audio data; and denoising the specific audio frame based on at least one preset audio signal to obtain denoised audio data.
[0088] Among them, the specific audio frame containing noise in the audio data may be one or more, which is not limited in the embodiments of the present application. The specific method of detecting the specific audio frame containing noise in the audio data may be: comparing each frame of audio of the audio data with the noise frequency of the noise library, and determining the audio frame that meets the noise frequency as the specific audio frame. In addition, the specific audio frame can be denoised based on at least one preset audio signal to obtain denoised audio data; wherein the preset audio signal can be represented as a gate function for filtering out noise in the audio. If there are multiple preset audio signals, then the specific audio frame is denoised based on at least one preset audio signal to obtain the denoised audio data. The specific implementation method may be: according to the noise frequency corresponding to the specific audio frame, a target preset audio signal suitable for the specific audio frame is selected from multiple preset audio signals, and the specific audio frame is denoised according to the target preset audio signal to obtain the denoised audio data.
[0089] It can be seen that implementing this optional embodiment can achieve audio denoising, improve the purity of the audio, and reduce the impact of noise on the accuracy of user voice segment detection, thereby achieving accurate discrimination between user voice segments and non-user voice segments.
[0090] In step S320, a key frame of each audio segment in the plurality of audio segments is determined according to a preset key frame rule.
[0091] The preset key frame rule can be used to define at least one selection condition for the key frame, and can also be used to define the number of key frames and the key frame positions in each audio segment. For example, the number of audio segments is 3, namely audio segment A, audio segment B, and audio segment C. Based on the preset key frame rule, key frame A can be determined from audio segment A, key frame B can be determined from audio segment B, and key frame C can be determined from audio segment C. In addition, optionally, if the preset key frame rule defines the number of key frames in each audio segment as N, then the specific implementation of determining the key frames of each audio segment in the multiple audio segments according to the preset key frame rule can be: determining N key frames from each audio segment respectively according to the preset key frame rule; wherein N is a positive integer. In addition, optionally, if the preset key frame rule defines the key frame position in each audio segment (such as the position of the first frame), then the specific implementation of determining the key frames of each audio segment in the multiple audio segments according to the preset key frame rule can be: determining the key frame position in the multiple audio segments according to the preset key frame rule; and determining the key frame corresponding to the key frame position as the key frame corresponding to the corresponding audio segment.
[0092] Optionally, determining a key frame of each audio segment in the plurality of audio segments according to a preset key frame rule includes: determining a first frame in each audio segment as a key frame of the corresponding audio segment according to the preset key frame rule.
[0093] In step S330, a voice segment including the user's voice is determined from the plurality of audio segments according to the key frames of the audio segments.
[0094] Among the multiple audio segments, at least one audio segment contains user voice.
[0095] As a specific implementation of step S330, a voice segment containing the user's voice is determined from multiple audio segments according to the key frames of each audio segment, including: byte-level segmentation of the key frames of each audio segment according to a preset number of bytes to obtain a set of byte groups corresponding to each key frame; wherein each byte group in the byte group set contains multiple bytes, and each byte group corresponds to the same number of bytes; and a voice segment containing the user's voice is determined from multiple audio segments according to the set of byte groups corresponding to each key frame.
[0096] The byte group set corresponding to each key frame may include multiple byte groups, and the number of bytes contained in each byte group meets the above-mentioned preset number of bytes. For example, the byte group set corresponding to key frame A may include 80 byte groups, and each byte group contains 2 bytes.
[0097] It can be seen that by implementing this optional embodiment, it is possible to determine whether an audio segment contains user voice at the byte level, which can make the determination result more accurate.
[0098] Furthermore, a voice segment containing the user's voice is determined from multiple audio segments according to the byte group sets corresponding to each key frame, including: reorganizing each byte group set at the frame level within the set to obtain reorganized frames corresponding to each audio segment; performing voice detection on the reorganized frames corresponding to each audio segment to obtain voice detection results corresponding to each audio segment; wherein the voice detection results are used to characterize the relationship between the target audio segment and the user's voice, and the target audio segment is any audio segment among the audio segments; and determining the voice segment containing the user's voice according to the voice detection results corresponding to each audio segment.
[0099] The method of reorganizing each byte group set at the frame level within the set to obtain the reorganized frames corresponding to each audio segment may specifically be: converting the byte groups in each byte group set into short integer pointers according to a preset rule to obtain short integer pointer sets corresponding to each byte group set, combining each short integer pointer set into a reorganized frame, and obtaining the reorganized frames corresponding to each audio segment. For example, the preset rule may be implemented as: short = (bytes[i*2+1]<<8)|(bytes[i*2]&0xFF).
[0100] Based on this, speech detection is performed on the reorganized frames corresponding to each audio segment to obtain the speech detection results corresponding to each audio segment, including: inputting the reorganized frames into the speech detection module, so that the speech detection module (such as the WebRtcVad detection module) detects whether each audio segment contains user voice based on the reorganized frames, that is, obtaining the speech detection results corresponding to each audio segment.
[0101] It can be seen that the implementation of this optional embodiment, compared with the prior art of directly converting the byte stream into a short integer pointer, can first combine the byte stream into byte groups, then convert the byte groups into short integer pointers, and then perform frame reorganization based on the short integer pointers, so as to avoid destroying the short-term stable state characteristics of the speech frame, and perform speech detection based on the short-term stable state of the speech signal, thereby improving the accuracy of speech detection.
[0102] Furthermore, the speech segments include a first category of speech segments and a second category of speech segments, and the speech segments containing the user's voice are determined according to the speech detection results corresponding to each audio segment, including: determining a previous audio segment that is temporally continuous with the target audio segment; if the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they contain the user's voice, then it is determined that there is no speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined to be speech segments containing the user's voice; if the speech detection result of the previous audio segment indicates that it does not contain the user's voice, and the speech detection result of the target audio segment indicates that it contains the user's voice, then it is determined that there is a speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined to be speech segments containing the user's voice.
[0103] In addition, the method further includes: if the voice detection result of the previous audio segment indicates that the user's voice is included, and the voice detection result of the target audio segment indicates that the user's voice is not included, it is determined that there is a voice state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as voice segments containing the user's voice. In this way, missed detection of the user's voice can be avoided.
[0104] In addition, it also includes: if the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they do not contain user voice, it is determined that there is no voice state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as audio segments that do not contain user voice.
[0105] It can be seen that the implementation of this optional embodiment can perform continuous speech detection based on multiple voice segments. Compared with performing speech detection on the entire audio file, it can reduce irrelevant calculations and improve the efficiency of distinguishing segments containing continuous speech. When applied to human-computer call scenarios, it can shorten the user's waiting time and improve the user's call experience.
[0106] In step S340, semantic recognition is performed on the voice segment, and a corresponding voice reply is matched based on the semantic recognition result.
[0107] The number of voice segments may be one or more, and the voice reply may consist of one or more sentences.
[0108] As a specific implementation of step S340, semantic recognition is performed on the voice segment, including: each time a voice segment is recognized, a step of performing semantic recognition on the voice segment is executed.
[0109] Alternatively, as a specific implementation of step S340, semantic recognition is performed on the speech segments, including: if there are temporally continuous speech segments, the temporally continuous speech segments are merged into speech to be recognized; the speech to be recognized is converted into text information, and semantic recognition is performed on the text information to obtain a semantic recognition result.
[0110] Among them, the temporally continuous speech segments can be understood as multiple temporally adjacent speech segments. For example, if speech segment 1 is 00:00-00:19, speech segment 2 is 00:20-00:29, and speech segment 3 is 00:30-00:39, then it can be determined that speech segment 1, speech segment 2, and speech segment 3 are temporally continuous. Furthermore, speech segment 1, speech segment 2, and speech segment 3 can be combined into the speech to be recognized, and the time period corresponding to the speech to be recognized can be 00:00-00:39.
[0111] Optionally, converting the speech to be recognized into text information includes: extracting speech features in the speech to be recognized, comparing the speech features with preset speech features, recognizing text corresponding to the speech features based on the preset speech features, and combining the recognized text into text information.
[0112] Furthermore, optionally, semantic recognition is performed on the text information to obtain a semantic recognition result, including: extracting a sub-vector of each word in the text information; identifying the degree of correlation between the words according to the order of the words and based on the correlation between the sub-vectors, and segmenting the text information according to the degree of correlation between the words; further identifying the part of speech of each word in the text information according to the segmentation result; and further generating a semantic recognition result according to the part of speech of each word and the text information.
[0113] It can be seen that the implementation of this optional embodiment can perform semantic recognition only on the voice segment containing the user's voice, which can effectively shorten the semantic recognition time and improve the semantic recognition efficiency.
[0114] See also Figure 5 , Figure 5 The structure diagram of the voice activity detection system according to an embodiment of the present application is schematically shown. The voice activity detection system 500 can be implemented on platforms such as Windows or Linux. In addition, the voice activity detection system 500 can be applied to a human-computer call scenario, and the received audio data can include the user's voice during the call. The session control protocol used during the call can be the SIP protocol, and the media transmission protocol used can be the RTP protocol.
[0115] The voice activity detection system 500 may specifically include: a voice reception preprocessing module 510 , a voice activity detection module 520 , an intelligent customer service module 530 , a voice synthesis module 540 , a voice recognition module 550 and a semantic analysis processing module 560 .
[0116] The voice receiving preprocessing module 510 is used to establish a communication session (session) and register a communication callback function corresponding to the communication session when detecting that a user calls a special service number (such as 10000) to enter the voice activity detection system 500, and initialize the communication parameters corresponding to the communication session, wherein the communication parameters may include: the initial voice activity state (such as 0 or 1), the minimum communication duration (min_speak_ms), the minimum silence duration (min_pause_ms), the maximum speaking duration (max_recording_ms), etc., which are not limited in the embodiment of the present application. The initial voice activity state can be used to characterize whether the preset initial audio contains / does not contain the user's voice.
[0117] The voice receiving preprocessing module 510 is also used to receive audio data that meets a preset coding format (such as PCMU (G.711u-law), PCMA (G.711A-law) or Opus) sent by the terminal device; when the audio data is received, the audio data is divided into multiple audio segments, and the key frame of each audio segment in the multiple audio segments is determined according to a preset key frame rule. The audio data can be divided into multiple audio segments based on the voice activity detection technology (Voice Activity Detection, VAD) Execution.
[0118] The voice receiving preprocessing module 510 is also used to monitor interruption events. When an interruption event is detected, a preset response plan can be selected from the preset response plans and the preset response plan can be executed; wherein, the interruption event is used to indicate that the intelligent customer service outputs audio to interrupt the user's speech during the user's speech.
[0119] The voice activity detection module 520 is configured to determine a voice segment containing the user's voice from among the multiple audio segments according to the key frames of each audio segment.
[0120] The speech recognition module 550 is used to merge temporally continuous speech segments into speech to be recognized and convert the speech to be recognized into text information; or, perform speech recognition on each speech segment to obtain text information of the corresponding speech segment.
[0121] The semantic analysis processing module 560 is used to perform semantic recognition on text information to obtain semantic recognition results.
[0122] The speech synthesis module 540 is used to generate a corresponding speech reply based on the semantic recognition result.
[0123] The intelligent customer service module 530 is used to output the above-mentioned voice reply.
[0124] It can be seen that implementation Figure 5 The system shown can segment the audio data and determine the audio segments containing the user's voice in the audio data based on the key frames in each audio segment, and then can perform semantic recognition on the audio segments containing the user's voice, thereby improving the utilization of computing resources and the efficiency of semantic recognition. In addition, it can also reduce labor costs and realize the intelligence and automation of telephone customer service.
[0125] See also Figure 6 , Figure 6 The flowchart of the voice activity detection method according to an embodiment of the present application is schematically shown. Figure 6 As shown, the voice activity detection method may include the following steps.
[0126] S600: When audio data is received, a specific audio frame containing noise is detected in the audio data, and denoising is performed on the specific audio frame based on at least one preset audio signal to obtain denoised audio data.
[0127] S602: Divide the audio data into multiple audio segments.
[0128] S604: Determine a key frame of each audio segment in the plurality of audio segments according to a preset key frame rule.
[0129] S606: Segment the key frames of each audio clip at byte level according to a preset number of bytes to obtain a set of byte groups corresponding to each key frame; wherein each byte group in the byte group set contains multiple bytes, and each byte group corresponds to the same number of bytes.
[0130] S608: Reorganize each byte group set at the frame level within the set to obtain a reorganized frame corresponding to each audio segment.
[0131] S610: Perform speech detection on the reconstructed frames corresponding to each audio segment to obtain speech detection results corresponding to each audio segment; wherein the speech detection results are used to characterize the relationship between the target audio segment and the user's voice, and the target audio segment is any audio segment among the audio segments.
[0132] S612: Determine a previous audio segment that is temporally continuous with the target audio segment.
[0133] S614: If the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they contain the user's voice, it is determined that there is no speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's voice.
[0134] S616: If the speech detection result of the previous audio segment indicates that it does not contain the user's voice, and the speech detection result of the target audio segment indicates that it contains the user's voice, it is determined that there is a speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's voice.
[0135] S618: If the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they do not contain user voice, it is determined that there is no speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined to be silent segments.
[0136] S620: If the speech detection result of the previous audio segment indicates that the user's speech is contained, and the speech detection result of the target audio segment indicates that the user's speech is not contained, it is determined that there is a speech state change between the previous audio segment and the target audio segment.
[0137] S622: Detect whether multiple audio segments in the audio data are all used as target audio segments for speech detection. If yes, execute step S624; if no, execute step S612.
[0138] S624: If there are temporally continuous speech segments, the temporally continuous speech segments are merged into the speech to be recognized.
[0139] S626: Convert the speech to be recognized into text information, and perform semantic recognition on the text information to obtain a semantic recognition result.
[0140] S628: Match the corresponding message to be replied according to the semantic recognition result and output the message to be replied.
[0141] It should be noted that steps S600 to S628 are Figure 3 The steps and their embodiments shown correspond to each other. For the specific implementation of steps S600 to S628, please refer to Figure 3 The steps and embodiments shown are not described in detail here.
[0142] It can be seen that implementation Figure 6 The method shown can segment the audio data and determine the voice segments containing the user's voice in the audio data based on the key frames in each audio segment, and then perform semantic recognition on the voice segments containing the user's voice, thereby improving the utilization of computing resources and the efficiency of semantic recognition. In addition, it can also reduce labor costs and realize the intelligence and automation of telephone customer service.
[0143] Furthermore, in this exemplary embodiment, a voice activity detection device is also provided. Figure 7As shown, the voice activity detection device 700 may include:
[0144] The audio segmentation unit 701 is used to segment the audio data into multiple audio segments when receiving the audio data;
[0145] A key frame determining unit 702, configured to determine a key frame of each audio segment in the plurality of audio segments according to a preset key frame rule;
[0146] A voice segment determining unit 703, configured to determine a voice segment containing a user's voice from a plurality of audio segments according to key frames of each audio segment;
[0147] The semantic recognition unit 704 is used to perform semantic recognition on the voice segment and match the corresponding voice reply based on the semantic recognition result.
[0148] The preset key frame rule is used to limit the number and position of key frames in each audio clip.
[0149] It can be seen that implementation Figure 7 The device shown can segment the audio data and determine the voice segments containing the user's voice in the audio data based on the key frames in each audio segment, and then can perform semantic recognition on the voice segments containing the user's voice, thereby improving the utilization of computing resources and the efficiency of semantic recognition. In addition, it can also reduce labor costs and realize the intelligence and automation of telephone customer service.
[0150] In an exemplary embodiment of the present application, the above-mentioned device further includes:
[0151] The denoising unit (not shown) is used to detect specific audio frames containing noise in the audio data before the audio segmentation unit 701 segments the audio data into multiple audio segments; denoise the specific audio frames based on at least one preset audio signal to obtain denoised audio data.
[0152] It can be seen that implementing this optional embodiment can achieve audio denoising, improve the purity of the audio, and reduce the impact of noise on the accuracy of user voice segment detection, thereby achieving accurate discrimination between user voice segments and non-user voice segments.
[0153] In an exemplary embodiment of the present application, the voice segment determining unit 703 determines a voice segment containing the user's voice from a plurality of audio segments according to key frames of each audio segment, including:
[0154] The key frames of each audio clip are segmented at the byte level according to a preset number of bytes to obtain a set of byte groups corresponding to each key frame; wherein each byte group in the set of byte groups contains a plurality of bytes, and each byte group corresponds to the same number of bytes;
[0155] A voice segment containing the user's voice is determined from the multiple audio segments according to the byte group set corresponding to each key frame.
[0156] It can be seen that by implementing this optional embodiment, it is possible to determine whether an audio segment contains user voice at the byte level, which can make the determination result more accurate.
[0157] In an exemplary embodiment of the present application, the voice segment determining unit 703 determines a voice segment containing the user's voice from multiple audio segments according to the byte group set corresponding to each key frame, including:
[0158] Reorganize each byte group set at the frame level within the set to obtain a reorganized frame corresponding to each audio clip;
[0159] Performing speech detection on the reconstructed frames corresponding to each audio segment to obtain speech detection results corresponding to each audio segment; wherein the speech detection results are used to characterize the relationship between the target audio segment and the user's speech, and the target audio segment is any audio segment among the audio segments;
[0160] The voice segment containing the user's voice is determined according to the voice detection results corresponding to each audio segment.
[0161] It can be seen that the implementation of this optional embodiment, compared with the prior art of directly converting the byte stream into a short integer pointer, can first combine the byte stream into byte groups, then convert the byte groups into short integer pointers, and then perform frame reorganization based on the short integer pointers, so as to avoid destroying the short-term stable state characteristics of the speech frame, and perform speech detection based on the short-term stable state of the speech signal, thereby improving the accuracy of speech detection.
[0162] In an exemplary embodiment of the present application, the voice segment includes a first type of voice segment and a second type of voice segment. The voice segment determination unit 703 determines the voice segment containing the user voice according to the voice detection results corresponding to each audio segment, including:
[0163] Determine a previous audio segment that is temporally continuous with the target audio segment;
[0164] If the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they contain the user's speech, it is determined that there is no speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's speech;
[0165] If the speech detection result of the previous audio segment indicates that it does not contain the user's voice, and the speech detection result of the target audio segment indicates that it contains the user's voice, it is determined that there is a speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's voice.
[0166] It can be seen that the implementation of this optional embodiment can perform continuous speech detection based on multiple voice segments. Compared with performing speech detection on the entire audio file, it can reduce irrelevant calculations and improve the efficiency of distinguishing segments containing continuous speech. When applied to human-computer call scenarios, it can shorten the user's waiting time and improve the user's call experience.
[0167] In an exemplary embodiment of the present application, the semantic recognition unit 704 performs semantic recognition on the speech segment including:
[0168] When there are temporally continuous speech segments, merging the temporally continuous speech segments into speech to be recognized;
[0169] The speech to be recognized is converted into text information, and semantic recognition is performed on the text information to obtain the semantic recognition result.
[0170] It can be seen that the implementation of this optional embodiment can perform semantic recognition only on the voice segment containing the user's voice, which can effectively shorten the semantic recognition time and improve the semantic recognition efficiency.
[0171] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0172] Since the various functional modules of the voice activity detection device of the exemplary embodiment of the present application correspond to the steps of the exemplary embodiment of the voice activity detection method described above, for details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the voice activity detection method described above in the present application.
[0173] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiment.
[0174] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0175] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0176] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0177] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary technical means in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0178] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A voice activity detection method, characterized in that: include: When receiving the audio data, dividing the audio data into a plurality of audio segments; Determining a key frame of each audio segment in the multiple audio segments according to a preset key frame rule; Determining a voice segment containing the user's voice from the multiple audio segments according to the key frames of the audio segments; If there are temporally continuous voice segments, the temporally continuous voice segments are merged into a voice to be recognized, the voice to be recognized is converted into text information, semantic recognition is performed on the text information to obtain a semantic recognition result, and a corresponding voice reply is matched based on the semantic recognition result; The step of determining a voice segment containing the user's voice from the plurality of audio segments according to the key frames of the audio segments includes: Determine a previous audio segment that is temporally continuous with the target audio segment; If the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they contain the user's speech, it is determined that there is no speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's speech; If the speech detection result of the previous audio segment indicates that it does not contain the user's voice, and the speech detection result of the target audio segment indicates that it contains the user's voice, it is determined that there is a speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's voice.
2. The method according to claim 1, characterized in that Before dividing the audio data into a plurality of audio segments, the method further includes: Detecting a specific audio frame containing noise in the audio data; The specific audio frame is subjected to denoising processing based on at least one preset audio signal to obtain denoised audio data.
3. The method according to claim 1, characterized in that The preset key frame rule is used to limit the number and position of key frames in each audio segment.
4. The method according to claim 1, characterized in that: Determining a voice segment containing the user's voice from the multiple audio segments according to the key frames of the audio segments includes: The key frames of each audio clip are segmented at the byte level according to a preset number of bytes to obtain a set of byte groups corresponding to each key frame; wherein each byte group in the set of byte groups contains a plurality of bytes, and each byte group corresponds to the same number of bytes; A voice segment containing the user's voice is determined from the multiple audio segments according to the byte group sets corresponding to the key frames.
5. The method according to claim 4, characterized in that Determining a voice segment containing the user's voice from the multiple audio segments according to the byte group sets corresponding to the key frames includes: Reorganize each byte group set at the frame level within the set to obtain reorganized frames corresponding to each of the audio clips; Performing speech detection on the recombined frames corresponding to the audio segments to obtain speech detection results corresponding to the audio segments; wherein the speech detection results are used to characterize the relationship between the target audio segment and the user's voice, and the target audio segment is any audio segment among the audio segments; The voice segment containing the user's voice is determined according to the voice detection results respectively corresponding to the audio segments.
6. A voice activity detection device, characterized in that: include: An audio segmentation unit, configured to segment the audio data into a plurality of audio segments when receiving the audio data; a key frame determining unit, configured to determine a key frame of each audio segment in the plurality of audio segments according to a preset key frame rule; a voice segment determining unit, configured to determine a voice segment containing a user's voice from the plurality of audio segments according to key frames of the audio segments; A semantic recognition unit, configured to merge the temporally continuous voice segments into a to-be-recognized voice if there are any, convert the to-be-recognized voice into text information, perform semantic recognition on the text information to obtain a semantic recognition result, and match a corresponding voice reply based on the semantic recognition result; The step of determining a voice segment containing the user's voice from the plurality of audio segments according to the key frames of the audio segments includes: Determine a previous audio segment that is temporally continuous with the target audio segment; If the speech detection result of the previous audio segment and the speech detection result of the target audio segment both indicate that they contain the user's speech, it is determined that there is no speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's speech; If the speech detection result of the previous audio segment indicates that it does not contain the user's voice, and the speech detection result of the target audio segment indicates that it contains the user's voice, it is determined that there is a speech state change between the previous audio segment and the target audio segment, and the target audio segment and the previous audio segment are determined as speech segments containing the user's voice.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
8. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 5 by executing the executable instructions.
Citation Information
Patent Citations
Voice endpoint detection and device, and computer device
CN107527630A