Audio segmentation method and device, electronic equipment and storage medium
By combining silent segment annotation and voiceprint feature combination, the problem of not being able to segment the audio of the business hall by customer in the existing technology has been solved, realizing audio segmentation by customer and improving the accuracy of service quality inspection and evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2022-09-29
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technology cannot segment the audio in the business hall according to the customer's business type, resulting in low effectiveness of service quality inspection and evaluation.
The silence segments in the two-channel audio are identified by marking the silence segments. The audio is segmented using common silence separators. The customer audio is then combined based on voiceprint features to obtain audio segmentation on a customer-by-customer basis.
It enables audio segmentation on a customer-by-customer basis, improving the accuracy of service quality inspection and evaluation, and providing assistance for service quality inspection and evaluation of different businesses.
Smart Images

Figure CN115719596B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information processing technology, and in particular to an audio segmentation method, apparatus, electronic device, and storage medium. Background Technology
[0002] In order to conduct quality inspection and evaluation of the service provided by the staff in the business hall, it is usually necessary to segment the audio of the business hall. However, considering that different customers may handle different services and different staff may be assigned to different services, the audio of the business hall needs to be segmented according to the order in which the customer handles the service, so as to realize the service quality inspection and evaluation of the staff corresponding to different services.
[0003] Currently, most audio segmentation solutions for business halls use audio pickup hardware for timed segmentation and audio uploading, but they cannot segment audio according to the customer, making it difficult to distinguish customers, resulting in low efficiency and failing to enable service quality inspection and evaluation of staff corresponding to different services. Summary of the Invention
[0004] This invention provides an audio segmentation method, apparatus, electronic device, and storage medium to address the shortcomings of existing technologies that can only perform timed segmentation and cannot distinguish the corresponding customer for audio, resulting in low efficiency of the segmentation method. This invention achieves audio segmentation based on the customer.
[0005] This invention provides an audio segmentation method, comprising:
[0006] Identify the two-channel audio to be segmented;
[0007] The first and second channels of the stereo audio are marked with silent segments to obtain the silent segments in the first channel audio and the silent segments in the second channel audio.
[0008] Based on the silence segments in the first channel audio and the silence segments in the second channel audio, a common silence separation point in the dual-channel audio is determined, and based on the common silence separation point, the first channel audio is segmented to obtain multiple first segmented audio segments;
[0009] The silence segments of each first segmented audio segment are removed to obtain each second segmented audio segment. Based on the voiceprint features of each second segmented audio segment, the customer audio is combined to obtain customer audio for each customer.
[0010] According to an audio segmentation method provided by the present invention, the step of combining customer audio based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer includes:
[0011] The second segmented audio segments are combined according to their order in the first channel audio to obtain the client combined audio.
[0012] Based on the similarity between the voiceprint features of adjacent second segmented audio segments in the customer combined audio, suspected audio endpoints in the customer combined audio are determined;
[0013] Based on the audio duration between adjacent suspected audio endpoints in the customer's combined audio and the preset noise duration, the suspected audio endpoints are filtered to obtain customer audio endpoints, and customer audio is determined on a customer-by-customer basis based on the customer audio endpoints.
[0014] According to an audio segmentation method provided by the present invention, determining suspected audio endpoints in the customer-combined audio based on the similarity between the voiceprint features of adjacent second-segmented audio segments in the customer-combined audio includes:
[0015] If the similarity between the voiceprint features of two adjacent second segmented audio segments in the customer combined audio is less than a preset similarity, then the audio combination point of the two adjacent second segmented audio segments is taken as the candidate audio endpoint of the customer combined audio.
[0016] Otherwise, the audio combination point of two adjacent second segmented audio segments is taken as the non-candidate audio endpoint of the client combined audio;
[0017] Based on the similarity between the voiceprint features of the second segmented audio segments corresponding to adjacent non-candidate audio endpoints in the customer's combined audio, and a preset similarity, the candidate audio endpoints are filtered to obtain the suspected audio endpoints.
[0018] According to an audio segmentation method provided by the present invention, the step of marking the first channel audio and the second channel audio in the dual-channel audio respectively to obtain the silent segments in the first channel audio and the silent segments in the second channel audio includes:
[0019] The frame energy of each audio frame in the first channel audio and the second channel audio is determined. Based on the frame energy of each audio frame and the energy threshold value, the silence detection state of each audio frame is determined. The energy threshold value is determined based on the corresponding channel audio.
[0020] Based on the number of audio frames contained in the audio windows of the first and second audio channels, and the silence detection status of each audio frame, the silence segments in the first and second audio channels are determined.
[0021] According to an audio segmentation method provided by the present invention, determining common silence separation points in the two-channel audio based on silence segments in the first channel audio and silence segments in the second channel audio includes:
[0022] Determine the mute endpoint of the mute segment in the first channel audio and the mute endpoint of the mute segment in the second channel audio.
[0023] A common silence endpoint is selected from the silence endpoints of the silence segments in the first channel audio and the silence endpoints of the silence segments in the second channel audio. The common silence endpoint corresponds to an audio frame within a silence segment in both the first channel audio and the second channel audio.
[0024] Based on the silence duration of the common silence endpoint in the corresponding channel audio and the preset silence duration, the common silence endpoint is filtered to obtain the common silence separation point.
[0025] According to an audio segmentation method provided by the present invention, the step of removing the silence segment from each first segmented audio segment to obtain each second segmented audio segment includes:
[0026] Each first segmented audio segment is subjected to a silent segment removal process to obtain each silent-removed audio segment.
[0027] Based on the audio duration of each silent cut-off audio segment and a preset audio duration, audio filtering is performed to obtain each second segmented audio segment.
[0028] According to an audio segmentation method provided by the present invention, the step of combining customer audio based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer unit further includes:
[0029] Based on the second channel audio, determine the standard audio endpoint;
[0030] Based on the standard audio endpoint and the client audio, determine the audio segmentation accuracy;
[0031] Adjust the client audio endpoint based on the audio segmentation accuracy.
[0032] The present invention also provides an audio segmentation device, comprising:
[0033] An audio determination unit is used to determine the two-channel audio to be segmented;
[0034] A silence annotation unit is used to annotate the first channel audio and the second channel audio in the dual-channel audio respectively to obtain the silence segments in the first channel audio and the silence segments in the second channel audio.
[0035] An audio segmentation unit is used to determine common silence separation points in the dual-channel audio based on silence segments in the first channel audio and silence segments in the second channel audio, and to segment the first channel audio based on the common silence separation points to obtain multiple first segmented audio segments.
[0036] The customer audio determination unit is used to remove the silence segment from each first segmented audio segment to obtain each second segmented audio segment, and to combine the customer audio based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer.
[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the audio segmentation method as described above.
[0038] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio segmentation method as described above.
[0039] The audio segmentation method, apparatus, electronic device, and storage medium provided by this invention determine common silence separation points in the dual-channel audio by identifying silence segments in the first and second channel audio obtained through silence segment annotation. Using these common silence separation points, the first channel audio is segmented, and silence segments are removed from the multiple first segmented audio segments to obtain second segmented audio segments. Customer audio is then combined based on the voiceprint characteristics of each second segmented audio segment to obtain customer audio on a customer-by-customer basis. This overcomes the shortcomings of traditional solutions that can only perform timed segmentation and cannot distinguish the corresponding customer, resulting in low efficiency. This invention achieves customer-by-customer audio segmentation, providing assistance for different service quality inspections and service evaluations. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating the audio segmentation method provided by the present invention;
[0042] Figure 2 This is an example diagram of the two-channel audio provided by the present invention;
[0043] Figure 3This is a schematic diagram of the process for determining customer audio provided by the present invention;
[0044] Figure 4 This is a schematic diagram of the process for determining a suspected audio endpoint provided by the present invention;
[0045] Figure 5 This is an example diagram of a suspected audio endpoint provided by the present invention;
[0046] Figure 6 This is a flowchart illustrating step 120 in the audio segmentation method provided by the present invention;
[0047] Figure 7 This is an example diagram of the silent segment provided by the present invention;
[0048] Figure 8 This is an example diagram of the silence detection status of each audio frame provided by the present invention;
[0049] Figure 9 This is a schematic diagram illustrating the process of determining the common silent separation point provided by the present invention;
[0050] Figure 10 This is a schematic diagram of the process for determining the second segmented audio segment provided by the present invention;
[0051] Figure 11 This is an example diagram of the second audio segmentation provided by the present invention;
[0052] Figure 12 This is a schematic diagram of the adjustment process of the client audio endpoint provided by the present invention;
[0053] Figure 13 This is a comparison diagram of the standard audio endpoint and the client audio endpoint provided by this invention;
[0054] Figure 14 This is a general framework diagram of the audio segmentation method provided by the present invention;
[0055] Figure 15 This is a schematic diagram of the audio segmentation device provided by the present invention;
[0056] Figure 16 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0058] Currently, audio segmentation in service halls typically uses audio pickup hardware. However, this hardware can only perform timed segmentation and audio upload, and cannot distinguish the customer corresponding to the audio. Therefore, this segmentation method is not very effective. Furthermore, since the start and end times of each customer's transaction are unknown, audio cannot be segmented according to the order of the customer's transactions, making it impossible to conduct service quality inspection and evaluation for staff corresponding to different transactions.
[0059] To address this issue, this invention provides an audio segmentation method. The method identifies silent segments in two-channel audio by marking them, segments the audio using common silent delimiters, removes silent segments from the segmented audio, and then combines customer audio using voiceprint features to obtain the customer's audio. This achieves customer-based audio segmentation, providing assistance for different service quality inspections and evaluations. Figure 1 This is a flowchart illustrating the audio segmentation method provided by the present invention, as shown below. Figure 1 As shown, the method includes:
[0060] Step 110: Determine the two-channel audio to be segmented;
[0061] Specifically, before performing audio segmentation, it is first necessary to determine the audio of the business hall to be segmented. In this case, the audio of the business hall is a two-channel audio. Figure 2 This is an example diagram of the two-channel audio provided by the present invention, such as... Figure 2 As shown, the dual-channel audio contains two mono audio channels. Specifically, the dual-channel audio provided by the service center may include customer-side audio and customer service-side audio. Figure 2 The top center is the first audio channel (customer-side audio), and the bottom center is the second audio channel (customer service-side audio).
[0062] The stereo audio to be segmented can be a segment of audio extracted from real-time recorded audio or a segment of audio selected from pre-recorded audio. For example, the duration and recording time of the stereo audio can be preset, and then a segment of audio with the set duration can be extracted from the real-time recorded audio as stereo audio, or an audio segment that meets the set duration at the set recording time can be selected from the pre-recorded audio as stereo audio.
[0063] After obtaining the two-channel audio to be segmented, it is necessary to perform channel separation to separate the audio from the two different channels contained therein, thereby obtaining two mono audios, namely the first channel audio and the second channel audio. It should be noted that the first channel audio and the second channel audio here can be customer-side audio and customer service-side audio, that is, customer channel audio and customer service channel audio, or it can be two different channels of audio separated from the acquired two-channel audio when the two parties are having a conversation. This embodiment of the invention does not make specific limitations on this.
[0064] In addition, the stereo audio to be segmented can be one or more. If there are multiple stereo audio tracks, audio segmentation needs to be performed on each stereo audio track to obtain the customer audio in each stereo audio track, which is based on the customer.
[0065] Step 120: Mark the silence segments in the first and second channels of the stereo audio respectively to obtain the silence segments in the first and second channels of the stereo audio.
[0066] Specifically, after obtaining the two-channel audio to be segmented, the first and second channels of the two-channel audio can be marked with silent segments to identify the silent segments in the first and second channels. The process of marking silent segments can be understood as the process of determining the silent and non-silent segments in the corresponding channel audio. A silent segment is an audio segment in which no one is speaking, while a non-silent segment is the opposite, representing an audio segment in which someone is speaking.
[0067] The determination of silent and non-silent segments depends on the silence detection status of each audio frame in the corresponding audio channel. The silence detection status here is used to indicate whether the corresponding audio frame is silent or not. It can be detected by Voice Activity Detection (VAD) technology. In other words, VAD can be used to perform silence detection on the first and second audio channels respectively to detect the silence points in the first and second audio channels. By combining the silence points and non-silent points, the silent segments in the first and second audio channels can be determined and labeled, thereby obtaining the silent segments in the first audio channel and the silent segments in the second audio channel.
[0068] It should be noted that VAD here can be either energy VAD or model VAD. Energy VAD can perform silence detection based on the frame energy of each audio frame in the corresponding channel audio, while model VAD can perform silence detection based on DNN (Deep Neural Network) features extracted from each audio frame.
[0069] As a preferred embodiment, to ensure the accuracy of the silence segment labeling process for the first and second channel audio, the VAD selected in this embodiment includes both energy VAD and model VAD. That is, silence detection can be performed by energy VAD and model VAD respectively. Then, by combining the silence detection status of each audio frame obtained by energy VAD and the silence detection status of each audio frame obtained by model VAD, the silence segment in the corresponding channel audio is determined. In this way, the silence segment labeling for the two mono audios is completed.
[0070] In this embodiment of the invention, the first and second channels of the dual-channel audio are marked with silence segments, which lays the foundation for subsequent silence segment removal. At the same time, it can also provide data support for determining common silence separation points, thus facilitating the audio segmentation process. In addition, the use of speech endpoint detection technologies at different levels and angles for silence segment marking ensures the accuracy of the silence segment marking process and improves the precision of the silence detection process.
[0071] Step 130: Based on the silence segments in the first channel audio and the silence segments in the second channel audio, determine the common silence separation points in the two-channel audio, and based on the common silence separation points, segment the first channel audio to obtain multiple first segmented audio segments.
[0072] Specifically, in step 120, after determining the silence segments in the first channel audio and the second channel audio, step 130 can be executed. Based on these silence segments, a common silence separation point is determined, and the first channel audio in the dual-channel audio is segmented according to this common silence separation point. This process specifically includes the following steps:
[0073] First, based on the silence segments in the first and second audio channels, common silence separation points can be determined. That is, based on the silence segments in the first and second audio channels, common silence separation points that are simultaneously in silence segments within the two-channel audio can be determined. Specifically, the silence endpoints of each silence segment in the corresponding audio channel can be determined. Then, based on the silence status of each silence endpoint in the other audio channel, common silence separation points can be selected from the silence endpoints. Here, the silence status indicates whether the audio frame corresponding to the silence endpoint is within the silence segment. In other words, the silence endpoints that are simultaneously in the silence segment are selected from the silence endpoints. These silence endpoints are the common silence endpoints, which can be directly used as common silence separation points.
[0074] Further filtering can be performed to obtain the final common silence separation point. That is, the common silence endpoints can be filtered using a preset silence duration to remove common silence endpoints whose silence duration does not meet the preset silence duration. This results in the filtered common silence endpoint, which is the common silence separation point in the desired two-channel audio. Here, the preset silence duration is the duration of the preset silence segment, which can be set according to the actual situation, such as 8 seconds, 10 seconds, 15 seconds, etc.
[0075] Then, the common silence dividing points in the two-channel audio can be used to segment the audio. That is, the first channel audio can be segmented based on the common silence dividing points in the two-channel audio to obtain multiple segmented audio segments. Since this segmentation is the first segmentation of the first channel audio, the multiple audio segments obtained can be called multiple first segmented audio segments. In other words, after segmenting the audio based on the common silence dividing points in the two-channel audio, multiple first segmented audio segments can be obtained.
[0076] In this embodiment of the invention, further screening based on the preset silence duration can ensure the accuracy of the common silence separation point selection process to the greatest extent, improve its selection precision, refine the audio segmentation process based on the common silence separation point, and provide key assistance for the audio segmentation process on a customer-by-customer basis.
[0077] Step 140: Remove the silent segments from each of the first segmented audio segments to obtain each of the second segmented audio segments. Combine the customer audio based on the voiceprint features of each of the second segmented audio segments to obtain customer audio for each customer.
[0078] Specifically, in step 130, after segmenting the first channel audio to obtain multiple first segmented audio segments, step 140 can be executed to remove the silence segments from each first segmented audio segment, and then the client audio is combined based on the second segmented audio segments after the silence segments are removed, so as to obtain client audio on a client-by-client basis. The specific process includes the following steps:
[0079] First, it is necessary to determine the silent segments in each first segmented audio segment. The silent segments can be determined by marking the silent segments in each first segmented audio segment. That is, VAD can be used to perform silence detection on each first segmented audio segment, and the silent segments in each first segmented audio segment can be determined based on the silence detection status of each audio frame.
[0080] Subsequently, the silent segments of each first segmented audio segment can be removed, that is, the silent segments in each first segmented audio segment can be removed. Since the segmentation here is a second segmentation for the first channel audio, the audio segments obtained by the removal can be called each second segmented audio segment.
[0081] Subsequently, based on the voiceprint features of each second segmented audio segment, customer audio can be combined to obtain customer audio per customer. Specifically, voiceprint features can be generated for each second segmented audio segment, and then the second segmented audio segments can be combined according to the similarity between the voiceprint features to obtain customer audio per customer. This process actually involves using the similarity between voiceprint features to perform adjacent segment comparison and cross-segment comparison to find audio segments with short noise from each second segmented audio segment. Through the recombination of audio segments, customer audio per customer is obtained.
[0082] The audio segmentation method provided by this invention determines common silence separation points in the dual-channel audio by identifying silence segments in the first and second channel audio obtained through silence segment annotation. Using these common silence separation points, the first channel audio is segmented, and silence segments are removed from the multiple first segmented audio segments to obtain second segmented audio segments. Customer audio is then combined based on the voiceprint features of each second segmented audio segment to obtain customer audio on a per-customer basis. This overcomes the shortcomings of traditional solutions that can only perform timed segmentation and cannot distinguish the corresponding customer, resulting in low efficiency. This method achieves customer-based audio segmentation, providing assistance for different service quality inspections and service evaluations.
[0083] Based on the above embodiments, Figure 3 This is a schematic diagram of the process for determining customer audio provided by the present invention, as shown below. Figure 3 As shown, customer audio is combined based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer, including:
[0084] Step 310: Combine the second segmented audio segments according to their order in the first channel audio to obtain the client combined audio.
[0085] Step 320: Based on the similarity between the voiceprint features of adjacent second segmented audio segments in the customer combined audio, determine the suspected audio endpoints in the customer combined audio.
[0086] Step 330: Based on the audio duration between adjacent suspected audio endpoints in the customer combined audio and the preset noise duration, the suspected audio endpoints are filtered to obtain the customer audio endpoints, and the customer audio is determined based on the customer audio endpoints.
[0087] Specifically, in step 140, the process of combining customer audio based on the voiceprint features of each second segmented audio to obtain customer audio on a per-customer basis may include:
[0088] Step 310: First, the second segmented audio segments can be combined according to their order in the first channel audio to obtain the customer combined audio. That is, the order of the second segmented audio segments in the first channel audio can be used as a reference to combine the audio segments to obtain the customer combined audio. In other words, the audio segments are spliced according to their order in the first channel audio to obtain the customer combined audio for multiple customers.
[0089] The order of each second segmented audio segment in the first channel audio can be determined by the recording time of each second segmented audio segment, or by the sequence number of each second segmented audio segment, or by the order of each second segmented audio segment obtained during the segmentation process. This embodiment of the invention does not make specific limitations in this regard.
[0090] Step 320: The voiceprint features of each second segmented audio segment in the customer's combined audio can be determined. Voiceprint features can be generated for each second segmented audio segment. Then, neighboring segments can be compared using the voiceprint features to determine the suspected audio endpoints from the audio combination points of each second segmented audio segment. This means that the similarity between the voiceprint features of all two adjacent second segmented audio segments in the customer's audio can be determined, and the suspected audio endpoints can be determined based on the similarity between these voiceprint features. The audio combination points here are the connection points (splicing points) of each second segmented audio segment in the customer's combined audio, and the suspected audio endpoints can be understood as the audio endpoints between customer audios of different customers that are initially determined.
[0091] Since voiceprints are unique, the similarity between voiceprint features can reflect the similarities and differences between customers corresponding to adjacent second-segment audio segments. In other words, if the similarity between voiceprint features reaches a preset similarity, it can be determined that the second-segment audio segments corresponding to these two are the audio of the same customer. Conversely, if the similarity between voiceprint features does not reach the preset similarity, it is determined that the second-segment audio segments corresponding to these two belong to different customers, that is, they correspond to different customers.
[0092] Specifically, in this embodiment of the invention, a cloud-based recognition engine can be used to generate voiceprint features for each of the segmented second audio segments. This recognition engine uses filter bank features to compare the voiceprint features of each of the generated second audio segments. That is, it can extract voiceprints from each of the input second audio segments and compare the voiceprint features of each of the extracted second audio segments to determine the similarity between them.
[0093] Step 330: Subsequently, the suspected audio endpoints in the customer-combined audio can be filtered using a preset noise duration to obtain the customer audio endpoints for each customer. Then, based on these customer audio endpoints, the customer audio for each customer can be determined. Specifically, the suspected audio endpoints can be filtered using the preset noise duration as a benchmark, based on the audio duration between adjacent suspected audio endpoints (the current suspected audio endpoint and the next suspected audio endpoint). That is, the current suspected audio endpoints in the suspected audio endpoints corresponding to the second segmented audio segment with an audio duration less than the preset noise duration are filtered out, thereby obtaining the customer audio endpoints for each customer. Then, based on these customer audio endpoints, the customer audio for each customer can be determined from the customer-combined audio, that is, the customer audio for each customer.
[0094] Correspondingly, if the audio duration between all pairs of adjacent suspected audio endpoints in the customer's combined audio is greater than or equal to the preset noise duration, then there is no need to filter out any suspected audio endpoints, and all suspected audio endpoints are taken as customer audio endpoints.
[0095] The preset noise duration is a pre-defined noise segment duration, which can be set according to actual conditions, such as 6 seconds, 8 seconds, or 10 seconds. Preferably, in this embodiment, the preset noise duration is set to 10 seconds. This means identifying a second segmented audio segment with an audio duration less than 10 seconds between adjacent suspected audio endpoints, then filtering out the earlier suspected audio endpoint from the two adjacent suspected audio endpoints corresponding to this second segmented audio segment, ultimately obtaining the customer's audio endpoint for each customer.
[0096] For example, if there are suspected audio endpoints T+1 and T+2 in the customer's audio combination, and the audio duration between T+1 and T+2 is 8 seconds, while the preset noise duration is 10 seconds, it can be determined that the audio duration between T+1 and T+2 is less than the preset noise duration. Therefore, the earlier suspected audio endpoint T+1 among the adjacent suspected audio endpoints can be filtered out to obtain the customer's audio endpoint T+2.
[0097] Based on the above embodiments, Figure 4 This is a schematic diagram of the process for determining suspected audio endpoints provided by the present invention, as shown below. Figure 4 As shown, step 320 includes:
[0098] Step 321: If the similarity between the voiceprint features of two adjacent second segmented audio segments in the customer combined audio is less than the preset similarity, then the audio combination point of the two adjacent second segmented audio segments is taken as the candidate audio endpoint of the customer combined audio.
[0099] Step 322, otherwise, take the audio combination point of two adjacent second segmented audio segments as the non-candidate audio endpoint of the client combined audio;
[0100] Step 323: Based on the similarity between the voiceprint features of the second segmented audio segments corresponding to adjacent non-candidate audio endpoints in the customer's combined audio, and the preset similarity, the candidate audio endpoints are filtered to obtain the suspected audio endpoints.
[0101] Specifically, in step 320, the process of determining the suspected audio endpoints in the customer combined audio based on the similarity between the voiceprint features of adjacent second segmented audio segments in the customer combined audio may include the following steps:
[0102] First, in step 321, if the similarity between the voiceprint features of two adjacent second segmented audio segments in the customer's combined audio is less than the preset similarity, that is, if the difference between the voiceprint features of these two adjacent second segmented audio segments is large, in other words, if the customers corresponding to these two adjacent second segmented audio segments are different, the audio combination point of these two adjacent second segmented audio segments can be used as a "suspected audio endpoint". However, considering that the "suspected audio endpoint" obtained at this time has not yet eliminated the interference of short noise, that is, there is still a situation where the audio combination point of short noise and customer audio is mistakenly used as a suspected audio endpoint due to multiple cuts caused by short noise. In order to avoid this influence, it is necessary to filter and screen it. Therefore, the "suspected audio endpoint" obtained at this time can be called a candidate audio endpoint.
[0103] Simultaneously, step 322 is executed. If the similarity between the voiceprint features of two adjacent second segmented audio segments in the customer's combined audio is greater than or equal to the preset similarity, that is, if the voiceprint features of these two adjacent second segmented audio segments are extremely similar, in other words, if these two adjacent second segmented audio segments correspond to the same customer, it can be determined that there are no suspected audio endpoints between these two adjacent second segmented audio segments. In other words, the audio combination point of these two adjacent second segmented audio segments is not included in the candidate category of suspected audio endpoints, that is, it is a non-candidate audio endpoint of the customer's combined audio.
[0104] Here, the preset similarity is a pre-set value used to determine the similarity between two second-segmented audio segments represented by the voiceprint features and corresponding to the same person. It can be set according to actual needs, for example, it can be 60%, 75%, 85%, etc., and the embodiments of the present invention do not specifically limit it. However, as a preferred embodiment of the present invention, the preset similarity is set to 60%.
[0105] Then, step 323 is executed, which can perform cross-segment comparison to filter candidate audio endpoints and obtain suspected audio endpoints. Specifically, it can be done by determining the later second segment of the two second segmented audio segments corresponding to each non-candidate audio endpoint in the customer's combined audio, then determining the similarity between the voiceprint features of the later second segmented audio segments corresponding to adjacent non-candidate audio endpoints, and using a preset similarity to filter candidate audio endpoints to obtain suspected audio endpoints. That is, filtering out candidate silent endpoints between two non-candidate silent endpoints corresponding to two second segmented audio segments whose similarity between voiceprint features is greater than or equal to the preset similarity, thereby obtaining suspected silent endpoints.
[0106] Correspondingly, if the similarity between the voiceprint features of the second segmented audio segments corresponding to adjacent non-candidate audio endpoints in the customer's combined audio is greater than or equal to the preset similarity, then there is no need to filter out any candidate audio endpoints, and all candidate audio endpoints are regarded as suspected audio endpoints.
[0107] The following example illustrates the process of identifying suspected audio endpoints:
[0108] Figure 5 This is an example diagram of a suspected audio endpoint provided by the present invention, such as... Figure 5 As shown, the customer's combined audio contains six second-segmented audio segments, namely audio segments one to six. The audio combination points of adjacent audio segments are T-1, T, T+1, T+2, and T+3, respectively. The similarity between the voiceprint features of audio segments one and two, three and four, and four and five is less than the preset similarity of 0.6. The similarity between the voiceprint features of audio segments two and three, and five and six is greater than the preset similarity of 0.6. Therefore, the audio combination point T-1 of audio segments one and two, the audio combination point T+1 of audio segments three and four, and the audio combination point T+2 of audio segments four and five can be identified as candidate audio endpoints of the customer's combined audio. The other audio combination points T and T+3, as well as the audio endpoint T-2, are identified as non-candidate audio endpoints.
[0109] Furthermore, the similarity between the voiceprint features of the second segmented audio segments (audio segment 3 and audio segment 6) corresponding to adjacent non-candidate audio endpoints T and T+3 is less than the preset similarity, so candidate audio endpoints T+1 and T+2 between non-candidate audio endpoints T and T+3 are retained; while the similarity between the voiceprint features of the second segmented audio segments (audio segment 1 and audio segment 3) corresponding to adjacent non-candidate audio endpoints T-2 and T is greater than the preset similarity, so candidate audio endpoint T-1 between non-candidate audio endpoints T-2 and T is filtered out, and finally, suspected audio endpoints T+1 and T+2 are obtained.
[0110] Based on the above embodiments, Figure 6 This is a flowchart illustrating step 120 of the audio segmentation method provided by the present invention, as shown below. Figure 6 As shown, step 120 includes:
[0111] Step 121: Determine the frame energy of each audio frame in the first and second audio channels. Based on the frame energy of each audio frame and the energy threshold value, determine the silence detection state of each audio frame. The energy threshold value is determined based on the corresponding audio channel.
[0112] Step 122: Based on the number of audio frames contained in the audio window of the first and second audio channels, and the silence detection status of each audio frame, determine the silence segments in the first and second audio channels.
[0113] Specifically, in step 120, the process of marking the first and second channels of the two-channel audio to obtain the silent segments in the first and second channels of the audio may include the following steps:
[0114] Step 121: First, silence detection can be performed using energy VAD to determine the frame energy of each audio frame in the first and second audio channels. Then, the silence and non-silent segments can be divided based on the frame energy of each audio frame. Specifically, the silence detection can be completed using the frame energy of each audio frame as a benchmark, combined with energy threshold values, through a state machine and smoothing process, thereby obtaining the silence detection state of each audio frame, i.e., determining whether each audio frame is silent or not. The energy threshold values can be determined based on the corresponding channel audio, i.e., based on the corresponding channel audio, the four energy threshold algorithms can be used to determine the four energy threshold values of the corresponding channel audio.
[0115] It is worth noting that, in this embodiment of the invention, in addition to using energy VAD to determine the silence detection state of each audio frame in the first and second channel audio as described above, the silence detection state of each audio frame can also be determined based on model VAD. That is, features can be extracted from each audio frame in the first and second channel audio using MLP (Multilayer Perceptron) to obtain DNN features of a specific dimension (75*11 dimensions). Then, the extracted DNN features of the specific dimension are input into the VAD model. The VAD model determines the silence detection state of each audio frame based on the DNN features of each audio frame, whether it is 0 or 1. Here, 0 represents Speech and 1 represents No Speech.
[0116] The VAD model can be trained using scientifically proportioned silent and non-silent corpora. The trained VAD model can also comprehensively consider the silence detection status of each audio frame in the corresponding audio channel, thereby directly determining the silent segment in the corresponding audio channel.
[0117] Step 122 determines the window length of the audio window in the corresponding channel audio. This window length can be understood as the number of audio frames contained within the audio window. Then, based on the number of audio frames contained in the audio windows of the first and second channel audio, and the silence detection status of each audio frame determined based on the energy VAD and model VAD, the silence segments in the first channel audio and the silence segments in the second channel audio are identified. Figure 7 This is an example diagram of the silent segment provided by the present invention.
[0118] The following example illustrates the process of marking silent segments:
[0119] Figure 8 This is an example diagram of the silence detection state of each audio frame provided by the present invention, such as... Figure 8 As shown, the audio window in the first or second channel audio contains 21 audio frames. In this case, the rule for determining the silence segment can be that the number of audio frames in the audio window with a silence detection state of 0 is greater than or equal to a preset number of frames. The preset number of frames can be 13, 15, 17, etc. In this embodiment of the invention, the preset number of frames is determined to be 15. That is, when the number of audio frames in the audio window with a silence detection state of 1 is greater than or equal to 15, it is determined to be Speech Start. Correspondingly, when the number of audio frames in the audio window with a silence detection state of 0 is greater than or equal to 15, it is determined to be Speech End.
[0120] It should be noted that during the process of marking silent segments, the length (size) of the audio window must remain unchanged, and its length can be set according to actual needs.
[0121] Based on the above embodiments, the frame energy of an audio frame can be calculated using the following formula:
[0122]
[0123] In the formula, x j E represents the amplitude of the sampling point. i N represents the frame energy, N represents the frame length, j represents the j-th audio frame, and C is the minimum threshold value, which is usually a constant. Setting C can prevent the frame energy from falling below 0.
[0124] The calculation of the four energy thresholds can be divided into the following four cases:
[0125] Firstly, since most of the audio in the corresponding channel is background noise, the energy value E of the background noise can be obtained through the first clustering method. Noise The formulas for calculating the four energy thresholds are: K i =E Noise +a i (i = 1, 2, 3, 4), select C1Centroid as E Noise ;
[0126] K1 = C1Centroid + 2.0
[0127] K2 = C1Centroid + 5.0
[0128] K3 = C1Centroid + 3.0
[0129] K4 = C1Centroid + 8.0
[0130] In the formula, C1Centroid represents the centroid of the first clustering method, and K1, K2, K3 and K4 are four energy threshold values.
[0131] Secondly, in the corresponding audio channel, human voices are more prominent, and the energy of the human voice segment is significantly higher than that of the noise segment. In this case, the energy value E of the background noise can be obtained through the second clustering method. Noise And the energy value E of the human voice segment Voice The formulas for calculating the four energy thresholds are: K i =E Noise +(E Voice -E Noise )*a i (i = 1, 2, 3, 4), select C2Centroid[0] as E Noise :
[0132] K1 = C2Centroid[0] + Value * 0.1
[0133] K2 = C2Centroid[0] + Value * 0.3
[0134] K3 = C2Centroid[0] + Value * 0.2
[0135] K4 = C2Centroid[0] + Value * 0.6
[0136] In the formula, C2Centroid[0] represents the centroid of the second clustering method, Value represents the loss function of the centroid, Value = C2Centroid[1] - C2Centroid[0], where C2Centroid[1] represents the expected value and C2Centroid[0] represents the actual value.
[0137] Third, the energy difference between the human voice segment and the noise segment in the corresponding audio channel is not significant. At this time, C1Centroid-C2Centroid[0]>Value*0.2, and the calculation formula for the four energy threshold values is: K i =E Noise +a i (i = 1, 2, 3, 4), select C2Centroid[0] as E Noise :
[0138] K1 = C2Centroid[0] + 1.0
[0139] K2 = C2Centroid[0] + 4.0
[0140] K3 = C2Centroid[0] + 2.0
[0141] K4 = C2Centroid[0] + 8.0
[0142] Fourth, when the above three conditions are not met, the formula for calculating the four energy thresholds is: K i =E Noise +a i (i = 1, 2, 3, 4), select C1Centroid as E Noise :
[0143] K1 = C1Centroid + 1.0
[0144] K2 = C1Centroid + 4.0
[0145] K3 = C1Centroid + 2.0
[0146] K4 = C1Centroid + 8.0
[0147] Based on the above embodiments, Figure 9 This is a schematic diagram illustrating the process of determining the common silent separation point provided by the present invention, as shown below. Figure 9 As shown, based on the silence segments in the first channel audio and the silence segments in the second channel audio, common silence separation points in the two-channel audio are determined, including:
[0148] Step 910: Determine the mute endpoint of the mute segment in the first audio channel and the mute endpoint of the mute segment in the second audio channel.
[0149] Step 920: Select a common silence endpoint from the silence endpoints of the silence segments in the first channel audio and the silence endpoints of the silence segments in the second channel audio. The audio frames corresponding to the common silence endpoints in the first channel audio and the second channel audio are both within the silence segment.
[0150] Step 930: Based on the silence duration of the common silence endpoint in the corresponding channel audio and the preset silence duration, filter the common silence endpoint to obtain the common silence separation point.
[0151] Specifically, the process of determining the common silence separation point in the two-channel audio based on the silence segments in the first channel audio and the silence segments in the second channel audio can include:
[0152] First, by executing step 910, the silence endpoints in the first audio channel and the second audio channel can be determined based on the silence segments in the first audio channel and the silence segments in the second audio channel, that is, the audio endpoints of each silence segment in the first audio channel and the second audio channel can be determined.
[0153] Then, step 920 is executed, selecting the corresponding audio frames from the silence endpoints of the silence segments in the first channel audio and the silence endpoints of the silence segments in the second channel audio. These silence endpoints are the common silence endpoints, which can also be understood as the audio frames corresponding to the common silence endpoints in both the first channel audio and the second channel audio being within the silence segment.
[0154] Subsequently, by executing step 930, common silence separation points can be obtained by filtering based on the common silence endpoints using a preset silence duration. Specifically, common silence endpoints can be filtered based on the silence duration of the corresponding silence segment in the corresponding channel audio and the preset silence duration to obtain common silence separation points. This can be achieved by filtering out common silence endpoints whose silence duration in the corresponding silence segment in the corresponding channel audio does not meet the preset silence duration, thereby obtaining the filtered common silence endpoints. These common silence endpoints are the desired common silence separation points in the two-channel audio.
[0155] Here, the preset silence duration is the duration of a pre-set silence segment, which can be set according to actual conditions, such as 8 seconds, 10 seconds, 15 seconds, etc. Preferably, in this embodiment of the invention, the preset silence duration is set to 10 seconds, that is, common silence endpoints with a silence duration greater than or equal to 10 seconds can be selected from the common silence endpoints as common silence separation points.
[0156] Based on the above embodiments, Figure 10 This is a schematic diagram of the process for determining the second segmented audio segment provided by the present invention, as shown below. Figure 10 As shown, the silence segments are removed from each of the first segmented audio segments to obtain each of the second segmented audio segments, including:
[0157] Step 1010: Perform silent segment removal on each of the first segmented audio segments to obtain each silent-removed audio segment;
[0158] Step 1020: Based on the audio duration of each silent cut-off audio segment and the preset audio duration, perform audio filtering to obtain each second segmented audio segment.
[0159] Specifically, the process of removing the silent segments from each of the first segmented audio segments to obtain each of the second segmented audio segments can include the following steps:
[0160] Step 1010: First, the multiple first-segment audio segments obtained by audio segmentation using common silence dividing points can be subjected to silence segment removal, that is, the silence segments in each first-segment audio segment can be removed to obtain the silence-removed audio segments, i.e., each silence-removed audio segment; wherein, the silence segments in each first-segment audio segment can be determined by silence segment annotation, the process of silence segment annotation has been explained in detail above, and will not be repeated here;
[0161] Step 1020: Subsequently, each silent segmented audio segment can be filtered using a preset audio duration to obtain each second segmented audio segment. Specifically, the audio duration of each silent segmented audio segment can be used as a benchmark to filter audio using a preset audio duration, filtering out silent segments whose audio duration does not reach the preset audio duration, thereby obtaining each filtered second segmented audio segment.
[0162] Figure 11 This is an example diagram of the second audio segmentation provided by the present invention, such as... Figure 11 As shown, segments 1, 2, 3, 5, 7, 9, and 11 are the audio segments obtained after the silent segments are removed. Among them, segments 1, 2, and 7 were filtered out during the audio filtering process because their audio duration was less than the preset audio duration. Segments 3, 5, 9, and 11, which were not filtered out, are the second segmented audio segments.
[0163] The preset audio duration here can be set according to the actual situation. For example, it can be 2 seconds, 3 seconds, 5 seconds, etc. In a preferred embodiment of the present invention, the preset audio duration is set to 2 seconds. That is, the silent cut audio with an audio duration of 2 seconds or more is selected from each silent cut audio segment as the second segmented audio. In other words, silent cut audio segments with an audio duration of less than 2 seconds are filtered out.
[0164] Based on the above embodiments, Figure 12 This is a schematic diagram of the adjustment process of the client audio endpoint provided by the present invention, as shown below. Figure 12 As shown, customer audio is combined based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer, which then includes:
[0165] Step 1210: Determine the standard audio endpoint based on the second channel audio;
[0166] Step 1220: Determine the audio segmentation accuracy based on the standard audio endpoint and client audio;
[0167] Step 1230: Adjust the client audio endpoint based on the audio segmentation accuracy.
[0168] Specifically, after obtaining customer audio data on a per-customer basis through the above steps, to ensure the accuracy of the customer audio, the audio segmentation process of the first channel audio can be verified to check its accuracy rate, thereby obtaining the audio segmentation accuracy rate. Based on this audio segmentation accuracy rate, the customer audio endpoints in the audio segmentation process can be corrected. This process can specifically include the following steps:
[0169] First, by executing step 1210, the standard audio segmentation point can be determined based on the second channel audio in the dual-channel audio. That is, the information contained in the second channel audio can be used to determine the standard audio segmentation point in the audio segmentation process. Specifically, since staff usually use common phrases or symbolic language to indicate the start or end of the business transaction when handling business for customers, such as "What business do you need?", "Your business has been completed", "Please rate my service", "You're welcome", "Welcome to visit again", "Take care", etc., the second channel audio can be turned on and the information contained therein can be used to determine the start and end times of each customer's business transaction. That is, the start and end times of the customer audio for each customer can be determined and manually marked to form the standard audio endpoints of the customer audio.
[0170] It should be noted that the start and end times should be marked in the format of "0:00:00". In other words, the start and end times of each customer's audio should be marked in the format of "0:00:00". The table below shows the standard audio endpoints for manual marking:
[0171]
[0172]
[0173] When manually annotating, in addition to marking the start and end times, you can also add notes on the general situation of the audio, such as the background noise in the audio, the noise segments in the audio, etc., thus obtaining audio notes.
[0174] Corresponding to the manually labeled standard audio endpoints, the customer audio endpoints of each customer determined through the above steps in this embodiment of the invention can be represented as shown in the table below:
[0175]
[0176] Then, step 1220 is executed, which can determine the audio segmentation accuracy rate based on the standard audio endpoints and customer audio. That is, the standard audio endpoints marked by humans and the customer audio endpoints of each customer determined in the audio segmentation process can be compared to obtain the audio segmentation accuracy rate of the audio segmentation process. This audio segmentation accuracy rate can reflect the accuracy rate of the audio segmentation process.
[0177] Figure 13 This is a comparison diagram of the standard audio endpoint and the client audio endpoint provided by this invention, as shown below. Figure 13 As shown, if the time difference between the standard audio endpoint and the client audio endpoint is within 10 seconds, it indicates that the client audio endpoint is accurately determined and can be marked in green to indicate that they are consistent. If the time difference between the standard audio endpoint and the client audio endpoint exceeds 2 minutes, it is confirmed that the client audio endpoint is incorrectly determined. In other words, the audio segmentation corresponding to the client audio endpoint is incorrect, so it can be marked in red to indicate an incorrect segmentation warning. If a standard audio endpoint exists but no corresponding client audio endpoint exists, it can be determined that the corresponding client audio endpoint is missing, that is, there is a missed segmentation in the audio segmentation process. In this case, it can be marked in yellow to indicate a missed segmentation prompt. Correspondingly, if a client audio endpoint exists but no corresponding standard audio endpoint exists, it can be determined that there is an extra client audio endpoint, that is, there is multiple segmentation in the audio segmentation process. In this case, it can be marked in purple to indicate a multiple segmentation prompt.
[0178] The audio segmentation accuracy can be determined based on the colors marked during the comparison process, i.e., audio segmentation accuracy = green / (green + red + yellow + purple) × 100%.
[0179] Then, step 1230 can be executed, adjusting the audio segmentation process according to the audio segmentation accuracy, so that the customer audio endpoints of each customer in the audio segmentation process can correspond to the standard audio endpoints marked by humans, and the time difference between the two is as small as possible. Even if the customer audio endpoints are infinitely close to the standard audio endpoints marked by humans, the verification of the audio segmentation process is completed, the audio segmentation accuracy is verified, and the feedback adjustment based on the audio segmentation accuracy is realized, ensuring the accuracy of the audio segmentation process.
[0180] Based on the above embodiments, Figure 14 This is a general framework diagram of the audio segmentation method provided by the present invention, as shown below. Figure 14 As shown, the overall process of audio segmentation includes the following steps:
[0181] First, determine the two-channel audio to be segmented;
[0182] Subsequently, the first and second channels of the two-channel audio can be marked with silence segments to obtain the silence segments in the first and second channels. Specifically, this can be done by determining the frame energy of each audio frame in the first and second channels, determining the silence detection state of each audio frame based on the frame energy and energy threshold value, where the energy threshold value is determined based on the corresponding channel audio, and determining the silence segments in the first and second channels based on the number of audio frames contained in the audio window in the first and second channels and the silence detection state of each audio frame.
[0183] Subsequently, based on the silence segments in the first and second channel audio, common silence separation points in the two-channel audio are determined. Specifically, this can be done by determining the silence endpoints of the silence segments in the first and second channel audio; selecting common silence endpoints from the silence endpoints of the silence segments in the first and second channel audio, where the corresponding audio frames of the common silence endpoints in both the first and second channel audio are within the silence segments; and filtering the common silence endpoints based on the silence duration of the corresponding silence segments in the corresponding channel audio and a preset silence duration to obtain common silence separation points.
[0184] Subsequently, based on the common silence dividing point, the first channel audio can be segmented to obtain multiple first segmented audio segments, and silence segments can be removed from each first segmented audio segment to obtain each second segmented audio segment. Specifically, this process can be as follows: silence segments can be removed from each first segmented audio segment to obtain each silence-removed audio segment; audio filtering can be performed based on the audio duration of each silence-removed audio segment and a preset audio duration to obtain each second segmented audio segment.
[0185] Finally, customer audio can be combined based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer. Specifically, the second segmented audio segments can be combined in the order of the first channel audio to obtain customer combined audio. Based on the similarity between the voiceprint features of adjacent second segmented audio segments in the customer combined audio, suspected audio endpoints in the customer combined audio are determined. Based on the audio duration between adjacent suspected audio endpoints in the customer combined audio and the preset noise duration, suspected audio endpoints are filtered to obtain customer audio endpoints. Based on the customer audio endpoints, customer audio per customer is determined.
[0186] The process of determining suspected audio endpoints in a customer-combined audio based on the similarity between the voiceprint features of adjacent second-segmented audio segments in the customer-combined audio may include the following steps: if the similarity between the voiceprint features of two adjacent second-segmented audio segments in the customer-combined audio is less than a preset similarity, then the audio combination point of the two adjacent second-segmented audio segments is taken as a candidate audio endpoint of the customer-combined audio; otherwise, the audio combination point of the two adjacent second-segmented audio segments is taken as a non-candidate audio endpoint of the customer-combined audio; based on the similarity between the voiceprint features of the second-segmented audio segments corresponding to the adjacent non-candidate audio endpoints in the customer-combined audio, and the preset similarity, the candidate audio endpoints are filtered to obtain suspected audio endpoints.
[0187] After that, the standard audio endpoint can be determined based on the second channel audio; the audio segmentation accuracy can be determined based on the standard audio endpoint and the client audio; and the client audio endpoint can be adjusted based on the audio segmentation accuracy.
[0188] It is worth noting that in the process of determining the suspected audio endpoints in the customer combined audio based on the similarity between the voiceprint features of adjacent second segmented audio segments in the customer combined audio, the voiceprint comparison can be a 1:1 scenario or a 1:N scenario, and this invention does not specifically limit this.
[0189] The 1:1 scenario is used to confirm the target, that is, to compare a certain voiceprint in two audio segments to confirm whether it is the same person's voiceprint. The comparison result is reflected in the form of a score. In other words, if the score is higher than the threshold, it is determined to be true, that is, the same person's original voiceprint; otherwise, it is false.
[0190] In a 1:1 scenario, there are also false alarm rates and missed alarm rates.
[0191] The false rejection rate (FRR) is the rate at which a correct target is rejected. Its calculation formula is shown below:
[0192]
[0193] False alarm rate, also known as false acceptance rate (FAR), refers to the number of times a non-target is allowed to pass through; in other words, the number of times a person impersonating a target is allowed to pass through. Its calculation formula is as follows:
[0194]
[0195] Both the false alarm rate and the missed alarm rate are related to the threshold. The higher the threshold, the higher the missed alarm rate, but the lower the false alarm rate; conversely, the lower the threshold, the lower the missed alarm rate and the higher the false alarm rate.
[0196] The 1:N scenario is used for large-scale data retrieval and recall. It uses a specific audio segment as a benchmark, compares its voiceprint features with those of other audio segments, and returns the N closest results, i.e., the Top N. If the Top N contains the actual target, the recall is considered successful; otherwise, it fails. The recall rate is calculated using the following formula:
[0197]
[0198] The method provided in this invention identifies common silence separation points in the dual-channel audio by marking silence segments in the first and second channel audio. Using these common silence separation points, the first channel audio is segmented, and silence segments are removed from the resulting multiple first segmented audio segments to obtain second segmented audio segments. Customer audio is then combined based on the voiceprint features of each second segmented audio segment to obtain customer audio for each customer. This overcomes the shortcomings of traditional solutions that only perform timed segmentation and cannot distinguish the corresponding customer, resulting in low efficiency. The method achieves customer-based audio segmentation, providing assistance for different service quality inspections and service evaluations.
[0199] The audio segmentation device provided by the present invention is described below. The audio segmentation device described below can be referred to in correspondence with the audio segmentation method described above.
[0200] Figure 15 This is a schematic diagram of the audio segmentation device provided by the present invention, as shown below. Figure 15 As shown, the device includes:
[0201] The audio determination unit 1510 is used to determine the two-channel audio to be segmented;
[0202] The silence labeling unit 1520 is used to label the first channel audio and the second channel audio in the dual-channel audio respectively to obtain the silence segments in the first channel audio and the silence segments in the second channel audio.
[0203] The audio segmentation unit 1530 is used to determine common silence separation points in the dual-channel audio based on silence segments in the first channel audio and silence segments in the second channel audio, and to segment the first channel audio based on the common silence separation points to obtain multiple first segmented audio segments.
[0204] The customer audio determination unit 1540 is used to remove the silence segment from each of the first segmented audio segments to obtain each of the second segmented audio segments, and to combine the customer audio based on the voiceprint features of each of the second segmented audio segments to obtain customer audio per customer.
[0205] The audio segmentation device provided by this invention determines common silence separation points in the dual-channel audio by identifying silence segments in the first and second channel audio obtained through silence segment annotation. Using these common silence separation points, the first channel audio is segmented, and silence segments are removed from the multiple first segmented audio segments to obtain second segmented audio segments. Customer audio is then combined based on the voiceprint characteristics of each second segmented audio segment to obtain customer audio on a per-customer basis. This overcomes the shortcomings of traditional solutions that can only perform timed segmentation and cannot distinguish the corresponding customer, resulting in low efficiency. This device achieves customer-based audio segmentation, providing assistance for different service quality inspections and service evaluations.
[0206] Based on the above embodiments, the customer audio determination unit 1540 is used for:
[0207] The second segmented audio segments are combined according to their order in the first channel audio to obtain the client combined audio.
[0208] Based on the similarity between the voiceprint features of adjacent second segmented audio segments in the customer combined audio, suspected audio endpoints in the customer combined audio are determined;
[0209] Based on the audio duration between adjacent suspected audio endpoints in the customer's combined audio and the preset noise duration, the suspected audio endpoints are filtered to obtain customer audio endpoints, and customer audio is determined on a customer-by-customer basis based on the customer audio endpoints.
[0210] Based on the above embodiments, the customer audio determination unit 1540 is used for:
[0211] If the similarity between the voiceprint features of two adjacent second segmented audio segments in the customer combined audio is less than a preset similarity, then the audio combination point of the two adjacent second segmented audio segments is taken as the candidate audio endpoint of the customer combined audio.
[0212] Otherwise, the audio combination point of two adjacent second segmented audio segments is taken as the non-candidate audio endpoint of the client combined audio;
[0213] Based on the similarity between the voiceprint features of the second segmented audio segments corresponding to adjacent non-candidate audio endpoints in the customer's combined audio, and a preset similarity, the candidate audio endpoints are filtered to obtain the suspected audio endpoints.
[0214] Based on the above embodiments, the silence labeling unit 1520 is used for:
[0215] The frame energy of each audio frame in the first channel audio and the second channel audio is determined. Based on the frame energy of each audio frame and the energy threshold value, the silence detection state of each audio frame is determined. The energy threshold value is determined based on the corresponding channel audio.
[0216] Based on the number of audio frames contained in the audio windows of the first and second audio channels, and the silence detection status of each audio frame, the silence segments in the first and second audio channels are determined.
[0217] Based on the above embodiments, the audio segmentation unit 1530 is used for:
[0218] Determine the mute endpoint of the mute segment in the first channel audio and the mute endpoint of the mute segment in the second channel audio.
[0219] A common silence endpoint is selected from the silence endpoints of the silence segments in the first channel audio and the silence endpoints of the silence segments in the second channel audio. The common silence endpoint corresponds to an audio frame within a silence segment in both the first channel audio and the second channel audio.
[0220] Based on the silence duration of the common silence endpoint in the corresponding channel audio and the preset silence duration, the common silence endpoint is filtered to obtain the common silence separation point.
[0221] Based on the above embodiments, the customer audio determination unit 1540 is used for:
[0222] Each first segmented audio segment is subjected to a silent segment removal process to obtain each silent-removed audio segment.
[0223] Based on the audio duration of each silent cut-off audio segment and a preset audio duration, audio filtering is performed to obtain each second segmented audio segment.
[0224] Based on the above embodiments, the device further includes a client audio endpoint adjustment unit, used for:
[0225] Based on the second channel audio, determine the standard audio endpoint;
[0226] Based on the standard audio endpoint and the client audio, determine the audio segmentation accuracy;
[0227] Adjust the client audio endpoint based on the audio segmentation accuracy.
[0228] Figure 16 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 16 As shown, the electronic device may include: a processor 1610, a communications interface 1620, a memory 1630, and a communication bus 1640, wherein the processor 1610, the communications interface 1620, and the memory 1630 communicate with each other through the communication bus 1640. The processor 1610 can call logic instructions in the memory 1630 to execute an audio segmentation method, which includes: determining the two-channel audio to be segmented; marking the first channel audio and the second channel audio in the two-channel audio with silence segments respectively, to obtain silence segments in the first channel audio and silence segments in the second channel audio; determining common silence separation points in the two-channel audio based on the silence segments in the first channel audio and the silence segments in the second channel audio, and segmenting the first channel audio based on the common silence separation points to obtain multiple first segmented audio segments; removing silence segments from each first segmented audio segment to obtain each second segmented audio segment; and combining customer audio based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer.
[0229] Furthermore, the logical instructions in the aforementioned memory 1630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0230] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the audio segmentation method provided by the above methods, the method comprising: determining the two-channel audio to be segmented; marking the first channel audio and the second channel audio in the two-channel audio with silence segments respectively, to obtain silence segments in the first channel audio and silence segments in the second channel audio; determining common silence separation points in the two-channel audio based on the silence segments in the first channel audio and silence segments in the second channel audio, and segmenting the first channel audio based on the common silence separation points to obtain a plurality of first segmented audio segments; removing silence segments from each first segmented audio segment to obtain each second segmented audio segment; and combining customer audio based on the voiceprint features of each second segmented audio segment to obtain customer audio on a customer-by-customer basis.
[0231] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the audio segmentation method provided by the above methods. The method includes: determining a two-channel audio to be segmented; marking silence segments in the first and second channels of the two-channel audio respectively to obtain silence segments in the first and second channels; determining common silence separation points in the two-channel audio based on the silence segments in the first and second channels, and segmenting the first channel audio based on the common silence separation points to obtain multiple first segmented audio segments; removing silence segments from each first segmented audio segment to obtain second segmented audio segments; and combining customer audio based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer.
[0232] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0233] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0234] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio segmentation method, characterized in that, include: Identify the two-channel audio to be segmented; The first and second channels of the stereo audio are marked with silent segments to obtain the silent segments in the first channel audio and the silent segments in the second channel audio. Based on the silence segments in the first channel audio and the silence segments in the second channel audio, a common silence separation point in the dual-channel audio is determined, and based on the common silence separation point, the first channel audio is segmented to obtain multiple first segmented audio segments; The silence segments are removed from each of the first segmented audio segments to obtain each of the second segmented audio segments. Based on the similarity between the voiceprint features of each of the second segmented audio segments, the second segmented audio segments are compared and combined to obtain customer audio for each customer. The step of determining the common silence separation point in the two-channel audio based on the silence segments in the first channel audio and the silence segments in the second channel audio includes: Determine the mute endpoint of the mute segment in the first channel audio and the mute endpoint of the mute segment in the second channel audio. A common silence endpoint is selected from the silence endpoints of the silence segments in the first channel audio and the silence endpoints of the silence segments in the second channel audio. The common silence endpoint corresponds to an audio frame within a silence segment in both the first channel audio and the second channel audio. Based on the common mute endpoints, the common mute separation points are determined.
2. The audio segmentation method according to claim 1, characterized in that, The method involves comparing and combining the second segmented audio segments based on the similarity of their voiceprint features to obtain customer audio for each customer, including: According to the order of each of the second segmented audio segments in the first channel audio, the second segmented audio segments are combined to obtain the customer combined audio; Based on the similarity between the voiceprint features of adjacent second segmented audio segments in the customer combined audio, suspected audio endpoints in the customer combined audio are determined; Based on the audio duration between adjacent suspected audio endpoints in the customer's combined audio and the preset noise duration, the suspected audio endpoints are filtered to obtain customer audio endpoints, and customer audio is determined on a customer-by-customer basis based on the customer audio endpoints.
3. The audio segmentation method according to claim 2, characterized in that, The step of determining suspected audio endpoints in the customer-combined audio based on the similarity between the voiceprint features of adjacent second-segmented audio segments in the customer-combined audio includes: If the similarity between the voiceprint features of two adjacent second segmented audio segments in the customer combined audio is less than a preset similarity, then the audio combination point of the two adjacent second segmented audio segments is taken as the candidate audio endpoint of the customer combined audio. Otherwise, the audio combination point of two adjacent second segmented audio segments is taken as the non-candidate audio endpoint of the client combined audio; Based on the similarity between the voiceprint features of the second segmented audio segments corresponding to adjacent non-candidate audio endpoints in the customer's combined audio, and a preset similarity, the candidate audio endpoints are filtered to obtain the suspected audio endpoints.
4. The audio segmentation method according to any one of claims 1 to 3, characterized in that, The step of marking the first and second channel audio in the dual-channel audio respectively with silence segments to obtain the silence segments in the first channel audio and the silence segments in the second channel audio includes: The frame energy of each audio frame in the first channel audio and the second channel audio is determined. Based on the frame energy of each audio frame and the energy threshold value, the silence detection state of each audio frame is determined. The energy threshold value is determined based on the corresponding channel audio. Based on the number of audio frames contained in the audio windows of the first and second audio channels, and the silence detection status of each audio frame, the silence segments in the first and second audio channels are determined.
5. The audio segmentation method according to any one of claims 1 to 3, characterized in that, The step of determining the common silence separation point based on the common silence endpoint includes: Based on the silence duration of the common silence endpoint in the corresponding channel audio and the preset silence duration, the common silence endpoint is filtered to obtain the common silence separation point.
6. The audio segmentation method according to any one of claims 1 to 3, characterized in that, The step of removing the silence segment from each of the first segmented audio segments to obtain each of the second segmented audio segments includes: Each first segmented audio segment is subjected to a silent segment removal process to obtain each silent-removed audio segment. Based on the audio duration of each silent cut-off audio segment and a preset audio duration, audio filtering is performed to obtain each second segmented audio segment.
7. The audio segmentation method according to any one of claims 1 to 3, characterized in that, The process of combining customer audio based on the voiceprint features of each second segmented audio segment to obtain customer audio per customer unit further includes: Based on the second channel audio, determine the standard audio endpoint; Based on the standard audio endpoint and the client audio, determine the audio segmentation accuracy; Adjust the client audio endpoint based on the audio segmentation accuracy.
8. An audio segmentation device, characterized in that, include: An audio determination unit is used to determine the two-channel audio to be segmented; A silence annotation unit is used to annotate the first channel audio and the second channel audio in the dual-channel audio respectively to obtain the silence segments in the first channel audio and the silence segments in the second channel audio. An audio segmentation unit is used to determine common silence separation points in the dual-channel audio based on silence segments in the first channel audio and silence segments in the second channel audio, and to segment the first channel audio based on the common silence separation points to obtain multiple first segmented audio segments. The customer audio determination unit is used to remove the silence segment from each first segmented audio segment to obtain each second segmented audio segment. Based on the similarity between the voiceprint features of each second segmented audio segment, the second segmented audio segments are compared and combined to obtain customer audio for each customer. The step of determining the common silence separation point in the two-channel audio based on the silence segments in the first channel audio and the silence segments in the second channel audio includes: Determine the mute endpoint of the mute segment in the first channel audio and the mute endpoint of the mute segment in the second channel audio. A common silence endpoint is selected from the silence endpoints of the silence segments in the first channel audio and the silence endpoints of the silence segments in the second channel audio. The common silence endpoint corresponds to an audio frame within a silence segment in both the first channel audio and the second channel audio. Based on the common mute endpoints, the common mute separation points are determined.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the audio segmentation method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the audio segmentation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
CN111131616A
CN111508498A