Method, apparatus and computer-readable storage medium for quickly extracting telephone channel data
By performing effective audio segmentation and voiceprint feature comparison of mono telephone channel data, audio data of each speaker is extracted clustered, solving the problem of low data acquisition efficiency in the prior art, and automatic fast and accurate data extraction is achieved.
Patent Information
- Application Number
- CN202111469131.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-03
AI Technical Summary
The prior art is difficult to quickly and efficiently extract audio data of each speaker from mono telephone channel data, resulting in low data acquisition efficiency and high cost.
By obtaining the effective audio in the channel data to be extracted, dividing it into several segments, and by comparing the voiceprint characteristics of the adjacent segments, the clip belonging to the first speaker and the speaker change segment are determined, and the effective audio is then clustered to obtain the effective audio of each speaker.
It realizes automatic and rapid extraction of telephone channel data without manual intervention, significantly improving data extraction efficiency and ensuring the accuracy of extracted audio data.
Smart Images

Figure CN114333842B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer information technology, and in particular to a method and apparatus for quickly extracting telephone channel data and a computer-readable storage medium. Background Art
[0002] The field of artificial intelligence usually requires a large amount of data for support. For example, training samples, input layer data, etc. all need to be preprocessed including collection, cleaning, classification, marking, construction, etc. These preprocessing tasks are mainly achieved manually. Therefore, whether it is the training of algorithm models or the data required for model applications, a large amount of human cost needs to be invested and it is time-consuming.
[0003] Similar problems will also be encountered in the application of artificial intelligence technology in telephone calls. Telephone channel data is mainly divided into two categories: one is stereo data, where the two channels store the audio of the speakers at both ends of the call respectively, and the other is mono data, where the audio of the two speakers is placed in the same channel. In the scenario of a two-person call, for the first case, the audio data of the two speakers can be directly extracted through channel separation. However, for the second case, currently only the method of manual marking can be used, that is, the audio is segmented, recognized and then synthesized to achieve the acquisition of the audio data of different speakers. And mono data accounts for a large proportion in actual applications. How to quickly extract the audio data of each speaker in mono data is an urgent problem to be solved. Summary of the Invention
[0004] In view of the above problems, an embodiment of this application provides a method for quickly extracting telephone channel data. The method includes the steps of: obtaining valid audio in the channel data to be extracted, where the channel data to be extracted is the channel data collected during a call of at least two people; segmenting the valid audio into a plurality of segments; determining the segments belonging to the first speaker and the speaker change segments by comparing the voiceprint features of adjacent segments; and clustering the segments belonging to the first speaker in the valid audio according to the segments of the first speaker and the change segments to obtain the valid audio of the first speaker.
[0005] In one embodiment, the step of segmenting the valid audio into a plurality of segments includes sequentially segmenting the valid audio based on the data length or time length to obtain each segment of a fixed length.
[0006] In one embodiment, there is partial identical data in two adjacent segments.
[0007] In one implementation, determining the segments belonging to the first speaker and the speaker change segments by comparing the voiceprint features of two adjacent segments includes: determining the voiceprint feature of the first speaker; sequentially identifying each of the segments based on a voiceprint recognition model to obtain the voiceprint features of each of the segments; slidingly comparing the voiceprint features of two adjacent segments before and after to determine whether the two segments belong to the same speaker, so as to determine the segments belonging to the first speaker and the segments not belonging to the first speaker among each of the segments; and determining the segments not belonging to the first speaker as the speaker change segments.
[0008] In one implementation, determining the voiceprint feature of the first speaker includes: pre-collecting the voice audio of the first speaker, and calculating the voiceprint feature of the first speaker based on the voiceprint recognition model for the voice audio; or determining the voiceprint feature corresponding to the first segment in the valid audio as the voiceprint feature of the first speaker.
[0009] In one implementation, clustering the segments belonging to the first speaker in the valid audio according to the segments of the first speaker and the change segments to obtain the valid audio of the first speaker includes: sequentially determining N consecutive segments from the segments of the first speaker as the basic clustering segments; sequentially obtaining other segments from the last position of the basic clustering segments as the new clustering segments until the proportion of the change segments in the basic clustering segments and the new clustering segments exceeds a preset ratio; and clustering the basic clustering segments and the new clustering segments to obtain the valid audio of the first speaker.
[0010] In one implementation, the method further includes determining the change segments as the segments belonging to the second speaker, and clustering the segments of the second speaker to obtain the valid audio of the second speaker.
[0011] In one implementation, the method further includes: determining the change segments as the segments belonging to other speakers; obtaining the voiceprint features of other speakers; determining the speaker to which each segment belongs; and respectively clustering the segments of each speaker to obtain the valid audio of each speaker.
[0012] Based on the same inventive concept, the present application also provides a device for quickly extracting telephone channel data, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for quickly extracting telephone channel data provided in the above embodiments.
[0013] In addition, the present application further provides a computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements the method for quickly extracting telephone channel data provided in the above embodiments.
[0014] Based on the method, apparatus and computer-readable storage medium for quickly extracting telephone channel data provided in the embodiments of the present application, automatic and quick extraction of telephone channel data can be achieved without manual intervention, significantly improving the data extraction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplary illustrations do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated. The drawings in the figures do not constitute a scale limitation.
[0016] Figure 1 Show the flowchart of the method for quickly extracting telephone channel data provided in Embodiment 1 of the present application;
[0017] Figure 2 Show the schematic diagram of the effective audio segmentation process in the method provided in Embodiment 1 of the present application;
[0018] Figure 3 Show the implementation method of step S103 in Embodiment 1 of the present application;
[0019] Figure 4 Show the schematic diagram of the sliding comparison process in Embodiment 1 of the present application;
[0020] Figure 5 Show the flowchart of the method for quickly extracting telephone channel data provided in Embodiment 2 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will elaborate on each embodiment of the present application with reference to the drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are provided to help the reader better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0022] In the method for quickly extracting telephone channel data provided in the embodiments of the present application, by effectively extracting audio from the telephone channel data to preliminarily filter out non-human voice noise data, the effective audio is segmented into fragments, and the fragments are identified to determine the fragments belonging to the first speaker and the changed fragments. Then, based on the fragments of the first speaker and the changed fragments, the fragments belonging to the first speaker in the effective audio are clustered to obtain the effective audio of the first speaker. The whole process does not require manual intervention and can be automatically completed quickly. The interference data is filtered out when extracting the effective audio belonging to the speaker, making the finally extracted audio data more accurate. The embodiments of the present application will be described in detail below in combination with specific application scenarios.
[0023] Embodiment 1
[0024] In this embodiment, the speakers in the channel data to be extracted are the first speaker and the second speaker. Based on the method provided in this embodiment, the audio belonging to the first speaker and the second speaker can be quickly and effectively extracted respectively. For details, please refer to Figure 1 , Figure 1 which shows the flowchart of the method for quickly extracting telephone channel data provided in Embodiment 1 of the present application. As Figure 1 shown, the method includes the steps:
[0025] S101, obtain the effective audio in the channel data to be extracted.
[0026] Specifically, the channel data to be extracted can be the mono-channel data collected during a two-person call. In this embodiment, the VAD (Voice Activity Detection) technology can be used to clean the data and retain the effective audio. Among them, the effective audio refers to the audio data containing speech. It can be understood that during a telephone call, non-speech audio data such as background noise and silent sounds are usually collected together, and these audio data will interfere with the human voice recognition. Therefore, by extracting the effective audio in the channel to be extracted, the non-speech noise data in the channel data to be extracted can be filtered to reduce the interference to subsequent recognition and ensure the recognition accuracy. It can be understood that the embodiments of the present application do not limit the technology for extracting effective audio, and it can be selected from the existing technologies according to actual needs.
[0027] S102, segment the effective audio into several fragments.
[0028] In one implementation, the valid audio can be sequentially segmented into segments of a fixed length according to the chronological order of the valid audio. Here, the fixed length can be measured in terms of time or data volume. Taking time as an example, the valid audio can be segmented into multiple segments with a duration of 1 second. Preferably, in order to improve the accuracy of the subsequent clustering results, an overlapping window can be set during the process of segmenting the segments, that is, there is some identical data in two adjacent segments. Specifically, reference can be made to Figure 2 In the example shown in, there is an overlap of 0.5 seconds in two adjacent segments. In this way, when clustering consecutive segments subsequently, since there is identical data in the two segments, the obtained clustering result will be more continuous and more real and accurate.
[0029] S103. By comparing the voiceprint features of adjacent segments, determine the segments belonging to the first speaker and the speaker change segments.
[0030] In this step, it is necessary to obtain the voiceprint features of each segment. In implementation, a voiceprint recognition model can be pre-constructed. The voiceprint recognition model can be trained by itself or an open-source model can be adopted. For example, a currently mainstream and verified reliable model such as ECAPA-TDNN or a model integrating multiple algorithms can be selected, and then a telephone channel dataset is constructed by itself for training. The model that meets the test expectations after training is used as the voiceprint recognition model. When the voiceprint recognition model is ready, each segment can be used as the input layer respectively, and the voiceprint features of each segment are obtained through the calculation and recognition of the voiceprint recognition model. Furthermore, the segments belonging to the first speaker are determined based on the voiceprint features. Specifically, reference can be made to Figure 3 shown Figure 3 Illustrates the implementation method of step S103 in the first embodiment of the present application. As Figure 3 shown, the method includes:
[0031] S301. Determine the voiceprint features of the first speaker.
[0032] According to different application scenarios, the method for determining the voiceprint features of the first speaker is different.
[0033] In an application scenario, the identity of the speaker is known. Then, the voiceprint features of the first speaker can be obtained by pre-collecting the voice audio of the first speaker and calculating the voice audio based on the above voiceprint recognition model.
[0034] In another application scenario, the identity of the speaker is unknown. Then, the voiceprint features corresponding to the first segment in the valid audio can be determined as the voiceprint features of the first speaker.
[0035] S302. Based on the voiceprint recognition model, sequentially identify each of the segments to obtain the voiceprint features of each of the segments.
[0036] S303. Slide and compare the voiceprint features of two adjacent segments before and after to determine whether the two segments belong to the same speaker, so as to determine the segments belonging to the first speaker and the segments not belonging to the first speaker among all the segments.
[0037] The specific way of sliding comparison is to sequentially compare the latter segment with the former segment according to the position order of each segment in the effective audio, determine whether the two segments belong to the same speaker, and mark the recognition result. The process of sliding comparison can refer to Figure 4 .
[0038] In one implementation, the method for determining the segments belonging to the first speaker may include comparing the voiceprint feature values of each segment with the voiceprint feature values of the first speaker. If the difference is within the first threshold, it can be determined that the segment belongs to the first speaker; otherwise, it does not belong to the first speaker.
[0039] In another implementation, since the application scenario of this embodiment is a two-person call scenario, if each segment does not belong to the first speaker, it belongs to the second speaker. In this way, it can be directly determined whether the difference between the voiceprint features of two adjacent segments is within the second threshold through sliding comparison. If so, it is determined that they belong to the same person; otherwise, it is determined as segments of different speakers. In this way, the change points of the speakers can be obtained and marked, so as to determine the segments belonging to the first speaker and the second speaker respectively.
[0040] It can be understood that the magnitudes of the first threshold and the second threshold can be set according to the characteristics of the voiceprint recognition model, and this application does not make any restrictions.
[0041] S304. Determine the segments not belonging to the first speaker as speaker change segments.
[0042] It can be understood that in the application scenario of this embodiment, the speaker change segments are the segments of the second speaker.
[0043] S104. Cluster the segments belonging to the first speaker in the effective audio according to the segments of the first speaker and the change segments to obtain the effective audio of the first speaker.
[0044] In the implementation, N consecutive segments (N is a positive integer greater than 1) can be determined from the segments of the first speaker in sequence as the basic clustering segments; then other segments are sequentially obtained from the last position of the basic clustering segments as new clustering segments until the proportion of the change segments in the basic clustering segments and the new clustering segments exceeds the preset proportion, and then the basic clustering segments and the new clustering segments are clustered to obtain the effective audio of the first speaker.
[0045] In one example, the valid audio is segmented into segments 1, 2, 3... 100. According to the above method, it is determined that segments 1 to 20, 30 to 35, and 37 to 50 are the segments of the first speaker. Assuming N is 10, then 1 - 10 can be obtained as the basic clustering segments first, and then 11 is obtained. The proportion of the changed segments in 1 - 11 is calculated to be 0. If the preset proportion is 10%, then segment 11 can be recorded as a new clustering segment, and segments 12 - 20 are judged in the same way in turn; when segment 21 is recorded, since 21 is a changed segment, the proportion of the changed segments in 1 - 21 is 4.7%, which has not exceeded the preset proportion, so segment 21 can be used as a new clustering segment and continue to judge backward until the proportion of the changed segments exceeds 10%, that is, until the 23rd segment. Thus, segments 11 - 22 can be used as new clustering segments, and then segments 1 - 22 can be clustered to obtain the valid audio of the first speaker. Then, the remaining segments can be clustered in the same way to obtain all the audio belonging to the first speaker in the valid audio.
[0046] In implementation, the clustering methods may include IAC, AHC, k - means, spectral clustering, etc., or based on the speaking characteristics of each person, extract features for clustering, etc.
[0047] S105, determine the changed segments as the segments belonging to the second speaker, and cluster the segments of the second speaker to obtain the valid audio of the second speaker.
[0048] It can be understood that in the scenario of a two - person call, the changed segments can be directly determined as the segments of the second speaker. Based on the same clustering method as above, the extraction of the valid audio of the second speaker can be realized.
[0049] It can be seen that based on the method provided in this embodiment, the channel data containing only two speakers can be quickly and effectively extracted automatically to obtain the valid audio of different speakers respectively, saving labor costs and improving work efficiency. Further, this method is applicable to scenarios where the speakers are known or unknown, with a wide and flexible application range. And by setting an overlapping window when segmenting, the continuity of the data can be ensured, thus providing a good data basis for recognition and clustering, making the extracted valid audio more real and accurate.
[0050] Embodiment 2
[0051] In this embodiment, the speakers in the channel data to be extracted can be more than two, that is, the first speaker and at least two other speakers. Based on the method provided in this embodiment, the audio belonging to each speaker can be quickly and effectively extracted respectively. For details, please refer to Figure 5 , Figure 5The flowchart of the method for quickly extracting telephone channel data provided in the second embodiment of the present application is shown. As Figure 5 shown, the method includes the steps:
[0052] S501, obtain the valid audio in the channel data to be extracted.
[0053] Among them, the channel data to be extracted in this embodiment can be mono-channel data containing more than two speakers, or stereo-channel data collected during a call between more than two people.
[0054] S502, divide the valid audio into several segments.
[0055] S503, determine the segments belonging to the first speaker and the speaker change segments by comparing the voiceprint features of adjacent segments.
[0056] S504, cluster the segments belonging to the first speaker in the valid audio according to the segments of the first speaker and the change segments to obtain the valid audio of the first speaker.
[0057] The specific implementation manners of the above steps S501 - S504 are the same as those of steps S101 - S104 in the first embodiment. For specific descriptions, reference can be made to the above, and details will not be repeated.
[0058] The difference between this embodiment and the above first embodiment lies in the processing of the change segments. It can be understood that the change segments contain segments of at least two speakers. Therefore, it is necessary to further confirm the segments belonging to each speaker in the change segments. The method includes:
[0059] S505, determine the change segments as the segments belonging to other speakers.
[0060] S506, obtain the voiceprint features of other speakers.
[0061] S507, determine the speaker to which each segment belongs.
[0062] S508, cluster the segments of each speaker respectively to obtain the valid audio of each speaker.
[0063] Specifically, according to different application scenarios, the methods for obtaining the voiceprint features of other people are also different. When the number and identities of other speakers are known, the segments of each speaker can be determined based on the voiceprint features of each speaker. The judgment method is the same as that of the first speaker, and details will not be repeated.
[0064] When the number and identities of other speakers are unknown, it is necessary to first determine the voiceprint features of the second speaker from the changed segments. That is, the voiceprint features corresponding to the first segment in the changed segments can be used as the voiceprint features of the second speaker, and then all the changed segments are compared and identified to determine the segments belonging to the second speaker in the changed segments. The specific method can refer to the method in step S503 correspondingly to divide the changed segments into segments of the second speaker and non-second speaker segments.
[0065] Furthermore, when there are N consecutive segments in the non-second speaker segments identified from the changed segments, then the segments of the third person, or even the fourth person, can continue to be identified from the changed segments based on the same method. It can be understood that if there are no N consecutive segments in the changed segments, there is no need to perform further identification processing on the changed segments, and they are directly identified as invalid data.
[0066] After obtaining the segments of each speaker, the segments can be clustered according to the speaker identity respectively to obtain the valid audio of each speaker. The specific clustering method can refer to the clustering method in Embodiment 1 and will not be elaborated here.
[0067] In this embodiment, a method for quickly extracting telephone channel data with more than two speakers is provided. It can realize the method of automatically and quickly extracting the valid audio corresponding to each speaker from the channel data when the number and identity of the speakers are known or unknown.
[0068] Based on the same inventive concept, an embodiment of the present application also provides a device for quickly extracting telephone channel data, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can execute the method for quickly extracting telephone channel data provided in the above embodiment.
[0069] The device for quickly extracting telephone channel data provided in this embodiment can process the received channel data to be extracted to automatically and quickly extract the valid audio of each speaker, and is applicable to all scenarios of two or more speakers, where the number and identity of the speakers are known or unknown. It not only saves labor costs and improves the extraction efficiency, but also can adapt to various application requirements.
[0070] In addition, another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiment is implemented.
[0071] Those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps in the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0072] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A method for quickly extracting telephone channel data, characterized in that, The method includes the steps of: Obtaining valid audio in the channel data to be extracted, where the channel data to be extracted is channel data collected during a call between at least two people, and the step of obtaining valid audio in the channel data to be extracted includes: cleaning the data using voice activity detection technology, retaining the audio data containing speech to filter out non-speech noise data; Dividing the valid audio into several segments; Determining the segments belonging to the first speaker and the speaker change segments by comparing the voiceprint features of adjacent segments; Clustering the segments belonging to the first speaker in the valid audio according to the segments of the first speaker and the change segments to obtain the valid audio of the first speaker, including: Sequentially determining N consecutive segments from the segments of the first speaker as the basic clustering segments; Sequentially obtaining other segments from the last position of the basic clustering segments as new clustering segments until the proportion of change segments in the basic clustering segments and the new clustering segments exceeds a preset ratio; Clustering the basic clustering segments and the new clustering segments to obtain the valid audio of the first speaker.
2. The method for quickly extracting telephone channel data according to claim 1, characterized in that The dividing the valid audio into several segments includes sequentially dividing the valid audio based on the data length or time length to obtain each segment of a fixed length.
3. The method for quickly extracting telephone channel data according to claim 2, characterized in that, There is partial identical data between two adjacent segments.
4. The method for quickly extracting telephone channel data according to claim 1, characterized in that, The determining the segments belonging to the first speaker and the speaker change segments by comparing the voiceprint features of two adjacent segments includes: Determining the voiceprint feature of the first speaker; Sequentially identifying each segment based on a voiceprint recognition model to obtain the voiceprint features of each segment; Slidingly comparing the voiceprint features of two adjacent segments before and after to determine whether the two segments belong to the same speaker, so as to determine the segments belonging to the first speaker and the segments not belonging to the first speaker in each segment; Determining the segments not belonging to the first speaker as the speaker change segments.
5. The method for quickly extracting telephone channel data according to claim 4, wherein The determining the voiceprint feature of the first speaker includes: Pre-collecting the voice audio of the first speaker, calculating the voice audio based on the voiceprint recognition model to obtain the voiceprint feature of the first speaker; or, Determining the voiceprint feature corresponding to the first segment in the valid audio as the voiceprint feature of the first speaker.
6. The method for quickly extracting telephone channel data according to claim 1, wherein, The method further includes determining the change segments as the segments belonging to the second speaker, and clustering the segments of the second speaker to obtain the valid audio of the second speaker.
7. The method for quickly extracting telephone channel data according to claim 1, wherein The method further includes: Determining the change segments as the segments belonging to other speakers; Obtaining the voiceprint features of other speakers; Determining the speaker to which each segment belongs; Respectively clustering the segments of each speaker to obtain the valid audio of each speaker.
8. A device for quickly extracting telephone channel data, characterized in that, Including: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method for quickly extracting telephone channel data according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for quickly extracting telephone channel data according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice data processing method and device, computer equipment and storage medium
CN111613231A