Voice processing method and device, electronic equipment and storage medium
Through screening, separation and clustering of speech techniques, the problem of extracting vocal clips from speech is solved, efficient and effective speech processing is achieved, and speech quality and accuracy are improved.
Patent Information
- Application Number
- CN202510150664.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively extract vocal clips from speech, especially in environments with abundant noise and invalid information.
By filtering candidate speeches based on speech quality, performing speech separation to obtain single-person voice clips, and clustering these clips to form a collection of voice clips, the target speech for each set is finally determined.
It realizes efficient screening of candidate speeches with good speech quality from the set of pending speeches, and successfully separates and extracts each person's target speech, improving the effectiveness and quality of the extracted speech.
Smart Images

Figure CN120183385A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and more particularly, to a voice processing method, apparatus, electronic device, and storage medium. Background Art
[0002] As a common data form, audio contains rich information, and voice carries various information for people to communicate with each other, making voice processing technology particularly important.
[0003] However, voice includes a large amount of noise and invalid information. Therefore, there is an urgent need for a voice processing means to extract effective human voice segments from the voice. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a voice processing method, apparatus, electronic device, and storage medium to provide a means for extracting effective human voice segments from voice.
[0005] In a first aspect, an embodiment of the present application provides a voice processing method, the method including: determining candidate voices from the set of voices to be processed according to the voice quality of each voice in the set of voices to be processed; performing voice separation on the candidate voices to obtain a plurality of single-person voice segments; each single-person voice segment includes the human voice of the same person; clustering the plurality of single-person voice segments to obtain at least one voice segment set; the single-person voice segments within a voice segment set all belong to the same person; determining a target voice corresponding to each voice segment set based on the single-person voice segments in each voice segment set.
[0006] In a second aspect, an embodiment of the present application provides a training apparatus for a voice processing model, the apparatus including: a first determination module for determining candidate voices from the set of voices to be processed according to the voice quality of each voice in the set of voices to be processed; a separation module for performing voice separation on the candidate voices to obtain a plurality of single-person voice segments; each single-person voice segment includes the human voice of the same person; a clustering module for clustering the plurality of single-person voice segments to obtain at least one voice segment set; the single-person voice segments within a voice segment set all belong to the same person; a second determination module for determining a target voice corresponding to each voice segment set based on the single-person voice segments in each voice segment set.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.
[0008] Fourthly, an embodiment of the present application provides a computer-readable storage medium, in which program codes are stored. When the program codes are run by a processor, the above-mentioned method is executed.
[0009] Fifthly, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the above-mentioned method.
[0010] A voice processing method, device, electronic device and storage medium provided by an embodiment of the present application. First, according to the voice quality of each voice in a set of voices to be processed, candidate voices are determined from the set of voices to be processed, so as to filter out voices with poor quality and obtain candidate voices with better quality. Then, the candidate voices are separated to obtain multiple single-person voice segments, and the multiple single-person voice segments are clustered to obtain at least one voice segment set, so as to determine a respective voice segment set for each person. Finally, a target voice is determined according to the single-person voice segments in each voice segment set. It can be seen that in the present application, the purpose of screening candidate voices with good quality from the set of voices to be processed is achieved first. Secondly, the target voice of each person is separated from the screened candidate voices, and the purpose of extracting the corresponding voice for each person is achieved. That is, through the method of the present application, the purpose of extracting the respective voices of each person is achieved, and the extracted voices are highly effective and have good quality. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, they can also
[0012] Figure 1 show a flowchart of a voice processing method proposed in an embodiment of the present application;
[0013] Figure 2 show Figure 1 a flowchart of step S140 in a corresponding embodiment in one embodiment;
[0014] Figure 3 show Figure 1 a flowchart of step S140 in a corresponding embodiment in another embodiment;
[0015] Figure 4 show Figure 1Flowchart of step S140 in a corresponding embodiment in yet another embodiment;
[0016] Figure 5 The block diagram of a voice processing device proposed in an embodiment of the present application is shown;
[0017] Figure 6 The structural block diagram of an electronic device proposed in an embodiment of the present application is shown;
[0018] Figure 7 The structural block diagram of a computer-readable storage medium provided in an embodiment of the present application is shown. Detailed implementation manners
[0019] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. According to the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0020] In the following description, the terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0022] Please refer to Figure 1 , Figure 1 The flowchart of a voice processing method proposed in an embodiment of the present application is shown. The method can be used for an electronic device, and the method includes:
[0023] S110. Determine candidate voices from the set of voices to be processed according to the voice quality of each voice in the set of voices to be processed.
[0024] In this embodiment, the set of voices to be processed may refer to a set including multiple voices. The voices in the set of voices to be processed may be transaction communication voices in a financial scenario, patient communication voices in a medical scenario, guiding communication voices in a transportation scenario, etc. The voices in the set of voices to be processed may have different uses. For example, the voices in the set of voices to be processed may be used for population classification of speakers (i.e., the people who speak the voices), voice content classification, voiceprint cleaning, voice content cleaning in specific scenarios, verification code cleaning, etc. Among them, specific scenarios may be, for example, a transaction scenario, a meeting scenario, etc. Cleaning means removing noise data.
[0025] First, voice collection can be performed for a certain scenario or multiple scenarios to obtain multiple voices. The multiple collected voices are summarized into a set of voices to be processed. Then, according to the voice quality of each voice in the set of voices to be processed, voices with higher quality are selected as candidate voices.
[0026] In this embodiment, the voice quality includes at least one of signal-to-noise ratio, effective human voice duration, average noise energy, and clipping ratio.
[0027] The signal-to-noise ratio refers to the ratio of the intensity of the received useful signal (i.e., the voice) to the intensity of the received interference signal (i.e., the noise in the voice). When the noise is strong and the human voice is weak, the signal-to-noise ratio is low. Therefore, a signal-to-noise ratio threshold can be set. When it is determined that the signal-to-noise ratio of the voice is higher than the signal-to-noise ratio threshold, the voice quality is high and it is used as a candidate voice. Among them, the signal-to-noise ratio threshold can be set based on requirements. For example, the signal-to-noise ratio threshold is 2.
[0028] The effective human voice duration refers to the duration during which the human voice exists in the voice. The larger the effective human voice duration, the more human voices there are in the voice, and the voice includes more human voices, thus possibly including more content. Therefore, a first duration threshold can be set. When the effective human voice duration of the voice is higher than the first duration threshold, it is determined that the voice quality is high and it is used as a candidate voice. Among them, the first duration threshold can be set based on requirements. For example, it is 5s.
[0029] The average noise energy refers to the average power or energy of the noise signal. The higher the average noise energy of the voice, the stronger the noise of the voice. The lower the average noise energy of the voice, the weaker the noise of the voice. Therefore, a noise energy threshold can be set. When the average noise energy of the voice is lower than the noise energy threshold, it means that the noise of the voice is weak, and it is determined that the voice quality is high and it is used as a candidate voice. Among them, the noise energy threshold can be set based on requirements. For example, the noise energy threshold is 3.
[0030] The clipping ratio refers to the ratio of the duration during which the signal amplitude in speech exceeds the linear range of the system to the speech duration. The larger the clipping ratio of the speech, the stronger the noise or human voice; the smaller the clipping ratio of the speech, the weaker the noise or human voice. Therefore, a clipping ratio threshold can be set. If the clipping ratio of the speech is higher than the clipping ratio threshold, it is determined that the speech quality is higher, and it is used as a candidate speech. Among them, the clipping ratio threshold can be set based on requirements. For example, the clipping ratio threshold is 0.3.
[0031] Of course, in some embodiments, multiple of signal-to-noise ratio, effective human voice duration, average noise energy, and clipping ratio can also be combined to determine the candidate speech. For example, select the speech with a signal-to-noise ratio higher than the signal-to-noise ratio threshold, an effective human voice duration higher than the first duration threshold, and an average noise energy lower than the noise energy threshold as the candidate speech. Another example is to select the speech with a signal-to-noise ratio higher than the signal-to-noise ratio threshold, an effective human voice duration higher than the first duration threshold, an average noise energy lower than the noise energy threshold, and a clipping ratio higher than the clipping ratio threshold as the candidate speech.
[0032] S120. Perform speech separation on the candidate speech to obtain multiple single-person speech segments.
[0033] Among them, each single-person speech segment includes the human voice of the same person. There can be multiple candidate speeches. For each candidate speech, speech separation is performed to obtain single-person speech segments, so that the human voice in each single-person speech segment belongs to the same person (which means there is only one speaker in this speech segment). In the following embodiments of this application, for the convenience of explanation and understanding, it is explained with only one candidate speech.
[0034] In this embodiment, the candidate speech can be cut into several sub-speeches - multiple single-person speech segments by a human voice separation module.
[0035] In some embodiments, the human voice separation module can extract all the voiceprint information in the candidate speech, and take the segments with the voiceprint information difference less than the voiceprint difference threshold as a single-person speech segment. Among them, the voiceprint difference threshold can be set based on requirements, and this application does not make a limitation. The difference in the voiceprint information of the speech segment being less than the voiceprint difference threshold means that the speaker in the speech segment is the same person, and it is determined that this speech segment is a single-person speech segment.
[0036] In still other embodiments, the human voice separation module can perform speech separation on the candidate speech through a human voice separation model to obtain multiple single-person speech segments. The human voice separation model can be a neural network model, which can be trained based on a sample speech including the human voices of multiple different speakers and the single-person sample speech segments after the sample speech is divided. The single-person sample speech segment refers to the speech segment of the same speaker in the sample speech.
[0037] S130. Cluster multiple single-person speech segments to obtain at least one speech segment set.
[0038] Among them, the single-person speech segments within one speech segment set all belong to the same person.
[0039] After the candidate speech is divided into multiple single-person speech segments, the same speaker may correspond to multiple speech segments. Therefore, clustering the multiple single-person speech segments to divide the single-person speech segments of each speaker into the same set, so as to obtain at least one speech segment set, and one speaker corresponds to one speech segment set.
[0040] In some embodiments, the tone or intonation of each single-person speech segment can be extracted, and the single-person speech segments with similar tone or intonation are divided into one speech segment set to achieve clustering of the multiple single-person speech segments.
[0041] In still other embodiments, the voiceprint information of each single-person speech segment can be determined; the single-person speech segments with the difference between the voiceprint information less than the voiceprint difference threshold are divided into one speech segment set to obtain at least one speech segment set. That is to say, the difference between the voiceprint information being less than the voiceprint difference threshold means that the speakers of the single-person speech segments are the same person, and they are divided into one speech segment set.
[0042] S140. Based on the single-person speech segments in each speech segment set, determine the target speech corresponding to each speech segment set.
[0043] After obtaining the speech segment set, for each speech segment set, the single-person speech segments in the speech segment set are spliced into one speech in chronological order as the target speech corresponding to the speech segment set. This target speech is the single-person speech of the speaker corresponding to this speech segment set (the single-person speech refers to the speech with the same speaker). Traverse each speech segment set to obtain the target speech of each speech segment set. Thus, the purpose of extracting the single-person speech of different speakers from the speech set to be processed is achieved.
[0044] Generally speaking, the target speech obtained according to the foregoing S110 - S140 can be used in the user group classification scenario. For example, distinguish the gender, occupation, preferences, etc. of the user according to the content in the target speech.
[0045] In this embodiment, candidate voices are determined from the set of voices to be processed according to the voice quality of each voice in the set of voices to be processed, so as to filter out voices with poor quality and obtain candidate voices with better voice quality. Then, the candidate voices are separated to obtain multiple single-person voice segments, and the multiple single-person voice segments are clustered to obtain at least one set of voice segments, so as to determine a respective set of voice segments for each person. Finally, a target voice is determined according to the single-person voice segments in each set of voice segments. It can be seen that in this application, the purpose of screening candidate voices with good voice quality from the set of voices to be processed is first achieved. Secondly, the respective target voices of each person are separated from the screened candidate voices, and the purpose of extracting the corresponding voice for each person is achieved. That is, through the method of this application, the purpose of extracting the respective voices of each person is achieved, and the extracted voices are highly effective and have good voice quality.
[0046] In some embodiments, S140 may include: determining second candidate single-person voice segments from the set of voice segments according to the voice quality of each single-person voice segment in the set of voice segments; determining the target voice corresponding to the set of voice segments according to the second candidate single-person voice segments in the set of voice segments.
[0047] Wherein, the voice quality here may also include at least one of signal-to-noise ratio, effective voice duration, average noise energy, and clipping ratio. Therefore, the step of screening candidate second candidate single-person voice segments according to the voice quality of each single-person voice segment is similar to the step of determining candidate voices from the set of voices to be processed as described above, and will not be elaborated here.
[0048] That is to say, for each set of voice segments, further voice screening is performed on the single-person voice segments in the set of voice segments to filter out some single-person voice segments with poor voice quality and obtain second candidate single-person voice segments with higher voice quality. Then, the second candidate single-person voice segments in the set of voice segments are spliced into a voice in chronological order as the target voice corresponding to the set of voice segments. The target voice is the single-person voice of the speaker corresponding to the set of voice segments, and the quality of the target voice determined at this time is further improved.
[0049] Of course, after further screening, the obtained target voice can also be used in the user group classification scenario, and the classification accuracy is higher at this time.
[0050] In some embodiments, as Figure 2 shown, S140 may further include:
[0051] S210. Determining first candidate single-person voice segments from the set of voice segments according to the number of characters in the text of each single-person voice segment in the set of voice segments and / or the effective voice duration of each single-person voice segment.
[0052] For each set of speech segments, count the number of characters in each single-person speech segment in the set of speech segments, and count the effective duration of the human voice in each single-person speech segment in the set of speech segments. Then, in combination with at least one of the number of characters in each single-person speech segment in the set of speech segments and the effective duration of the human voice in each single-person speech segment, screen out the effective single-person speech segments as the first candidate single-person speech segments.
[0053] In some embodiments, determine, from the set of speech segments, single-person speech segments with the number of characters higher than a character count threshold as the first candidate single-person speech segments. Among them, the character count threshold can be, for example, 3. The higher the number of characters, the lower the possibility that the characters included in the single-person speech segment are filler words, the more effective content the single-person speech segment includes, and the more effective the single-person speech segment is. On the contrary, the lower the number of characters, the higher the possibility that the characters included in the single-person speech segment are filler words, the less effective content the single-person speech segment includes, and the less effective the single-person speech segment is. Therefore, the first candidate single-person speech segments can be screened by the character count threshold.
[0054] In still some other embodiments, determine, from the set of speech segments, single-person speech segments with the effective duration of the human voice higher than a second duration threshold as the first candidate single-person speech segments. Among them, the second duration threshold can be set based on requirements. Since the single-person speech segments are obtained by dividing the candidate speech, generally, the second duration threshold is less than the first duration threshold.
[0055] The higher the effective duration of the human voice, the more human voice the single-person speech segment includes, the more effective content the single-person speech segment includes, and the more effective the single-person speech segment is. On the contrary, the lower the effective duration of the human voice, the less human voice the single-person speech segment includes, the less effective content the single-person speech segment includes, and the less effective the single-person speech segment is. Therefore, the first candidate single-person speech segments can be screened by the second duration threshold.
[0056] In yet some other embodiments, for each single-person speech segment in the set of speech segments, determine the character frequency of the single-person speech segment according to the number of characters in the single-person speech segment and the effective duration of the human voice in the single-person speech segment; determine the first candidate single-person speech segments from the set of speech segments according to the character frequencies of the respective single-person speech segments in the set of speech segments.
[0057] It can be to calculate the ratio of the number of characters in the single-person speech segment and the effective duration of the human voice in the single-person speech segment as the character frequency of the single-person speech segment, and then screen out single-person speech segments with the character frequency higher than a first character frequency threshold as the first candidate single-person speech segments. Among them, the first character frequency threshold can be, for example, 0.5.
[0058] The higher the word frequency, the higher the frequency of the human voice in a single-person speech segment, the lower the possibility that the single-person speech segment includes filler words, the more effective content the single-person speech segment includes, and the more effective the single-person speech segment is. Conversely, the lower the word frequency, the lower the frequency of the human voice in the single-person speech segment, the higher the possibility that the single-person speech segment includes filler words, the less effective content the single-person speech segment includes, and the less effective the single-person speech segment is. Therefore, the first candidate single-person speech segments can be screened by a first word frequency threshold.
[0059] S220. Determine the target speech corresponding to the speech segment set according to the first candidate single-person speech segments in the speech segment set.
[0060] It can be directly splicing the first candidate single-person speech segments in chronological order into a speech as the target speech corresponding to the speech segment set.
[0061] Of course, in some embodiments, the first candidate single-person speech segments with higher speech quality can also be screened according to the speech quality of the first candidate single-person speech segments as the third candidate single-person speech segments, and then the third candidate single-person speech segments screened from the same speech segment set are spliced into a speech in chronological order as the target speech corresponding to the speech segment set.
[0062] In this embodiment, the first candidate single-person speech segments are also screened by the number of characters in the single-person speech segment and the effective duration of the human voice, and the target speech is determined according to the first candidate single-person speech segments, further filtering out ineffective single-person speech segments with few characters, low effective duration of the human voice or low word frequency, and improving the quality of the obtained target speech.
[0063] Above, the target speech determined in this embodiment can be used in the voiceprint cleaning scenario to determine high-accuracy voiceprint information for different users respectively.
[0064] In some embodiments, as Figure 3 shown, S140 may include:
[0065] S310. For each single-person speech segment in the speech segment set, determine the voice timestamp of each human voice segment in the single-person speech segment and the text timestamp of each character in the single-person speech segment.
[0066] There may be pauses between each word spoken by the speaker. Therefore, for each single-person speech segment, the human voice included is not completely continuous. Each continuous human voice part in the single-person speech segment is regarded as a human voice segment, and the start time of the human voice segment can be determined as the human voice timestamp of the human voice segment. Among them, when the speech interval between two adjacent words is lower than the interval threshold, the segment including these two adjacent words is determined as a continuous human voice part. For example, the interval threshold can be 0.5s.
[0067] At the same time, for each word in the single-person speech segment, determine the moment when each word is spoken by the speaker as the word timestamp of each word.
[0068] S320. According to the human voice timestamp of each human voice segment and the word timestamp of each word, perform word alignment processing on each human voice segment to obtain the words corresponding to each human voice segment.
[0069] For any human voice segment in the same speech segment set, the words belonging to the word timestamp with the same human voice timestamp as that of this human voice segment can be used as the words corresponding to this human voice segment, or the words belonging to the word timestamp with a time difference less than the time difference threshold (the time difference threshold is 0.3s, for example) between the same human voice timestamp as that of this human voice segment can be used as the words corresponding to this human voice segment. In this way, by traversing all human voice segments, the words corresponding to each human voice segment are obtained.
[0070] S330. Determine the word frequency of each human voice segment according to the duration of each human voice segment and the number of words of the text corresponding to each human voice segment; determine candidate human voice segments from the single-person speech segment according to the word frequencies of the human voice segments in the single-person speech segment.
[0071] For each human voice segment, the ratio of the number of words of the text corresponding to this human voice segment to the duration of this human voice segment can be calculated as the word frequency of the human voice segment. Then, obtain the human voice segments with a word frequency higher than the second word frequency threshold as the candidate human voice segments corresponding to the single-person speech segment. The second word frequency threshold can be 0.3, for example.
[0072] The higher the word frequency, the higher the frequency of the human voice in the human voice segment, the lower the possibility that the human voice segment includes filler words, the more effective content the human voice segment includes, and the more effective the human voice segment is. On the contrary, the lower the word frequency, the lower the frequency of the human voice in the human voice segment, the higher the possibility that the human voice segment includes filler words, the less effective content the human voice segment includes, and the less effective the human voice segment is. Therefore, the candidate human voice segments can be screened through the second word frequency threshold.
[0073] S340. Determine the target voice corresponding to the speech segment set according to the candidate human voice segments in the speech segment set.
[0074] For each set of speech segments, the candidate human voice segments determined from the set of speech segments can be spliced into a single speech in chronological order as the target speech corresponding to the set of speech segments.
[0075] Of course, in some embodiments, candidate human voice segments with higher voice quality can also be screened according to the voice quality of the candidate human voice segments as effective human voice segments, and then the effective human voice segments screened from the same set of speech segments are spliced into a single speech in chronological order as the target speech corresponding to the set of speech segments.
[0076] It is worth mentioning that after determining the first candidate single-person speech segment according to the steps of the foregoing S210, the first candidate single-person speech segment can be used as the single-person speech segment in S310, and then the single-person speech segment is processed according to the steps of S310-S340 to determine candidate human voice segments from the first candidate single-person speech segment, and then according to the steps of S340, the target speech is obtained based on the candidate human voice segments. Of course, after determining the candidate human voice segments here, candidate human voice segments with higher voice quality can also be screened according to the voice quality of the candidate human voice segments as effective human voice segments, and then the effective human voice segments screened from the same set of speech segments are spliced into a single speech in chronological order as the target speech corresponding to the set of speech segments. In this way, candidate human voice segments with poor voice quality can be further filtered out, and the voice quality of the determined target speech can be improved.
[0077] In this embodiment, the text and the human voice in the single-person speech segment are also aligned to screen candidate human voice segments, further filtering out invalid human voice segments including filler words with low word frequencies and improving the quality of the obtained target speech.
[0078] In addition, after determining the first candidate single-person speech segment, the human voice segments in the first candidate single-person speech segment can be further screened to screen candidate human voice segments, further filtering out invalid human voice segments with low word frequencies and greatly improving the quality of the obtained target speech.
[0079] In this embodiment, the quality of the determined target speech is higher, and the target speech can be used in the cleaning scenario of voice verification codes to accurately identify verification code information in the target speech.
[0080] In some embodiments, as Figure 4 shown, S140 may include:
[0081] S410. For each single-person speech segment in the speech segment set, perform sentence alignment processing on the single-person speech segment according to the punctuation timestamp of the punctuation in the single-person speech segment and the text timestamp of the text in the single-person speech segment, to obtain at least one candidate sentence corresponding to the single-person speech segment.
[0082] Speech recognition can be performed on the single-person speech segment to obtain the punctuation in the single-person speech segment, and the moment when the punctuation in the single-person speech segment is located is used as the punctuation timestamp of the punctuation in the single-person speech segment.
[0083] The text timestamp of the text in the single-person speech segment may include the text timestamp of each text in the single-person speech segment, so as to divide the text included in the single-person speech segment into multiple text segments through the punctuation timestamp of the punctuation in the single-person speech segment, and implement sentence alignment processing on the single-person speech segment. At this time, the sentences formed by the text segments between two adjacent punctuations, the text segment before the first punctuation, and the text segment before the last punctuation are respectively used as candidate sentences. Of course, if a punctuation is set at the end of the single-person speech segment, the sentences formed by the text segments between two adjacent punctuations and the text segment before the first punctuation are respectively used as candidate sentences.
[0084] For example, if the number of punctuations recognized in the single-person speech segment is 3 (no punctuation is set at the end of the single-person speech segment), then 4 text segments are obtained, and the texts in each text segment are arranged in the order of the text timestamps in sequence to form a candidate sentence. At this time, four candidate sentences are obtained.
[0085] S420. Screen target sentences from the candidate sentences according to the semantic fluency of each candidate sentence.
[0086] In this application, the perplexity, BLEU (Bilingual Evaluation Understudy) value, METEOR (Metric for Evaluation of Translation with Explicit Ordering) value, etc. of each candidate sentence can be determined as the semantic fluency of the candidate sentence.
[0087] Obtain the candidate sentences with semantic fluency higher than the fluency threshold as the target sentences, where the fluency threshold can be set based on requirements and is not limited in this application.
[0088] S430. Determine the target speech corresponding to the speech segment set according to the target sentences corresponding to the speech segment set.
[0089] For each single-person speech segment in the speech segment set, according to the target sentence determined in the single-person speech segment, the corresponding partial segment in the single-person speech segment is used as the target human voice segment. Then, all the target human voice segments determined in all the single-person speech segments in the speech segment set are obtained, and then they are spliced into a voice in the chronological order of the target human voice segments as the target voice corresponding to the speech segment set.
[0090] In some embodiments, it is also possible to continue to screen the result segments with higher speech quality according to the speech quality of the target human voice segments, and then splice the screened result segments in the speech segment set in the chronological order of time into a voice as the target voice corresponding to the speech segment set.
[0091] In some embodiments, after determining the first candidate single-person speech segment according to the steps of the foregoing S210, the first candidate single-person speech segment is used as the single-person speech segment in S410, and the steps of S410-S430 are continued to determine the target human voice segment, and then based on the target human voice segment, the target voice is obtained. Of course, after determining the target human voice segment here, it is also possible to continue to screen the result segments with higher speech quality according to the speech quality of the target human voice segment, and then splice the screened result segments in the same speech segment set in the chronological order of time into a voice as the target voice corresponding to the speech segment set.
[0092] In still other embodiments, after determining the first candidate single-person speech segment according to the steps of the foregoing S210, the first candidate single-person speech segment is used as the single-person speech segment in S310, and then the single-person speech segment is processed according to the steps of S310-S340 to determine the candidate human voice segment from the first candidate single-person speech segment. After that, the candidate human voice segment is used as the single-person speech segment in S410, and the steps of S410-S430 are continued to determine the target human voice segment, and then based on the target human voice segment, the target voice is obtained. Of course, after determining the target human voice segment here, it is also possible to continue to screen the result segments with higher speech quality according to the speech quality of the target human voice segment, and then splice the screened result segments in the same speech segment set in the chronological order of time into a voice as the target voice corresponding to the speech segment set.
[0093] In this embodiment, the semantic fluency of the single-person speech segment is also detected to filter out the segments with poor semantic fluency to obtain the target human voice segment, thereby further filtering out the human voice segments with poor semantic fluency and improving the semantic fluency of the obtained target voice.
[0094] In this embodiment, the quality of the determined target voice is relatively high, and the target voice can be used in the data cleaning scenario to accurately identify various data contents in the target voice.
[0095] Please refer to Figure 5 , Figure 5 which shows a block diagram of a voice processing device proposed in an embodiment of the present application. The device 500 includes:
[0096] A first determination module 510, configured to determine candidate voices from the set of voices to be processed according to the voice quality of each voice in the set of voices to be processed;
[0097] A separation module 520, configured to perform voice separation on the candidate voices to obtain multiple single-person voice segments; each single-person voice segment includes the voice of the same person;
[0098] A clustering module 530, configured to cluster the multiple single-person voice segments to obtain at least one voice segment set; the single-person voice segments within a voice segment set all belong to the same person;
[0099] A second determination module 540, configured to determine the target voice corresponding to each voice segment set based on the single-person voice segments in each voice segment set.
[0100] Optionally, the second determination module 540 is further configured to determine a first candidate single-person voice segment from the voice segment set according to the number of characters in each single-person voice segment in the voice segment set and / or the effective voice duration of each single-person voice segment; and determine the target voice corresponding to the voice segment set according to the first candidate single-person voice segment in the voice segment set.
[0101] Optionally, the second determination module 540 is further configured to, for each single-person voice segment in the voice segment set, determine the word frequency of the single-person voice segment according to the number of characters in the single-person voice segment and the effective voice duration of the single-person voice segment; and determine the first candidate single-person voice segment from the voice segment set according to the word frequencies of the respective single-person voice segments in the voice segment set.
[0102] Optionally, the second determination module 540 is further configured to, for each single-person voice segment in the voice segment set, determine the voice timestamp of each voice segment in the single-person voice segment and the text timestamp of each character in the single-person voice segment; perform text alignment processing on each voice segment according to the voice timestamp of each voice segment and the text timestamp of each character to obtain the text corresponding to each voice segment; determine the word frequency of each voice segment according to the duration of each voice segment and the number of characters in the text corresponding to each voice segment; determine the candidate voice segments from the single-person voice segment according to the word frequencies of the respective voice segments in the single-person voice segment; and determine the target voice corresponding to the voice segment set according to the candidate voice segments in the voice segment set.
[0103] Optionally, the second determination module 540 is further configured to, for each single-person speech segment in the speech segment set, perform sentence alignment processing on the single-person speech segment according to the punctuation timestamp of the punctuation marks in the single-person speech segment and the text timestamp of the text in the single-person speech segment, so as to obtain at least one candidate sentence corresponding to the single-person speech segment; screen a target sentence from the candidate sentences according to the semantic fluency of each candidate sentence; and determine the target speech corresponding to the speech segment set according to the target sentences corresponding to the speech segment set.
[0104] Optionally, the second determination module 540 is further configured to determine a second candidate single-person speech segment from the speech segment set according to the speech quality of each single-person speech segment in the speech segment set; and determine the target speech corresponding to the speech segment set according to the second candidate single-person speech segment in the speech segment set.
[0105] Optionally, the clustering module 530 is further configured to determine the voiceprint information of each of the multiple single-person speech segments; and divide the single-person speech segments with the difference between the voiceprint information being less than the voiceprint difference threshold into a speech segment set, so as to obtain at least one speech segment set.
[0106] It should be noted that the device embodiments in this application correspond to the foregoing method embodiments. The specific principles in the device embodiments can be referred to the content in the foregoing method embodiments, and will not be elaborated here.
[0107] Figure 6 The structural block diagram of an electronic device proposed in an embodiment of this application is shown. The electronic device is used to execute the speech processing method according to the embodiment of this application. As Figure 6 shown, the electronic device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 into the random access memory (RAM) 1203, such as executing the method in the foregoing embodiment. In the RAM 1203, various programs and data required for system operation are also stored. The CPU 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. The input / output (I / O) interface 1205 is also connected to the bus 1204.
[0108] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1210 as needed so that a computer program read therefrom is installed into the storage section 1208 as needed.
[0109] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, various functions defined in the system of the present application are executed.
[0110] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0111] Reference Figure 7 , Figure 7 shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable storage medium 600, and this program code can be called by a processor to execute the method described in the above method embodiment.
[0112] The computer-readable storage medium 600 can be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 600 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 600 has a storage space for program code 610 that executes any of the method steps in the above-described method. These program codes can be read out from or written into one or more computer program products. The program code 610 can be compressed in an appropriate form, for example.
[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that executes the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0114] The units described in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0115] As another aspect, the present application also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments; or can exist separately without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in any of the above embodiments is implemented.
[0116] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method in any of the above embodiments.
[0117] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0118] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0119] After considering the specification and practicing the disclosed embodiments here, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application. It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A speech processing method, characterized in that: The method comprises: Determining candidate voices from the set of voices to be processed according to the voice quality of each voice in the set of voices to be processed; Performing voice separation on the candidate voice to obtain a plurality of single-person voice segments, each of which includes the voice of the same person; Clustering the multiple single-person voice segments to obtain at least one voice segment set; the single-person voice segments in a voice segment set all belong to the same person; Based on the single-speaker speech segments in each of the speech segment sets, the target speech corresponding to each of the speech segment sets is determined.
2. The method according to claim 1, characterized in that The determining, based on the single-person speech segments in each speech segment set, the target speech corresponding to each speech segment set comprises: Determine a first candidate single-person voice segment from the voice segment set according to the number of characters in each single-person voice segment in the voice segment set and / or the effective duration of the human voice of each single-person voice segment; According to the first candidate single-person voice segment in the voice segment set, a target voice corresponding to the voice segment set is determined.
3. The method according to claim 2, characterized in that The step of determining a first candidate single-person voice segment from the voice segment set according to the number of characters in each single-person voice segment in the voice segment set and / or the effective duration of human voice in each single-person voice segment comprises: For each single-person voice segment in the voice segment set, determine the word frequency of the single-person voice segment according to the number of characters in the single-person voice segment and the effective duration of the human voice in the single-person voice segment; According to the word frequencies of the individual speech segments in the speech segment set, a first candidate individual speech segment is determined from the speech segment set.
4. The method according to claim 1, characterized in that The determining, based on the single-person speech segments in each speech segment set, the target speech corresponding to each speech segment set comprises: For each single-person voice segment in the voice segment set, determining a voice timestamp of each vocal segment in the single-person voice segment and a text timestamp of each text in the single-person voice segment; According to the vocal timestamp of each vocal segment and the text timestamp of each text, performing text alignment processing on each vocal segment to obtain the text corresponding to each vocal segment; Determine the word frequency of each vocal segment according to the duration of each vocal segment and the number of words corresponding to each vocal segment; Determining candidate vocal segments from the single-person voice segments according to the word frequencies of the vocal segments in the single-person voice segments; According to each candidate vocal segment in the speech segment set, a target speech corresponding to the speech segment set is determined.
5. The method according to claim 1, characterized in that: The determining, based on the single-person speech segments in each speech segment set, the target speech corresponding to each speech segment set comprises: For each single-person voice segment in the voice segment set, according to the punctuation timestamps of the punctuation marks in the single-person voice segment and the text timestamps of the text in the single-person voice segment, sentence alignment processing is performed on the single-person voice segment to obtain at least one candidate sentence corresponding to the single-person voice segment; Selecting a target sentence from each of the candidate sentences according to the semantic fluency of each of the candidate sentences; According to each target sentence corresponding to the speech segment set, a target speech corresponding to the speech segment set is determined.
6. The method according to claim 1, characterized in that The determining, based on the single-person speech segments in each speech segment set, the target speech corresponding to each speech segment set comprises: Determining a second candidate single-person voice segment from the voice segment set according to the voice quality of each single-person voice segment in the voice segment set; According to the second candidate single-person voice segment in the voice segment set, a target voice corresponding to the voice segment set is determined.
7. The method according to any one of claims 1 to 6, characterized in that The speech quality includes at least one of a signal-to-noise ratio, an effective duration of human voice, an average noise energy, and a clipping ratio.
8. A speech processing device, characterized in that: The device comprises: A first determination module, configured to determine candidate voices from the set of voices to be processed according to the voice quality of each voice in the set of voices to be processed; A separation module, used for performing voice separation on the candidate voice to obtain a plurality of single-person voice segments; each single-person voice segment includes the voice of the same person; A clustering module, used for clustering the multiple single-person voice segments to obtain at least one voice segment set; the single-person voice segments in a voice segment set all belong to the same person; The second determination module is used to determine the target speech corresponding to each of the speech segment sets based on the single-person speech segments in each of the speech segment sets.
9. An electronic device, characterized in that: include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.