Audio processing method, device, processor and system

By automatically identifying and registering unknown characters in audio processing, the problem of early registration of voiceprint role separation in the prior art is solved, the ease of use and applicable scenarios are improved, and the separation of voiceprint roles without early registration is achieved.

CN115440229BActive Publication Date: 2025-08-22BEIJING SINOVOICE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211119654.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2025-08-22
Estimated Expiration
2042-09-14

AI Technical Summary

Technical Problem

In the prior art, the separation of voiceprint roles requires the speaker's voiceprint to be registered in advance, resulting in poor ease of use and high preparation costs.

Method used

By obtaining audio clips, the voiceprint recognition model is used for identification. According to the duration of the audio clips and the relationship between the recognition score and the threshold, unknown characters are automatically identified and registered in the voiceprint recognition model library to achieve the separation of voiceprint characters without early registration.

Benefits of technology

It realizes automatic identification of unknown characters in audio, improves the ease of use and applicable scenarios of voiceprint character separation, and reduces the cost of preparation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440229B_ABST
    Figure CN115440229B_ABST
Patent Text Reader

Abstract

The present application provides an audio processing method, device, processor, and system. The method includes: obtaining at least one audio clip and performing voiceprint recognition on the at least one audio clip using a voiceprint recognition model to obtain a first recognition result; obtaining the highest recognition score in the first recognition result when the first recognition result indicates that the at least one audio clip is a non-target silence clip and the duration of the at least one audio clip is greater than or equal to a first duration threshold; determining that the role corresponding to the at least one audio clip is an unknown role when the audio duration of the at least one audio clip is greater than or equal to a second duration threshold and the highest recognition score is less than the score threshold; and registering the unknown role in a voiceprint recognition model library. This method achieves the technical effect of voice role separation through an unknown role separation algorithm, solving the technical problem of requiring the speaker's voiceprint to be registered in advance when performing role separation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to an audio processing method, device, processor, and system. Background Art

[0002] Currently, both voice conferencing systems and transcript interrogation systems require a role separation system to separate the voices of multiple speakers and perform voice transcription or speaker role display based on the separation results.

[0003] However, current role separation technology usually requires the speaker's voiceprint to be registered in advance when performing voiceprint role separation. In actual application scenarios, it has poor usability and high preparation costs.

[0004] The above information disclosed in the background technology section is only used to enhance the understanding of the background technology of the technology described in this article. Therefore, the background technology may contain certain information that does not form the prior art known in this country to those skilled in the art. Summary of the Invention

[0005] The main purpose of this application is to provide an audio processing method, device, processor and system to solve the problem in the prior art that the speaker's voiceprint needs to be registered in advance when performing voiceprint role separation.

[0006] According to one aspect of an embodiment of the present invention, there is provided an audio processing method, comprising: obtaining at least one audio clip, and performing voiceprint recognition on the at least one audio clip using a voiceprint recognition model to obtain a first recognition result; when the first recognition result characterizes that the at least one audio clip is a non-target silent clip and the duration of the at least one audio clip is greater than or equal to a first duration threshold, obtaining a highest recognition score in the first recognition result; when the audio duration of the at least one audio clip is greater than or equal to a second duration threshold and the highest recognition score is less than the score threshold, determining that a role corresponding to the at least one audio clip is an unknown role, and the second duration threshold is greater than the first duration threshold; and registering the unknown role in the voiceprint recognition model library.

[0007] Optionally, when the audio duration of the at least one audio segment is greater than or equal to a second duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio segment is an unknown character includes: a first determination step, when the audio duration of the at least one audio segment is less than the second duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio segment is a candidate unknown character; a second determination step, obtaining a subsequent audio segment of the at least one audio segment to obtain a first updated audio segment, and performing the voiceprint recognition on the first updated audio segment to obtain a second recognition result, and when the second recognition result represents the When the audio duration of the first updated audio segment is greater than or equal to the second duration threshold and the highest recognition score is greater than the score threshold, the candidate unknown character is updated to a known character; when the second recognition result indicates that the audio duration of the first updated audio segment is greater than or equal to the second duration threshold and the highest recognition score is less than or equal to the score threshold, the candidate unknown character is updated to the unknown character; when the second recognition result indicates that the audio duration of the first updated audio segment is less than the second duration threshold, the second determination step is repeated until it is determined that the character corresponding to the first updated audio segment is the known character or the unknown character.

[0008] Optionally, when the audio duration of the at least one audio segment is less than the second duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio segment is a candidate unknown character includes: when the audio duration of the at least one audio segment is less than the first duration threshold and the highest recognition score is less than the first score threshold, determining that the character corresponding to the at least one audio segment is the candidate unknown character; when the audio duration of the at least one audio segment is greater than or equal to the first duration threshold and less than a third duration threshold and the highest recognition score is greater than or equal to the first score threshold and less than a second score threshold, determining that the character corresponding to the at least one audio segment is the candidate unknown character, the first duration threshold is less than the third duration threshold, and the first score threshold is less than the second score threshold; when the audio duration of the at least one audio segment is greater than or equal to the third duration threshold and less than the second duration threshold and the highest recognition score is greater than or equal to the second score threshold and less than the third score threshold, determining that the character corresponding to the at least one audio segment is the candidate unknown character, the third duration threshold is less than the second duration threshold, and the second score threshold is less than the third score threshold.

[0009] Optionally, when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio clip is an unknown character includes: when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio clip is the unknown character; when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is greater than or equal to the score threshold, determining that the character corresponding to the at least one audio clip is a known character.

[0010] Optionally, obtaining at least one audio segment and performing voiceprint recognition on the at least one audio segment using a voiceprint recognition model to obtain the first recognition result includes: a third determination step of, if the first recognition result indicates that the at least one audio segment is the target silence segment, obtaining a duration of the at least one audio segment; if the duration of the at least one audio segment is greater than a fourth duration threshold, determining that the at least one audio segment is the target silence segment and the corresponding role is null; a fourth determination step of, if the duration of the at least one audio segment is less than or equal to the fourth duration threshold, obtaining a subsequent audio segment of the at least one audio segment to obtain a second updated audio segment, performing voiceprint recognition on the second updated audio segment to obtain a third recognition result; if the third recognition result indicates that the duration of the second updated audio segment is greater than the fourth duration threshold, determining that the at least one audio segment is the target silence segment; and if the third recognition result indicates that the audio duration of the second updated audio segment is less than or equal to the fourth duration threshold, repeating the fourth determination step until the second updated audio segment is determined to be the target silence segment or the non-target silence segment.

[0011] Optionally, when the duration of the at least one audio segment is greater than the fourth duration threshold, determining the at least one audio segment as the target silent segment includes: when the duration of the at least one audio segment is greater than the second duration threshold, determining the at least one audio segment as the target silent segment; when the duration of the at least one audio segment is less than or equal to the second duration threshold and greater than the fourth duration threshold, determining the at least one audio segment as the target silent segment, and the second duration threshold is greater than the fourth duration threshold.

[0012] Optionally, the method further includes: a fifth determination step, when the historical role is not empty, determining whether the historical role is the same as the current role, wherein the current role is the role corresponding to the current at least one audio clip, and the historical role is the role corresponding to the audio clip before the at least one audio clip; a sixth determination step, when the historical role is the same as the current role, determining that no role switch has occurred; a seventh determination step, when the historical role is not the same as the current role, determining whether the duration of the at least one audio clip is greater than or equal to a second duration threshold, and when the duration of the at least one audio clip is greater than or equal to the second duration threshold, determining that the role switch has occurred; when the duration of the at least one audio clip is less than the second duration threshold, obtaining a subsequent audio clip of the at least one audio clip to obtain a third updated audio clip, and repeating the fifth determination step to the seventh determination step at least once in sequence until it is determined that the role switch has occurred or not, and during the repeated execution, the current role is the role corresponding to the third updated audio clip.

[0013] According to another aspect of an embodiment of the present invention, an audio processing device is provided, comprising: a first acquisition unit for acquiring at least one audio segment, and performing voiceprint recognition on the at least one audio segment using a voiceprint recognition model to obtain a first recognition result; a second acquisition unit for acquiring the highest recognition score in the first recognition result when the first recognition result indicates that the at least one audio segment is a non-target silent segment and the duration of the at least one audio segment is greater than or equal to a first duration threshold; a determination unit for determining that a character corresponding to the at least one audio segment is an unknown character, and the second duration threshold is greater than the first duration threshold, when the audio duration of the at least one audio segment is greater than or equal to a second duration threshold and the highest recognition score is less than a score threshold; and a registration unit for registering the unknown character into the voiceprint recognition model library.

[0014] According to another aspect of the present application, a processor is provided, which is used to run a program, wherein the program executes any one of the processing methods when running.

[0015] According to another aspect of the present application, an audio processing system is provided, which includes: a voiceprint recognition system, one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of the processing methods.

[0016] In an embodiment of the present invention, voiceprint recognition is adopted to achieve the purpose of automatically identifying multiple unknown characters in the audio through an unknown character separation algorithm, thereby achieving the technical effect of voice character separation, and further solving the technical problem of usually having to register the speaker's voiceprint in advance when performing character separation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings that constitute part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application. In the drawings:

[0018] Figure 1 A schematic flow chart of an embodiment of an audio processing method according to the present application is shown;

[0019] Figure 2 A schematic diagram showing a flow chart of unknown character detection according to another embodiment of an audio processing method of the present application is shown;

[0020] Figure 3 A schematic diagram of a silence detection process according to another embodiment of an audio processing method of the present application is shown;

[0021] Figure 4 A schematic diagram of a role switching process according to another embodiment of an audio processing method of the present application is shown;

[0022] Figure 5 A schematic diagram of an embodiment of an audio processing device according to the present application is shown. DETAILED DESCRIPTION

[0023] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] For ease of description, some nouns or terms involved in the embodiments of the present application are explained below:

[0027] Voiceprint recognition: A technology that extracts the speaker's voice characteristics and speech content information to automatically verify the speaker's identity;

[0028] Automatic speech recognition technology: a technology that converts human speech into text.

[0029] As mentioned in the background technology, in the prior art, when voiceprint role separation is performed, the speaker's voiceprint is usually registered in advance before voiceprint role separation can be performed. In order to solve the above problem, a typical embodiment of the present application provides an audio processing method, device, processor and system.

[0030] Figure 1 Flowchart of the audio processing method according to the embodiment of the present application. Figure 1 As shown, the method includes the following steps:

[0031] Step S101: obtaining at least one audio segment, and performing voiceprint recognition on the at least one audio segment using a voiceprint recognition model to obtain a first recognition result.

[0032] In the process of obtaining audio clips, it can be obtained through any one or more of the following methods: ordinary microphone audio collection, array microphone audio collection, hand-in-hand microphone audio collection, computer speakers, mobile phone recording and network audio collection, so that the voice collection mechanism has diversity and can be applied to different scenarios.

[0033] The at least one audio segment mentioned above may be one audio segment or multiple audio segments. In different application scenarios, the number of audio segments may be different.

[0034] The voiceprint recognition model of the present application can be any feasible voiceprint recognition model in the prior art, specifically a template model or a random model. The template model is a non-parametric model that compares the training feature parameters with the test feature parameters, and the distortion between the two is used as the similarity. For example, the VQ (vector quantization) model and the DTW (dynamic time warping) model are vector quantization models. The VQ model generates a codebook through clustering and quantization methods, quantizes and encodes the test data during recognition, and uses the degree of distortion as the judgment criterion. The DTW model compares the input feature vector sequence to be recognized with the feature vector extracted during training and performs recognition through the optimal path matching method. The random model is a parametric model that uses a probability density function to simulate the speaker. The training process is used to predict the parameters of the probability density function, and the matching process is completed by calculating the similarity of the test sentences of the corresponding model. For example, the GMM model is a Gaussian mixture model, which is one of the most effective and commonly used models in text-independent speaker recognition. The HMM model is a hidden Markov model, which is a statistical model used to describe a Markov process with hidden unknown parameters. More specifically, the voiceprint recognition model includes a training phase and a testing phase. The training phase includes four parts: training speech, feature extraction, model training, and model library; the testing phase includes test speech, feature extraction, and scoring judgment.

[0035] Step S102: if the first recognition result indicates that at least one audio segment is a non-target silence segment and the duration of the at least one audio segment is greater than or equal to a first duration threshold, obtaining the highest recognition score in the first recognition result;

[0036] The above-mentioned first recognition result at least includes information indicating whether the at least one audio segment is a non-target silent segment, the duration information of the at least one audio segment, and the corresponding recognition score of the at least one audio segment. In the above steps, when the at least one audio segment is a non-target silent segment, it means that the at least one audio segment is not a silent segment. That is to say, the highest recognition score is obtained only when the at least one audio segment is not a silent segment. Because, when the at least one audio segment is a silent segment, the at least one silent segment does not correspond to any character, so it does not involve the recognition of unknown characters. In addition, if the duration of the at least one audio segment is too short, less than the first duration threshold, the determined character may be inaccurate. Therefore, another prerequisite for obtaining the highest recognition score is that the duration of the at least one audio segment is greater than or equal to the first duration threshold.

[0037] In addition, the highest recognition score mentioned above refers to the highest score obtained after matching at least one audio clip with a role in the model library in the voiceprint recognition model.

[0038] Step S103: If the audio duration of at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, determine that the character corresponding to the at least one audio clip is an unknown character, and the second duration threshold is greater than the first duration threshold;

[0039] In the above steps, if the audio clip is too short, it is impossible to accurately determine whether the corresponding character is an unknown character. Therefore, only when the audio duration of at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, it is considered that there is no character corresponding to the audio in the model library, that is, the character corresponding to the audio is considered to be an unknown character.

[0040] Step S104: register the unknown character into the voiceprint recognition model library.

[0041] In the above audio processing method, by comparing the duration of the audio clip and the relationship between the highest recognition score and the corresponding threshold, it can be determined whether the character corresponding to the audio is an unknown character. If it is determined to be an unknown character, the unknown character is registered in the voiceprint recognition model library. In this way, there is no need to register in advance, and voiceprint role separation can be performed later, thereby solving the problem in the existing technology that requires advance registration to perform voiceprint role separation. Compared with the existing solution that requires advance registration, this solution is easier to use and has a wider range of applicable scenarios.

[0042] In a specific embodiment of the present application, on the basis of including the above steps S101 to S104, the specific step S103 is further refined, and the step specifically includes: step S1031, a first determination step, when the audio duration of the above at least one audio segment is less than the above duration threshold and the above highest recognition score is less than the above score threshold, determining that the role corresponding to the above at least one audio segment is a candidate unknown role, that is, first determining at least one audio segment that may be an unknown role; step S1032, a second determination step, obtaining a subsequent audio segment of the above at least one audio segment to obtain a first updated audio segment, and performing the above voiceprint recognition on the above first updated audio segment to obtain a second recognition result, when the above second recognition result indicates that the audio duration of the above first updated audio segment is greater than or equal to the above duration threshold And when the above-mentioned highest recognition score is greater than the above-mentioned score threshold, the above-mentioned candidate unknown role is updated to a known role. When the above-mentioned second recognition result indicates that the audio duration of the above-mentioned first updated audio segment is greater than or equal to the above-mentioned duration threshold and the above-mentioned highest recognition score is less than or equal to the above-mentioned score threshold, the above-mentioned candidate unknown role is updated to the above-mentioned unknown role. That is, after determining the possible unknown role, subsequent audio segments are obtained, and the subsequent audio segments are recognized. According to the recognition results of the subsequent audio segments, it is determined whether the candidate unknown role is an unknown role; step S1033, when the above-mentioned second recognition result indicates that the audio duration of the above-mentioned first updated audio segment is less than the above-mentioned duration threshold, the above-mentioned second determination step is repeated until it is determined that the role corresponding to the above-mentioned first updated audio segment is the above-mentioned known role or the above-mentioned unknown role. In this method, by first determining the candidate unknown role and then obtaining the subsequent audio segment for determination, the result of determining whether it is an unknown role is more accurate.

[0043] In one embodiment of the present application, based on the above steps S1031 to S1033, the above step S1031 is further refined. Figure 2 is a flow chart of unknown role detection according to an embodiment of the present application, such as Figure 2As shown, the method includes the following steps: when the audio duration of the at least one audio clip is less than the duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio clip is a candidate unknown character, including: when the audio duration of the at least one audio clip is less than the first duration threshold and the highest recognition score is less than the first score threshold, determining that the character corresponding to the at least one audio clip is the candidate unknown character; when the audio duration of the at least one audio clip is greater than or equal to the first duration threshold and less than a third duration threshold and the highest recognition score is greater than or equal to the first score threshold and less than a second score threshold, determining that the character corresponding to the at least one audio clip is the candidate unknown character, the first duration threshold is less than the third duration threshold, and the first score threshold is less than the second score threshold; when the audio duration of the at least one audio clip is greater than or equal to the third duration threshold and less than the second duration threshold and the highest recognition score is greater than or equal to the second score threshold and less than the third score threshold, determining that the character corresponding to the at least one audio clip is the candidate unknown character, the third duration threshold is less than the second duration threshold, and the second score threshold is less than the third score threshold.

[0044] In order to improve the sensitivity of model recognition in the above steps, a variety of different duration thresholds and score thresholds are set. This application sets three duration thresholds and three score thresholds. In the above steps, by comparing the relationship between the duration of the audio clip and the highest recognition score and different corresponding thresholds, the candidate unknown role, that is, the possible unknown role, can be determined first, and then wait for subsequent determination, which can improve the accuracy of the recognition model. Therefore, there are three situations in which the role corresponding to the audio can be considered as a candidate unknown role, that is, a possible unknown role. The first situation: the audio duration of at least one audio clip is less than the first duration threshold and the above-mentioned highest recognition score is less than the first score threshold; the second situation: the audio duration of at least one audio clip is greater than or equal to the above-mentioned first duration threshold and less than the third duration threshold and the highest recognition score is greater than or equal to the above-mentioned first score threshold and less than the second score threshold; the third situation: the audio duration of at least one audio clip is greater than or equal to the above-mentioned third duration threshold and less than the second duration threshold and the above-mentioned highest recognition score is greater than or equal to the above-mentioned second score threshold and less than the third score threshold. In one embodiment of the present application, on the basis of steps S101 to S104, step S103 is further refined, which specifically includes: when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio clip is the unknown character, that is, when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, it can be determined that the character corresponding to the at least one audio clip is an unknown character; when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is greater than or equal to the score threshold, determining that the character corresponding to the at least one audio clip is a known character, that is, when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is greater than the score threshold, it can be determined that the character corresponding to the at least one audio clip is a known character.

[0045] Through the above steps, it is possible to determine whether the character corresponding to at least one audio clip is known or unknown, thus enabling the recognition of unknown characters. After determining whether the audio clip corresponds to a known or unknown character, automatic speech recognition technology can be used to convert the speech into text or machine translation technology can be used. Furthermore, the audio character name and the audio-to-text conversion results can be displayed, enabling functions such as comparing the audio with the processed results, selecting audio clips, and modifying the audio character name.

[0046] In the actual voiceprint recognition process, the audio can be silenced to detect the silent segments in the audio. Only when the audio is not silent can the subsequent role separation and role switching recognition be continued, so as to improve the efficiency of audio recognition. In one embodiment of the present application, based on the above steps S101 to S104, the above step S101 is refined. Figure 3 Flowchart of silence detection according to an embodiment of the present application. Figure 3 As shown, the method includes the following steps: step S1011, a third determination step, in which, when the first recognition result indicates that the at least one audio segment is the target silent segment, the duration of the at least one audio segment is obtained, and when the duration of the at least one audio segment is greater than a fourth duration threshold, the at least one audio segment is determined to be the target silent segment, and the corresponding role is empty, that is, it is considered that the at least one audio segment is silent and has no corresponding role; step S1012, a fourth determination step, in which, when the duration of the at least one audio segment is less than or equal to the fourth duration threshold, a subsequent audio segment of the at least one audio segment is obtained to obtain a second updated audio segment, and the at least one audio segment is updated. The second updated audio segment undergoes voiceprint recognition to obtain a third recognition result. If the third recognition result indicates that the duration of the updated audio segment is greater than the fourth duration threshold, the at least one audio segment is determined to be the target silent segment. That is, after the silent segment is determined, subsequent audio segments are obtained and identified. Based on the recognition results of the subsequent audio segments, whether the segment is silent or not is determined. In step S1013, if the third recognition result indicates that the duration of the second updated audio segment is less than or equal to the fourth duration threshold, the fourth determination step is repeated until the second updated audio segment is determined to be the target silent segment or the non-target silent segment. In this method, by first determining the target silent segment and then obtaining subsequent audio segments for determination, the result of determining whether the target silent segment is obtained is more accurate.

[0047] In one embodiment of the present application, based on the above steps S1011 to S1013, step S1011 is further refined, specifically including: if the duration of the at least one audio segment is greater than the fourth duration threshold, determining the at least one audio segment as the target silence segment, including: if the duration of the at least one audio segment is greater than the third duration threshold, determining the at least one audio segment as the target silence segment, that is, after the target silence segment is determined, if the duration is greater than the third duration threshold, it can be determined as the target silence segment; if the duration of the at least one audio segment is less than or equal to the third duration threshold and greater than the fourth duration threshold, determining the at least one audio segment as the target silence segment, the third duration threshold is greater than the fourth duration threshold, that is, after the target silence segment is determined, if the duration is less than or equal to the third duration threshold and greater than the fourth duration threshold, it can be determined as the target silence segment. Through the above steps, the setting of the fourth duration threshold can eliminate silence caused by speaker pauses in the audio, making the voiceprint recognition result more accurate.

[0048] In actual use, there is also a situation where the speaker changes, so it is necessary to perform role switching detection on the audio. In another embodiment of the present application, based on the above steps S101 to S104, Figure 4 is a flowchart of role switching according to an embodiment of the present application, such as Figure 4As shown, the method includes the following steps: a fifth determination step, when the historical role is not empty, determining whether the historical role is the same as the current role, wherein the current role is the role corresponding to the current at least one audio clip, and the historical role is the role corresponding to the audio clip before the at least one audio clip, that is, first determining whether the role corresponding to the at least one audio clip is the same as the role corresponding to the previous audio clip; a sixth determination step, when the historical role is the same as the current role, determining that no role switch occurs, that is, when at least one audio clip is the same as the role corresponding to the previous audio clip, no role switch occurs; a seventh determination step, when the historical role is not the same as the current role, determining whether the duration of the at least one audio clip is greater than or equal to a third duration threshold, and when the duration of the at least one audio clip is greater than or equal to a third duration threshold, If the third duration threshold is met, the role switch is determined to have occurred. That is, if it is determined that the role corresponding to at least one audio segment is different from that corresponding to the previous audio segment, and the duration satisfies the third duration threshold, a role switch is determined to have occurred. If the duration of the at least one audio segment is less than the third duration threshold, a subsequent audio segment of the at least one audio segment is obtained to obtain a third updated audio segment. The fifth to seventh determination steps are repeated at least once until it is determined that the role switch has occurred or not. During the repeated execution, the current role is the role corresponding to the third updated audio segment. That is, after determining that a role switch has occurred, a subsequent audio segment is obtained, and the subsequent audio segment is identified. Based on the identification result of the subsequent audio segment, whether a role switch has occurred is determined. In this method, by first determining whether a role switch has occurred and then determining it by obtaining the subsequent audio segment, the result of determining whether a role switch has occurred is more accurate. This method can accurately identify whether a role switch has occurred in audio.

[0049] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0050] The present application also provides an audio processing device. It should be noted that the audio processing device of the present application can be used to execute the audio processing method provided in the present application. The following describes the audio processing device provided in the present application.

[0051] Figure 5 Schematic diagram of an audio processing device according to an embodiment of the present application. Figure 5 As shown, the device includes:

[0052] The first acquiring unit 10 is configured to acquire at least one audio segment and perform voiceprint recognition on the at least one audio segment using a voiceprint recognition model to obtain a first recognition result.

[0053] In the process of obtaining audio clips, it can be obtained through any one or more of the following methods: ordinary microphone audio collection, array microphone audio collection, hand-in-hand microphone audio collection, computer speakers, mobile phone recording and network audio collection, so that the voice collection mechanism has diversity and can be applied to different scenarios.

[0054] The at least one audio segment mentioned above may be one audio segment or multiple audio segments. In different application scenarios, the number of audio segments may be different.

[0055] The voiceprint recognition model of the present application can be any feasible voiceprint recognition model in the prior art, specifically a template model or a random model. The template model is a non-parametric model that compares the training feature parameters with the test feature parameters, and the distortion between the two is used as the similarity. For example, the VQ (vector quantization) model and the DTW (dynamic time warping) model are vector quantization models. The VQ model generates a codebook through clustering and quantization devices, quantizes and encodes the test data during recognition, and uses the degree of distortion as the judgment criterion. The DTW model compares the input feature vector sequence to be recognized with the feature vector extracted during training, and performs recognition through an optimal path matching device. The random model is a parametric model that uses a probability density function to simulate the speaker. The training process is used to predict the parameters of the probability density function, and the matching process is completed by calculating the similarity of the test sentences of the corresponding model. For example, the GMM model is a Gaussian mixture model, which is one of the most effective and commonly used models in text-independent speaker recognition. The HMM model is a hidden Markov model, which is a statistical model used to describe a Markov process with hidden unknown parameters. More specifically, the voiceprint recognition model includes a training phase and a testing phase. The training phase includes four parts: training speech, feature extraction, model training, and model library; the testing phase includes test speech, feature extraction, and scoring judgment.

[0056] The second obtaining unit 20 is configured to obtain a highest recognition score in the first recognition result when the first recognition result indicates that the at least one audio segment is a non-target silence segment and the duration of the at least one audio segment is greater than or equal to a first duration threshold;

[0057] The above-mentioned first recognition result at least includes information indicating whether the at least one audio segment is a non-target silent segment, the duration information of the at least one audio segment, and the corresponding recognition score of the at least one audio segment. In the above-mentioned unit, when the at least one audio segment is a non-target silent segment, it indicates that the at least one audio segment is not a silent segment. That is to say, the highest recognition score is obtained only when the at least one audio segment is not a silent segment. Because, when the at least one audio segment is a silent segment, the at least one silent segment does not correspond to any character, so it does not involve the recognition of unknown characters. In addition, if the duration of the at least one audio segment is too short, less than the first duration threshold, the determined character may be inaccurate. Therefore, another prerequisite for obtaining the highest recognition score is that the duration of the at least one audio segment is greater than or equal to the first duration threshold.

[0058] In addition, the highest recognition score mentioned above refers to the highest score obtained after at least one audio clip is matched with a role in the model library in the voiceprint recognition model.

[0059] The determining unit 30 is configured to determine that the character corresponding to the at least one audio segment is an unknown character, if the audio duration of the at least one audio segment is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, and the second duration threshold is greater than the first duration threshold;

[0060] In the above unit, if the audio clip is too short, it is impossible to accurately determine whether the corresponding character is an unknown character. Therefore, only when the audio duration of at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, it is considered that there is no character corresponding to the audio in the model library, that is, the character corresponding to the audio is considered to be an unknown character.

[0061] The registration unit 40 is used to register the unknown character into the library of the voiceprint recognition model.

[0062] In the above-mentioned audio processing device, by comparing the duration of the audio clip and the relationship between the highest recognition score and the corresponding threshold, it can be determined whether the character corresponding to the audio is an unknown character. If it is determined to be an unknown character, the unknown character is registered in the voiceprint recognition model library. In this way, there is no need to register in advance, and voiceprint role separation can be performed later, thereby solving the problem in the existing technology that requires advance registration to perform voiceprint role separation. Compared with the existing solution that requires advance registration, this solution is easier to use and has a wider range of applicable scenarios.

[0063] In a specific embodiment of the present application, on the basis of including the above-mentioned first acquisition unit, second acquisition unit, determination unit and registration unit, the specific determination unit is further refined, and the determination unit includes a first determination module, a second determination module and a third determination module, wherein the first determination module is used to determine that the role corresponding to the above-mentioned at least one audio segment is a candidate unknown role when the audio duration of the above-mentioned at least one audio segment is less than the above-mentioned duration threshold and the above-mentioned highest recognition score is less than the above-mentioned score threshold, that is, first determine at least one audio segment that may be an unknown role; the second determination module is used to obtain a subsequent audio segment of the above-mentioned at least one audio segment to obtain a first updated audio segment, and perform the above-mentioned voiceprint recognition on the above-mentioned first updated audio segment to obtain a second recognition result, and the above-mentioned second recognition result indicates that the audio duration of the above-mentioned first updated audio segment is greater than or equal to When the duration threshold is met and the highest recognition score is greater than the score threshold, the candidate unknown character is updated to a known character. When the second recognition result indicates that the audio duration of the first updated audio segment is greater than or equal to the duration threshold and the highest recognition score is less than or equal to the score threshold, the candidate unknown character is updated to the unknown character. That is, after determining the possible unknown character, subsequent audio segments are obtained, and the subsequent audio segments are recognized. It is determined whether the candidate unknown character is an unknown character based on the recognition results of the subsequent audio segments. The third determination module is used to repeatedly execute the second determination module when the second recognition result indicates that the audio duration of the first updated audio segment is less than the duration threshold, until it is determined that the character corresponding to the first updated audio segment is the known character or the unknown character. In this device, by first determining the candidate unknown character and then obtaining the subsequent audio segment for determination, the result of determining whether the character is an unknown character is more accurate.

[0064] In one embodiment of the present application, on the basis of including the above-mentioned first determination module, the second determination module and the third determination module, the specific first determination module is further refined, and the module specifically includes a first determination submodule, a second determination submodule and a third determination submodule, wherein the first determination submodule is used to determine that the role corresponding to the above-mentioned at least one audio segment is the above-mentioned candidate unknown role when the audio duration of the above-mentioned at least one audio segment is less than the first duration threshold and the above-mentioned highest recognition score is less than the first score threshold; the second determination submodule is used to determine that the role corresponding to the above-mentioned at least one audio segment is the above-mentioned candidate unknown role when the audio duration of the above-mentioned at least one audio segment is greater than or equal to the above-mentioned first duration threshold and less than the third duration threshold and the above-mentioned highest recognition score is greater than or equal to When the above-mentioned first score threshold is greater than or equal to the above-mentioned third score threshold and is less than the second score threshold, the character corresponding to the above-mentioned at least one audio clip is determined to be the above-mentioned candidate unknown character, the above-mentioned first duration threshold is less than the above-mentioned third duration threshold, and the above-mentioned first score threshold is less than the above-mentioned second score threshold; the third determination submodule is used to determine that when the audio duration of the above-mentioned at least one audio clip is greater than or equal to the above-mentioned third duration threshold and less than the second duration threshold and the above-mentioned highest recognition score is greater than or equal to the above-mentioned second score threshold and less than the third score threshold, the character corresponding to the above-mentioned at least one audio clip is determined to be the above-mentioned candidate unknown character, the above-mentioned third duration threshold is less than the above-mentioned second duration threshold, and the above-mentioned second score threshold is less than the above-mentioned third score threshold.

[0065] In order to improve the sensitivity of model recognition, the above-mentioned unit sets a variety of different duration thresholds and score thresholds. The present application sets three duration thresholds and three score thresholds. In the above-mentioned device, by comparing the relationship between the duration of the audio clip and the highest recognition score and different corresponding thresholds, the candidate unknown role, i.e., the possible unknown role, can be determined first, and then wait for subsequent determination, which can improve the accuracy of the recognition model. Therefore, there are three situations in which the role corresponding to the audio can be considered as a candidate unknown role, i.e., a possible unknown role. The first situation is: the audio duration of at least one audio clip is less than the first duration threshold and the highest recognition score is less than the first score threshold; the second situation is: the audio duration of at least one audio clip is greater than or equal to the first duration threshold and less than the third duration threshold and the highest recognition score is greater than or equal to the first score threshold and less than the second score threshold; the third situation is: the audio duration of at least one audio clip is greater than or equal to the third duration threshold and less than the second duration threshold and the highest recognition score is greater than or equal to the second score threshold and less than the third score threshold.

[0066] In one embodiment of the present application, on the basis of including the above-mentioned first acquisition unit, second acquisition unit, determination unit and registration unit, the specific determination unit is further refined, and the unit specifically includes: a fourth determination module, which is used to determine that the character corresponding to the at least one audio clip is the unknown character when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, that is, when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, the character corresponding to the at least one audio clip can be determined as an unknown character; a fifth determination module, which is used to determine that the character corresponding to the at least one audio clip is a known character when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is greater than or equal to the score threshold, that is, when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is greater than the score threshold, the character corresponding to the at least one audio clip can be determined as a known character.

[0067] The above units can be used to determine whether the character corresponding to at least one audio clip is known or unknown, thus enabling the recognition of unknown characters. After determining whether the audio corresponds to a known or unknown character, automatic speech recognition technology can be used to convert the speech into text or machine translation technology can be used. Furthermore, the audio character name and the audio-to-text conversion results can be displayed, enabling functions such as comparing the audio with the processed results, selecting audio clips, and modifying the audio character name.

[0068] In the actual voiceprint recognition process, silence detection can be performed on the audio, the purpose of which is to identify silent segments in the audio, and only when the audio is non-silent will the subsequent role separation and role switching recognition be continued, so as to improve the efficiency of audio recognition. In one embodiment of the present application, on the basis of including the above-mentioned first acquisition unit, the second acquisition unit, the determination unit and the registration unit, the specific first acquisition unit is refined, and the first acquisition unit includes a sixth determination module, a seventh determination module and an eighth determination module, wherein the sixth determination module is used to obtain the duration of the above-mentioned at least one audio segment when the above-mentioned first recognition result characterizes that the above-mentioned at least one audio segment is the above-mentioned target silent segment, and when the duration of the above-mentioned at least one audio segment is greater than the fourth duration threshold, it is determined that the above-mentioned at least one audio segment is the above-mentioned target silent segment, and the corresponding role is empty, that is, it is considered that the at least one audio segment is silent and has no corresponding role; the seventh determination module is used to obtain the duration of the above-mentioned at least one audio segment when the duration of the above-mentioned at least one audio segment is less than or equal to the above-mentioned fourth duration threshold. The subsequent audio segment of the at least one audio segment is used to obtain a second updated audio segment, and the voiceprint recognition is performed on the second updated audio segment to obtain a third recognition result. When the third recognition result indicates that the duration of the updated audio is greater than the fourth duration threshold, the at least one audio segment is determined to be the target silent segment. That is, after the silent segment is determined, the subsequent audio segment is obtained and recognized. The silent segment or the non-silent segment is determined based on the recognition result of the subsequent audio segment. The eighth determination module is used to repeatedly execute the fourth determination module when the third recognition result indicates that the audio duration of the second updated audio segment is less than or equal to the fourth duration threshold, until the second updated audio segment is determined to be the target silent segment or the non-target silent segment. In this device, by first determining the target silent segment and then obtaining the subsequent audio segment for determination, the result of determining whether it is the target silent segment is more accurate.

[0069] In one embodiment of the present application, based on the sixth determination module, the seventh determination module, and the eighth determination module, the sixth determination module is further refined to include: a fourth determination submodule and a fifth determination submodule, wherein the fourth determination submodule is configured to determine that the at least one audio segment is the target silence segment if the duration of the at least one audio segment is greater than a third duration threshold, i.e., after the target silence segment is determined, if the duration is greater than the third duration threshold, it can be determined as the target silence segment; and the fifth determination submodule is configured to determine that the at least one audio segment is the target silence segment if the duration of the at least one audio segment is less than or equal to the third duration threshold and greater than the fourth duration threshold, i.e., after the target silence segment is determined, if the duration is less than or equal to the third duration threshold and greater than the fourth duration threshold, it can be determined as the target silence segment. Through the above device, the setting of the fourth duration threshold can eliminate silence caused by speaker pauses in the audio, thereby making the voiceprint recognition result more accurate.

[0070] During actual use, there are also situations where the speaker changes, so it is necessary to perform role switching detection on the audio. In another embodiment of the present application, based on the above-mentioned first acquisition unit, second acquisition unit, determination unit and registration unit, it also includes a ninth determination module, a tenth determination module, an eleventh determination module and a twelfth determination module, wherein the ninth determination module is used to determine whether the historical role is the same as the current role when the historical role is not empty, wherein the above-mentioned current role is the role corresponding to the current at least one audio segment, and the above-mentioned historical role is the role corresponding to the audio segment before the above-mentioned at least one audio segment, that is, first determine whether the role corresponding to the at least one audio segment is the same as the role corresponding to the previous audio segment; the tenth determination module is used to determine that no role switching has occurred when the above-mentioned historical role is the same as the above-mentioned current role, that is, when the role corresponding to at least one audio segment is the same as the role corresponding to the previous audio segment, no role switching has occurred; the eleventh determination module is used to determine whether the duration of the above-mentioned at least one audio segment is greater than when the above-mentioned historical role is not the same as the above-mentioned current role. or equal to a third duration threshold, and if the duration of the at least one audio segment is greater than or equal to the third duration threshold, determining that the role switch has occurred, that is, if it is determined that the role corresponding to the at least one audio segment is different from the previous audio segment, and the duration satisfies the third duration threshold, determining that a role switch has occurred; a twelfth determination module is configured to, if the duration of the at least one audio segment is less than the third duration threshold, obtain a subsequent audio segment of the at least one audio segment to obtain a third updated audio segment, and repeat the ninth to eleventh determination modules at least once until it is determined that the role switch has occurred or not. During the repeated execution, the current role is the role corresponding to the third updated audio segment, that is, after determining that a role switch has occurred, a subsequent audio segment is obtained, and the subsequent audio segment is identified, and whether a role switch has occurred is determined based on the identification result of the subsequent audio segment. In this device, by first determining whether a role switch has occurred and then obtaining a subsequent audio segment to determine it, the result of determining whether a role switch has occurred is more accurate. With this device, accurate identification of whether a role switch has occurred in audio can be achieved.

[0071] The above-mentioned audio processing device includes a processor and a memory. The above-mentioned first acquisition unit, second acquisition unit, determination unit and registration unit are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.

[0072] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the kernel parameters can be adjusted to achieve the recognition of the unknown character corresponding to the audio.

[0073] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0074] An embodiment of the present invention provides a processor, which is used to run a program, wherein the audio processing method is executed when the program is run.

[0075] An embodiment of the present invention provides a device, comprising a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, at least the following steps are performed:

[0076] Step S101: obtaining at least one audio clip, and performing voiceprint recognition on the at least one audio clip using a voiceprint recognition model to obtain a first recognition result;

[0077] Step S102: if the first recognition result indicates that the at least one audio segment is a non-target silence segment and the duration of the at least one audio segment is greater than or equal to a first duration threshold, obtaining the highest recognition score in the first recognition result;

[0078] Step S103: If the audio duration of the at least one audio clip is greater than or equal to the duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio clip is an unknown character;

[0079] Step S104: register the unknown character into the library of the voiceprint recognition model.

[0080] The devices in this article can be servers, PCs, PADs, mobile phones, etc.

[0081] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program for initializing at least the following method steps:

[0082] Step S101: obtaining at least one audio clip, and performing voiceprint recognition on the at least one audio clip using a voiceprint recognition model to obtain a first recognition result;

[0083] Step S102: if the first recognition result indicates that the at least one audio segment is a non-target silence segment and the duration of the at least one audio segment is greater than or equal to a first duration threshold, obtaining the highest recognition score in the first recognition result;

[0084] Step S103: If the audio duration of the at least one audio clip is greater than or equal to the duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio clip is an unknown character;

[0085] Step S104: register the unknown character into the library of the voiceprint recognition model.

[0086] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0087] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the above-mentioned units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0088] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0089] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0090] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods of each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0091] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0092] 1) In the audio processing method of the present application, at least one audio clip is first obtained, and a voiceprint recognition model is used to perform voiceprint recognition on the at least one audio clip to obtain a first recognition result. Then, when the first recognition result indicates that the at least one audio clip is a non-target silent clip and the duration of the at least one audio clip is greater than or equal to the first duration threshold, the highest recognition score in the first recognition result is obtained; when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, it is determined that the character corresponding to the at least one audio clip is an unknown character, and the second duration threshold is greater than the first duration threshold; finally, the unknown character is registered in the voiceprint recognition model library. By setting the recognition score threshold and the audio duration threshold and judging the audio, it is possible to detect unknown characters appearing in the audio, and register the detected unknown characters in the voiceprint recognition model library.

[0093] 2) In the audio processing device of the present application, the first acquisition unit is used to acquire at least one audio segment, and use a voiceprint recognition model to perform voiceprint recognition on the at least one audio segment to obtain a first recognition result. The at least one audio segment can be one audio segment or multiple audio segments. In different application scenarios, the number of audio segments may be different. The second acquisition unit is used to obtain the highest recognition score in the first recognition result when the first recognition result characterizes that the at least one audio segment is a non-target silent segment and the duration of the at least one audio segment is greater than or equal to the first duration threshold; when the at least one audio segment is a non-target silent segment, it indicates that the at least one audio segment is not a silent segment, that is, the highest recognition score is obtained when the at least one audio segment is not a silent segment. In addition, if the duration of the at least one audio segment is too short, less than the first duration threshold, the determined role may be inaccurate. Therefore, another prerequisite for obtaining the highest recognition score is that the duration of the at least one audio segment is greater than or equal to the first duration threshold. The determination unit is used to determine that the role corresponding to the at least one audio clip is an unknown role when the audio duration of the at least one audio clip is greater than or equal to the second duration threshold and the highest recognition score is less than the score threshold, and the second duration threshold is greater than the first duration threshold; if the audio clip is too short, it is impossible to accurately determine whether the corresponding role is an unknown role. The registration unit is used to register the above-mentioned unknown role to the library of the above-mentioned voiceprint recognition model. In the above-mentioned audio processing device, by comparing the relationship between the duration of the audio clip and the highest recognition score and the corresponding threshold, it can be determined whether the role corresponding to the audio is an unknown role. When it is determined to be an unknown role, the unknown role is registered in the model library of voiceprint recognition. In this way, there is no need to register in advance, and voiceprint role separation can be performed later, thereby solving the problem in the prior art that prior registration is required for voiceprint role separation. Compared with the prior art solution that requires prior registration, this solution is easier to use and has a wider range of applicable scenarios.

[0094] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Persons skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. An audio processing method, characterized in that: include: Obtaining at least one audio clip, and performing voiceprint recognition on the at least one audio clip using a voiceprint recognition model to obtain a first recognition result; If the first recognition result indicates that the at least one audio segment is a non-target silence segment and the duration of the at least one audio segment is greater than or equal to a first duration threshold, obtaining a highest recognition score in the first recognition result; If the audio duration of the at least one audio clip is greater than or equal to a second duration threshold and the highest recognition score is less than a score threshold, determining that the character corresponding to the at least one audio clip is an unknown character, and the second duration threshold is greater than the first duration threshold; Registering the unknown character into the model library of the voiceprint recognition; The method further includes: when the audio duration of the at least one audio segment is less than the second duration threshold and greater than or equal to the first duration threshold, A first determining step, in which, when the audio duration of the at least one audio segment is less than the second duration threshold and the highest recognition score is less than the score threshold, the character corresponding to the at least one audio segment is determined to be a candidate unknown character; A second determination step is to obtain a subsequent audio segment of the at least one audio segment to obtain a first updated audio segment, and perform the voiceprint recognition on the first updated audio segment to obtain a second recognition result. If the second recognition result indicates that the audio duration of the first updated audio segment is greater than or equal to the second duration threshold and the highest recognition score is greater than the score threshold, the candidate unknown role is updated to a known role. If the second recognition result indicates that the audio duration of the first updated audio segment is greater than or equal to the second duration threshold and the highest recognition score is less than or equal to the score threshold, the candidate unknown role is updated to the unknown role. If the second recognition result indicates that the audio duration of the first updated audio segment is less than the second duration threshold, the second determination step is repeated until it is determined that the character corresponding to the first updated audio segment is the known character or the unknown character.

2. The method according to claim 1, characterized in that The method further comprises: When the audio duration of the at least one audio segment is less than a first duration threshold and the highest recognition score is less than a first score threshold, determining that the character corresponding to the at least one audio segment is the candidate unknown character; When the audio duration of the at least one audio segment is less than the second duration threshold and the highest recognition score is less than the score threshold, determining that the character corresponding to the at least one audio segment is a candidate unknown character includes: Determining that the character corresponding to the at least one audio clip is the candidate unknown character if the audio duration of the at least one audio clip is greater than or equal to the first duration threshold and less than a third duration threshold, and the highest recognition score is greater than or equal to the first score threshold and less than a second score threshold, the first duration threshold is less than the third duration threshold, and the first score threshold is less than the second score threshold; When the audio duration of the at least one audio clip is greater than or equal to the third duration threshold and less than the second duration threshold and the highest recognition score is greater than or equal to the second score threshold and less than the third score threshold, the character corresponding to the at least one audio clip is determined to be the candidate unknown character, the third duration threshold is less than the second duration threshold, and the second score threshold is less than the third score threshold.

3. The method according to claim 1, characterized in that The method further comprises: When the audio duration of the at least one audio segment is greater than or equal to the second duration threshold and the highest recognition score is greater than or equal to the score threshold, it is determined that the character corresponding to the at least one audio segment is a known character.

4. The method according to claim 1, wherein Acquiring at least one audio clip and performing voiceprint recognition on the at least one audio clip using a voiceprint recognition model to obtain a first recognition result includes: A third determining step, if the first recognition result indicates that the at least one audio segment is the target silent segment, obtaining a duration of the at least one audio segment; if the duration of the at least one audio segment is greater than a fourth duration threshold, determining that the at least one audio segment is the target silent segment and the corresponding role is null; a fourth determining step, if the duration of the at least one audio segment is less than or equal to the fourth duration threshold, obtaining a subsequent audio segment of the at least one audio segment to obtain a second updated audio segment, performing the voiceprint recognition on the second updated audio segment to obtain a third recognition result, and determining the at least one audio segment as the target silence segment if the third recognition result indicates that the duration of the second updated audio segment is greater than the fourth duration threshold; When the third recognition result indicates that the audio duration of the second updated audio segment is less than or equal to the fourth duration threshold, the fourth determination step is repeated until it is determined that the second updated audio segment is the target silent segment or the non-target silent segment.

5. The method according to claim 4, characterized in that When the duration of the at least one audio segment is greater than the fourth duration threshold, determining the at least one audio segment as the target silence segment includes: When the duration of the at least one audio segment is greater than a second duration threshold, determining the at least one audio segment as the target silence segment; If the duration of the at least one audio segment is less than or equal to the second duration threshold and greater than the fourth duration threshold, the at least one audio segment is determined to be the target silence segment, and the second duration threshold is greater than the fourth duration threshold.

6. The audio processing method according to claim 1, characterized in that: Also includes: A fifth determination step, if the historical role is not empty, determining whether the historical role is the same as the current role, wherein the current role is the role corresponding to the current at least one audio segment, and the historical role is the role corresponding to the audio segment before the at least one audio segment; A sixth determining step, in a case where the historical role is the same as the current role, determining that no role switch has occurred; a seventh determining step, if the historical role is different from the current role, determining whether a duration of the at least one audio segment is greater than or equal to a second duration threshold, and determining that the role switch occurs if the duration of the at least one audio segment is greater than or equal to the second duration threshold; When the duration of the at least one audio segment is less than the second duration threshold, a subsequent audio segment of the at least one audio segment is obtained to obtain a third updated audio segment, and the fifth determination step to the seventh determination step are repeated at least once in sequence until it is determined that the role switch occurs or does not occur. During the repeated execution, the current role is the role corresponding to the third updated audio segment.

7. An audio processing device, characterized in that: The processing device comprises: a first acquiring unit, configured to acquire at least one audio clip, and perform voiceprint recognition on the at least one audio clip using a voiceprint recognition model to obtain a first recognition result; a second acquiring unit, configured to acquire a highest recognition score in the first recognition result, if the first recognition result indicates that the at least one audio segment is a non-target silence segment and the duration of the at least one audio segment is greater than or equal to a first duration threshold; a determining unit, configured to determine that the character corresponding to the at least one audio clip is an unknown character if the audio duration of the at least one audio clip is greater than or equal to a second duration threshold and the highest recognition score is less than a score threshold, and the second duration threshold is greater than the first duration threshold; The device is further configured to, when the audio duration of the at least one audio segment is less than the second duration threshold and greater than or equal to the first duration threshold, performing a first determining step, determining that the character corresponding to the at least one audio segment is a candidate unknown character when the audio duration of the at least one audio segment is less than the second duration threshold and the highest recognition score is less than the score threshold; Executing a second determination step, obtaining a subsequent audio segment of the at least one audio segment to obtain a first updated audio segment, and performing the voiceprint recognition on the first updated audio segment to obtain a second recognition result; if the second recognition result indicates that the audio duration of the first updated audio segment is greater than or equal to the second duration threshold and the highest recognition score is greater than the score threshold, updating the candidate unknown role to a known role; if the second recognition result indicates that the audio duration of the first updated audio segment is greater than or equal to the second duration threshold and the highest recognition score is less than or equal to the score threshold, updating the candidate unknown role to the unknown role; If the second recognition result indicates that the audio duration of the first updated audio segment is less than the second duration threshold, repeatedly performing the second determining step until it is determined that the character corresponding to the first updated audio segment is the known character or the unknown character; A registration unit is used to register the unknown role into the model library of the voiceprint recognition.

8. A processor, characterized in that: The processor is configured to run a program, wherein the program executes the method according to any one of claims 1 to 6 when running.

9. An audio processing system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image acquisition control method and device and acquisition terminal

    CN110505399A

  • Automatic identity recognition method based on voiceprint information of speaker

    CN113113022A

  • Utterance cutting and dividing system and method therefor

    JP2022071960A