An audio separation and script violation reminding method, device and computer equipment
By processing the audio feature vector twice and utilizing a speaker prediction model and an audio clustering model, the problem of low speaker identification accuracy in existing technologies is solved, and a higher audio separation accuracy is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, the trained speaker prediction model only processes the current audio segment, resulting in low speaker identification accuracy and audio segmentation accuracy.
After the audio feature vector is processed in the first step, the predicted speaker identifier is determined by the trained speaker prediction model. Then, the target predicted speaker identifier is corrected in the second step by the trained audio clustering model to obtain the target speaker identifier.
This improved the accuracy of speaker identification, which in turn improved the accuracy of audio separation.
Smart Images

Figure CN116631432B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to a method, device, and computer equipment for audio separation and speech violation alerts. Background Technology
[0002] Currently, when performing speech segmentation on audio to be separated, most methods only utilize a trained speaker prediction model to process each audio segment in the audio to be separated, obtaining a corresponding predicted speaker identifier. Based on the same predicted speaker identifier, the corresponding audio segments are then merged to obtain the target audio corresponding to each predicted speaker identifier. However, because the trained speaker prediction model only obtains the predicted speaker identifier based on the current audio segment, the accuracy of the determined predicted speaker identifier is low, resulting in low accuracy in audio segmentation.
[0003] Improving the accuracy of audio segmentation is a problem that urgently needs to be solved in existing technologies. Summary of the Invention
[0004] To address the problems in the prior art, embodiments of this specification provide an audio separation and speech violation alert method, apparatus, and computer device. After performing a first processing on the audio feature vector obtained from the audio segment to obtain a predicted speaker identifier, a second correction processing is performed on the audio feature vector corresponding to the target predicted speaker identifier in the predicted speaker identifier, to obtain the target speaker identifier. This achieves correction of the predicted speaker identifier, improving the accuracy of speaker identifier determination, and thus improving the accuracy of audio separation.
[0005] To solve the above-mentioned technical problems, the specific technical solution in this specification is as follows:
[0006] On the one hand, the embodiments of this specification provide an audio separation method, including,
[0007] Speech segmentation is performed on the audio to be separated to obtain multiple audio segments;
[0008] For each audio segment, feature extraction is performed to obtain an audio feature vector corresponding to the audio segment;
[0009] The trained speaker prediction model is used to process the audio feature vector to determine the predicted speaker identifier corresponding to each audio segment;
[0010] The trained audio clustering model is used to process the audio feature vector corresponding to the target predicted speaker identifier to obtain the target speaker identifier corresponding to the audio feature vector. The target predicted speaker identifier is determined from the predicted speaker identifier.
[0011] The audio segments corresponding to the same target speaker identifier are merged to obtain the target audio corresponding to each target speaker identifier.
[0012] Furthermore, the audio to be separated is subjected to speech segmentation, resulting in multiple audio segments, which further include:
[0013] Based on the audio segmentation model, the audio to be separated is processed to obtain multiple pre-audio segments;
[0014] Determine whether each of the pre-audio segments includes speech information; and
[0015] If it is determined that the pre-audio segment includes the speech information, then the pre-audio segment is determined to be the audio segment.
[0016] Furthermore, the training method for this trained speaker prediction model further includes,
[0017] Each sample audio segment is labeled, and a label corresponding to each sample audio segment is determined. The label is the identifier of the real speaker corresponding to the sample audio segment.
[0018] A preset speaker prediction model is used to process the sample audio feature vector corresponding to each sample audio segment to obtain the corresponding sample speaker identifier; and
[0019] Based on the difference between the sample speaker identifier and the label, the preset speaker prediction model is trained to obtain the trained speaker prediction model.
[0020] Furthermore, the determination of the speaker identifier for this target prediction further includes:
[0021] For each predicted speaker identifier corresponding to a given time moment, determine whether the predicted speaker identifier for the previous moment corresponding to the moment preceding the given time moment and the predicted speaker identifier for the moment following the given time moment are both different; and
[0022] If it is determined that the current predicted speaker identifier is consistent with the previous predicted speaker identifier and / or the next predicted speaker identifier, then multiple predicted speaker identifiers that are consecutively identical to the current predicted speaker identifier are identified as the target predicted speaker identifier.
[0023] Further, after determining, for each time-corresponding predicted speaker identifier, whether the predicted speaker identifier of the previous moment corresponding to the moment before the current moment and the predicted speaker identifier of the moment after the current moment are both different, the process includes:
[0024] If it is determined that the current predicted speaker identifier is inconsistent with both the previous and next predicted speaker identifiers, then the current predicted speaker identifier is determined to be the target speaker identifier.
[0025] Furthermore, the trained audio clustering model includes a trained feature weight set and a trained feature prediction model, and the training method of the trained audio clustering model further includes,
[0026] Using a preset set of feature weights, the sample audio feature vectors corresponding to multiple consecutive sample sub-audio segments are processed to obtain feature values corresponding to each sample sub-audio segment.
[0027] A preset feature prediction model is used to process each feature value to obtain a prediction result identifier corresponding to each sample sub-audio; and
[0028] By utilizing the difference between the predicted result identifier and the sub-label corresponding to the sample sub-audio segment, the preset feature weight set and the preset feature prediction model are trained to obtain the trained feature weight set and the trained feature prediction model, thus obtaining the trained audio clustering model.
[0029] The preset feature weight set includes multiple weight data. The weight data corresponding to each sample sub-audio segment is determined based on the time of each sample sub-audio segment and the time corresponding to each weight data.
[0030] Furthermore, the weight data corresponding to each time point is determined based on an exponential forgetting model.
[0031] On the other hand, the embodiments of this specification also provide a method for reminding users of script violations, including,
[0032] From the audio to be separated, determine the target audio corresponding to each target speaker identifier;
[0033] Based on the acoustic model, each target audio is identified to obtain text information corresponding to each target audio.
[0034] The trained illegal language recognition model is used to process the text information to determine whether it is illegal; and
[0035] If the text information is determined to be in violation of regulations, the target speaker identifier corresponding to the text information is identified as the user identifier to be reminded of the violation, in order to issue a violation reminder.
[0036] Wherein, the determination of the target audio corresponding to each target speaker identifier from the audio to be separated is determined according to any of the audio separation methods described above.
[0037] Furthermore, the trained illegal language recognition model is used to process the text information to determine whether it is illegal text information. This further includes...
[0038] Each piece of text information is verified using a pre-set violation corpus. If the text information matches the target violation prediction in the pre-set violation corpus, the text information is determined to be the violation text information.
[0039] On the other hand, embodiments of this specification also provide an audio separation device, including,
[0040] The segmentation unit is used to perform speech segmentation on the audio to be separated, resulting in multiple audio segments;
[0041] The feature extraction unit is used to extract features for each audio segment to obtain an audio feature vector corresponding to the audio segment.
[0042] The first processing unit is used to process the audio feature vector using the trained speaker prediction model to determine the predicted speaker identifier corresponding to each audio segment;
[0043] The second processing unit uses a trained audio clustering model to process the audio feature vector corresponding to the target predicted speaker identifier, obtaining the target speaker identifier corresponding to the audio feature vector, wherein the target predicted speaker identifier is determined from the predicted speaker identifier; and
[0044] The merging unit is used to merge the audio segments corresponding to the same target speaker identifier to obtain target audio corresponding to each target speaker identifier.
[0045] On the other hand, embodiments of this specification also provide a speech violation reminder device, including,
[0046] The first determining unit is used to determine the target audio corresponding to each target speaker identifier from the audio to be separated;
[0047] The recognition unit is configured to recognize each target audio based on an acoustic model, and obtain text information corresponding to each target audio; and
[0048] The third processing unit is used to process the text information using the trained violation language recognition model to determine whether the text information is violation text information; and
[0049] The second determining unit is used to, when determining that the text information is illegal text information, determine the target speaker identifier corresponding to the text information as the user identifier to be reminded of the violation, so as to issue a violation reminder.
[0050] The first determining unit is determined based on the audio separation device.
[0051] On the other hand, embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.
[0052] On the other hand, embodiments of this specification also provide a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-described method.
[0053] On the other hand, embodiments of this specification also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the above-described method.
[0054] Using the embodiments of this specification, upon receiving audio to be separated, speech segmentation is performed to obtain multiple audio segments; features are extracted from each audio segment to obtain an audio feature vector corresponding to that segment; a trained speaker prediction model is used to process the audio feature vectors to determine a predicted speaker identifier for each audio segment; a trained audio clustering model is used to process the audio feature vectors corresponding to the target predicted speaker identifiers to obtain the target speaker identifiers, which are determined from the predicted speaker identifiers; and audio segments corresponding to the same target speaker identifier are merged to obtain the target audio corresponding to each target speaker identifier. After the first processing of the audio feature vectors obtained from the audio segments to obtain the predicted speaker identifiers, a second correction processing is performed on the audio feature vectors corresponding to the target predicted speaker identifiers within the predicted speaker identifiers to obtain the target speaker identifier. Thus, a second processing of the predicted speaker identifiers is implemented, improving the accuracy of speaker identifier determination and consequently improving the accuracy of audio separation. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 The diagram shown is a schematic representation of an implementation system for an audio separation method according to an embodiment of this specification.
[0057] Figure 2 The diagram shown is a flowchart of an audio separation method according to an embodiment of this specification;
[0058] Figure 3 The diagram shown is a flowchart of an audio segment determination method according to an embodiment of this specification;
[0059] Figure 4 The diagram shown is a flowchart of a method for determining the speaker identifier in target prediction according to an embodiment of this specification;
[0060] Figure 5 The diagram shown is a flowchart of a training method for a trained audio clustering model according to an embodiment of this specification.
[0061] Figure 6 The diagram shown is a flowchart of a method for reminding users of prohibited speech patterns according to an embodiment of this specification.
[0062] Figure 7 The diagram shown is a structural schematic of an audio separation device according to an embodiment of this specification.
[0063] Figure 8 The diagram shown is a structural schematic of a speech violation reminder device according to an embodiment of this specification;
[0064] Figure 9 This is a schematic diagram of the structure of a computer device according to an embodiment of this specification.
[0065] [Explanation of Labels in the Attached Image]
[0066] 101. User terminal;
[0067] 102. Server;
[0068] 710. Segmentation unit;
[0069] 720. Feature extraction unit;
[0070] 730. First processing unit;
[0071] 740. Second processing unit;
[0072] 750. Merging Units;
[0073] 810. First Determined Unit;
[0074] 820. Identification unit;
[0075] 830. Third processing unit;
[0076] 840. Second Determined Unit;
[0077] 902. Computer equipment;
[0078] 904. Processing equipment;
[0079] 906. Storage resources;
[0080] 908. Drive mechanism;
[0081] 910. Input / Output Module;
[0082] 912. Input devices;
[0083] 914. Output devices;
[0084] 916. Presentation equipment;
[0085] 918. Graphical User Interface;
[0086] 920. Network interface;
[0087] 922. Communication link;
[0088] 924. Communication bus. Detailed Implementation
[0089] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0090] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0091] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0092] Figure 1 The diagram shown is a schematic of an implementation system for an audio separation method according to an embodiment of this specification. The system may include a user terminal 101 and a server 102. The user terminal 101 and the server 102 communicate with each other through a network. The network may include a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof, and is connected to a website, user equipment (e.g., a computing device), and a backend system. The user can send the audio to be separated to the server 102 through the user terminal 101. The server 102 performs speech segmentation on the received audio to be separated, obtaining multiple audio segments; extracts features for each audio segment, obtaining an audio feature vector corresponding to the audio segment; processes the audio feature vector using a trained speaker prediction model to determine the predicted speaker identifier corresponding to each audio segment; processes the audio feature vector corresponding to the target predicted speaker identifier using a trained audio clustering model, obtaining the target speaker identifier corresponding to the audio feature vector, and the target predicted speaker identifier is determined from the predicted speaker identifier; and merges the audio segments corresponding to the same target speaker identifier to obtain the target audio corresponding to each target speaker identifier, and sends the target audio to the user terminal 101.
[0093] Furthermore, the server 102 can also identify each target audio based on an acoustic model to obtain text information corresponding to each target audio; use the trained illegal speech recognition model to process the text information to determine whether the text information is illegal text information; and if the text information is determined to be illegal text information, determine the target speaker identifier corresponding to the text information as the identifier of the illegal user to be reminded, so as to send the illegal reminder information to the user terminal 101 to remind the user of the illegality.
[0094] Alternatively, server 102 may be a node of a cloud computing system (not shown in the figure), or each server 102 may be a separate cloud computing system comprising multiple computers interconnected by a network and operating as a distributed processing system.
[0095] In an optional embodiment, the user terminal 101 may include electronic devices, including but not limited to smartphones, data acquisition devices, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, and other similar electronic devices. Optionally, the operating system running on the electronic device may include, but is not limited to, Android, iOS, Linux, Windows, etc.
[0096] In addition, it should be noted that, Figure 1 The example shown is merely one application environment provided in this manual. In actual applications, it may include multiple user terminals 101, and this manual does not impose any restrictions.
[0097] like Figure 2 The diagram shows a flowchart of an audio separation method according to an embodiment of this specification. The audio separation process is depicted in this figure, but may include more or fewer steps based on conventional or non-inventive methods. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible order. In actual system or device products, the method can be executed sequentially or in parallel according to the embodiment or the accompanying drawings. Specifically, as shown... Figure 2 As shown, the method may include:
[0098] S210 performs speech segmentation on the audio to be separated, resulting in multiple audio segments;
[0099] S220, perform feature extraction for each audio segment to obtain the audio feature vector corresponding to the audio segment;
[0100] S230: The trained speaker prediction model is used to process the audio feature vector to determine the predicted speaker identifier corresponding to each audio segment.
[0101] S240, using the trained audio clustering model to process the audio feature vector corresponding to the target predicted speaker identifier, to obtain the target speaker identifier corresponding to the audio feature vector;
[0102] S250, merge the audio segments corresponding to the same target speaker identifier to obtain the target audio corresponding to each target speaker identifier.
[0103] Using the embodiments of this specification, upon receiving audio to be separated, speech segmentation is performed to obtain multiple audio segments; features are extracted from each audio segment to obtain an audio feature vector corresponding to that segment; a trained speaker prediction model is used to process the audio feature vectors to determine a predicted speaker identifier for each audio segment; a trained audio clustering model is used to process the audio feature vectors corresponding to the target predicted speaker identifiers to obtain the target speaker identifiers, which are determined from the predicted speaker identifiers; and audio segments corresponding to the same target speaker identifier are merged to obtain the target audio corresponding to each target speaker identifier. After the first processing of the audio feature vectors obtained from the audio segments to obtain the predicted speaker identifiers, a second correction processing is performed on the audio feature vectors corresponding to the target predicted speaker identifiers within the predicted speaker identifiers to obtain the target speaker identifier. Thus, a second processing of the predicted speaker identifiers is implemented, improving the accuracy of speaker identifier determination and consequently improving the accuracy of audio separation.
[0104] According to one embodiment of this specification, the audio to be separated is speech that needs to be separated according to different speakers. For example, when the speech of two people talking is taken as the audio to be separated, the final separation result is two target audios, each target audio corresponding to a corresponding speaker, and each target audio includes at least one audio segment.
[0105] Upon receiving the audio to be separated, the audio is segmented into multiple audio segments. Specifically, any method for segmenting speech can be used to segment the audio, such as segmenting by speech frames or by duration.
[0106] A feature extraction model is used to extract features from each separated audio segment, resulting in a corresponding audio feature vector. Specifically, this feature extraction model can be, for example, any trained neural network model capable of extracting feature vectors from speech segments. This audio feature vector is a feature vector that indicates the speaker.
[0107] Using sample audio segments and the corresponding labels, a preset speaker prediction model is trained to obtain the trained speaker prediction model. The label is the real speaker label corresponding to the sample audio segment.
[0108] The trained speaker prediction model is used to process the audio feature vectors to determine the predicted speaker identifiers corresponding to each feature vector. For example, audio segments include A, B, C, and D, and the corresponding audio feature vectors for these segments are x, y, z, and t, respectively. The trained speaker prediction model is then used to process x, y, z, and t separately to obtain the corresponding predicted speaker identifiers F, F, E, and W. Therefore, it can be determined that F spoke about A and B, E spoke about C, and W spoke about D.
[0109] After processing the audio feature vectors once, the target speaker identifier that meets the preset conditions is determined from the predicted speaker identifiers. The trained audio clustering model is then used to process the audio feature vectors corresponding to the target speaker identifiers again to obtain the final target speaker identifier. Specifically, meeting the preset conditions can be, for example, a predicted speaker identifier corresponding to an audio segment not at the first moment. During the training phase, the trained audio clustering model predicts based on multi-latented audio feature vectors, where each latent state is associated with an audio feature vector. Simultaneously, the audio feature vectors are sorted according to the order of the speaking moments to obtain an audio feature vector sequence. Based on this sequence, a latent state sequence is obtained. In each prediction during training, the audio feature vectors used are the audio feature vectors of all speaking moments preceding the current speaking moment (i.e., all previous latent states). This achieves prediction based on the audio feature vectors of all previous speaking moments, improving the accuracy of speaker identifier prediction. If the preset conditions allow for a predicted speaker identifier corresponding to an audio segment not at the first moment, the predicted speaker identifier corresponding to the first moment is determined as the target speaker identifier.
[0110] After determining the target speaker identifier corresponding to each audio segment, audio segments corresponding to the same speaker identifier are merged to obtain the target audio corresponding to each target speaker identifier. Specifically, for example, consecutive audio segments corresponding to the same speaker identifier are directly merged, while non-consecutive audio segments corresponding to the same speaker identifier are merged at intervals to obtain the target audio corresponding to the target speaker identifier. For example, audio segments A, B, C, D, and T correspond to target speaker identifiers F, F, E, W, and F, respectively. Therefore, it can be determined that F spoke AB and T, E spoke C, and W spoke D.
[0111] According to another embodiment of this specification, the training method of the trained speaker prediction model includes: labeling each sample audio segment to determine a label corresponding to each sample audio segment, wherein the label is the real speaker identifier corresponding to the sample audio segment; processing the sample audio feature vector corresponding to each sample audio segment using a preset speaker prediction model to obtain the corresponding sample speaker identifier; and training the preset speaker prediction model based on the difference between the sample speaker identifier and the label to obtain the trained speaker prediction model.
[0112] Specifically, based on the differences between the speaker identifiers and labels in the sample, a preset speaker prediction model is trained to obtain a trained speaker prediction model. This model can be obtained by processing the speaker identifiers and labels in the sample using a first loss function to obtain a first loss function value. The preset speaker prediction model is then trained based on this first loss function value to obtain a trained speaker prediction model that satisfies a first training condition. This first training condition can be, for example, the number of training iterations and / or the convergence of a sequence composed of the first loss function values.
[0113] Figure 3 The diagram shows a flowchart of an audio segment determination method according to an embodiment of this specification. This diagram depicts a process for determining an audio segment, but based on conventional or non-creative labor, it may include more or fewer operational steps. Specifically, as shown... Figure 3 As shown, the method may include:
[0114] S311, based on the audio segmentation model, processes the audio to be separated to obtain multiple pre-audio segments;
[0115] S312, determine whether the pre-audio segment includes speech information;
[0116] S313, if it is determined that the pre-audio segment includes speech information, the pre-audio segment is determined to be an audio segment;
[0117] S314, if it is determined that the pre-audio segment does not contain speech information, delete the pre-audio segment.
[0118] Using the embodiments in this specification, when performing speech segmentation on audio to be separated, a fine segmentation granularity ensures that each segmented audio segment corresponds to a different speaker. However, a finer segmentation granularity results in multiple blank audio segments, i.e., audio segments without any speaker. If speaker identification prediction is also performed on these blank audio segments, it leads to a waste of resources.
[0119] According to another embodiment of this specification, an audio segmentation model is used to process the audio to be separated to obtain multiple pre-audio segments. Specifically, the audio to be separated is segmented at a preset granularity to obtain multiple pre-audio segments. The preset granularity can be any granularity that ensures that the segmented audio segments correspond to different speakers.
[0120] After obtaining multiple pre-audio segments, for each pre-audio segment, it is determined whether the pre-audio segment contains speech information. Specifically, it is determined whether the pre-audio segment contains a sound of a preset volume or a sound corresponding to a preset timbre. If at least one of these is determined to exist, the pre-audio segment is determined to contain speech information.
[0121] If a pre-audio segment is determined to contain speech information, it is identified as an audio segment for speaker identification prediction; if it is determined not to contain speech information, the pre-audio segment is deleted. This reduces resource waste and increases the speed of audio separation.
[0122] Figure 4 The diagram shows a flowchart of a method for determining a target predicted speaker identifier according to an embodiment of this specification. While the process of determining a target predicted speaker identifier is described in this figure, it may include more or fewer operational steps based on conventional or non-inventive labor. Specifically, as shown... Figure 4 As shown, the method may include:
[0123] S441, for the current predicted speaker identifier corresponding to each time point, determine whether the predicted speaker identifier of the previous moment corresponding to the previous time point and the predicted speaker identifier of the next moment corresponding to the next time point are both different.
[0124] S442, if it is determined that the current predicted speaker identifier is consistent with the predicted speaker identifier of the previous moment and / or the predicted speaker identifier of the next moment, multiple predicted speaker identifiers that are consecutively identical with the current predicted speaker identifier are identified as target predicted speaker identifiers;
[0125] S443, if it is determined that the current predicted speaker identifier is inconsistent with the previous predicted speaker identifier and the next predicted speaker identifier, the current predicted speaker identifier shall be determined as the target speaker identifier.
[0126] According to another embodiment of this specification, if two consecutive audio segments spoken by different speakers are predicted to have the same predicted speaker identifier, it will lead to incorrect merging of the final target audio and a low accuracy rate. Therefore, further prediction processing is required for consecutive identical predicted speaker identifiers.
[0127] For each time-time corresponding to the current predicted speaker identifier, it is checked whether the predicted speaker identifier of the previous moment corresponding to the previous time and the predicted speaker identifier of the next moment corresponding to the next time are all different. If they are not different, multiple consecutively identical predicted speaker identifiers are identified as target predicted speaker identifiers for this current predicted speaker identifier. Then, the trained audio clustering model is used to process the target predicted speaker identifier again.
[0128] If all are determined to be different, then the predicted speaker identifier is determined as the target speaker identifier. This is because the subsequent processing after erroneous segmentation is relatively simpler than the subsequent processing after erroneous merging. Therefore, at this point, only the audio segments that would lead to erroneous merging can be processed to reasonably reduce the waste of resources.
[0129] Figure 5 The diagram shows a flowchart of a training method for a trained audio clustering model according to an embodiment of this specification. This diagram describes the training process of the trained audio clustering model, but based on conventional or non-creative work, it may include more or fewer operational steps. Specifically, as shown... Figure 5 As shown, the method may include:
[0130] S510: Using a preset set of feature weights, the sample audio feature vectors corresponding to multiple consecutive sample sub-audio segments are processed to obtain the feature values corresponding to each sample sub-audio segment.
[0131] S520 uses a preset feature prediction model to process each feature value and obtains the prediction result identifier corresponding to each sample sub-audio.
[0132] S530 utilizes the difference between the prediction result identifier and the sub-label corresponding to the sample sub-audio segment to train a preset feature weight set and a preset feature prediction model, thereby obtaining a trained feature weight set and a trained feature prediction model, and thus obtaining a trained audio clustering model.
[0133] Using the embodiments of this specification, preset feature weights corresponding to multiple audio feature vectors are trained to make predictions based on audio feature vectors from all previous speech moments, thereby improving the accuracy of speaker identification prediction.
[0134] According to another embodiment of this specification, the trained audio clustering model includes a trained feature weight set and a trained feature prediction model, which may be, for example, a trained fully connected layer. Before the audio feature vectors enter the fully connected layer, they are processed using a preset feature weight set to ensure that when predicting the speaker identifier for each audio feature vector, all previous latent states (audio feature vectors) are considered.
[0135] The preset feature weight set includes multiple weight data. During each training iteration, the number of weight data is automatically determined. For example, for each sample audio segment, its position within a plurality of consecutive input sample audio segments is determined (e.g., which position). This position is then used to determine the number of weight data. For example, if there are eight consecutive input sample audio segments, and the third sample audio segment has a position of three, then the number of weight data corresponding to that sample audio segment is also three. Further, based on the time (position) of each sample audio segment and the time (position) corresponding to each weight data, the weight data corresponding to each sample audio segment is determined. For example, based on the same time, the weight data corresponding to each sample audio segment is determined; this weight data is determined based on that time (position). For example, there are eight consecutive sample audio segments in the input. The third sample audio segment corresponds to the third weight data in the preset feature weight set. The preset feature weight set includes eight weight data. The weight data is determined based on at least one of the following: the current time of the weight data, the duration between the current time and the first time, and the position. For example, the earlier the position, the smaller the weight data.
[0136] Using a preset feature weight set, the sample audio feature vectors corresponding to multiple consecutive sample sub-audio segments are processed to obtain feature values corresponding to each sample sub-audio segment. Specifically, each sample sub-audio segment is processed to obtain a sample audio feature vector; then, using the weight data in the preset feature weight set, the corresponding sample audio feature vector is processed to obtain corresponding sub-feature values, which constitute the feature values. The feature values include multiple sub-feature values.
[0137] A pre-defined feature prediction model is used to process each feature value to obtain a prediction result identifier corresponding to each sample sub-audio segment. Then, a second loss function is used to process this prediction result identifier and the sub-label corresponding to the sample sub-audio segment to obtain a second loss function value. The sub-label is the identifier of the real speaker corresponding to the sample sub-audio segment. Based on this second loss function value, the pre-defined feature weight set and the pre-defined feature prediction model are trained to obtain a trained feature weight set and a trained feature prediction model that satisfy the second training condition, thus obtaining a trained audio clustering model.
[0138] According to another embodiment of this specification, the weight data corresponding to each time step is determined based on an exponential forgetting model.
[0139] For example, one could input at least one of the following into an exponential forgetting model: the current time, the duration between the current time and the first time, and the position of the current time. This would yield weighted data corresponding to the current time. Specifically, for example, the larger the current time, the larger the weighted data; the longer the duration between the current time and the first time, the larger the weighted data; and the larger the position, the larger the weighted data. The exponential forgetting model could be a function corresponding to the forgetting curve.
[0140] Figure 6 The diagram shown is a flowchart of a method for alerting users to script violations according to an embodiment of this specification. While the process of alerting users to script violations is described in this diagram, it may include more or fewer operational steps based on conventional or non-creative work. Specifically, as shown... Figure 6 As shown, the method may include:
[0141] S610, determine the target audio corresponding to each target speaker identifier from the audio to be separated;
[0142] S620, based on an acoustic model, identifies each target audio and obtains the text information corresponding to each target audio.
[0143] S630 uses a trained illegal language recognition model to process text information to determine whether the text information is illegal.
[0144] S640, if it is determined that the text information is illegal, the target speaker identifier corresponding to the text information is identified as the user identifier to be reminded of the violation, so as to issue a violation reminder.
[0145] Using the embodiments described in this specification, currently, due to the low accuracy of determining the target audio corresponding to each target speaker identifier when separating audio to be separated, the accuracy of violation alerts is also low. This application addresses this by performing a first processing on the audio feature vector obtained from the audio segment to obtain the predicted speaker identifier, and then performing a second correction processing on the audio feature vector corresponding to the target predicted speaker identifier within the predicted speaker identifier to obtain the target speaker identifier. This achieves a second processing of the predicted speaker identifier, improving the accuracy of speaker identifier determination, and thus improving the accuracy of audio separation. Based on the target audio corresponding to each target speaker identifier with higher accuracy, violation alerts are then issued, further improving the accuracy of violation alerts.
[0146] According to another embodiment of this specification, based on such Figures 2-5Any of the methods can be used to determine the target audio corresponding to each target speaker identifier from the audio to be separated.
[0147] The acoustic model is a trained neural network model that can recognize any audio and obtain the corresponding text information. The training process of this acoustic model is similar to that of the trained speaker prediction model, and will not be described in detail here.
[0148] The trained violation-related text recognition model can be any binary classification model. This model can determine whether the text to be judged is violation-related. The training process of this model is similar to that of existing binary classification models and will not be described in detail here.
[0149] The trained violation speech recognition model is used to process the text information. If the text information is determined to be violation-related, the target speaker identifier corresponding to the text information is identified as the user identifier to be alerted for violation, and a violation alert is issued. If the text information is determined not to be violation-related, the next text information is evaluated. The method of issuing the violation alert can be arbitrarily determined; for example, when a violation alert is determined to be required, a prompt indicating the violation speech can be given, or speech interference can be used.
[0150] According to another embodiment of this specification, the text information is processed using a trained illegal speech recognition model to determine whether the text information is illegal text information. This includes: verifying each text information using a preset illegal speech corpus, and determining that the text information is illegal text information if it is determined that the text information matches the target illegal speech corpus in the preset illegal speech corpus.
[0151] Multiple violation corpora are pre-set in a pre-defined violation corpus. Each violation corpus consists of text information that requires violation notification. Violation corpora can be added or deleted at any time in this pre-defined violation corpus.
[0152] A text similarity recognition model is used to process each piece of text and each piece of non-compliant text separately. For each piece of text, if the similarity between it and a target piece of non-compliant text is greater than or equal to a preset condition, the text is determined to be non-compliant. The text similarity recognition model can be any function that can determine the similarity between two pieces of text, such as the cosine similarity function.
[0153] Figure 7 The diagram shown is a structural schematic of an audio separation device according to an embodiment of this specification. Figure 7 As shown, including,
[0154] The segmentation unit 710 is used to perform speech segmentation on the audio to be separated, resulting in multiple audio segments;
[0155] The feature extraction unit 720 is used to extract features for each audio segment to obtain an audio feature vector corresponding to the audio segment;
[0156] The first processing unit 730 is used to process the audio feature vector using the trained speaker prediction model to determine the predicted speaker identifier corresponding to each audio segment.
[0157] The second processing unit 740 uses a trained audio clustering model to process the audio feature vector corresponding to the target predicted speaker identifier, obtaining the target speaker identifier corresponding to the audio feature vector. The target predicted speaker identifier is determined from the predicted speaker identifier; and
[0158] The merging unit 750 is used to merge audio segments that correspond to the same target speaker identifier to obtain target audio corresponding to each target speaker identifier.
[0159] Since the principle of the above-mentioned device in solving the problem is similar to that of the above-mentioned method, the implementation of the above-mentioned device can refer to the implementation of the above-mentioned method, and the repeated parts will not be described again.
[0160] Figure 8 The diagram shown is a structural schematic of a speech violation reminder device according to an embodiment of this specification. Figure 8 As shown, including,
[0161] The first determining unit 810 is used to determine the target audio corresponding to each target speaker identifier from the audio to be separated;
[0162] Recognition unit 820 is used to recognize each target audio based on an acoustic model, and obtain text information corresponding to each target audio; and
[0163] The third processing unit 830 is used to process text information using the trained illegal speech recognition model to determine whether the text information is illegal; and
[0164] The second determining unit 840 is used to, when determining that the text information is illegal, identify the target speaker identifier corresponding to the text information as the user identifier to be reminded of the violation, so as to issue a violation reminder.
[0165] Among them, the first determining unit 810 is based on Figure 7 The device is determined.
[0166] Since the principle of the above-mentioned device in solving the problem is similar to that of the above-mentioned method, the implementation of the above-mentioned device can refer to the implementation of the above-mentioned method, and the repeated parts will not be described again.
[0167] like Figure 9 This is a schematic diagram of the structure of a computer device according to an embodiment of this specification. The apparatus in this specification can be the computer device in this embodiment, performing the methods described in this specification. The computer device 902 may include one or more processing devices 904, such as one or more central processing units (CPUs), each processing unit implementing one or more hardware threads. The computer device 902 may also include any storage resource 906 for storing information of any kind, such as code, settings, data, etc. Non-limitingly, for example, the storage resource 906 may include any one or more combinations of: any type of RAM, any type of ROM, flash memory, hard disk, optical disk, etc. More generally, any storage resource can use any technology to store information. Further, any storage resource can provide volatile or non-volatile retention of information. Further, any storage resource can represent a fixed or removable component of the computer device 902. In one case, when the processing device 904 executes associated instructions stored in any storage resource or combination of storage resources, the computer device 902 can perform any operation of the associated instructions. The computer device 902 also includes one or more drive mechanisms 908 for interacting with any storage resource, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.
[0168] Computer device 902 may also include an input / output module 910 (I / O) for receiving various inputs (via input device 912) and providing various outputs (via output device 914). A specific output mechanism may include a presentation device 916 and an associated graphical user interface (GUI) 918. In other embodiments, the input / output module 910 (I / O), input device 912, and output device 914 may be omitted, and the device may function solely as a computer device within a network. Computer device 902 may also include one or more network interfaces 920 for exchanging data with other devices via one or more communication links 922. One or more communication buses 924 couple the components described above together.
[0169] Communication link 922 can be implemented in any way, such as via a local area network (LAN), a wide area network (WAN) (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 922 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0170] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0171] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.
[0172] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0174] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0175] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0176] The above specific embodiments further illustrate the purpose, technical solutions, and beneficial effects of this specification. It should be understood that the above are merely specific embodiments of this specification and are not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.
Claims
1. An audio separation method, characterized by, Comprise: Voice segmentation is carried out on the audio to be separated, and a plurality of audio segments are obtained; Feature extraction is carried out on each of the audio segments, and an audio feature vector corresponding to the audio segment is obtained; The trained speaker prediction model is used to process the audio feature vector to determine a predicted speaker identifier corresponding to each of the audio segments; The trained audio clustering model is used to process the audio feature vector corresponding to the target predicted speaker identifier to obtain a target speaker identifier corresponding to the audio feature vector, wherein the target predicted speaker identifier is determined from the predicted speaker identifier; and The audio segments corresponding to the same target speaker identifier are merged to obtain a target audio corresponding to each of the target speaker identifiers; The determination of the target predicted speaker identifier comprises: For the current predicted speaker identifier corresponding to each time, it is judged whether the previous moment predicted speaker identifier corresponding to the previous time of the time and the next moment predicted speaker identifier corresponding to the next time of the time are both different; and In the case where it is determined that the current predicted speaker identifier is consistent with the previous moment predicted speaker identifier and / or the next moment predicted speaker identifier, a plurality of predicted speaker identifiers which are continuously the same as the current predicted speaker identifier are determined as the target predicted speaker identifier.
2. The method of claim 1, wherein, The voice segmentation is carried out on the audio to be separated, and a plurality of audio segments are obtained, comprising: Based on the audio segmentation model, the audio to be separated is processed to obtain a plurality of pre-audio segments; It is judged whether each of the pre-audio segments includes voice information; and In the case where it is determined that the pre-audio segment includes the voice information, the pre-audio segment is determined as the audio segment.
3. The method of claim 1, wherein, The training method of the trained speaker prediction model comprises: Each sample audio segment is labeled to determine a label corresponding to each sample audio segment, wherein the label is a real speaker identifier corresponding to the sample audio segment; A preset speaker prediction model is used to process a sample audio feature vector corresponding to each of the sample audio segments to obtain a corresponding sample speaker identifier; and Based on the difference between the sample speaker identifier and the label, the preset speaker prediction model is trained to obtain the trained speaker prediction model.
4. The method of claim 1, wherein, After the current predicted speaker identifier corresponding to each time is determined, it is judged whether the previous moment predicted speaker identifier corresponding to the previous time of the time and the next moment predicted speaker identifier corresponding to the next time of the time are both different, further comprising: In the case where it is determined that the current predicted speaker identifier is inconsistent with the previous moment predicted speaker identifier and the next moment predicted speaker identifier, the current predicted speaker identifier is determined as the target speaker identifier.
5. The method of claim 1, wherein, The trained audio clustering model comprises a trained feature weight set and a trained feature prediction model, and the training method of the trained audio clustering model comprises: The preset feature weight set is used to process sample audio feature vectors corresponding to a plurality of continuous sample sub-audio segments, to obtain feature values corresponding to each of the sample sub-audio segments; Each of the feature values is processed by using a preset feature prediction model to obtain a prediction result identifier corresponding to each of the sample sub-audio segments; and The preset feature weight set and the preset feature prediction model are trained by using the difference between the prediction result identifier and a sub-label corresponding to the sample sub-audio segment, to obtain a trained feature weight set and a trained feature prediction model, so as to obtain the trained audio clustering model, The preset feature weight set includes a plurality of weight data, and the weight data corresponding to each of the sample sub-audio segments is determined according to the time of each sample sub-audio segment and the time corresponding to each of the weight data.
6. The method of claim 5, wherein, The weight data corresponding to each of the times is determined based on an exponential forgetting model.
7. A script violation alerting method, characterized by, It includes: From the audio to be separated, determine the target audio corresponding to each target speaker identifier; Based on the acoustic model, each of the target audio is identified to obtain the text information corresponding to each of the target audio; The trained violation rhetoric identification model is used to process the text information to determine whether the text information is a violation text information; And In the case where the text information is determined to be a violation text information, the target speaker identifier corresponding to the text information is determined to be a to-be-reminded violation user identifier, so as to perform a violation reminder, The method for determining the target audio corresponding to each target speaker identifier from the audio to be separated is determined according to any one of claims 1-6.
8. The method of claim 7, wherein, The trained violation rhetoric identification model is used to process the text information to determine whether the text information is a violation text information, which includes: Each of the text information is verified by using a preset violation corpus, and in the case where the text information is determined to match a target violation corpus in the preset violation corpus, the text information is determined to be the violation text information.
9. An audio separating apparatus, characterized by comprising: It includes: The segmentation unit is used for speech segmentation on the audio to be separated to obtain a plurality of audio segments; The feature extraction unit is used for feature extraction on each of the audio segments to obtain an audio feature vector corresponding to the audio segment; The first processing unit is used for processing the audio feature vector by using the trained speaker prediction model to determine a predicted speaker identifier corresponding to each of the audio segments; The second processing unit is used for processing the audio feature vector corresponding to a target predicted speaker identifier by using the trained audio clustering model to obtain a target speaker identifier corresponding to the audio feature vector, wherein the target predicted speaker identifier is determined from the predicted speaker identifier; and The merging unit is used for merging the audio segments corresponding to the same target speaker identifier to obtain a target audio corresponding to each of the target speaker identifiers. The second processing unit is further configured to: for a current predicted speaker identifier corresponding to each time point, determine whether a previous moment predicted speaker identifier corresponding to a previous time point of the time point and a next moment predicted speaker identifier corresponding to a next time point of the time point are both different; and in a case where the current predicted speaker identifier is consistent with the previous moment predicted speaker identifier and / or the next moment predicted speaker identifier, determine a plurality of predicted speaker identifiers that are continuously same as the current predicted speaker identifier as the target predicted speaker identifier.
10. A script violation alerting apparatus, characterized by, Comprise: A first determining unit configured to determine target audio corresponding to each target speaker identifier from audio to be separated; An identifying unit configured to identify each target audio based on an acoustic model to obtain text information corresponding to each target audio; And A third processing unit configured to process the text information by using a trained rule-violating rhetoric identification model to determine whether the text information is rule-violating text information; And A second determining unit configured to determine, in a case where the text information is rule-violating text information, a target speaker identifier corresponding to the text information as a rule-violating user identifier to be reminded, so as to perform rule violation reminding. The first determining unit is determined by the apparatus according to claim 9.
11. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is run by the processor to execute the method of any one of claims 1-8.
13. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the method according to any one of claims 1-8. The computer program / instruction is executed by the processor to implement the method according to any one of claims 1-8.
Citation Information
Patent Citations
Speaker clustering method for distributed microphone
CN102074236A
Audio signal processing method and related product
CN110111808A