Speech processing method, speech processing apparatus, electronic device, and storage medium
By classifying and denoising voice data, voice segments with individual user identifiers are generated, solving the problem of low accuracy of voice-to-text in multi-person speaking scenarios and achieving more accurate text generation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUANGZHOU SHIYUAN ELECTRONICS CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-07-23
AI Technical Summary
In multi-person speaking scenarios, existing speech processing models generate speech text containing errors, resulting in low text accuracy.
By classifying the voice data to be processed into events, identifying voice segments with individual user identifiers, and generating standard text based on these voice segments, combined with noise reduction and dereverberation processing, accurate text content is generated using speech recognition technology.
It improves the accuracy of voice-to-text translation, avoids text errors caused by multiple people speaking simultaneously, enhances the accuracy of event classification results, and ensures the clarity and accuracy of the text.
Smart Images

Figure CN2026071691_23072026_PF_FP_ABST
Abstract
Description
Speech processing method, speech processing apparatus, electronic device, and storage medium
[0001] The present application claims priority from the Chinese patent application No. 2025100774640 filed on January 17, 2025, and entitled "Speech processing method, speech processing apparatus, electronic device, and storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of speech technology, and in particular to a speech processing method, a speech processing apparatus, an electronic device, and a storage medium in the field of speech technology. BACKGROUND
[0003] In order to facilitate recording of the speech content of a speaker in a multi-person speech scene such as a meeting, teaching, and court trial site, the speaker's speech can be input into a speech processing model, and the speaker's speech is converted into text by the speech processing model to generate speech text corresponding to the speaker's speech.
[0004] However, when the speech processing model generates speech text corresponding to the speaker's speech, due to the complexity and diversity of the speaker's speech, there may be errors in the generated speech text, which reduces the accuracy of the speech text.
[0005] Therefore, when generating text corresponding to speech data, how to ensure the accuracy of the text is a problem that needs to be solved at present. SUMMARY
[0006] The present application provides a speech processing method, a speech processing apparatus, an electronic device, and a storage medium, which can ensure the accuracy of the text when generating text corresponding to speech data.
[0007] In a first aspect, the present application provides a speech processing method, comprising: obtaining speech data to be processed; determining a first speech segment in the speech data to be processed; wherein the first speech segment represents a plurality of continuous frames of speech data in the speech data to be processed; performing event classification on the first speech segment to obtain a classification result; determining a target speech segment based on the classification result; wherein the user identifier corresponding to the target speech segment is a single user identifier; and generating a target standard text based on the target speech segment.
[0008] In the embodiment of the present application, when the speech data to be processed is obtained, the speech data of the continuous multiple frames in the speech data to be processed is classified by events, the speech segment (i.e., the target speech segment) in which the user identifier is a single user identifier in the speech data to be processed is determined, and the standard text (i.e., the target standard text) is generated from the target speech segment. Since the user identifier of the target speech segment is a single user identifier, it indicates that there is only the speech of one user in the target speech segment. Therefore, when the standard text is generated from the target speech segment, the problem of the inaccuracy of the standard text caused by the multiple-person speech data corresponding to the multiple users speaking at the same time in the speech data to be processed can be avoided, thereby ensuring the accuracy of the text when the text corresponding to the speech data is generated.
[0009] In addition, by classifying the speech data of the continuous multiple frames in the speech data to be processed by events, the problem that the duration of one frame of speech data is too short to reflect the event type corresponding to the frame of speech data can be avoided, thereby improving the accuracy of the event classification result. Furthermore, on the basis of the more accurate event classification result, the accuracy of the text is further ensured.
[0010] In combination with the first aspect and the above implementation manners, in some implementation manners of the first aspect, the above event classification of the first speech segment to obtain the classification result comprises: determining the similarity of the first speech segment to each of the plurality of preset events; determining whether the maximum similarity in the plurality of similarities is greater than or equal to a preset similarity; and in the case that the maximum similarity is greater than or equal to the preset similarity, determining the event type corresponding to the maximum similarity as the classification result.
[0011] In the embodiment of the present application, when the maximum similarity in the similarity of the first speech segment to each of the plurality of preset events is greater than or equal to the preset similarity, the similarity of the first speech segment to the classification result can be ensured to the greatest extent, and the problem that the maximum similarity is small in similarity to the classification result and causes classification error can be avoided, thereby improving the accuracy of the event classification result of the first speech segment. Furthermore, on the basis of the more accurate event classification result of the first speech segment, the accuracy of the text is further ensured.
[0012] In combination with the first aspect and the above implementation manners, in some implementation manners of the first aspect, the above determination of the target speech segment based on the classification result comprises: in the case that the classification result is the event type corresponding to a single user identifier, determining the first speech segment as the target speech segment.
[0013] In the embodiment of the present application, in the case that the classification result is the event type corresponding to the single user identifier, it is illustrated that there is only one user's speech in the first speech segment, and the first speech segment can be determined as the target speech segment, thereby avoiding the problem that in the case that the classification result is the event type corresponding to the speech of multiple users, the multi-person speech data corresponding to the speech of multiple users is determined as the target speech segment, and thereby the accuracy of the target speech segment is improved. Furthermore, on the basis of the more accurate target speech segment, the accuracy of the text is further ensured.
[0014] With reference to the first aspect and the implementation manners above, in some implementation manners of the first aspect, the speech processing method further includes: pre-processing the speech data to be processed to obtain pre-processed speech data; wherein the pre-processing includes de-noise and de-reverberation processing; and the determining the first speech segment in the speech data to be processed includes: acquiring the first speech segment in the pre-processed speech data.
[0015] In the embodiment of the present application, by performing de-noise and de-reverberation processing on the speech data to be processed, the noise and reflected sound of the speech data to be processed can be reduced, and the pre-processed speech data is more accurate, and thereby the first speech segment in the pre-processed speech data can also be more accurate. Furthermore, on the basis of the more accurate first speech segment, the accuracy of the text is further ensured.
[0016] With reference to the first aspect and the implementation manners above, in some implementation manners of the first aspect, the generating the target standard text based on the target speech segment includes: performing speech recognition on the target speech segment to generate first text content corresponding to the target speech segment; and generating a first event type identifier corresponding to the target speech segment, and a second event type identifier corresponding to a second speech segment in the speech data to be processed; wherein the second speech segment is a speech segment in the speech data to be processed except the target speech segment; and generating the target standard text based on the first text content, the first event type identifier and the second event type identifier.
[0017] In the embodiment of the present application, when the corresponding text content (i.e. the first text content) is generated by performing speech recognition on the target speech segment, the event type identifier corresponding to the target speech segment (i.e. the first event type identifier) and the event type identifier corresponding to the speech segment (i.e. the second speech segment) in the speech data to be processed except the target speech segment (i.e. the second event type identifier) are also generated, so that the target standard text generated based on the first text content, the first event type identifier and the second event type identifier has the event type identifier of the target speech segment and the event type identifier of the non-target speech segment in addition to the text content corresponding to the target speech segment, and thereby the final generated standard text is clearer, and the accuracy of the text is further ensured.
[0018] With the first aspect and the above implementation manners, in some implementation manners of the first aspect, the text content of the target user is included in the first text content; the speech recognition on the target speech segment to generate the first text content corresponding to the target speech segment includes: obtaining, based on the target speech segment, an acoustic feature of each frame of speech data and a first acoustic feature of each speech segment, wherein a time length of each speech segment is longer than a time length of each frame of speech data; the first acoustic feature includes a plurality of acoustic features of each frame; determining a target acoustic feature of the target user based on the first acoustic feature; determining target speech data of the target user in each frame of speech data based on the acoustic feature of each frame and the target acoustic feature; and performing speech recognition on the target speech data to obtain the text content of the target user.
[0019] In the embodiments of the present application, when the target speech segment is obtained, the acoustic feature of each frame of speech data corresponding to the target speech segment can be obtained, and the acoustic feature of a speech segment with a time length longer than that of each frame of speech data (i.e., the first acoustic feature). Then, the acoustic feature of the target user (i.e., the target acoustic feature) is determined based on the first acoustic feature, so that each frame of speech data is classified based on the target acoustic feature and the acoustic feature of each frame, and the speech data (i.e., the target speech data) of the target user in each frame of speech data is obtained. Compared with the classification result of the speech data obtained by clustering in the related art, the speech data to be classified is classified at the frame level based on the target acoustic feature of the target user and the acoustic feature of each frame, which avoids the problem of classification error of the speech data by clustering directly. At the same time, the granularity of classification of the speech data to be classified is reduced by frame-level classification, which classifies each frame of speech data more carefully and avoids the problem of classification error of the speech segment with a long time length, thereby improving the accuracy of classification of the speech data.
[0020] In addition, since the time length of the speech segment corresponding to the first acoustic feature is longer than that of each frame of speech data corresponding to the acoustic feature of each frame, the first acoustic feature includes more acoustic information, so that the target acoustic feature determined based on the first acoustic feature is more accurate. Based on the more accurate target acoustic feature of the target user, each frame of speech data can be classified more accurately based on the target acoustic feature and the acoustic feature of each frame, thereby further improving the accuracy of classification of the speech data.
[0021] Optionally, in the speech segment with a long time length, there may be speech data of the target user and speech data of a non-target user. If the speech segment is classified into the speech data of the target speaker, the speech data of the non-target user will also be classified into the speech data of the target speaker, thereby causing an error in classification of the speech segment.
[0022] In conjunction with the first aspect, in certain implementations of the first aspect, the determination of the target user's target voiceprint feature based on the first voiceprint feature includes:
[0023] Clustering is performed on the first voiceprint feature to obtain multiple voiceprint feature sets; the target cluster center of the target voiceprint feature set is determined; wherein, the target voiceprint feature set is any one of the multiple voiceprint feature sets; the target cluster center is determined as the target voiceprint feature of the first target user; wherein, the first target user is the user corresponding to the target voiceprint feature set.
[0024] In this embodiment, when clustering the first voiceprint features, multiple voiceprint feature sets are obtained, and the cluster center of each voiceprint feature set is determined. Each cluster center is then identified as the target voiceprint feature corresponding to a user (i.e., the first target user). By clustering the first voiceprint features, unknown speech data can be classified, improving the generalization of speech data processing and further enhancing the accuracy of speech data classification.
[0025] In combination with the first aspect and the above implementation methods, in some implementation methods of the first aspect, the above-mentioned determination of the target speech data corresponding to the target user in each frame of speech data based on the voiceprint features of each frame and the target voiceprint features includes: determining the similarity value between the voiceprint features of each frame and multiple target voiceprint features; and determining the speech data corresponding to the largest similarity value among the multiple similarity values as the target speech data of the first target user.
[0026] In this embodiment of the application, by comparing the similarity between the voiceprint features of each frame and multiple target voiceprint features, the voice data corresponding to the maximum similarity value is determined as the voice data of a target user (i.e., the target voice data). This can ensure the matching between the target user and the determined voice data to the greatest extent, thereby further improving the accuracy of voice data classification.
[0027] Secondly, this application provides a voice processing apparatus configured in an electronic device, comprising: an acquisition module for acquiring voice data to be processed; a determination module for determining a first voice segment in the voice data to be processed; wherein the first voice segment represents multiple consecutive frames of voice data in the voice data to be processed; a classification module for classifying the first voice segment into events to obtain a classification result; a processing module for determining a target voice segment based on the classification result; wherein the user identifier corresponding to the target voice segment is a single user identifier; and a generation module for generating target standard text based on the target voice segment.
[0028] Thirdly, this application provides an electronic device including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the electronic device to perform the methods described in the first aspect or any possible implementation thereof.
[0029] Fourthly, this application provides a computer program product comprising: computer program code, which, when run on a computer, causes the computer to perform the method described in the first aspect or any possible implementation thereof.
[0030] Fifthly, this application provides a computer-readable storage medium storing computer program code that, when executed on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0031] In this embodiment, when the voice data to be processed is obtained, event classification can be performed on multiple consecutive frames of voice data to determine the voice segment with a single user identifier (i.e., the target voice segment). Then, standard text (i.e., the target standard text) is generated from this target voice segment. Since the target voice segment has a single user identifier, it indicates that only one user is speaking in the target voice segment. Therefore, when generating standard text from this target voice segment, the problem of errors in the standard text caused by multiple users speaking simultaneously in the voice data to be processed, thus reducing the accuracy of the standard text, can be avoided. This ensures the accuracy of the text when generating the text corresponding to the voice data.
[0032] Furthermore, by classifying events from multiple consecutive frames of audio data in the processed audio data, the problem of a single frame being too short to reflect the corresponding event type can be avoided, thus improving the accuracy of the event classification results. In turn, based on more accurate event classification results, the accuracy of the text can be further ensured. Attached Figure Description
[0033] Figure 1 is a schematic diagram of a speech processing scenario using related technologies.
[0034] Figure 2 is a schematic diagram of the structure of the speech processing model provided in the embodiment of this application.
[0035] Figure 3 is a schematic diagram of the structure of the noise reduction and reverberation reduction module provided in the embodiment of this application.
[0036] Figure 4 is a schematic diagram of the structure of the event detection module provided in an embodiment of this application.
[0037] Figure 5 is a schematic diagram of the speaker classification module provided in an embodiment of this application.
[0038] Figure 6 is a schematic diagram of the standard text provided in the embodiments of this application.
[0039] Figure 7 is a schematic flowchart of a speech processing method provided in an embodiment of this application.
[0040] Figure 8 is a schematic diagram of the structure of the voice processing device provided in the embodiment of this application.
[0041] Figure 9 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. Detailed Implementation
[0042] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0043] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0044] For example, in multi-person speaking scenarios such as meetings, teaching, telephone monitoring, radio programs, and court hearings, there may be a need to record and analyze the speakers' speech in order to transcribe the speech into text.
[0045] For ease of explanation, this application uses a meeting scenario as an example, as shown in Figure 1.
[0046] Figure 1 is a schematic diagram of a speech processing scenario using related technologies.
[0047] For example, as shown in Figure 1, Figure 1 includes user A, user B, user C, voice data 110 corresponding to user A, voice data 120 corresponding to user B, voice data 130 corresponding to user C, environmental noise data (hereinafter referred to as "noise data") 140, microphone 150, and electronic device 160.
[0048] The microphone 150 can be used to collect voice data 110, voice data 120, voice data 130 and noise data 140, and transmit the collected voice data 110, voice data 120, voice data 130 and noise data 140 to the electronic device 160 through a wireless network or wired connection. When the electronic device 160 receives the voice data 110, voice data 120, voice data 130 and noise data 140, it can process the voice data 110, voice data 120, voice data 130 and noise data 140 through its own configured voice processing model to obtain the voice text corresponding to the voice data 110, voice data 120 and voice data 130.
[0049] For example, as neural network technology matures, speech data can be processed using speech processing models to obtain corresponding speech text, thus reducing labor costs.
[0050] Speech processing models can perform automatic echo cancellation (AEC), automatic noise suppression (ANS), and automatic gain control (AGC) on speech data captured by microphones. This ensures that the processed speech data meets the needs of voice communication, making it clearer for people at a distance from the speaker. Furthermore, it removes interference information from the speech data, making the speech text generated by the speech processing model more accurate. AEC, ANS, and AGC can be collectively referred to as "3A technology."
[0051] Because the processed voice data may contain voice data from multiple people speaking simultaneously (which can be called "overlapping voice data"), the presence of overlapping voice data can easily lead to errors in the voice text, reduce the accuracy of the voice text, and affect the user experience.
[0052] It should be noted that the electronic device can be a smart device configured with a voice processing model, including but not limited to: personal computers, tablets, handheld devices, in-vehicle devices, wearable devices, computing devices, or other processing devices connected to a wireless modem. The electronic device may have different names in different networks, such as: user equipment, access electronic device, user unit, user station, mobile station, mobile station, remote station, remote electronic device, mobile device, user electronic device, electronic device, wireless communication device, user agent or user device, cellular phone, cordless phone, electronic device in a 5G network or future evolved network, etc. The embodiments in this application are not limited in scope.
[0053] Therefore, in order to solve the problem of low text accuracy, this application proposes a speech processing method, a speech processing device, an electronic device, and a storage medium.
[0054] The speech processing method provided in the embodiments of this application will be described in detail below with reference to Figures 2 to 7.
[0055] Figure 2 is a schematic diagram of the structure of the speech processing model provided in the embodiment of this application.
[0056] For example, as shown in Figure 2, the speech processing model 200 may include an echo cancellation module 210, a beamforming module 220, a noise reduction and de-reverberation module 230, a gain control module 240, an event detection module 250, a speaker classification module 260, and a speech recognition module 270. All speech data collected by the microphone (e.g., speech data 110, speech data 120, speech data 130, and noise data 140) are used as input, and standard text is used as output.
[0057] The echo cancellation module 210, beamforming module 220, noise reduction and de-reverberation module 230, and gain control module 240 process all voice data acquired by the microphone in real time, and can store the processed voice data along with corresponding video and images for later review and analysis. The event detection module 250, speaker classification module 260, and speech recognition module 270 can process all voice data acquired by the microphone in real time or not; this embodiment does not limit this process.
[0058] The echo cancellation module 210 can be used to eliminate the echo generated when the far-end sound is played through the speaker and then picked up again by the microphone after receiving all the voice data collected by the microphone, thereby ensuring the clarity of the far-end speaker's voice and obtaining echo-free voice data. The echo-free voice data is then sent to the beamforming module 220. All voice data can be multi-channel voice data.
[0059] The beamforming module 220 can be used to focus and enhance the speech signal in a specific direction when receiving echo-canceling speech data, while suppressing noise and interference from other directions, making the focused speech data clearer. It also sends the focused speech data to the denoising and de-reverberation module 230.
[0060] The noise reduction and dereverberation module 230 can be used to remove noise and reverberation from the received focused speech data to obtain the user's speech data (e.g., speech data 110, speech data 120, and speech data 130). It also sends the user's speech data to the gain control module 240.
[0061] Figure 3 is a schematic diagram of the structure of the noise reduction and reverberation reduction module provided in the embodiment of this application.
[0062] For example, as shown in Figure 3, the noise reduction and de-reverberation module 230 may include an encoding module 201 (Encoder), a decoding module 202 (Decoder), and an intermediate layer Long Short-Term Memory Network (LSTM) 203 that serves as a connector. Specifically, each frame of focused speech data undergoes a short-time Fourier transform to obtain the complex spectrum corresponding to each frame of focused speech data. The real and imaginary parts of the complex spectrum are used as input, and the resulting mask matrix is used as the output. Multiplying the mask matrix by the complex spectrum yields the complex spectrum corresponding to the output speech data. Performing a short-time Fourier transform on the complex spectrum of the output speech data then yields the user's speech data after noise reduction and de-reverberation.
[0063] It should be understood that consecutive 5ms, 15ms, 10ms or 25ms can be considered as a frame, and the duration of each frame is extremely short. This application does not limit this.
[0064] The encoding module 201 may include multiple encoders with different numbers of channels, and the number of channels in each encoder gradually increases. Referring to Figure 3, the encoding module 201 consists of four encoders, which are arranged in ascending order of the number of channels, with the encoders having fewer channels at the beginning and those having more channels at the end. The number of channels in the four encoders increases from 2 to 16 to 32 to 64 to 128, so that deeper speech features can be captured by the encoders as the number of channels gradually increases. Each encoder may consist of a convolutional layer, a normalization layer, and an activation layer. The kernel size in the convolutional layer is (3×3), and the stride of the kernel is (2, 1), that is, the frequency domain stride is 2 and the time domain stride is 1.
[0065] For example, "2→16" can indicate that the encoder has 2 input channels and 16 output channels; "16→32" can indicate that the encoder has 16 input channels and 32 output channels, and so on, without further explanation.
[0066] The decoding module 202 may include multiple decoders with different numbers of channels, and the number of channels in each decoder gradually increases. Referring to Figure 3, the decoding module 202 consists of 4 decoders, which are arranged in descending order of the number of channels, with the decoders having larger channels first and the decoders having smaller channels last. The number of channels in the 4 decoders increases from 128 to 64 to 32 to 16 to 2, so as to enable more comprehensive decoding of the encoded speech features as the number of channels gradually increases, ultimately obtaining an output mask matrix of the same size as the input. Each decoder may consist of a transposed convolution, a normalization layer, and an activation layer. The kernel size in the transposed convolution is (3×3), and the stride of the kernel is (2, 1), that is, the frequency domain stride is 2 and the time domain stride is 1.
[0067] For example, "128→64" can indicate that the encoder has 128 input channels and 64 output channels; "64→32" can indicate that the encoder has 64 input channels and 32 output channels, and so on.
[0068] Furthermore, considering the real-time nature of voice data processing, the convolutional layers in the encoder and the transposed convolutions in the decoder can employ causal convolutions with a processing latency of 1 frame, for example, 25ms.
[0069] The LSTM203 can use the output of the last encoder as input and its own output as input to the first decoder. It can combine the currently encoded speech features with historical information from previous frames to make predictions, so that the predicted speech features are correlated with each other.
[0070] Alternatively, in the denoising and de-reverberation module 230, the LSTM can be removed, and the output of the last encoder can be directly used as the input of the first decoder without affecting the final mask matrix.
[0071] Referring to Figure 3, there are skip connections between each encoder and each decoder, making the architecture of the denoising and de-reverberation module 230 a U-Net architecture. This allows low-level speech features to be directly passed to the corresponding decoding stage, and each encoded speech feature can be used to guide the corresponding decoding process to preserve more fine-grained details in the speech features.
[0072] Optionally, when training the denoising and dereverberation module 230, the sample speech data can be preprocessed by simultaneously adding noise and reverberation to obtain noisy and reverberated sample speech data. This noisy and reverberated sample speech data can then be input as ground truth to train the denoising and dereverberation module 230, enabling the trained module to perform denoising and dereverberation on speech data. The data preprocessing of the sample speech data and the training process of the denoising and dereverberation module 230 can be referred to as a "data-driven approach."
[0073] The gain control module 240 can adjust the intensity or amplitude of the user's voice data upon receipt, ensuring that the gained voice data maintains an appropriate volume level in different scenarios. It also converts the gained voice data from multi-channel voice data into single-channel voice data. The single-channel gained voice data (hereinafter referred to as "gained voice data") is then sent to the event detection module 250.
[0074] The event detection module 250 can detect the event type corresponding to the received amplified speech data and send the detected event type and the corresponding speech data to the speaker classification module 260. The event type can include, but is not limited to, single-person speaking events, multi-person speaking events, and no-speaking events. A single-person speaking event can represent an event type where only one user speaks at the same time; a multi-person speaking event can represent an event type where two or more users speak at the same time; and a no-speaking event can represent an event type where no user speaks. Furthermore, when the amplified speech data contains speech data corresponding to multi-person speaking events and no-speaking events, the event type corresponding to that speech data (e.g., multi-person speaking events and no-speaking events) can be sent to the speaker classification module 260, but the speech data itself can be sent instead. Alternatively, when the amplified speech data contains speech data corresponding to a single-person speaking event, both the speech data and the corresponding event type (i.e., single-person speaking event) can be sent to the speaker classification module 260.
[0075] Figure 4 is a schematic diagram of the structure of the event detection module provided in an embodiment of this application.
[0076] For example, as shown in Figure 4, the event detection module 250 may include a first convolutional layer 204, a second convolutional layer 205, ..., an nth convolutional layer 206, and a second LSTM 207. Specifically, the inverse Mel-frequency spectrum (FilterBank, Fbank) of the gained speech data for each frame is extracted to obtain an 80-dimensional Fbank. All Fbanks corresponding to the gained speech data (which can be denoted as "80-dimensional × T frames") are used as input, and the event type and the corresponding speech data are used as output.
[0077] Furthermore, since a single frame is relatively short, it may not be able to reflect the event category it corresponds to. Therefore, in order to accurately reflect the type of the amplified speech data, multiple consecutive frames of speech data can be assigned to the same event type by default. For example, four consecutive frames of speech data might correspond to the event type of a single person speaking.
[0078] Each convolutional layer in the event detection module 250 can consist of a convolutional layer, a pooling layer, and an activation layer. The kernel size in each convolutional layer is (3×3), and the stride of the bottom two kernels is (2, 2), meaning the frequency domain stride is 2 and the time domain stride is 2, equivalent to downsampling by a factor of 4 in both the feature and time dimensions. Multiple convolutional layers can extract local features from the amplified speech data, helping to identify specific speaking events (i.e., event types). The recognition results are then sent to the second LSTM207.
[0079] The second LSTM207 can be used to consider the influence between the voice data of the preceding and following frames when the event type is received, so as to obtain the final event type.
[0080] The speaker classification module 260 can be used to perform speaker recognition on the speech data of a single speaker when receiving speech data of a single speaker, and obtain the speaker's tag information (e.g., name, number, and code) corresponding to each frame of speech data. The tag information and the speech data of the single speaker corresponding to the tag information are then sent to the speech recognition module 270.
[0081] Figure 5 is a schematic diagram of the speaker classification module provided in an embodiment of this application.
[0082] For example, as shown in Figure 5, the speaker classification module 260 may include a speech segmentation module 211, a voiceprint module 212, a clustering module 213, and a voiceprint determination module 214 (also referred to as the "voiceprint determination module"). The speech data of a single speaker is used as input, and the classification result of the speech data is used as output.
[0083] The speech segmentation module 211 can be used to segment the speech data of a single person's speech according to a fixed duration (e.g., 2 seconds) upon receiving speech data of a single person's speech, thereby segmenting at least one speech segment. Furthermore, it can send the segmented speech segment to the voiceprint model 212.
[0084] The voiceprint module 212 can be used to extract voiceprint features (which can be called "segment-level voiceprint features") included in at least one segmented speech segment upon receiving the segmented speech segment, and send the extracted segment-level voiceprint features to the clustering module 213. It also extracts frame-level voiceprint features (D×T dimension) corresponding to each frame of speech data in the at least one speech segment. Finally, it sends the extracted frame-level voiceprint features to the voiceprint determination module 214. The voiceprint features may include, but are not limited to, spectral features, temporal features, frequency features, pitch, stress, vocalization order, and prosody.
[0085] Optionally, the voiceprint features of continuous 10ms (which can be called "frame duration") of speech data can be used as frame-level voiceprint features; or, the voiceprint features of continuous 5ms and 15ms of speech data can be used as frame-level voiceprint features, that is, the duration corresponding to the frame-level voiceprint features is extremely short. This application does not limit this.
[0086] For example, the voiceprint module 212 may employ a residual network structure to output at least one D-dimensional feature corresponding to a speech segment. For instance, the voiceprint module 212 may employ a stacked Transformer structure.
[0087] The clustering module 213 can be used to perform cluster analysis on the received segment-level voiceprint features, grouping similar voiceprint features into one category, and using the speech data corresponding to each category of voiceprint features as the voiceprint features of the same speaker (e.g., the voiceprint features of user A). Furthermore, after performing cluster analysis on the segment-level voiceprint features, the cluster center corresponding to each category of voiceprint features can be obtained, and each cluster center can be sent to the voiceprint determination module 214.
[0088] The voiceprint determination module 214 can be used to, upon receiving frame-level voiceprint features and each cluster center, separately judge whether the frame-level voiceprint features are the same as or similar to each cluster center. When the frame-level voiceprint feature is the voiceprint feature of user A, the frame speech data corresponding to the frame-level voiceprint feature can be classified as the speech data of user A. For example, after obtaining the speaking probabilities of users A, B, and C corresponding to the frame-level voiceprint feature, and smoothing the speaking probabilities in the time dimension to eliminate the probability of occasional short-term speakers, the user with the highest speaking probability among users A, B, and C is determined as the speaker corresponding to the frame-level voiceprint feature.
[0089] For example, the voiceprint determination module 214 can perform attention mechanism calculation on the speaker voiceprint features S*N*D (where S represents the number of speaker categories, N represents the number of representative voiceprints retained in each category, and D represents the feature dimension) corresponding to each speech segment obtained by the clustering module 213, to obtain the voiceprint representation S*D after feature fusion, and then copy the S×D dimensional feature into T copies to obtain (S×D)×T dimensional features. The (S×D)×T dimensional features are then concatenated with the D×T dimensional features to obtain ((S+1)×D)×T dimensional features.
[0090] Figure 6 is a schematic diagram of the standard text provided in the embodiments of this application.
[0091] For example, as shown in Figure 6, the voice data start and end times, event type, speaker's tag information, and corresponding text records (i.e., the voice text) are output to the user.
[0092] For example, the event type corresponding to the voice data from 10:00 to 10:10 is a single-person speaking event, the speaker is ID1, and the text record corresponding to the voice data from 10:00 to 10:10 is also included.
[0093] For example, the event type corresponding to the voice data from 10:11 to 10:12 is a single-person speaking event, the speaker is ID2, and the text record corresponding to the voice data from 10:11 to 10:12 is also included.
[0094] For example, if the event type corresponding to the voice data from 10:13 to 10:20 is a multi-person speaking event, only the corresponding event type will be output, without displaying the corresponding speaker's label information or the text record corresponding to the voice data from 10:13 to 10:20.
[0095] For example, the event type corresponding to the voice data from 10:21 to 10:25 is "no one is speaking," and only the corresponding event type is output.
[0096] Figure 7 is a schematic flowchart of a speech processing method provided in an embodiment of this application. This method can be executed by the electronic device 160 in Figure 1. Furthermore, this method can be executed within a pre-trained speech processing model 200.
[0097] For example, as shown in FIG7, the method 700 includes S710-S730:
[0098] S710 acquires the voice data to be processed.
[0099] For example, in a multi-person speaking scenario (as shown in Figure 1, in order to facilitate the recording of participants' speeches, all audio data in the meeting scenario (which can be referred to as "voice data to be processed") can be collected through a microphone.
[0100] All audio data may include, but is not limited to, the voice data and noise data of each participant.
[0101] S720, determine the first speech segment in the speech data to be processed; wherein, the first speech segment represents multiple consecutive frames of speech data in the speech data to be processed.
[0102] For example, to avoid the problem that a single frame of audio data is too short to reflect the event type corresponding to that frame, multiple consecutive frames of audio data (which can be called "first audio segments") can be identified in the audio data to be processed.
[0103] S730 performs event classification on the first speech segment and obtains the classification result.
[0104] For example, when the first voice segment is obtained, the first voice segment can be input to the event detection module, and the event detection module can classify the first voice segment into events to obtain the event classification result corresponding to the first voice segment (e.g., single person speaking event, multiple people speaking event, or no one speaking event).
[0105] In one possible implementation, the above-mentioned event classification of the first speech segment to obtain the classification result includes: determining the similarity between the first speech segment and each of the multiple preset events; determining whether the maximum similarity among the multiple similarities is greater than or equal to the preset similarity; and determining the event type corresponding to the maximum similarity as the classification result if the maximum similarity is greater than or equal to the preset similarity.
[0106] Among them, multiple preset events may include, but are not limited to, single-person speaking events, multi-person speaking events, and no-person speaking events.
[0107] For example, the first speech segment can be input into the event detection module to obtain the similarity between the first speech segment and single-person speaking events, multi-person speaking events, and no-person speaking events, respectively. For instance, the similarity between the first speech segment and a single-person speaking event can be recorded as "similarity A", the similarity between the first speech segment and a multi-person speaking event can be recorded as "similarity B", and the similarity between the first speech segment and a no-person speaking event can be recorded as "similarity C".
[0108] Furthermore, the similarities A, B, and C can be sorted from largest to smallest or smallest to largest to obtain the maximum similarity among them. For example, if similarity A > similarity B > similarity C, then similarity A can be determined as the maximum similarity. Alternatively, if similarity B > similarity C > similarity A, then similarity B can be determined as the maximum similarity.
[0109] When obtaining the maximum similarity, in order to avoid the problem that the first speech segment does not match the event type due to the maximum similarity being too small, it can be determined whether the maximum similarity is greater than or equal to the preset similarity (e.g., 95%).
[0110] When the maximum similarity (e.g., similarity A is 99%) is ≥ 95%, it can be said that the first speech segment matches the single-person speaking event corresponding to similarity A. Therefore, the single-person speaking event corresponding to similarity A can be identified as the event classification result of the first speech segment.
[0111] Alternatively, if the maximum similarity (e.g., similarity A is 55%) is less than 5%, it indicates that the first speech segment does not match the single-person speaking event corresponding to similarity A, and the event type corresponding to the first speech segment does not belong to single-person speaking events, multi-person speaking events, or no-person speaking events. In this case, a reminder message can be generated to alert the user that there is an unknown event type in the first speech segment, enabling the user to promptly train and update the event detection model.
[0112] It should be noted that the preset similarity can be calibrated according to the training requirements of the event detection model. The preset similarity can be 95%, 98%, or 92%, etc., and this application embodiment does not limit it.
[0113] In this embodiment, when the maximum similarity among the similarities between the first speech segment and each of the multiple preset events is greater than or equal to a preset similarity, the similarity between the first speech segment and the classification result can be ensured to the greatest extent possible, while also avoiding the problem of classification errors caused by a small maximum similarity between the two events. This improves the accuracy of the event classification result of the first speech segment. Furthermore, based on the more accurate event classification result of the first speech segment, the accuracy of the text is further ensured.
[0114] S740 determines the target speech segment based on the classification results.
[0115] The user identifier corresponding to the target speech segment is a single user identifier.
[0116] For example, when the event classification result of the first speech segment is obtained, the speech data corresponding to only one user identifier in the speech data to be processed (which can be called the "target speech segment") can be determined by the event classification result of the first speech segment.
[0117] For example, the first speech segments in the speech data to be processed are: speech data A, speech data B, and speech data C.
[0118] If the event classification result corresponding to voice data A is the speaking time of a single person, that is, there is only one user identifier corresponding to voice data A, then voice data A can be used as the target voice segment.
[0119] If the event classification result corresponding to voice data B is a multi-person speaking event, that is, if there are two or more user identifiers corresponding to voice data B, then voice data B is not the target voice segment (it can be called a "non-target voice segment").
[0120] The event classification result corresponding to voice data C is "no one speaking," meaning the user identifier for voice data C is 0, indicating that there is no user speech in voice data C. Therefore, voice data C can be considered a non-target voice segment. This can be understood as follows: a target voice segment is a voice segment in the voice data to be processed where only one user speaks at any given time; while a non-target voice segment is a voice segment in the voice data to be processed where two or more users speak at any given time, or a voice segment in the voice data to be processed where no user speaks.
[0121] Optionally, the above-mentioned determination of the target speech segment based on the classification result includes: when the classification result is the event type corresponding to a single user identifier, the first speech segment is determined as the target speech segment.
[0122] For example, when determining the event type corresponding to the first voice segment, it can be determined whether the event type is the event type corresponding to a single user identifier (i.e., a single-person speaking event). If the event type corresponding to the first voice segment is a single-person speaking event, the first voice segment can be identified as the aforementioned target voice segment.
[0123] Alternatively, if the event type corresponding to the first speech segment is a multi-person speaking event or a no-person speaking event, it indicates that the first speech segment is not the aforementioned target speech segment.
[0124] In this embodiment, when the classification result corresponds to an event type with a single user identifier, it indicates that the first speech segment contains only one user's speech. Therefore, the first speech segment can be identified as the target speech segment. This avoids the problem of misidentifying multi-user speech data when the classification result corresponds to an event type with multiple users speaking simultaneously, thus improving the accuracy of the target speech segment. Furthermore, based on the improved accuracy of the target speech segment, the accuracy of the text is further ensured.
[0125] Optionally, the speech data to be processed is preprocessed to obtain preprocessed speech data; wherein, the preprocessing includes noise reduction and dereverberation reduction; determining the first speech segment in the speech data to be processed includes: acquiring the first speech segment in the preprocessed speech data.
[0126] Preprocessing may include echo cancellation, beamforming, noise and reverberation removal, and gain control on the speech data to be processed.
[0127] For example, the voice data to be processed can be input to the echo cancellation module, the echo cancellation module can eliminate the echo in the voice data to be processed to obtain the echo-free voice data, and the echo-free voice data can be sent to the beamforming module.
[0128] When the beamforming module receives echo-canceling speech data, it can focus and enhance the speech signal in a specific direction in the echo-canceling speech data, while suppressing noise and interference in other directions, making the focused speech data clearer, and sending the focused speech data to the noise reduction and de-reverberation module.
[0129] When the noise reduction and dereverberation module receives the focused speech data, it can perform noise reduction and dereverberation on the focused speech data to obtain the noise reduction and dereverberation speech data, and send the noise reduction and dereverberation speech data to the gain control module.
[0130] When the gain control module receives denoised and dereverberated speech data, it can adjust the intensity and amplitude of the denoised and dereverberated speech data to obtain the gained speech data. It can also convert the gained speech data from multi-channel speech data into single-channel speech data to obtain single-channel gained speech data (which can be called "preprocessed speech data").
[0131] For example, multiple consecutive frames of audio data in the preprocessed audio data can be identified as the first audio segment mentioned above.
[0132] In this embodiment, by performing noise reduction and de-reverberation processing on the speech data to be processed, the noise and reflected sound in the speech data can be reduced, making the preprocessed speech data more accurate. This, in turn, makes the first speech segment in the preprocessed speech data more accurate. Furthermore, based on the increased accuracy of the first speech segment, the accuracy of the text is further ensured.
[0133] S750 generates target standard text based on target speech segments.
[0134] For example, when a target speech segment is identified in the speech data to be processed, the corresponding standard text (which can be called the "target standard text") can be obtained from the target speech segment, such as the meeting text corresponding to the meeting scenario in Figure 1. Furthermore, the obtained target standard text can be displayed as shown in Figure 6.
[0135] Optionally, when the target standard text is obtained, it can be input into a pre-trained meeting summary model, which can then generate a corresponding meeting summary based on the target standard text.
[0136] In method 700 as shown in Figure 7, when the speech data to be processed is obtained, event classification can be performed on multiple consecutive frames of speech data to determine the speech segment with a single user identifier (i.e., the target speech segment). Then, standard text (i.e., the target standard text) is generated from this target speech segment. Since the target speech segment has a single user identifier, it indicates that only one user is speaking in the target speech segment. Therefore, when generating standard text from this target speech segment, the problem of multiple users speaking simultaneously in the speech data to be processed, which could lead to errors in the standard text and reduce its accuracy, can be avoided. This ensures the accuracy of the text when generating the text corresponding to the speech data. Furthermore, by performing event classification on multiple consecutive frames of speech data to be processed, the problem of a single frame of speech data being too short to reflect the event type corresponding to that frame can be avoided, thus improving the accuracy of the event classification results. Furthermore, based on more accurate event classification results, the accuracy of the text is further ensured.
[0137] In one possible implementation, the above-mentioned generation of target standard text based on target speech segments includes: performing speech recognition on the target speech segments to generate first text content corresponding to the target speech segments; and generating a first event type identifier corresponding to the target speech segments and a second event type identifier corresponding to a second speech segment in the speech data to be processed; wherein the second speech segment is a speech segment in the speech data to be processed other than the target speech segment; and generating target standard text based on the first text content, the first event type identifier, and the second event type identifier.
[0138] For example, when a target speech segment is identified in the speech data to be processed, the target speech segment can be input into the speech recognition module. The speech recognition module performs speech recognition on the target speech segment to obtain the text corresponding to the target speech segment (i.e., the aforementioned speech text, which can be referred to as the "first text content").
[0139] Furthermore, upon obtaining the event type of the target speech segment (i.e., a single-person speaking event), an identifier corresponding to that single-person speaking event (which can be called the "first event type identifier") can be generated.
[0140] Furthermore, when the target speech segment in the speech data to be processed is determined, the event detection module can identify speech segments other than the target speech segment (which can be called "second speech segments", i.e. the non-target speech segments mentioned above) in the speech data to be processed, and perform event classification on the second speech segments to obtain the event classification results corresponding to the second speech segments (e.g., multiple people speaking event or no one speaking event).
[0141] Furthermore, upon obtaining the event type of the second speech segment (e.g., a multi-person speaking event or a no-person speaking event), an identifier corresponding to the multi-person speaking event or the no-person speaking event (which can be called a "second event type identifier") can be generated.
[0142] Furthermore, after obtaining the first text content and the first event type identifier corresponding to the target speech data in the target speech segment, and the second event type identifier corresponding to the second speech segment, the aforementioned target standard text can be generated by combining the first text content, the first event type identifier, and the second event type identifier (as shown in Figure 6).
[0143] The first event type identifier is different from the second event type identifier, and the first event type identifier and the second event type identifier can be represented by one or more of the following: subtitles, numbers, codes and images. This application embodiment does not limit this.
[0144] In this embodiment of the application, when generating corresponding text content (i.e., first text content) by performing speech recognition on the target speech segment, an event type identifier (i.e., first event type identifier) corresponding to the target speech segment can also be generated, as well as an event type identifier (i.e., second event type identifier) corresponding to the speech segment (i.e., second speech segment) in the speech data to be processed, excluding the target speech segment. This ensures that the target standard text generated by the first text content, the first event type identifier, and the second event type identifier contains not only the text content corresponding to the target speech segment but also the event type identifiers of the target speech segment and the non-target speech segment, making the final generated standard text clearer and further ensuring the accuracy of the text.
[0145] Optionally, the first text content mentioned above includes the text content of the target user; the above-mentioned speech recognition of the target speech segment to generate the first text content corresponding to the target speech segment includes: obtaining the voiceprint features of each frame corresponding to each frame of speech data, and the first voiceprint features of each speech segment based on the target speech segment; wherein, the duration of each speech segment is greater than the duration of each frame of speech data; the first voiceprint features include multiple voiceprint features of each frame; determining the target voiceprint features of the target user based on the first voiceprint features; determining the target speech data corresponding to the target user in each frame of speech data based on the voiceprint features of each frame and the target voiceprint features; and performing speech recognition on the target speech data to obtain the text content of the target user.
[0146] The target user can refer to different speakers in the target speech segment, such as user A and user B, and does not refer to a specific user.
[0147] For example, since the target speech segment may include speech data from different users, and the speech data from different users do not overlap at the same time, in order to determine the speech data corresponding to different speech segments of different users in the target speech segment, the target speech segment can be input into a speech segmentation module for segmentation to obtain at least one speech segment. Furthermore, at least one speech segment is input into a voiceprint module, which extracts segment-level voiceprint features (which can be called "first voiceprint features") and frame-level voiceprint features (which can be called "each frame voiceprint features") from the segmented speech segments. The segment-level voiceprint features include multiple frame-level voiceprint features.
[0148] Optionally, when extracting the voiceprint features of the segmented speech segments using the voiceprint module, the voiceprint features corresponding to each frame of speech data in the segmented speech segments can be extracted according to the frame duration (e.g., 10ms) (which can be called "frame-by-frame voiceprint features," i.e., the aforementioned frame-level voiceprint features). The number of voiceprint features per frame is at least one.
[0149] Optionally, the voiceprint features of each segmented speech fragment can be directly extracted using a voiceprint module, and the extracted voiceprint features of the speech fragments can be used as the first voiceprint features. The number of first voiceprint features is at least one.
[0150] Optionally, since the first voiceprint feature is a segment-level voiceprint feature, including at least two frames of voiceprint features, it can be said that the first voiceprint feature contains more voiceprint information than each frame of voiceprint features. Therefore, when determining the target voiceprint feature of the target user through the first voiceprint feature, a more accurate target voiceprint feature can be obtained. Furthermore, the duration of the speech segment corresponding to the first voiceprint feature (which can be referred to as "each speech segment") is greater than the duration of the speech data in each frame. For example, the duration of each speech segment is 1 second, and the duration of each frame of speech data is 10 ms.
[0151] Furthermore, when the first voiceprint feature is extracted through the voiceprint module, the first voiceprint feature can be input into the clustering module, and the clustering module can perform cluster analysis on the first voiceprint feature to obtain the clustering result.
[0152] In one possible implementation, determining the target voiceprint feature of a target user based on the first voiceprint feature includes: performing clustering processing on the first voiceprint feature to obtain multiple voiceprint feature sets; determining the target cluster center of the target voiceprint feature set; wherein the target voiceprint feature set is any one of the multiple voiceprint feature sets; and determining the target cluster center as the target voiceprint feature of the first target user; wherein the first target user is the user corresponding to the target voiceprint feature set.
[0153] For example, the clustering module can first perform clustering processing on at least one first voiceprint feature, grouping first voiceprint features with high similarity into one category. After the clustering is completed, multiple voiceprint feature sets and the cluster center corresponding to each voiceprint feature set in the multiple voiceprint feature sets can be obtained.
[0154] For example, after clustering, we obtain feature set A and feature set B, and the cluster center a corresponding to feature set A and the cluster center b corresponding to feature set B.
[0155] Furthermore, when multiple voiceprint feature sets are obtained, the cluster center (or target cluster center) corresponding to any one of the multiple voiceprint feature sets (which can be called the "target voiceprint feature set") can be determined. That is, each voiceprint feature set in the multiple voiceprint feature sets can be used as the target voiceprint feature set.
[0156] For example, when feature set A is the target voiceprint feature set, cluster center a is the target cluster center; and when feature set B is the target voiceprint feature set, cluster center b is the target cluster center.
[0157] For example, when determining the target cluster center corresponding to the target voiceprint feature set, the target cluster center can be identified as the target voiceprint feature of the user (which can be referred to as the "first target user") corresponding to the target voiceprint feature set. It should be understood that one target voiceprint feature set corresponds to one first target user, and each target voiceprint feature set corresponds to a different first target user.
[0158] For example, if the target user corresponding to feature set A is user A, then cluster center a can be determined as the target voiceprint feature corresponding to user A. Similarly, if the target user corresponding to feature set B is user B, then cluster center b can be determined as the target voiceprint feature corresponding to user B.
[0159] In this embodiment, when clustering the first voiceprint features, multiple voiceprint feature sets are obtained, and the cluster center of each voiceprint feature set is determined. Each cluster center is then identified as the target voiceprint feature corresponding to a user (i.e., the first target user). By clustering the first voiceprint features, unknown speech data can be classified, improving the generalization of speech data processing and further enhancing the accuracy of speech data classification.
[0160] Furthermore, after obtaining the voiceprint features of each frame of speech data in the target speech segment and the target voiceprint features corresponding to the target user, the frame voiceprint features belonging to the target user in each frame voiceprint feature can be determined by using the voiceprint features of each frame and the target voiceprint features. Then, by using the timestamp information carried in the frame voiceprint features of the target user, the speech data corresponding to the frame voiceprint features of the target user can be determined in the target speech segment, and this speech data is identified as the speech data corresponding to the target user (which can be called "target speech data").
[0161] Once the target user's voice data is obtained, the voice data can be fed into a speech recognition model for speech recognition to generate the corresponding text (i.e., the target user's text content).
[0162] In this embodiment, upon acquiring a target speech segment, the speaker features corresponding to each frame of speech data in the target speech segment, as well as the speaker features (i.e., the first speaker features) of speech segments whose duration is longer than that of each frame of speech data, can be obtained. Then, the speaker features corresponding to the target user (i.e., the target speaker features) are determined using the first speaker features. These target speaker features, along with the speaker features of each frame, are used to classify each frame of speech data, thus obtaining the speech data corresponding to the target user in each frame of speech data (i.e., the target speech data). Compared to related technologies that directly obtain the classification results of speech data through clustering, this solution performs frame-level classification of the speech data to be classified using the target user's target speaker features and the speaker features of each frame, avoiding the problem of incorrect classification of speech data directly through clustering. Simultaneously, frame-level classification reduces the granularity of the classification of the speech data to be classified, allowing for more detailed classification of each frame of speech data, avoiding classification errors for longer speech segments, thereby improving the accuracy of speech data classification.
[0163] Furthermore, since the duration of the speech segment corresponding to the first voiceprint feature is longer than the duration of each frame of speech data corresponding to each frame of voiceprint features, the first voiceprint feature can include more voiceprint information. Therefore, the target voiceprint feature determined by the first voiceprint feature can be more accurate. Based on the more accurate target voiceprint feature corresponding to the target user, the target voiceprint feature and each frame of voiceprint features can be used to perform more accurate frame-level classification of each frame of speech data, thereby further improving the accuracy of speech data classification.
[0164] Optionally, the above-mentioned determination of the target speech data corresponding to the target user in each frame of speech data based on the voiceprint features of each frame and the target voiceprint features includes: determining the similarity value between the voiceprint features of each frame and multiple target voiceprint features; and determining the speech data corresponding to the largest similarity value among the multiple similarity values as the target speech data of the first target user.
[0165] For example, when obtaining the voiceprint features of each frame of the target speech segment and the target voiceprint features corresponding to multiple first target users (e.g., user A and user B), the similarity of each frame of voiceprint features with the target voiceprint features of multiple first target users can be compared to obtain multiple similarity values of each frame of voiceprint features after the similarity comparison.
[0166] Furthermore, the multiple similarity values are sorted from largest to smallest or smallest to largest to obtain the largest similarity value. The speech data corresponding to the speaker features of the frame with the largest similarity value is determined as the target speech data of the user with the largest similarity value among multiple first target users. Each frame of speech data corresponding to the speaker features has a corresponding first target user.
[0167] For example, each frame of voiceprint features includes frame voiceprint feature 1 and frame voiceprint feature 2, with the first target users being user A and user B. Comparing the similarity of frame voiceprint feature 1 with the target voiceprint features of both user A and user B, we find that the similarity value A between frame voiceprint feature 1 and user A's target voiceprint feature is 95%, and the similarity value B between frame voiceprint feature 1 and user B's target voiceprint feature is 63%. Therefore, the maximum similarity between similarity value A and similarity value B is determined to be similarity value A. Since the first target user corresponding to similarity value A is user A, and the corresponding frame voiceprint feature is frame voiceprint feature 1, the frame speech data corresponding to frame voiceprint feature 1 can be identified as the target speech data for user A.
[0168] Furthermore, a similarity comparison is performed between frame voiceprint feature 2 and the target voiceprint features of both user A and user B. The similarity value C between frame voiceprint feature 2 and user A's target voiceprint feature is 55%, and the similarity value D between frame voiceprint feature 2 and user B's target voiceprint feature is 93%. Therefore, the maximum similarity between similarity value C and similarity value D is determined to be similarity value D. Since the first target user corresponding to similarity value D is user B, and the corresponding frame voiceprint feature is frame voiceprint feature 2, the frame speech data corresponding to frame voiceprint feature 2 can be identified as user B's target speech data.
[0169] Optionally, the frame speech data corresponding to the frame speech feature can be determined from the speech data to be classified by using the timestamp information carried in the frame speech feature.
[0170] In this embodiment of the application, by comparing the similarity between the voiceprint features of each frame and multiple target voiceprint features, the voice data corresponding to the maximum similarity value is determined as the voice data of a target user (i.e., the target voice data). This can ensure the matching between the target user and the determined voice data to the greatest extent, thereby further improving the accuracy of voice data classification.
[0171] Optionally, when obtaining the voiceprint features of each frame of the target speech segment and the target voiceprint features corresponding to the target user, the voiceprint features of each frame and the target voiceprint features can be first input into the aforementioned pre-trained voiceprint determination module (which can be called the "pre-trained target module"). After inputting the voiceprint features of each frame and the target voiceprint features into the pre-trained target module, the target voiceprint features can be fused through the attention layer in the pre-trained target module to obtain the fused voiceprint features.
[0172] Furthermore, in the pre-trained target module, the speech data corresponding to the target user (i.e., the aforementioned target speech data) can be determined by using the fused voiceprint features and the voiceprint features of each frame. Moreover, since the fused voiceprint features can include more voiceprint information of the target user, more accurate frame-level classification can be performed when classifying each frame of speech data using the fused voiceprint features corresponding to the target user and the voiceprint features of each frame, thereby further improving the accuracy of speech data classification.
[0173] Optionally, when obtaining the fused voiceprint features of each frame of the target speech segment and the corresponding voiceprint features of the target user, the fusion layer in the pre-trained voiceprint determination module can be used to concatenate the voiceprint features of each frame with the fused voiceprint features to obtain the concatenated voiceprint features. Furthermore, since concatenating the voiceprint features of each frame with the fused voiceprint features of the target user improves the classification performance of each frame of speech data corresponding to each frame of voiceprint features, the accuracy of speech data classification is further improved.
[0174] For example, each frame of voiceprint features includes frame voiceprint feature 1, frame voiceprint feature 2, and frame voiceprint feature 3. Frame voiceprint feature 1, frame voiceprint feature 2, and frame voiceprint feature 3 can be concatenated with the fused voiceprint features to obtain the concatenated voiceprint feature 1 corresponding to frame voiceprint feature 1, the concatenated voiceprint feature 2 corresponding to frame voiceprint feature 2, and the concatenated voiceprint feature 3 corresponding to frame voiceprint feature 3.
[0175] Furthermore, the concatenated voiceprint features can be input into the decoder in a pre-trained voiceprint determination module for decoding, resulting in decoded voiceprint features. These decoded features can then be used to determine whether the corresponding frame's voiceprint features are identical or similar to the target user's target voiceprint features. Additionally, decoding the concatenated voiceprint features reduces the difficulty of classifying them, making classification easier and thus improving the efficiency of speech data classification.
[0176] Among them, the frame voiceprint feature corresponding to the decoded voiceprint feature can represent the frame voiceprint feature that has not been concatenated with the fused voiceprint feature. For example, if the frame voiceprint feature that has not been concatenated with the fused voiceprint feature is frame voiceprint feature 1, then the frame voiceprint feature corresponding to the decoded voiceprint feature is frame voiceprint feature 1.
[0177] Optionally, when obtaining the decoded voiceprint features through the decoder, the decoded voiceprint features can be input into the fully connected layer in the pre-trained voiceprint determination module for dimensionality reduction and mapping processing. The spatial dimension of the decoded voiceprint features is reduced by the dimensionality reduction processing to obtain low-dimensional voiceprint features (e.g., 1-dimensional voiceprint features).
[0178] Furthermore, when obtaining low-dimensional voiceprint features through fully connected layers, these features can be mapped to obtain corresponding mapping values (which can be called "first mapping values"). Using the first mapping value, it can be determined whether the frame voiceprint features corresponding to the low-dimensional voiceprint features are the same as the target voiceprint features. By mapping the decoded voiceprint features, the classification ability of the decoded voiceprint features can be enhanced, making different categories of voiceprint features more separated in the feature space, thereby further improving the accuracy of speech data classification.
[0179] To avoid errors in speech classification caused by excessively large or small first mapping values, the first mapping value can be input into the activation layer of a pre-trained speaker recognition module for normalization. This normalizes the first mapping value to a set range (e.g., (0,1)) to obtain the normalized first mapping value (which can be called the "target mapping value"). By normalizing the first mapping value of the decoded speaker features, the normalized mapping value is kept within a set range, avoiding abnormal data such as excessively large or small first mapping values that could lead to errors in speech data classification, thus further improving the accuracy of speech data classification.
[0180] For example, the first mapping value before normalization is 8. After normalizing 8 to (0,1), the normalized first mapping value can be obtained as 0.8.
[0181] Furthermore, when obtaining the target mapping value through the activation layer, it is determined whether the target mapping value is within a preset range. If the target mapping value is within the preset range, it can be determined that the frame voiceprint feature (e.g., frame voiceprint feature 1) spliced with the fused voiceprint feature is the same as the target voiceprint feature, that is, frame voiceprint feature 1 is the second voiceprint feature.
[0182] Alternatively, when the target mapping value is not within the preset range, it can be determined that the frame voiceprint feature 1 spliced with the fused voiceprint feature is different from the target voiceprint feature.
[0183] The preset range can be set by setting the range (0,1). When the preset range is (0.5,1), it can be said that frame voiceprint feature 1 is the same as the target voiceprint feature.
[0184] For example, the target mapping value corresponding to frame voiceprint feature 1 is 0.7, which is in (0.5,1), indicating that frame voiceprint feature 1 is the same as the target voiceprint feature.
[0185] For example, the target mapping value corresponding to frame voiceprint feature 2 is 0.3, which is not in (0.5,1), indicating that frame voiceprint feature 2 is different from the target voiceprint feature.
[0186] It should be noted that the setting range can be (0,1) or (0,10), and this application embodiment does not limit it.
[0187] It should be understood that the above examples are provided to help those skilled in the art understand the embodiments of this application, and are not intended to limit the embodiments of this application to the specific values or scenarios exemplified. Those skilled in the art can obviously make various equivalent modifications or variations based on the above examples, and such modifications or variations also fall within the scope of the embodiments of this application.
[0188] The speech processing method provided by the embodiments of this application has been described in detail above with reference to Figures 1 to 7; the device embodiments of this application will be described in detail below with reference to Figures 8 and 9. It should be understood that the device in the embodiments of this application can execute the various methods of the foregoing embodiments of this application, that is, the specific working process of the various products below can be referred to the corresponding process in the foregoing method embodiments.
[0189] Figure 8 is a schematic diagram of the structure of the voice processing device provided in the embodiment of this application.
[0190] For example, as shown in FIG8, the device 800 is configured in an electronic device and includes: an acquisition module 810 for acquiring voice data to be processed; a determination module 820 for determining a first voice segment in the voice data to be processed; wherein the first voice segment represents multiple consecutive frames of voice data in the voice data to be processed; a classification module 830 for classifying the first voice segment into events to obtain a classification result; a processing module 840 for determining a target voice segment based on the classification result; wherein the user identifier corresponding to the target voice segment is a single user identifier; and a generation module 850 for generating target standard text based on the target voice segment.
[0191] In one possible implementation, the classification module 830 is specifically used to: determine the similarity between the first speech segment and each of the multiple preset events; determine whether the maximum similarity among the multiple similarities is greater than or equal to a preset similarity; and if the maximum similarity is greater than or equal to the preset similarity, determine the event type corresponding to the maximum similarity as the classification result.
[0192] In one possible implementation, the processing module 840 is specifically used to: determine the first speech segment as the target speech segment when the classification result is an event type corresponding to a single user identifier.
[0193] In one possible implementation, the determining module 820 is further configured to: preprocess the speech data to be processed to obtain preprocessed speech data; wherein, the preprocessing includes noise reduction and reverberation reduction processing; specifically, the determining module 820 is configured to: acquire the first speech segment in the preprocessed speech data.
[0194] In one possible implementation, the generation module 850 is specifically used to: perform speech recognition on the target speech segment to generate first text content corresponding to the target speech segment; and generate a first event type identifier corresponding to the target speech segment and a second event type identifier corresponding to a second speech segment in the speech data to be processed; wherein the second speech segment is a speech segment in the speech data to be processed other than the target speech segment; and generate target standard text based on the first text content, the first event type identifier and the second event type identifier.
[0195] In one possible implementation, the generation module 850 is specifically used to: obtain the voiceprint features of each frame of voice data corresponding to each frame, and the first voiceprint feature of each voice segment, based on the target voice segment; wherein the duration of each voice segment is greater than the duration of each frame of voice data; the first voiceprint feature includes multiple voiceprint features of each frame; determine the target voiceprint feature of the target user based on the first voiceprint feature; determine the target voice data corresponding to the target user in each frame of voice data based on the voiceprint features of each frame and the target voiceprint feature; and perform speech recognition on the target voice data to obtain the text content of the target user.
[0196] In one possible implementation, the generation module 850 is specifically used to: perform clustering processing on the first voiceprint features to obtain multiple voiceprint feature sets; determine the target cluster center of the target voiceprint feature set; wherein, the target voiceprint feature set is any one of the multiple voiceprint feature sets; and determine the target cluster center as the target voiceprint feature of the first target user; wherein, the first target user is the user corresponding to the target voiceprint feature set.
[0197] In one possible implementation, the generation module 850 is specifically used to: determine the similarity value between the voiceprint feature of each frame and multiple target voiceprint features; and determine the speech data corresponding to the largest similarity value among the multiple similarity values as the target speech data of the first target user.
[0198] It should be noted that the aforementioned device 800 is embodied in the form of a functional module. The term "module" here can be implemented in software and / or hardware, without specific limitations.
[0199] For example, a "module" can be a software program, hardware circuit, or a combination of both that implements the above functions. Hardware circuits may include application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors) and memory for executing one or more software or firmware programs, integrated logic circuits, and / or other suitable components that support the described functions.
[0200] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0201] Figure 9 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application.
[0202] For example, as shown in FIG9, the electronic device 900 includes a memory 910 and a processor 920, wherein the memory 910 stores executable program code 911, and the processor 920 is used to call and execute the executable program code 911 to perform a voice processing method.
[0203] This application can divide electronic devices into functional modules based on the above method examples. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into a single processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and represents only one logical functional division; in actual implementation, there may be other division methods.
[0204] When each functional module is divided according to its corresponding function, the electronic device may include: an acquisition module, a determination module, a classification module, a processing module, and a generation module, etc. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0205] The electronic device provided in this application is used to execute the above-described voice processing method, and thus can achieve the same effect as the above-described implementation method.
[0206] When using integrated units, the electronic device may include a processing module and a storage module. The processing module is used to control and manage the operation of the electronic device. The storage module is used to support the execution of relevant program code and data by the electronic device.
[0207] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits shown in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and microprocessors, etc., and the storage module may be a memory.
[0208] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in the foregoing embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs (Digital Video Discs), CD-ROMs (Compact Disc Read-Only Memory), microdrives, magneto-optical disks, ROMs (Read-Only Memory), RAMs (Random Access Memory), EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable Programmable Read Only Memory), DRAMs (Dynamic Random Access Memory), VRAMs (Video Random Access Memory), flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0209] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a speech processing method as described in the above embodiments.
[0210] In addition, the electronic device provided in the embodiments of this application may specifically be a chip, component or module. The electronic device may include a connected processor and a memory. The memory is used to store instructions. When the electronic device is running, the processor may call and execute the instructions to make the chip execute a voice processing method in the above embodiments.
[0211] The electronic devices, computer-readable storage media, computer program products or chips provided in this application are all used to perform the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0212] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0213] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0214] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A speech processing method, characterized in that, Applied to electronic devices, the method includes: Acquire the voice data to be processed; Identify a first speech segment in the speech data to be processed; wherein the first speech segment represents multiple consecutive frames of speech data in the speech data to be processed; The first audio segment is classified into events to obtain the classification results; Based on the classification results, a target speech segment is determined; wherein, the user identifier corresponding to the target speech segment is a single user identifier; Based on the target speech segment, generate the target standard text.
2. The method according to claim 1, characterized in that, The process of classifying the first speech segment into events to obtain classification results includes: Determine the similarity between the first speech segment and each of the multiple preset events; Determine whether the maximum similarity among multiple similarity values is greater than or equal to a preset similarity value; If the maximum similarity is greater than or equal to the preset similarity, the event type corresponding to the maximum similarity is determined as the classification result.
3. The method according to claim 1, characterized in that, The step of determining the target speech segment based on the classification result includes: If the classification result corresponds to the event type of the single user identifier, the first voice segment is identified as the target voice segment.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The speech data to be processed is preprocessed to obtain preprocessed speech data; wherein, the preprocessing includes noise reduction and reverberation reduction. Determining the first speech segment in the speech data to be processed includes: Obtain the first speech segment from the preprocessed speech data.
5. The method according to any one of claims 1 to 3, characterized in that, The step of generating target standard text based on the target speech segment includes: Speech recognition is performed on the target speech segment to generate the first text content corresponding to the target speech segment; and, Generate a first event type identifier corresponding to the target speech segment and a second event type identifier corresponding to a second speech segment in the speech data to be processed; wherein, the second speech segment is a speech segment in the speech data to be processed other than the target speech segment; The target standard text is generated based on the first text content, the first event type identifier, and the second event type identifier.
6. The method according to claim 5, characterized in that, The first text content includes text content from the target user; the step of performing speech recognition on the target speech segment to generate the first text content corresponding to the target speech segment includes: Based on the target speech segment, the voiceprint features corresponding to each frame of speech data are obtained, as well as the first voiceprint feature of each speech segment; wherein, the first voiceprint feature includes multiple voiceprint features per frame. Based on the first voiceprint feature, the target voiceprint feature of the target user is determined; Based on the voiceprint features of each frame and the target voiceprint features, the target voice data corresponding to the target user in each frame of voice data is determined; The target speech data is subjected to speech recognition to obtain the text content of the target user.
7. The method according to claim 6, characterized in that, The step of determining the target user's target voiceprint features based on the first voiceprint feature includes: Clustering is performed on the first voiceprint feature to obtain multiple voiceprint feature sets; Determine the target cluster center of the target voiceprint feature set; wherein, the target voiceprint feature set is any one of the multiple voiceprint feature sets; The target cluster center is determined as the target voiceprint feature of the first target user; wherein, the first target user is the user corresponding to the target voiceprint feature set.
8. The method according to claim 7, characterized in that, The step of determining the target voice data corresponding to the target user in each frame of voice data based on the voiceprint features of each frame and the target voiceprint features includes: Determine the similarity value between the voiceprint features of each frame and multiple target voiceprint features; The voice data corresponding to the largest similarity value among multiple similarity values is determined as the target voice data of the first target user.
9. A voice processing device, characterized in that, Configured in an electronic device, the device includes: The acquisition module is used to acquire the voice data to be processed. A determining module is used to determine a first speech segment in the speech data to be processed; wherein the first speech segment represents multiple consecutive frames of speech data in the speech data to be processed; The classification module is used to classify the first speech segment into events and obtain the classification result; The processing module is used to determine the target speech segment based on the classification result; wherein the user identifier corresponding to the target speech segment is a single user identifier; The generation module is used to generate target standard text based on the target speech segment.
10. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 8.