Voice annotation method and device, computer equipment and storage medium

By combining speech overlap detection and sound source separation technologies with speaker recognition and automatic speech recognition, the problem of speech annotation in multi-speaker overlapping speech scenarios is solved, achieving efficient and accurate speech annotation and improving the utilization efficiency and automation of speech data.

CN121963765APending Publication Date: 2026-05-01PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, there is a lack of automated speech annotation methods in multi-speaker overlapping speech scenarios, which leads to inconsistencies between speech and text supervision, affecting the quality and efficiency of automatic speech recognition and speech synthesis. In addition, manual annotation is costly and time-consuming.

Method used

A collaborative processing mechanism of speech overlap detection, speaker recognition and sound source separation is adopted. Overlap detection distinguishes overlapping segments from non-overlapping segments, constructs a speaker set, and uses a sound source separator to separate sound sources. Combined with automatic speech recognition and speech-text splicing alignment, a speech annotation set is generated.

Benefits of technology

It enables automated and high-precision annotation of overlapping speech from multiple speakers without human intervention, improving the accuracy and efficiency of speech annotation, reducing the workload of manual annotation, and ensuring the continuity of the timeline and the robustness of speech annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963765A_ABST
    Figure CN121963765A_ABST
Patent Text Reader

Abstract

The invention discloses a voice annotation method and device, computer equipment and a storage medium, belongs to the technical field of artificial intelligence, and is applied to voice data annotation of science and technology finance or health medical scenes. According to the method, a speech overlap detection, speaker recognition and sound source separation cooperative processing mechanism is introduced into a speech annotation process, and overlap detection is firstly utilized to distinguish overlapped segments and non-overlapped segments, so that recognition confusion caused when multiple persons speak at the same time is avoided; a speaker set is constructed based on the non-overlapping segments, speaker prior information is provided for sound source separation of the overlapping segments, and the separation quality and speaker consistency are improved. According to the method, automatic voice recognition and text splicing alignment are performed on each separated sound source branch and non-overlapping segments, so that the manual annotation and segmentation workload can be reduced while the continuity of a time axis is ensured, a uniform voice annotation result is generated, automatic and high-precision annotation of multi-speaker and strong-overlapping voice data is realized, and the voice annotation efficiency is improved. And the voice annotation efficiency and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a speech annotation method, device, computer equipment, and storage medium. Background Technology

[0002] In existing research on dialogue speech processing, speech datasets are typically obtained through manual transcription. However, in natural dialogue scenarios, it is common for two or more speakers to speak simultaneously, leading to overlapping speech signals. These overlapping segments often correspond to only a single text transcription, failing to clearly label the boundaries and content of each speaker's speech. This results in inconsistency between speech and text supervision, consequently affecting the training quality and efficiency of downstream tasks such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and speech dialogue systems. For example, in ASR training, the model cannot correctly learn the interaction patterns between speakers, leading to an increased error rate in recognition results; in TTS, the lack of accurate overlap annotation prevents the model from synthesizing natural multi-speaker dialogue. Furthermore, manually annotating overlapping regions segment by segment is costly and time-consuming, making it unsuitable for constructing large-scale datasets.

[0003] For example, when processing historical consultation data related to financial or medical products, the lack of professional annotators to manually annotate overlapping areas segment by segment, coupled with the significant manpower and time required for the annotation process, makes it difficult to quickly convert such historical consultation data containing complex overlapping speech into high-quality annotated data. When using this unannotated data to train an ASR model, the model is prone to problems such as text content confusion and speaker misidentification when recognizing overlapping speech segments in customer consultations or product introductions. This affects the accuracy of tasks such as semantic understanding and intent analysis, and fails to provide accurate datasets for applications such as optimizing intelligent customer service systems and mining customer needs in the financial or medical fields.

[0004] Therefore, how to automatically detect, separate, and label overlapping speech without human intervention has become a key problem that urgently needs to be solved in the field of speech processing. Summary of the Invention

[0005] The purpose of this application is to provide a speech annotation method, apparatus, computer device, and storage medium to solve the problem of automatically detecting, separating, and annotating overlapping speech without human intervention.

[0006] To address the aforementioned technical problems, this application provides a speech annotation method, employing the following technical solution: A speech annotation method, comprising: Acquire the speech data to be labeled, and perform speech overlap detection on the speech data to be labeled to obtain overlapping and non-overlapping segments in the speech data to be labeled; Speaker identification is performed on non-overlapping segments in the annotated speech data, and a speaker set is constructed based on the speaker identification results; Based on the speaker set, a preset sound source separator is used to separate the sound sources of overlapping segments in the unannotated speech data, resulting in several sound source branch outputs; Automatic speech recognition and speech-text concatenation and alignment are performed on the outputs of the several sound sources respectively to obtain the first speech annotation set; Automatic speech recognition and speech-text splicing alignment are performed on the non-overlapping segments in the speech data to be labeled to obtain a second speech annotation set; Based on the first and second speech annotation sets, a speech annotation set for the speech data to be annotated is generated.

[0007] To address the aforementioned technical problems, this application also provides a voice annotation device, which employs the following technical solution: A voice annotation device, characterized in that it comprises: The overlap detection module is used to acquire the speech data to be labeled and to perform speech overlap detection on the speech data to be labeled, so as to obtain overlapping segments and non-overlapping segments in the speech data to be labeled. The speaker recognition module is used to identify speakers in non-overlapping segments of the unannotated speech data and to construct a speaker set based on the speaker recognition results. The sound source separation module is used to separate overlapping segments in the unannotated speech data based on the speaker set using a preset sound source separator, and obtain several sound source branch outputs; The first speech annotation module is used to perform automatic speech recognition and speech-text splicing and alignment on the outputs of the plurality of sound sources respectively, so as to obtain the first speech annotation set; The second speech annotation module is used to automatically recognize speech and align speech and text in the non-overlapping segments of the speech data to be annotated, so as to obtain a second speech annotation set. The speech annotation integration module is used to generate a speech annotation set for the speech data to be annotated based on the first speech annotation set and the second speech annotation set.

[0008] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution: A computer device includes a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps of the speech annotation method as described in any of the preceding claims.

[0009] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below: A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the speech annotation method as described in any one of the above descriptions.

[0010] Compared with the prior art, the embodiments of this application have the following main advantages: This application discloses a speech annotation method, apparatus, computer device, and storage medium, belonging to the field of artificial intelligence technology, and applied to speech data annotation in fintech or healthcare scenarios. This application effectively improves the accuracy and completeness of speech annotation in complex scenarios by introducing a collaborative processing mechanism of speech overlap detection, speaker recognition, and sound source separation into the speech annotation process. First, overlap detection is used to distinguish overlapping and non-overlapping segments, avoiding recognition confusion caused by direct speech recognition when multiple people are speaking simultaneously. Then, a speaker set is constructed based on non-overlapping segments, providing prior speaker information for sound source separation of overlapping segments, improving separation quality and speaker consistency. By performing automatic speech recognition and text splicing alignment on each separated sound source path and non-overlapping segment, the workload of manual annotation and segmentation can be reduced while ensuring temporal continuity, lowering the probability of mismatch. Finally, a unified speech annotation result is generated based on the first and second speech annotation sets, achieving automated and high-precision annotation of multi-speaker, highly overlapping speech data, improving speech annotation efficiency and robustness. Attached Figure Description

[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 An exemplary system architecture diagram is shown, in which this application can be applied; Figure 2 A flowchart of one embodiment of the speech annotation method according to this application is shown; Figure 3 A flowchart illustrating another specific embodiment of the speech annotation method according to this application is shown; Figure 4 A schematic diagram of one embodiment of the speech annotation device according to this application is shown; Figure 5A schematic diagram of another specific embodiment of the speech annotation device according to this application is shown; Figure 6 A schematic diagram of the structure of one embodiment of a computer device according to this application is shown. Detailed Implementation

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0020] It should be noted that the speech annotation method provided in this application is generally executed by a server / terminal device, and correspondingly, the speech annotation device is generally set in the server / terminal device.

[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative; the system can have any number of terminal devices, networks, and servers depending on implementation needs.

[0022] Continue to refer to Figure 2-3 A flowchart of an embodiment of the speech annotation method according to this application is shown. The speech annotation method includes the following steps: S201, acquire the speech data to be labeled, and perform speech overlap detection on the speech data to be labeled to obtain overlapping segments and non-overlapping segments in the speech data to be labeled; Specifically, the audio data to be labeled first undergoes acoustic feature preprocessing, including but not limited to framing the original audio waveform, windowing, short-time Fourier transform, and logarithmic amplitude compression of the spectrogram to generate a frame-level acoustic feature sequence. Then, these acoustic features are input into an overlap detection model, which can be a time-series classification model based on a bidirectional recurrent neural network, convolutional temporal network, or Transformer structure, used to estimate the probability of whether each frame of audio belongs to an overlapping speech interval. The model uses a Sigmoid or Softmax activation function in the output layer to generate a frame-level overlap probability distribution. Further, by smoothing, dilating, and closing the probabilities of consecutive frames, combined with minimum duration constraints and silence filtering strategies, the detection results are transformed from frame-level labels into time interval-level sets of overlapping and non-overlapping segments.

[0023] For example, in a Q&A system for elderly health consultations, when the consultant and the health advisor speak simultaneously to supplement information, such as the consultant asking "How many times a day should this medicine be taken?" while the advisor responds "Taking it before meals is more effective," the system can identify the overlapping segment of 0.5-1.2 seconds through the aforementioned overlap detection process. The consultant's voice starts at 0.5 seconds and the advisor's voice starts at 0.7 seconds, with a 0.5-second overlap between the two in the 0.7-1.2 second range. After model calculation, this range is marked as an overlapping segment, while the first half of the consultant's question (0-0.5 seconds) and the advisor's complete reply (1.2 seconds later) are classified as non-overlapping segments.

[0024] In one specific embodiment of this application, acoustic feature extraction is performed on the speech data to be labeled, and acoustic features are extracted for each frame. (For example, log-mel, energy, short-time Fourier amplitude, and speaker embedding, etc.), the overlap probability is obtained after modeling with a context encoder (BiLSTM / Transformer). The specific principle is as follows:

[0025] In the formula, This represents the probability that at time frame t, there are two or more speakers speaking simultaneously. This represents the speech feature vector (encoded feature) of time frame t, which is usually derived from log-mel features, STFT spectrum, or context-coded output after BiLSTM / Transformer. , These are parameters belonging to the linear transformation layer (or fully connected layer), used to map input features to the latent space; This represents a non-linear activation function, which can be ReLU, GELU, Tanh, or other deep network activation units; , These are parameters belonging to the linear transformation of the output layer. This represents the Sigmoid function.

[0026] This formula represents the application of a two-layer neural network structure to perform linear transformation and nonlinear activation on the encoded frame-level speech features, and outputs the temporal frame-level overlapping speech probability through the Sigmoid function. , These are the hidden layer parameters. , For output layer parameters, Enhance features for the current frame context. For activation function, This is a probability mapping function.

[0027] S202, perform speaker identification on non-overlapping segments in the labeled speech data, and construct a speaker set based on the speaker identification results; Specifically, speaker recognition processing is performed on the non-overlapping segments obtained in step S201 to avoid interference from overlapping speech on the embedded features. First, speaker representation features are extracted from each non-overlapping speech segment, including Mel-frequency cepstral coefficients, speaker perceptual spectrum features, or time-frequency mask features after speech enhancement. Then, these features are input into a speaker embedding extraction model, such as an x-vector, ECAPA-TDNN, or ResNet architecture model, to generate an embedding vector representation with global session discrimination capabilities. Next, similarity calculations and clustering operations are performed on the embedding space, such as spectral clustering based on cosine similarity, hierarchical clustering, or adaptive graph clustering methods, to cluster speaker instances belonging to the same session, thereby forming a session-level speaker set. In this set, each speaker entry is bound to a corresponding embedding vector, a vocalization interval index, and a speech source identifier.

[0028] S203, based on the speaker set, uses a preset sound source separator to separate the overlapping segments in the speech data to be labeled, and obtains several sound source branch outputs; Specifically, in this step, for the overlapping segments detected in S201, feature enhancement and speaker prior constraint injection operations are first performed. The system extracts time-domain or frequency-domain mixed features from the overlapping audio segments and, combined with the speaker set obtained in S202, injects the embedding vectors of the corresponding candidate speakers as conditional inputs or control signals into the mask prediction module or feature decoder inside the sound source separation model. Subsequently, the enhanced features are input to a preset sound source separator, which can be Conv-TasNet, DPRNN, TF-GridNet, or a multi-mask prediction network. The model generates K sound source branch signals and corresponding time mask matrices at the output. To improve track stability, the system further calculates the similarity between each branch signal and each embedding vector in the speaker set, and assigns identity constraints and performs mask corrections to the tracks based on the similarity ranking results, completing the re-estimation and consistency assignment of the sound source signals, and finally outputting several sound source branch outputs that satisfy the mixed consistency constraints.

[0029] The source splitter generates several source-specific outputs, each corresponding to a specific speaker in the speaker set. The duration of each output signal is consistent with the original overlapping segment, ensuring accurate time alignment between speech recognition and annotation. For example, when the overlapping segment contains two speakers, the source splitter will output two source-specific outputs, each corresponding to an independent speech signal from one speaker within the overlapping time period. Each output signal retains only the speech component of its corresponding speaker, effectively suppressing interference from other speakers.

[0030] The separation loss of a sound source separator can be defined as:

[0031] In the formula, The input mixed speech signal is the original audio that is superimposed from multiple speakers; This represents the output of the kth sound source branch obtained from the model separation, which belongs to the network prediction result; This shows the signal after all the separated orbits have been re-overlaid. This represents the mixing consistency loss, used to constrain the superposition of separated tracks to be close to the original signal, and to avoid energy drift introduced during the separation process. This indicates the magnitude spectrum extraction operator; This represents the k-th reference source speech; This represents the difference between the amplitude spectrum of the separated speech and the amplitude spectrum of the reference speech. This represents the spectral amplitude consistency loss, used to constrain the separated speech to approximate the reference source signal in the frequency domain. It is the weighting coefficient.

[0032] S204, perform automatic speech recognition and speech-text splicing and alignment on the outputs of the plurality of sound sources respectively to obtain the first speech annotation set.

[0033] Specifically, please refer to Figure 3 Step S204 specifically includes sub-steps S241-S245, wherein: S241, Automatic speech recognition is performed on the outputs of several sound source channels respectively to obtain the candidate text corresponding to each sound source channel output; Specifically, in this step, automatic speech recognition processing is performed on each audio source branch output from S203. First, speech recognition feature sequences are extracted for each audio signal, including filter bank energy spectrum, log-Mel spectrum, or acoustic coding vectors required by the end-to-end model. Then, these features are input into the recognition model, which can be an end-to-end speech recognition network based on CTC, RNN-Transducer, or attention mechanism. The recognition model performs sequence modeling and character or word-level probability decoding on the entire branch speech, outputting the corresponding acoustic posterior probability distribution, possible text sequence candidates, and word-level confidence information. Simultaneously, the system retains intermediate layer alignment information or temporal distribution information to facilitate timestamp extraction and text alignment calculation, and binds and stores the currently recognized candidate text set with each audio source branch output.

[0034] During processing, the Automatic Speech Recognition (ASR) model preprocesses the segmented speech signals, performing noise suppression, endpoint detection, and speech enhancement to improve recognition accuracy. For the output candidate text, the system not only provides the optimal decoding result but also retains an N-best candidate list.

[0035] S242, perform permutation invariance training on several source branch outputs and several candidate texts respectively, and determine the first target text corresponding to each source branch output; Specifically, in this step, several source outputs and the candidate text set obtained in S204 are taken as input, and a permutation-invariant training mechanism is used to perform multi-permutation matching calculations. First, the system obtains the corresponding acoustic posterior probability sequence for each output and constructs alignment path constraints based on the character sequences of the candidate texts. Then, using a preset loss function, such as CTC loss, edit distance loss, or transduction loss, the matching loss is calculated for each output and each candidate text combination to obtain pairwise loss matrices. Next, the system enumerates all permutation and combination relationships between outputs and candidate texts, sums and counts the overall losses corresponding to different permutations and combinations, and selects the permutation with the smallest combination loss as the optimal matching result for the current sample. Finally, based on the optimal permutation result, the first target text corresponding to each source output is determined, and this mapping relationship is stored at the output level.

[0036] The Permutation Invariant Training (PIT) mechanism effectively addresses the uncertainty of the path order and speaker identity after overlapping segment separation by removing the fixed permutation constraint between the source path outputs and candidate texts. For example, when overlapping segments are separated into two paths and automatic speech recognition generates two candidate texts, PIT calculates the matching loss for four combinations: path A-text 1, path A-text 2, path B-text 1, and path B-text 2. It then selects the combination with the minimum total loss (e.g., path A-text 2 and path B-text 1) as the optimal pairing, ensuring that each path output corresponds to the semantically most matching first target text. In its implementation, the system constructs a path-text matching loss matrix and uses the Hungarian algorithm or a greedy search algorithm to find the minimum weight matching path, achieving precise binding between paths and text.

[0037] The overlapping regions are trained using PIT-CTC / Transducer, with the loss function set as follows:

[0038] In the formula, This represents the ASR training loss function in overlapping speech scenarios, specifically designed for multi-speaker overlapping speech recognition, and adopts the PIT permutation invariant loss form. This indicates the number of speakers participating in the current speech segment / the number of sound source separation output channels; This represents the set of all permutations of K elements, also known as the permutation space / matching search space; This represents the CTC (Connectionist Temporal Classification) loss function, which is mainly used for speech recognition sequence modeling and alignment learning without frame-level annotation; It is the ASR predicted acoustic posterior probability sequence of the output of the k-th sound source branch; This represents the target text sequence corresponding to the k-th speech path specified by the permutation mapping. It should be noted that non-overlapping segments are calculated using the standard single-stream CTC / CE loss.

[0039] S243, obtain the first timestamp corresponding to each sound source branch output, and splice and align several sound source branch outputs and several first target texts based on the first timestamp to generate initial annotation labels, wherein the initial annotation labels are pseudo annotation labels; Specifically, in this step, firstly, based on the alignment information or time boundary estimation results output by the automatic speech recognition model, the start time, end time, and character-level or word-level alignment timestamps of each sound source branch output at the corresponding position in the first target text are extracted, forming a first timestamp sequence that corresponds one-to-one with the text sequence. Subsequently, using the timestamps as a reference, the system maps several sound source branch outputs to a unified time axis, segments and indexes time intervals, and synchronously segments the text sequence according to the timestamps, constructing a time-level mapping relationship between speech segments and text segments. When multiple branches appear in the same time interval, the system performs cross-branch conflict detection and time segment alignment processing, splicing, trimming, and uniformly marking multiple text and speech segments to form continuous and compact pseudo-annotation structure entries on the time axis. Finally, the text content, time boundary, and branch identifier corresponding to each time interval are organized into initial annotation labels and output.

[0040] S244, calculate the confidence level of the initial label and filter based on the initial label to obtain the first label; Specifically, in this step, for each initial annotation generated by S206, the system extracts relevant confidence evaluation elements from multiple dimensions, including the mean probability of candidate text recognition, speech quality indicators of source-path speech, consistency scores between text and candidate permutations, and cross-model output consistency statistics. Subsequently, these elements are combined to form a confidence feature vector, and a weighted scoring rule or confidence evaluation model is used for comprehensive calculation to obtain segment-level or label-level confidence scores. Further, for multiple initial annotations from the same time region or the same speech segment, the system performs cross-consistency comparisons, calculates the consistency strength between different outputs, and uses the consistency index as an additional constraint in the confidence judgment. Finally, based on the confidence score threshold and consistency constraints, the system filters the initial annotations, retaining those that meet the confidence conditions and marking them as the first annotation; the remaining annotations are either downgraded or cached.

[0041] S245, based on several source branch outputs, several first target texts and first annotation labels, construct a first speech annotation set; Specifically, in this step, the system, based on the first annotation label, combines the audio segments output from the sound source branch associated with each label, the first target text sequence, and the corresponding time boundary information to form structured annotation entries. Subsequently, based on the mapping result between the branch number and the speaker set, the system binds the corresponding speaker identity and track source information to each annotation entry, thereby establishing a multi-dimensional data association between speech segments, text content, timestamps, and identity labels. For multiple annotation entries belonging to the same sound source branch and with consecutive time, the system performs segment splicing processing according to the timestamp order, and semantically splices and integrates the broken text at cross-segment boundaries while maintaining the original text boundary structure. Next, the system performs data integrity and format consistency checks on the spliced ​​annotation sequence, filtering out abnormal annotation entries with time conflicts, duplicate annotations, or missing text segments. Finally, all annotation sequences that meet the conditions are summarized according to the branch dimension to form the first speech annotation set.

[0042] S205, perform automatic speech recognition and speech-text splicing and alignment on the non-overlapping segments in the speech data to be labeled to obtain a second speech annotation set.

[0043] For details, please refer to the following: Figure 3 Step S205 specifically includes sub-steps S251-S253, wherein: S251, Automatic speech recognition is performed on non-overlapping segments in the speech data to be labeled to obtain the second target text; Specifically, in this step, automatic speech recognition processing is performed only on the non-overlapping segments obtained in step S201. First, acoustic recognition features are extracted from each non-overlapping segment, including spectrograms, Mel-frequency filter features, or input feature representations of an end-to-end speech recognition model. Then, these features are input into the same or independently configured automatic speech recognition model to decode the non-overlapping speech sequence, outputting the corresponding second target text and confidence information. Considering that non-overlapping speech does not involve source conflict or separation uncertainty, the model in this step can adopt a conventional single-stream recognition path, preserving alignment information and temporal boundary indices. Furthermore, the system maintains the correspondence between non-overlapping segments and speaker sets, ensuring consistent labeling of speech segments in the identity dimension.

[0044] S252, obtain the second timestamp of the non-overlapping segment in the speech data to be labeled, and splice and align the non-overlapping segment in the speech data to be labeled and the second target text based on the second timestamp to generate the second label; Specifically, in this step, firstly, based on the alignment information output by the automatic speech recognition model for non-overlapping segments, the start and end times and alignment trajectories of each character (or word) in the second target text are extracted to construct a second timestamp sequence corresponding to the text sequence. Subsequently, the system segments and maps the non-overlapping speech segments according to their timestamps, and performs time-level concatenation mapping between the segmented speech segments and their corresponding text segments in the second target text, forming a pairing relationship between speech and text segments. Based on this, the system organizes multiple mapped segments into structured annotation entries according to chronological order, and performs sequential concatenation processing on time-continuous but text-structure-related segments to generate a time-continuous second annotation label sequence. Simultaneously, the system combines speaker set information to consistently bind the annotation labels to the corresponding speaker identities, ensuring that the index association relationships between different data source dimensions remain consistent.

[0045] S253, construct a second speech annotation set based on non-overlapping segments, second target text, and second annotation labels in the speech data to be annotated; Specifically, in this step, the system constructs a second speech annotation set corresponding to non-overlapping speech segments based on the second annotation tags. First, the system combines each second annotation tag with its corresponding non-overlapping speech segment, second target text, and timestamp information to form a structured annotation entry, which is then indexed and managed according to chronological order and segment number. Subsequently, for time-continuous annotation entries belonging to the same speaker or the same dialogue turn, the system performs speech segment-level and text-level concatenation and merging processing to form a longer-scale annotation sequence. Next, the system performs format and integrity verification on the merged annotation sequence, removing abnormal entries with overlapping time intervals, text conflicts, or missing annotations, and marking suspicious entries. Finally, all verified annotation sequences are grouped and summarized to form the second speech annotation set.

[0046] S206, Based on the first speech annotation set and the second speech annotation set, generate a speech annotation set for the speech data to be annotated.

[0047] Specifically, in this step, the system receives a first speech annotation set from S245 and a second speech annotation set from S253, and merges the two sets under the constraints of a unified timeline, a unified speaker index, and a unified annotation format. First, based on timestamp information, the system maps the two annotation sets to a timeline within the same session, and performs conflict detection and boundary correction operations on annotation entries with time overlap or boundary overlap. Then, without changing the original tag source attributes, the system performs priority fusion or segment retention processing on annotation entries in overlapping areas to ensure the continuity and integrity of the annotation sequence in the time dimension. Next, based on the speaker set's identity index, the system groups and organizes annotation entries from different sets according to their identities, thus forming a consistent and hierarchically clear set of speech annotation entries. Finally, the system outputs the merged annotation entries as the target speech annotation set according to the session dimension or data sample dimension, maintaining an index-level binding relationship with the original speech data.

[0048] During the system training and deployment phase, the separator and speaker embedding are pre-trained first; then, a single-stream ASR is pre-trained; multi-stream joint fine-tuning is performed using pseudo-labels; finally, end-to-end fine-tuning is performed (if a unified model is used); overlapping segments are sampled by confidence weight, with high-confidence segments having a larger training weight; during inference, frame-level Povl(t) is used as a gating switch, and if the threshold is exceeded, multi-stream separation and multi-path decoding are enabled to save computational resources; for low-latency scenarios, "detection + soft assignment" can be used instead of explicit separation, that is, the hybrid posterior is soft-assigned based on the speaker posterior, and then multi-head decoding is performed.

[0049] Specifically, during the system training phase, the source separator and speaker embedding model are first pre-trained independently. The source separator is trained using speech data containing a large amount of real reverberation and noise scenarios to learn effective spatial-temporal feature separation capabilities and output source-splitting signals with a high signal-to-noise ratio. The speaker embedding model is trained using a large-scale speaker speech database to extract fixed-dimensional vectors that characterize individual acoustic features. Next, the single-stream automatic speech recognition (ASR) model is pre-trained on standard clean speech corpora to give it basic speech-to-text capabilities. After completing the above pre-training, the system uses the pseudo-annotated data generated in the previous steps (i.e., the high-confidence portions in the first and second speech annotation sets) to jointly fine-tune the separator, speaker embedding model, and multi-stream ASR model. During this process, for overlapping speech segments, the system performs weighted sampling based on their confidence calculated during pseudo-annotation generation, assigning greater training weights to high-confidence overlapping segments to guide the model to focus more on speech recognition accuracy in complex scenarios. If the system adopts an end-to-end unified model architecture, after the joint fine-tuning, another round of end-to-end overall fine-tuning will be carried out to further optimize the collaborative performance between the various modules of the model.

[0050] During the inference and deployment phase, the system introduces a frame-level overlap detection probability, Povl(t), as a gating mechanism. When the real-time calculated Povl(t) exceeds a preset threshold, it is determined that there is speaker overlap in the current speech frame. The source separation module is then activated for multi-stream separation processing, and a multi-path decoding process is triggered to obtain the candidate text corresponding to each path. When Povl(t) is below the threshold, it is determined to be single-speaker speech, and a single-stream ASR model is directly used for decoding. This effectively saves computational resources and power consumption while ensuring recognition accuracy. For low-latency application scenarios (such as real-time conference transcription and real-time caption generation), the system can adopt a lightweight strategy of "detection + soft allocation" instead of performing computationally intensive explicit source separation steps. Specifically, the system first calculates the frame-level speaker posterior probability of the mixed speech using a speaker activity detection model. Then, it performs soft-assignment on the acoustic posterior probability of the mixed speech based on this speaker posterior probability, that is, it distributes the mixed posterior probability proportionally to each active speaker to form multiple virtual "segmented posteriors". Finally, it performs multi-head decoding on these virtual segmented posteriors to significantly reduce end-to-end processing latency and meet real-time requirements while sacrificing a small amount of accuracy.

[0051] In the above embodiments, this application can automatically annotate conversational speech data containing overlapping speech without human intervention. It achieves a collaborative fusion of multiple speech processing technologies, including speech overlap detection, speaker recognition, source separation, automatic speech recognition, and permutation-invariant training, thereby simultaneously completing speech separation, text alignment, and structured speech segment annotation generation within the same processing flow. By constructing a speaker set in non-overlapping segments and using it as a separation and alignment constraint, the source outputs of overlapping segments are guided to be consistently assigned to the corresponding text content, allowing the speech trajectories of different speakers to be distinguished and decomposed along the timeline. Simultaneously, through timestamp splicing, cross-path alignment merging, and confidence filtering, the automatically generated pseudo-annotation labels possess high reliability in terms of temporal boundary consistency, text matching stability, and annotation data integrity. This transforms speech data, which originally only possesses a single, overall transcribed text, into a structured speech annotation set that can be directly used for multi-speaker speech recognition, speech synthesis, and speech modeling training, improving the efficiency of speech data utilization and the degree of automation in overlapping speech scenarios.

[0052] Specifically, for example, in processing historical consultation corpora for financial products, when overlapping speech segments exist in the consultation corpus, such as when customers and customer service personnel speak simultaneously, causing speech signal superposition, the speech annotation method of this application is adopted. The system first decomposes the customer's and customer service personnel's speech signals into two independent speech source outputs using sound source separation technology. Subsequently, based on the alignment information of the automatic speech recognition model, the timestamp of each branch at the corresponding position in the target text (such as the transcribed text of the consultation dialogue record) is extracted, constructing a time-level mapping relationship between speech segments and text segments.

[0053] During the conflict detection phase, the system detected a temporal overlap between the customer's question, "What is the annualized rate of return for this financial product?" and the customer service representative's simultaneous explanation, "Its return is calculated on a daily interest basis." Therefore, the system performed cross-segment time segment alignment processing, precisely splicing and trimming the two text and audio segments. For example, the system segmented and indexed the customer's question (00:01:23-00:01:28) and the customer service representative's explanation (00:01:25-00:01:30) along the timeline, forming continuous pseudo-annotation structure entries. Next, the system calculated the confidence level of the initial annotation labels based on dimensions such as text recognition probability and audio quality metrics. Assuming the customer's audio segment had a mean recognition probability of 0.92 and an audio quality metric (such as signal-to-noise ratio) of 25dB, while the corresponding values ​​for the customer service representative's audio segment were 0.90 and 23dB, both exceeding the set confidence threshold of 0.85, they were selected as the first annotation labels. Subsequently, the system combined these labels with the corresponding audio segments and text content to construct the first audio annotation set. For non-overlapping segments, such as the voice segment 00:01:10-00:01:22 in which the customer states their personal needs, the system directly generates the second target text and timestamp through the single-stream ASR model to construct the second voice annotation set.

[0054] Finally, when merging the two sets, the system integrates the overlapping customer and customer service annotation entries, as well as the non-overlapping customer monologue entries, according to a unified timeline and speaker index (customer is labeled S1, customer service is labeled S2), forming a complete structured speech annotation set. This set clearly records the monologue content of S1 from 00:01:10 to 00:01:22, the questions from 00:01:23 to 00:01:28, and the explanation content of S2 from 00:01:25 to 00:01:30, achieving accurate differentiation and annotation of the speech trajectories of different speakers in overlapping speech scenarios. This corpus can be used to train multi-speaker speech recognition models or dialogue intent analysis models in the financial customer service field.

[0055] Furthermore, based on the speaker set, a preset source separator is used to separate overlapping segments in the annotated speech data, resulting in several source-path outputs. This process specifically includes: Acoustic features are extracted from overlapping segments to obtain spectral feature sequences corresponding to time frames; Based on the speaker feature vectors in the speaker set, feature enhancement processing is performed on the spectral feature sequence to obtain enhanced speech features; The enhanced speech features are input into a preset sound source separator. The sound source separator is used to separate and calculate the enhanced speech features to obtain several initial sound source outputs corresponding to overlapping segments and corresponding time masks. Calculate the similarity between the output of each initial sound source and the feature vector of each speaker to obtain the candidate speaker set corresponding to each output of the initial sound source. The initial sound source branch outputs are assigned and reconstructed based on identity constraints. For the initial sound source branch outputs that do not meet the identity consistency requirements, mask correction and feature re-estimation are performed to obtain several sound source branch outputs.

[0056] In this embodiment, before performing source separation on overlapping segments, the system first performs short-time framing and window function processing on the audio to be processed based on the time-frequency structure characteristics of the speech signal, and performs short-time Fourier transform on each frame signal to extract a spectral feature sequence containing spectral amplitude and phase components. Based on this, by introducing the speaker set obtained in the non-overlapping segments, the embedding vectors or deep representation features corresponding to each speaker in the set are used as external prior conditions, and jointly modeled or concatenated with the spectral features. This allows for discriminative enhancement of each potential sound source track at the feature level, resulting in an enhanced speech feature set with speaker-constrained attributes. Subsequently, the enhanced speech features are input into a preset source separator. The source separator can employ a separation model based on time-domain mask estimation, frequency-domain amplitude recovery, or convolutional temporal network structure. By performing sequence modeling and mask prediction on the enhanced features, it outputs several initial source-splitting signals and their corresponding time mask matrices. Next, based on the feature vectors in the speaker set, the system calculates the similarity score between each initial sound source branch output and each candidate speaker, and establishes a matching relationship between the candidate speaker set and the branch outputs based on the score results. If some branch outputs are found to have inconsistencies in identity or insufficient similarity confidence with any candidate speaker, the system further performs local correction, re-estimation, or reallocation processing on the time mask and speech features corresponding to the branch output to improve the consistency and stability of the branch outputs in the identity dimension, ultimately forming several sound source branch outputs that meet the speaker constraint requirements.

[0057] Through the above steps, the sound source separation process can perform track differentiation and stable identity allocation under the speaker's prior constraints, thereby improving the distinguishability and consistency of the sound source splitting output under overlapping speech conditions.

[0058] Furthermore, the step of performing permutation-invariant training on several source outputs and several candidate texts to determine the first target text corresponding to each source output specifically includes: Several sound source outputs are input into a preset automatic speech recognition model to obtain the acoustic posterior probability sequence and the recognition confidence information of the candidate text corresponding to each sound source output. Based on the candidate text, the acoustic posterior probability sequence is aligned with the candidate text sequence for calculation, and the matching loss value between each sound source branch output and each candidate text is obtained according to the preset loss function. Using the matching loss value as the evaluation metric, we enumerate all permutations and combinations between several source outputs and several candidate texts, and calculate the combination loss corresponding to each permutation and combination. The target permutation and combination relationship corresponding to the minimum permutation and combination loss is determined from the permutation loss, and the target permutation and combination relationship is used as the optimal matching result of permutation-invariant training. Based on the optimal matching result, the candidate text corresponding to each sound source branch output is determined, and the candidate text is determined as the first target text corresponding to the sound source branch output.

[0059] In this embodiment, to address the uncertainty in the correspondence between the sound source separation output and the text sequence, after performing automatic speech recognition on several sound source branch outputs, the recognized candidate texts are not directly used as fixed results. Instead, a permutation-invariant training mechanism is introduced to jointly match and model the multiple outputs and multiple text candidates. Specifically, each sound source branch output is first independently recognized to obtain the acoustic posterior probability distribution sequence corresponding to that branch signal, while retaining the candidate text sequences and their confidence information generated during the recognition process. Subsequently, based on the candidate texts as reference sequences, each acoustic posterior probability sequence is aligned with the candidate texts one by one. By using preset loss functions such as concatenation time classification loss, edit distance loss, or transduction loss, the matching loss value between each sound source branch output and each candidate text is calculated, and a complete matching loss matrix is ​​constructed accordingly. Next, the system enumerates the permutation and combination relationships between all sound source branch outputs and candidate texts, and performs a weighted summation of the matching losses under each combination to obtain the combination loss under different permutation and combination conditions. By comparing the combined losses, the arrangement corresponding to the minimum combined loss is determined as the optimal matching relationship. This allows for the automatic selection of the most suitable candidate text for each sound source branch output without pre-constraining the output order. Finally, based on the optimal arrangement, the candidate text corresponding to each sound source branch output is determined as the first target text and bound to that output.

[0060] Through the above steps, adaptive matching between multi-channel speech and multiple candidate texts can be achieved without pre-setting the relationship between the output of the multi-channel speech and the order of the text, thus avoiding recognition ambiguity caused by order uncertainty and improving the accuracy and stability of text allocation.

[0061] Further, the steps of obtaining the first timestamp corresponding to each sound source branch output, and concatenating and aligning several sound source branch outputs and several first target texts based on the first timestamp to generate initial annotation labels specifically include: Based on the alignment results calculated by the alignment calculation, the start and end times of each speech segment in the output of each sound source branch are extracted to generate the first timestamp sequence corresponding to the first target text. Based on the first timestamp sequence, the outputs of several sound sources are divided into time boundaries on a unified time axis to obtain a set of branched speech segments in units of time segments; The first target text is segmented into text segments according to the first timestamp sequence, and each text segment is concatenated and associated with the corresponding split speech segment at the time level to form a mapping relationship between speech and text segments; When multiple sound source outputs correspond to the same time interval, based on the time overlap ratio and text coverage, multiple speech segments and text segments within the same time interval are aligned, merged, and spliced. Based on the mapping relationship of the processed speech-text segments, structured annotation entries containing time boundary information, sound source branch numbers, and the first target text content are generated, and these structured annotation entries are used as initial annotation labels.

[0062] In this embodiment, after matching the source branch outputs with the first target text, to ensure that the text content can form a structured binding relationship with the corresponding speech segments on the timeline, the system further performs timestamp extraction and alignment splicing processing on each source branch output based on the alignment data generated during automatic speech recognition. Specifically, the system first parses the alignment trajectory information at the character, word, or segment level from the recognition alignment results, extracts the start and end times of each speech segment in the branch output, and constructs a first timestamp sequence corresponding one-to-one with each text unit of the first target text. Subsequently, using a unified conversation timeline as a reference, several source branch outputs are segmented according to the timestamps to form a set of branch speech segments in units of time segments, and the time segments between different branches are aligned and mapped using a time index. Next, the system synchronously segments the first target text according to the timestamp sequence, and splices and associates each text segment with its corresponding branch speech segment at the time level to form a mapping relationship between speech and text segments. If overlapping coverage is detected between different output channels within the same time interval, the system further performs merging, trimming, or splicing operations on multiple speech and text segments within that time interval based on indicators such as the time overlap ratio and text content coverage. This ensures the consistency and integrity of the speech-text mapping relationship within the time interval. Finally, the system organizes the processed speech-text segment mapping relationship into structured annotation entries, recording the time boundary, channel number, and first target text content in each entry. This entry is then output as the initial annotation label corresponding to that time segment.

[0063] Through the above steps, a refined mapping relationship is formed between speech segments and text content on a unified time axis, thereby achieving clear segmentation and structured annotation output of multi-channel speech in the time dimension.

[0064] Furthermore, when multiple sound source outputs correspond to the same time interval, the steps of aligning, resolving, and concatenating multiple speech segments and text segments within the same time interval based on the time overlap ratio and text coverage specifically include: Obtain the speech segments and time intervals corresponding to the output of each sound source within the same time interval, and calculate the time overlap ratio of each speech segment based on the intersection duration of the speech segments and time intervals. Each text segment within the time interval is segmented and mapped according to the first timestamp, and the corresponding text coverage is calculated based on the effective coverage length of the text segment within the time interval. Based on the time overlap ratio and text coverage, multiple speech segments and text segments within the same time interval are jointly scored, and the speech segments and text segments are prioritized according to the joint score. Based on the priority ranking results, the dominant speech segment and its corresponding dominant text segment are determined, and the non-dominant speech segments are subjected to boundary trimming compensation processing. After completing the boundary clipping and compensation process, the processed speech segments and corresponding text segments within the same time interval are merged and spliced ​​to generate a merged and aligned result that is time-continuous, text-consistent, and has clear branching labels. The merged and aligned result is then used as the initial annotation sub-label corresponding to the time interval.

[0065] In this embodiment, when multiple sound source outputs simultaneously cover the same time interval, the system does not simply replace any one output with the result of any one output. Instead, it performs refined conflict resolution and merging processing on the multiple speech segments and their corresponding text segments based on a joint metric of time and text dimensions. Specifically, the system first obtains all sound source output segments within the same time interval based on time index information, and calculates the intersection duration of each speech segment within that time interval by combining the start and end times of the time interval, thus obtaining the time overlap ratio, which is used to characterize the main coverage of the output segments within that time period. Subsequently, the system segments and maps the text segments within the time interval according to the first timestamp sequence, and calculates the effective coverage length of each text segment within the time interval to form a corresponding text coverage index. Based on this, the system comprehensively scores the time overlap ratio and text coverage, using it as a joint evaluation criterion to measure the dominance of speech and text segments within that time interval, and prioritizes the multiple speech and text segments accordingly. Furthermore, for the dominant speech segment and its corresponding text segment in the sorting results, the system retains them as the dominant content of that time interval. For other non-dominant speech segments, boundary trimming or compensation is performed while ensuring temporal continuity, so that they form a consistent temporal boundary with the dominant segment. Finally, after completing the boundary adjustment, the system performs a merge and concatenation operation on the processed speech segments and their corresponding text segments. While keeping the branch identification information unchanged, it generates a merged and aligned result that is temporally continuous, has consistent text content, and a clear affixation. This result is then output and stored as the initial annotation sub-label corresponding to that time interval.

[0066] Through the above steps, joint conflict resolution and consistent splicing of speech and text can be achieved in multi-path overlapping time segments, thereby improving the completeness and structured expression of the annotation results within the time interval.

[0067] Further, the steps of calculating the confidence level of the initial annotation labels and filtering based on the initial annotation labels to obtain the first annotation labels specifically include: For each initial label, obtain the recognition confidence information, speech quality evaluation index and text matching score of the corresponding sound source output, and construct the confidence feature vector corresponding to the initial label; Based on a pre-set reliability assessment model, the confidence feature vector is comprehensively calculated to obtain the segment-level confidence score of the initial labeled label. Based on the segment-level confidence score, the initial annotation labels are compared with different generated results for the same speech segment across samples for cross-sample consistency, and a multi-result consistency index is calculated. Based on segment-level confidence scores and multiple result consistency indicators, the reliability of the initial label is judged, and the target initial label that meets the preset screening conditions is determined and the target initial label is determined as the first label.

[0068] In this embodiment, after generating the initial labels, to prevent low-quality pseudo-labels caused by sound source separation deviations, unstable speech recognition, or text matching errors from entering the training and label set construction process, the system further performs confidence calculation and screening on the initial labels. Specifically, for each initial label, the system first extracts quality assessment elements from multiple related data sources, including recognition confidence information output from the automatic speech recognition process, signal-to-noise ratio of the corresponding speech signals output from the sound source branching, speech clarity or residual crosstalk intensity, and text-side consistency indicators such as text matching scores and permutation stability scores generated during permutation-invariant training and alignment calculation. These indicators are then combined according to preset feature dimensions to construct a confidence feature vector describing the reliability of the label. Subsequently, the system inputs this confidence feature vector into a preset confidence assessment model or performs comprehensive calculation according to a weighted scoring strategy to obtain a segment-level confidence score for label-level judgment. Furthermore, for multiple initial annotation labels generated for the same speech segment under different processing paths or different model versions, the system performs cross-sample consistency comparison and calculates a multi-result consistency index to reflect the consistency of the annotation results of the speech segment from different perspectives. Finally, the system makes a reliability judgment based on the segment-level confidence score and the multi-result consistency index, and determines the initial annotation labels that meet the preset screening threshold and stability conditions as the first annotation labels, while removing or downgrading labels that do not meet the conditions.

[0069] By following the steps above, multidimensional confidence metrics can be used to screen for pseudo-labeling results, thereby improving the reliability and stability of labeled data entering the label set construction stage.

[0070] Furthermore, the step of constructing a first speech annotation set based on several source branch outputs, several first target texts, and first annotation labels specifically includes: Obtain the source branch output, first target text, and time boundary information corresponding to each first label to form a three-dimensional data entry containing speech segments, text content, and timestamps; Based on the branch identifiers corresponding to the output of the sound source branch and the matching results with the speaker set, the corresponding speaker identity information is associated with the three data entries to obtain the labeled entries with speaker identifiers. For multiple annotation entries belonging to the same sound source branch and being temporally continuous, the segments are spliced ​​and merged based on the timestamp order to form a semantically continuous speech annotation sequence; The quality of the merged speech annotation sequence is checked, and annotation segments with time conflicts, missing text, or confidence scores below the threshold are removed to obtain the target annotation sequence that meets the quality constraints. The target annotation sequences corresponding to each sound source branch are aggregated and stored to form a first speech annotation set consisting of multiple annotation sequences with speaker identifiers, text content and time information.

[0071] In this embodiment, after completing the confidence screening of the initial annotation labels and determining the first annotation label, the system further uses the first annotation label as the core information source to perform structured reorganization of its associated speech and text data to construct a first speech annotation set for training or data management. Specifically, the system first extracts the corresponding sound source branch output, the first target text, and the time boundary information from each first annotation label, and binds the speech segment, text content, and timestamp according to the index relationship to form a three-dimensional data entry, so that speech and text have a one-to-one structural expression in the time dimension. Subsequently, based on the branch number of the sound source branch output and the matching result with the speaker set, the system attaches the corresponding speaker identity to the three-dimensional data entry to generate a structured annotation entry with speaker annotation attributes. On this basis, for multiple annotation entries belonging to the same sound source branch and being temporally continuous or adjacent, the system performs segment splicing and merging processing according to the timestamp order to form a speech annotation sequence with a longer time span and semantic continuity. Next, the system performs quality checks on the merged labeled sequences, removing or correcting labeled segments with time interval conflicts, missing text content, abnormal boundaries, or confidence levels below a set threshold, ensuring that the data entering the set meets preset quality constraints. Finally, the system groups and centrally stores the target labeled sequences corresponding to each sound source branch according to the branch index and speaker identifier, thus forming a first speech labeled set composed of multiple labeled sequences with speaker identity, text content, and time information.

[0072] Through the above steps, the filtered pseudo-annotated entries can be uniformly integrated into a structured annotation set, realizing the multi-dimensional association and organization of speech, text and speaker information.

[0073] In this embodiment, the speech annotation method runs on an electronic device (e.g., Figure 1 The server shown can receive instructions or acquire data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods.

[0074] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned voice data to be labeled, the aforementioned voice data to be labeled can also be stored in a node of a blockchain.

[0075] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0076] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0077] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0078] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0079] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0080] Further reference Figure 4 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a speech annotation device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0081] like Figure 4 As shown, the speech annotation device 400 described in this embodiment is characterized by comprising: The overlap detection module 401 is used to acquire speech data to be labeled and perform speech overlap detection on the speech data to be labeled to obtain overlapping segments and non-overlapping segments in the speech data to be labeled. The speaker recognition module 402 is used to identify speakers in non-overlapping segments of the speech data to be labeled, and to construct a speaker set based on the speaker recognition results. The sound source separation module 403 is used to separate the sound sources of overlapping segments in the speech data to be labeled based on the speaker set and using a preset sound source separator to obtain several sound source branch outputs. The first speech annotation module 404 is used to perform automatic speech recognition and speech-text splicing and alignment on the outputs of the plurality of sound sources respectively, so as to obtain the first speech annotation set; In this embodiment, please refer to Figure 5The first speech annotation module 404 includes a first speech recognition submodule 441, a permutation-invariant training submodule 442, a first splicing and alignment submodule 443, a label filtering submodule 444, and a first speech annotation submodule 445, wherein: The first speech recognition submodule 441 is used to perform automatic speech recognition on several sound source branch outputs respectively, and obtain the candidate text corresponding to each sound source branch output; The permutation-invariant training submodule 442 is used to perform permutation-invariant training on several source branch outputs and several candidate texts respectively, and determine the first target text corresponding to each source branch output; The first splicing and alignment submodule 443 is used to obtain the first timestamp corresponding to each sound source branch output, and splice and align several sound source branch outputs and several first target texts based on the first timestamp to generate initial annotation labels, wherein the initial annotation labels are pseudo annotation labels; The label filtering submodule 444 is used to calculate the confidence level of the initial label and filter based on the initial label to obtain the first label; The first speech annotation submodule 445 is used to construct a first speech annotation set based on several source branch outputs, several first target texts and first annotation labels; The second speech annotation module 405 is used to automatically recognize speech and align speech and text in the non-overlapping segments of the speech data to be annotated, so as to obtain a second speech annotation set. In this embodiment, please continue to refer to Figure 5 The second speech annotation module 405 includes a second speech recognition submodule 451, a second splicing and alignment submodule 452, and a second speech annotation submodule 453, wherein: The second speech recognition submodule 451 is used to automatically recognize non-overlapping segments in the speech data to be labeled, and obtain the second target text. The second splicing and alignment submodule 452 is used to obtain the second timestamp of the non-overlapping segments in the speech data to be labeled, and splice and align the non-overlapping segments in the speech data to be labeled and the second target text based on the second timestamp to generate the second label. The second speech annotation submodule 453 is used to construct a second speech annotation set based on non-overlapping segments in the speech data to be annotated, the second target text, and the second annotation label; The speech annotation integration module 406 is used to generate a speech annotation set for the speech data to be annotated based on the first speech annotation set and the second speech annotation set.

[0082] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 6 , Figure 6 This is a basic structural block diagram of the computer device in this embodiment.

[0083] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are interconnected via a system bus. It should be noted that only the computer device 6 with memory 61, processor 62, and network interface 63 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0084] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0085] The memory 61 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 6. Of course, the memory 61 may also include both the internal storage unit and its external storage device of the computer device 6. In this embodiment, the memory 61 is typically used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions for speech annotation methods. In addition, the memory 61 can also be used to temporarily store various types of data that have been output or will be output.

[0086] In some embodiments, the processor 62 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 62 is typically used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to execute computer-readable instructions stored in the memory 61 or to process data, for example, to execute computer-readable instructions for the speech annotation method.

[0087] The network interface 63 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 6 and other electronic devices.

[0088] This application also provides an implementation method, namely, a computer device including a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the speech annotation method described above.

[0089] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the speech annotation method described above.

[0090] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0091] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0092] It should be noted that the software tools or components not belonging to this company that appear in the various embodiments of this application are merely illustrative examples and do not represent actual use.

[0093] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A speech annotation method, characterized in that, include: Acquire speech data to be labeled, and perform speech overlap detection on the speech data to be labeled to obtain overlapping segments and non-overlapping segments in the speech data to be labeled; Speaker identification is performed on non-overlapping segments in the speech data to be labeled, and a speaker set is constructed based on the speaker identification results; Based on the speaker set, a preset sound source separator is used to separate the overlapping segments in the speech data to be labeled, resulting in several sound source branch outputs; Automatic speech recognition and speech-text concatenation and alignment are performed on the outputs of the several sound sources respectively to obtain the first speech annotation set; Automatic speech recognition and speech-text splicing alignment are performed on the non-overlapping segments in the speech data to be labeled to obtain a second speech annotation set; Based on the first speech annotation set and the second speech annotation set, a speech annotation set for the speech data to be annotated is generated.

2. The speech annotation method as described in claim 1, characterized in that, The step of performing automatic speech recognition and speech-text concatenation and alignment on the outputs of the plurality of sound sources to obtain the first speech annotation set specifically includes: Automatic speech recognition is performed on the outputs of the various sound sources to obtain candidate text corresponding to each output of the sound source; Permutation invariance training is performed on the several source branch outputs and several candidate texts respectively to determine the first target text corresponding to each source branch output; Obtain the first timestamp corresponding to each of the sound source branch outputs, and concatenate and align the several sound source branch outputs and the several first target texts based on the first timestamp to generate initial annotation labels, wherein the initial annotation labels are pseudo annotation labels; Calculate the confidence level of the initial label and filter based on the initial label to obtain the first label; Based on the output of the plurality of sound sources, the plurality of first target texts, and the first annotation labels, a first speech annotation set is constructed; The step of using a preset sound source separator to separate overlapping segments in the speech data to be labeled, based on the speaker set, to obtain several sound source-directed outputs, specifically includes: Acoustic features are extracted from the overlapping segments to obtain spectral feature sequences corresponding to time frames; Based on the speaker feature vectors in the speaker set, the spectral feature sequence is subjected to feature enhancement processing to obtain enhanced speech features; The enhanced speech features are input into a preset sound source separator, and the sound source separator is used to separate and calculate the enhanced speech features to obtain several initial sound source branch outputs corresponding to the overlapping segments and the corresponding time masks. The similarity between the initial sound source branch outputs of several paths and the feature vectors of each speaker is calculated respectively to obtain the candidate speaker set corresponding to each initial sound source branch output; The initial sound source branch outputs of several channels are subjected to identity constraint allocation and reconstruction processing. For the initial sound source branch outputs that do not meet the identity consistency, mask correction and feature re-estimation are performed to obtain the several sound source branch outputs.

3. The speech annotation method as described in claim 2, characterized in that, The step of performing permutation-invariant training on the plurality of sound source branch outputs and the plurality of candidate texts respectively, and determining the first target text corresponding to each of the sound source branch outputs, specifically includes: The outputs of the various sound sources are input into a preset automatic speech recognition model to obtain the acoustic posterior probability sequence corresponding to each sound source output and the recognition confidence information of the candidate text. Based on the candidate text, the acoustic posterior probability sequence is aligned with the candidate text sequence, and the matching loss value between each sound source branch output and each candidate text is calculated according to the preset loss function. Using the matching loss value as the evaluation index, enumerate all permutations and combinations between the output of the plurality of sound sources and the plurality of candidate texts, and calculate the combination loss corresponding to each permutation and combination; The target permutation and combination relationship corresponding to the minimum permutation and combination loss is determined from the combined loss, and the target permutation and combination relationship is used as the optimal matching result of permutation-invariant training; Based on the optimal matching result, the candidate text corresponding to each sound source branch output is determined, and the candidate text is determined as the first target text corresponding to the sound source branch output.

4. The speech annotation method as described in claim 3, characterized in that, The step of obtaining the first timestamp corresponding to each of the sound source branch outputs, and concatenating and aligning the plurality of sound source branch outputs and the plurality of first target texts based on the first timestamp to generate initial annotation labels specifically includes: Based on the alignment result calculated, the start and end times of each speech segment in the output of each sound source branch are extracted to generate a first timestamp sequence corresponding to the first target text. Based on the first timestamp sequence, the outputs of the several sound sources are divided into time boundaries on a unified time axis to obtain a set of split speech segments in units of time segments; The first target text is segmented into text segments according to the first timestamp sequence, and each text segment is concatenated with the corresponding split speech segment at the time level to form a mapping relationship between speech and text segments; In the case where multiple sound source outputs correspond to the same time interval, the multiple speech segments and text segments within the same time interval are aligned, merged and spliced ​​based on the time overlap ratio and text coverage. Based on the mapping relationship of the processed speech-text segments, a structured annotation entry containing time boundary information, sound source branch number and first target text content is generated, and the structured annotation entry is used as the initial annotation label.

5. The speech annotation method as described in claim 4, characterized in that, The step of aligning, reconciling, and splicing multiple speech segments and text segments within the same time interval based on the time overlap ratio and text coverage when multiple sound source outputs correspond to the same time interval specifically includes: Obtain the speech segments and time intervals corresponding to the output of each sound source in the same time interval, and calculate the time overlap ratio of each speech segment based on the intersection duration of the speech segments and the time intervals. Each text segment within the time interval is segmented and mapped according to the first timestamp, and the corresponding text coverage is calculated based on the effective coverage length of the text segment within the time interval. Based on the time overlap ratio and the text coverage, multiple speech segments and text segments within the same time interval are jointly scored, and the speech segments and text segments are prioritized according to the joint score. Based on the priority ranking results, the dominant speech segment and its corresponding dominant text segment are determined, and the non-dominant speech segments are subjected to boundary trimming compensation processing. After completing the boundary clipping and compensation process, the processed speech segments and corresponding text segments within the same time interval are merged and spliced ​​to generate a merged and aligned result that is time-continuous, text-consistent, and has clear branching identifiers. The merged and aligned result is then used as the initial annotation sub-label corresponding to the time interval.

6. The speech annotation method as described in claim 2, characterized in that, The step of calculating the confidence level of the initial label and filtering based on the initial label to obtain the first label specifically includes: For each initial label, obtain the recognition confidence information, speech quality evaluation index and text matching score of the corresponding sound source branch output, and construct the confidence feature vector corresponding to the initial label; Based on a pre-set reliability assessment model, the confidence feature vector is comprehensively calculated to obtain the segment-level confidence score of the initial labeled label. Based on the segment-level confidence score, the initial annotation labels are compared with different generated results for the same speech segment across samples for cross-sample consistency, and a multi-result consistency index is calculated. Based on the segment-level confidence score and the multi-result consistency index, the reliability of the initial label is judged, the target initial label that meets the preset screening conditions is determined, and the target initial label is determined as the first label.

7. The speech annotation method as described in claim 2, characterized in that, The step of constructing a first speech annotation set based on the outputs of the plurality of sound sources, the plurality of first target texts, and the first annotation labels specifically includes: Obtain the source branch output, first target text, and time boundary information corresponding to each of the first annotation labels to form a three-dimensional data entry containing speech segments, text content, and timestamps; Based on the branch identifier corresponding to the output of the sound source branch and the matching result with the speaker set, the speaker identity information corresponding to the three data entries is associated with the three data entries to obtain the labeled entries with speaker identifiers. For multiple annotation entries that belong to the same sound source branch and are temporally continuous, the segments are spliced ​​and merged based on the timestamp order to form a semantically continuous speech annotation sequence. The quality of the merged speech annotation sequence is checked, and annotation segments with time conflicts, missing text, or confidence scores below the threshold are removed to obtain the target annotation sequence that meets the quality constraints. The target annotation sequences corresponding to each sound source branch are aggregated and stored to form a first speech annotation set consisting of multiple annotation sequences with speaker identifiers, text content and time information.

8. A voice annotation device, characterized in that, include: An overlap detection module is used to acquire speech data to be labeled and to perform speech overlap detection on the speech data to be labeled to obtain overlapping segments and non-overlapping segments in the speech data to be labeled. The speaker recognition module is used to identify speakers in non-overlapping segments of the speech data to be labeled, and to construct a speaker set based on the speaker recognition results. The sound source separation module is used to separate the overlapping segments in the speech data to be labeled using a preset sound source separator based on the speaker set, and to obtain several sound source branch outputs. The first speech annotation module is used to perform automatic speech recognition and speech-text splicing and alignment on the outputs of the plurality of sound sources respectively, so as to obtain the first speech annotation set; The second speech annotation module is used to automatically recognize speech and align speech and text in the non-overlapping segments of the speech data to be annotated, so as to obtain a second speech annotation set. The speech annotation integration module is used to generate a speech annotation set for the speech data to be annotated based on the first speech annotation set and the second speech annotation set.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the speech annotation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech annotation method as described in any one of claims 1 to 7.