A Real-Time Voice Interaction and Incremental Backfilling Method and System for Interrogation Transcripts

CN122575375APending Publication Date: 2026-08-14LISHUI CITY PUBLIC SECURITY BUREAU +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本发明提供了一种面向询问笔录的实时语音交互与增量回填方法及系统,用于解决询问笔录场景中答案文本与原始语音片段无法精确绑定导致笔录无法复核、缺乏对说话人角色的自动归属与持久化记录、以及新角色注册后已生成笔录无法关联新角色信息的技术问题

Benefits of technology

第一,本发明通过在笔录生成过程中将句末稳定转录段对应的语音片段与当前目标问题标识和说话人角色标签绑定为四元组,并基于答案补丁建立语音回溯索引项,实现了笔录中每句答案文本与原始语音片段的精确绑定。与现有技术中非结构化的角色发言区间无法与预设问题槽位关联且不支持答案溯源的方案相比,使笔录中的每一个答案字段均可通过语音回放控件回溯至原始语音片段,消除了转写错误或断章取义无法在事后审查中发现的隐患。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575375A_ABST
    Figure CN122575375A_ABST
Patent Text Reader

Abstract

This invention discloses a real-time voice interaction and incremental backfilling method and system for interrogation transcripts, relating to the field of voice interaction technology. The method includes: acquiring a real-time voice stream and segmenting it into several voice segments; extracting voiceprint feature vectors, performing similarity comparisons, and determining speaker role labels; binding the voice segment corresponding to the stable transcription segment at the end of the sentence with the current target question identifier and the speaker role label to generate a four-tuple; establishing a voice backfilling index; and generating a voice playback control for each incremental backfilling answer field, which, when triggered by the user, plays back the corresponding voice segment according to the index. This invention solves the technical problems in interrogation transcript scenarios where the answer text and the original voice segment cannot be accurately bound, leading to the inability to verify the transcript; the lack of automatic attribution and persistent recording of speaker roles; and the inability to associate newly generated transcripts with new role information after new role registration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice interaction technology, specifically to a real-time voice interaction and incremental backfilling method and system for inquiry transcripts. Background Technology

[0002] Existing voice interaction and speaker annotation methods, such as CN113421563A, achieve speaker annotation in meeting scenarios through voiceprint separation and role resetting. However, their output is an unstructured role speaking range, which cannot be associated with preset question slots and does not support answer tracing. Another example is CN121393423A, which uses wake words to extract voiceprints for voice enhancement to improve interaction accuracy. However, its voiceprints are temporarily bound at the session level and become invalid when the interaction ends, without involving the persistent storage and retrieval of structured transcripts.

[0003] In interrogation record scenarios, existing methods generally have the following shortcomings: they cannot accurately bind the answer text to the original audio segment, resulting in the record being unverifiable; they lack automatic attribution and persistent recording of the speaker's role; and when a new role is subsequently registered, the generated record cannot be associated with the new role's information.

[0004] Therefore, there is an urgent need for a method that can achieve answer and voice tracing and automatic role assignment during the transcript generation process. Summary of the Invention

[0005] This invention provides a real-time voice interaction and incremental backfilling method and system for interrogation transcripts, which solves the technical problems in interrogation transcript scenarios, such as the inability to accurately bind the answer text and the original voice segment, resulting in the transcript being unverifiable; the lack of automatic attribution and persistent recording of speaker roles; and the inability to associate the generated transcript with the new role information after the new role is registered.

[0006] In a first aspect, the present invention provides a real-time voice interaction and incremental backfilling method for interrogation transcripts, the method comprising: In the process of generating transcripts based on closed-loop voice interaction and incremental backfilling, real-time voice streams are acquired and segmented into several voice segments, and conversation timing confidence parameters are configured for the voice segments to verify the compliance of the transcript question-and-answer timing. Extract the voiceprint feature vector of each speech segment, compare it with the centroid of at least one registered role, and determine the speaker role label corresponding to the speech segment; Obtain the stable transcript segment detected during the transcript generation process, bind the speech segment corresponding to the stable transcript segment to the current target question identifier and the speaker role label, and generate a quadruple containing a session identifier, question identifier, role label and speech segment reference; When the transcript generation process generates an answer patch based on the stable transcript segment at the end of the sentence, a speech backtracking index item is established based on the question identifier and answer text of the answer patch, combined with the speech segment reference in the quadruple and the transcript conversation temporal confidence parameter. In the transcript display interface, a voice playback control is generated for each incrementally filled answer field. When the user triggers the voice playback control, the corresponding voice segment is played back according to the voice backtracking index item and the transcript session temporal confidence parameter.

[0007] Secondly, the present invention also provides a real-time voice interaction and incremental backfilling system for interrogation transcripts, the system comprising: The real-time voice segmentation module is used to acquire real-time voice streams and segment them into several voice segments during the transcript generation process based on closed-loop voice interaction and incremental backfilling, and to configure conversation timing confidence parameters for the voice segments to verify the compliance of the transcript question-and-answer timing. The voiceprint role determination module extracts the voiceprint feature vector of each speech segment, compares it with the centroid of at least one registered role, and determines the speaker role label corresponding to the speech segment. The quadruple binding module obtains the stable transcript segment detected during the transcript generation process, binds the speech segment corresponding to the stable transcript segment to the current target question identifier and the speaker role label, and generates a quadruple containing a session identifier, question identifier, role label and speech segment reference. The speech retrospective index module, when generating an answer patch based on the stable transcript segment at the end of the sentence during the transcript generation process, establishes a speech retrospective index item based on the question identifier and answer text of the answer patch, combined with the speech segment reference in the quadruple and the transcript conversation temporal confidence parameter; The voice playback control module generates a voice playback control for each incrementally filled answer field in the transcript display interface. When the user triggers the voice playback control, the corresponding voice segment is played back according to the voice backtracking index item and the transcript session temporal confidence parameter.

[0008] One or more technical solutions provided in this invention have at least the following technical effects or advantages: First, this invention achieves precise binding between each answer text in the transcript and the original speech segment by binding the speech segment corresponding to the stable transcription segment at the end of the sentence with the current target question identifier and speaker role label as a four-tuple during the transcript generation process, and establishing a speech backtracking index based on the answer patch. Compared with the existing technology where unstructured role speech intervals cannot be associated with preset question slots and do not support answer tracing, this invention allows each answer field in the transcript to be traced back to the original speech segment through the speech playback control, eliminating the potential risks of transcription errors or misinterpretations that cannot be detected in post-review.

[0009] Secondly, this invention achieves automatic determination and persistent recording of the roles of the questioner and the questionee by performing role registration at the beginning of the inquiry and continuously comparing the similarity of voice segments with the centroids of registered roles and automatically assigning speaker role tags thereafter. Compared with the existing technology where voiceprints are temporarily bound at the session level and expire when the interaction ends, this invention allows the role information of the same speaker to be reused in different sessions, reducing the operational burden of re-registering roles for each inquiry.

[0010] Third, this invention provides a manual review interface to receive manually corrected role types, adds the corrected voiceprint feature vector to the corresponding role's voiceprint database, and recalculates the role's centroid. After accumulating a preset number of corrections, it automatically and adaptively adjusts the similarity threshold for the corresponding role. Compared to existing technologies where existing records cannot be associated with new role information when registering new roles, this invention achieves incremental backfilling of new role information and dynamic optimization of the judgment threshold, ensuring that the accuracy of role attribution continuously improves with increasing usage. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating a real-time voice interaction and incremental backfilling method for interrogation transcripts provided in an embodiment of the present invention. Figure 2 This is a logic diagram of a real-time voice interaction and incremental backfilling method for interrogation transcripts provided in an embodiment of the present invention; Figure 3 This is an example data table diagram of the voice backtracking index item of a real-time voice interaction and incremental backfilling method for interrogation transcripts provided in an embodiment of the present invention; Figure 4This is a schematic diagram of a real-time voice interaction and incremental backfilling structure for interrogation transcripts provided in an embodiment of the present invention; The diagram shows: Real-time voice segmentation module 11, voiceprint role determination module 12, quadruple binding module 13, voice backtracking index module 14, and voice playback control module 15. Detailed Implementation

[0013] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0014] Example 1, as Figure 1 As shown, this invention provides a flowchart illustrating a real-time voice interaction and incremental backfilling method for interrogation transcripts; as... Figure 2 As shown, this invention provides a scheme logic diagram for a real-time voice interaction and incremental backfilling method for interrogation transcripts. The method includes: S100: In the process of generating transcripts based on closed-loop voice interaction and incremental backfilling, real-time voice streams are acquired and segmented into several voice segments, and conversation timing confidence parameters are configured for the voice segments to verify the compliance of the transcript question-and-answer timing. In interrogation recording scenarios, the interrogator and the interrogated person speak alternately in the same space, with their voices appearing alternately on the timeline. The recording system needs to segment the real-time acquired mixed speech stream into several speech segments. Each speech segment should completely contain a continuous speech by a single speaker to facilitate subsequent voiceprint feature extraction and speaker role determination. If the speech segmentation granularity is too coarse, the speech of two speakers may be mixed in one segment, resulting in mixed voiceprint features and making it impossible to accurately determine the speaker role; if the segmentation granularity is too fine, a complete sentence is split into multiple segments, which increases the fragmentation of subsequent processing.

[0015] Step S100 provided in this embodiment of the invention includes: Speech streams from different directions are acquired through at least two microphone channels, and the speech streams from the direction of the questioner and the direction of the questioned person are separated based on sound source localization.

[0016] The specific implementation method is as follows: At least two microphones are deployed in the inquiry area, one facing the inquirer's seat and the other facing the inquiree's seat. The microphone spacing is set according to the actual layout of the inquiry area, for example, 1 to 2 meters. The two microphone channels synchronously acquire audio signals from their respective directions, with a sampling rate of, for example, 16kHz and a sampling precision of 16bit. Since the inquirer and the inquiree are located at different positions on the microphone array, there is a difference in the arrival time of the two speakers' sound sources to the two microphones, i.e., the arrival time difference. By calculating the cross-correlation function between the two microphone channels, the time delay value corresponding to the peak value of the cross-correlation function is extracted, which is the arrival time difference between the two channels. Based on the arrival time difference and the microphone spacing, the azimuth angle of the sound source is estimated. The sound source with its azimuth angle facing the inquirer's seat is identified as the inquirer's speech, and the sound source with its azimuth angle facing the inquiree's seat is identified as the inquiree's speech, generating separate speech streams for the inquirer and inquiree directions.

[0017] During sound source localization, when the sound source azimuth angle is detected within the preset range of the interrogator's azimuth angle, the speech signal in the current time period is marked as the speech stream in the interrogator's direction; when the sound source azimuth angle is detected within the preset range of the questioner's azimuth angle, the speech signal in the current time period is marked as the speech stream in the questioner's direction; when the sound source azimuth angle is outside both ranges, the speech signal in the current time period is marked as ambient noise. Through continuous tracking of sound source localization, when the sound source azimuth angle is detected to switch from the interrogator's direction to the questioner's direction or vice versa, the speech stream is segmented at the switching moment, and each segmented speech segment contains only the continuous speech signal of a single speaker.

[0018] Step S100 provided in this embodiment of the invention further includes: Speech segments are segmented based on silence boundaries and semantic boundaries by using a single-channel microphone combined with speech activity detection and speaker change detection.

[0019] The specific implementation method is as follows: In deployment scenarios with only a single-channel microphone, silent segments are first eliminated through voice activity detection. Voice activity detection distinguishes between speech segments, silent segments, and noise segments by analyzing the energy, zero-crossing rate, or spectral characteristics of the audio signal. Specifically, the acquired continuous audio signal is processed frame by frame, with each frame lasting, for example, 20 milliseconds, and the frame shifted by 10 milliseconds. Short-time energy and zero-crossing rate are calculated for each frame signal. When the short-time energy exceeds a preset energy threshold and the zero-crossing rate exceeds a preset zero-crossing rate threshold, the frame is determined to be a speech frame; when the short-time energy of multiple consecutive frames is lower than the energy threshold, it is determined to be a silent segment. Continuous speech frames are merged into speech segments, and silent segments are used as segmentation boundaries to initially divide the audio stream into several speech segments.

[0020] Then, speaker change detection is performed on the segmented speech segments. Speaker change detection determines whether a speech segment contains the voices of two different speakers; if so, further segmentation is needed at the change point. For each speech segment, acoustic features are extracted using Mel-frequency cepstral coefficients (MFCCs). MFCCs are widely used acoustic features in speech recognition and speaker recognition, and their extraction process simulates the human ear's perception of different frequencies. The human ear has a higher ability to distinguish low-frequency sounds than high-frequency sounds. MFCCs map linear frequencies to the Mel-frequency scale by using denser filters in the low-frequency region and sparser filters in the high-frequency region, making the feature representation closer to the auditory perception of the human ear. The specific extraction steps are: performing a Fourier transform on each frame of speech signal to obtain the spectrum; passing the spectrum through a set of triangular filters distributed according to the Mel-frequency scale to obtain the logarithmic energy value of each filter; performing a discrete cosine transform on the logarithmic energy value of each filter, and taking the first few coefficients as the MFCC feature vector for that frame.

[0021] Within a speech segment, a sliding window of a preset length is used, for example, 0.5 seconds, with a sliding step of 0.1 seconds. Within each sliding window, the mean vector of the Mel-frequency cepstral coefficient eigenvectors of all speech frames within that window is calculated, serving as a statistical representation of the acoustic features within that window. This mean vector integrates acoustic feature information from multiple frames within the window, making it more stable than single-frame features and smoothing out abnormal fluctuations in individual frames.

[0022] For two adjacent windows, calculate the cosine distance between their mean vectors. Cosine distance is a method to measure the difference between two vectors in orientation space. It is calculated as follows: calculate the inner product of the two vectors, divide by the product of their magnitudes, and the quotient is the cosine similarity between the two vectors; subtract the cosine similarity from 1, and the difference is the cosine distance. The cosine distance ranges from 0 to 2. A smaller value indicates that the two vectors are closer in direction, meaning the acoustic features of the two windows are more similar; a larger value indicates a greater difference in direction, meaning the acoustic features have changed significantly. The characteristic of cosine distance is that it is insensitive to the absolute magnitude of the vectors, focusing only on the difference in direction, thus exhibiting strong robustness to changes in the loudness of the speaker's voice. Different speakers have different vocal tract shapes, vocalization methods, and speech rate habits, resulting in different distribution directions of their speech signals in the acoustic feature space. When the speech from the two windows comes from the same speaker, the direction of the mean vector remains relatively consistent in the acoustic feature space, and the cosine distance is small. When the speech from the two windows comes from different speakers, the acoustic feature space direction of the first speaker deviates significantly from that of the second speaker, the direction of the mean vector shifts in the feature space, and the cosine distance increases significantly.

[0023] When the cosine distance between adjacent windows exceeds a preset change threshold, that location is determined to be a speaker switching point. The change threshold is set to distinguish between the natural fluctuations in the same speaker's voice and the differences in acoustic features between different speakers. The cosine distance of the same speaker's voice at different times will fluctuate slightly due to changes in intonation and speech rate, but usually does not exceed the range of 0.3 to 0.4. The acoustic feature differences between different speakers are significant, and the cosine distance usually exceeds 0.5 to 0.7. The change threshold is set to a reasonable value between these two, for example, 0.5. When the cosine distance exceeds 0.5, it indicates that the change in the direction of acoustic features between windows has exceeded the normal fluctuation range of the same speaker, and there is a high probability that a speaker switch has occurred. At the beginning of that window, the speech segment is divided into two independent segments. To avoid misjudgment caused by brief voice changes or environmental noise interference, a continuous multi-point confirmation strategy can be adopted, for example, triggering segmentation only when the cosine distance between two consecutive adjacent windows exceeds the change threshold.

[0024] After segmentation by speech activity detection and speaker change detection, speech segments are further refined by incorporating semantic boundaries. Semantic boundaries refer to the natural pauses in a sentence, typically corresponding to a drop in intonation at the end of the sentence and a longer pause. By detecting the fundamental frequency change trend and pause duration within the speech segment, a semantic boundary is identified when the fundamental frequency drops to the end-of-sentence characteristic range and is accompanied by a silent segment exceeding a preset end-of-sentence pause threshold. If the semantic boundary falls within the speech segment and is close to its end, it is used as the final segmentation point, and the portion following the semantic boundary is moved to the next speech segment.

[0025] While segmenting audio clips, a conversation timing confidence parameter is configured for each clip. This parameter is used to verify the compliance of the question-and-answer sequence in the interrogation transcript. In an interrogation transcript scenario, the speaking order of the interrogator and the interrogated party should adhere to the "one question, one answer" timing standard, with the interrogator asking the question first and the interrogated party answering later. Overlapping or out-of-order question-and-answer sessions should be flagged during transcript review. The timing confidence is calculated by determining whether the alternating speaking rule of "interrogator and interrogated party" is met based on the speaker role labels of the current and previous audio clips. If the current voice segment is labeled as the respondent and the previous voice segment was labeled as the questioner, the timing is normal, and a higher timing confidence value is used. If the current voice segment has the same role label as the previous voice segment (i.e., the two consecutive segments belong to the same role), or if the current voice segment is labeled as the questioner while the previous voice segment was labeled as the respondent, it indicates possible interruption, interruption, or timing discrepancies, and a lower timing confidence value is used. The reduced timing confidence is used to alert reviewers to any abnormalities in the question-and-answer timing of the voice segment during subsequent voice recall index creation and playback.

[0026] The following technical effects were achieved through this step: First, in a dual-microphone deployment scenario, this step separates speech streams from different directions by sound source localization. It utilizes the natural isolation between the physical locations of the questioner and the questioned to achieve speech separation. Compared with existing technologies that only use voiceprint features for speaker separation, this method has a higher accuracy rate when the sound source localization accuracy meets the requirements and the directional information is used directly as the segmentation basis.

[0027] Second, in a single-microphone deployment scenario, this step employs a three-tiered progressive processing approach: speech activity detection to remove silent segments, speaker change detection to locate voiceprint transition points, and semantic boundary refinement to refine segmentation points. This process refines the segmentation of speech segments layer by layer, from energy boundaries to voiceprint boundaries and then to semantic boundaries, thereby improving the segmentation quality of speech segments in terms of speaker independence and semantic integrity.

[0028] S200: Extract the voiceprint feature vector of each speech segment, compare it with the centroid of at least one registered role, and determine the speaker role label corresponding to the speech segment; The above steps segment the real-time audio stream into several speaker-independent, semantically complete audio segments. Each audio segment contains a continuous audio signal emitted by a single speaker. In an interrogation recording scenario, the recording system needs to automatically determine the speaker identity of each audio segment, i.e., whether the segment was emitted by the interrogator or the interrogated. Accurate speaker role determination is the key foundation for subsequently binding the answer text with the speaker identity and realizing record verification. If the role determination is incorrect, the question-and-answer attribution relationship recorded in the record will be confused, leading to a broken chain of evidence.

[0029] Step S200 provided in this embodiment of the invention includes: Before extracting the speaker feature vector for each speech segment, preprocessing of the speech segments is also included: The preprocessing includes resampling, amplitude normalization, and removal of beginning and end silences; The voiceprint feature vector is extracted through a deep voiceprint model, which includes a speaker embedding network based on a residual network.

[0030] The specific implementation method is as follows: First, each segmented speech piece undergoes preprocessing. The first preprocessing operation is resampling, which resamples speech pieces from different sampling rate sources to a preset standard sampling rate, such as 16kHz. The purpose of resampling is to unify the temporal resolution of all speech pieces, ensuring that subsequent speaker feature extraction operates under a unified sampling rate parameter and avoiding feature bias caused by inconsistent sampling rates.

[0031] The second preprocessing operation is amplitude normalization. The maximum absolute value of the amplitude at all sampling points within the speech segment is taken as the normalization benchmark. The amplitude of each sampling point is then divided by this benchmark, normalizing the amplitude range to between -1 and 1. The purpose of amplitude normalization is to eliminate the influence of speaker volume, distance from the microphone, and microphone gain differences on voiceprint features, allowing the voiceprint model to focus on differences in vocal tract shape and vocalization method, rather than differences in volume.

[0032] The third preprocessing step is to remove beginning and ending silences. Silence detection is performed at both ends of the speech segment, and the silent segments from the beginning to the first speech frame and from the end to the last speech frame are removed. Beginning and ending silences do not contain any speaker's vocal information; removing them reduces the computational load for subsequent voiceprint feature extraction and avoids interference with voiceprint feature aggregation.

[0033] Then, the preprocessed speech segment is input into a residual network-based speaker embedding network to extract fixed-dimensional voiceprint feature vectors. The residual network-based speaker embedding network is a deep neural network model whose core structure consists of multiple stacked residual blocks. The residual block introduces a direct connection from input to output on top of the traditional convolutional layer, adding the input features element-wise with the convolutionally transformed features before outputting the result. This direct connection allows gradients to be directly propagated to shallower layers during backpropagation, effectively alleviating the vanishing gradient problem common in deep network training. This allows the network to be stacked deeper to extract more abstract and discriminative voiceprint features.

[0034] Each residual block contains convolutional layers, batch normalization layers, and activation function layers. The convolutional layers extract local time-frequency features, the batch normalization layers normalize the output of the convolutional layers to stabilize the training process, and the activation function introduces a non-linear transformation. Multiple residual blocks are connected sequentially from low to high layers. Low-layer residual blocks extract short-time spectral features of the speech signal, such as formant structure and fundamental frequency features; middle-layer residual blocks extract phoneme-level articulation features, such as vowel and consonant transition patterns; and high-layer residual blocks extract global vocal tract shape and vocalization habit features related to speaker identity. The pooling layer at the end of the network statistically pools the frame-level feature vector output from the last residual block in the time dimension. Before statistical pooling, the speech tags obtained from the preprocessing stage's speech activity detection are used to remove frames marked as silent or non-speech. Only frames with speech tags are calculated for mean and standard deviation. The mean and standard deviation are concatenated into a fixed-dimensional vector, which is the speaker signature feature vector for that speech segment. The dimension of the voiceprint feature vector is preset by the network structure, such as 128 or 256 dimensions, and is independent of the duration of the speech segment, providing a unified input dimension for comparing the centroid similarity of characters.

[0035] Step S200 provided in this embodiment of the invention further includes: At the beginning of the inquiry, role registration is performed: if a pre-registered voiceprint database of the inquirer exists, the matching voice segment is marked as the inquirer role; otherwise, the voiceprint feature vector of the first long voice segment of the inquired object that meets the preset criteria is set as the centroid of the inquired object role. Subsequent audio segments are assigned to corresponding roles by calculating their similarity to the centroids of registered roles. Segments with similarity greater than a first preset threshold are assigned to the corresponding roles, while segments with similarity less than the first preset threshold are created as new roles and marked as awaiting manual confirmation.

[0036] The specific implementation method is as follows: First, role registration is performed at the beginning of the query phase. The purpose of role registration is to establish a role centroid benchmark for subsequent similarity comparison. Role registration is performed in the following priority order.

[0037] The first priority is matching with the pre-registered questioner's voiceprint database. The system checks if a pre-registered questioner's voiceprint database exists. This database is pre-entered during system initialization and contains voiceprint feature vectors of several commonly used questioners. If such a database exists, after the questioning begins, the segmented speech fragments are compared one by one with the voiceprint feature vectors of each questioner in the database. The comparison method involves calculating the cosine similarity between the speech fragment's voiceprint feature vector and each voiceprint feature vector in the database, and taking the pre-registered questioner corresponding to the highest similarity as the matching result. If this highest similarity exceeds a preset matching threshold, the corresponding speech fragment is labeled as the questioner's role, and the matched voiceprint feature vector is set as the centroid of the questioner's role.

[0038] The second priority is to use the azimuth angle of the sound source localization for auxiliary determination. If there is no pre-registered interrogator voiceprint database, or if no voice segment reaches the matching threshold after matching, and dual-channel azimuth angle information based on sound source localization is available in the current deployment environment, then the role is determined by combining physical orientation. Voice segments with their azimuth angle in the direction of the interrogator's seat are assigned to the interrogator role, and the voiceprint feature vector of that voice segment is set as the centroid of the interrogator role. At the same time, voice segments with their azimuth angle in the direction of the person being questioned's seat are assigned to the person being questioned role.

[0039] The third priority is the analysis of a temporary benchmark based on the first speech segment. If there is no pre-registered interrogator voiceprint database and sound source localization information is unavailable, the voiceprint feature vector of the first speech segment is set as a temporary benchmark at the beginning of the interrogation, and its similarity with the voiceprint feature vectors of subsequent speech segments is checked. If the similarity of multiple consecutive subsequent segments exceeds the role identity threshold, the first speech segment is confirmed as the interrogator role, and its voiceprint feature vector is set as the centroid of the interrogator role. The role identity threshold is determined based on the statistical distribution of similarity among different speech segments of the same speaker using the voiceprint model.

[0040] Regardless of the priority method used to determine the centroid of the interrogator's role, the voiceprint feature vector of the first long speech segment of the interviewee that meets the preset criteria is set as the centroid of the interviewee's role. The preset criteria include: the speech segment duration exceeds the minimum effective duration threshold, for example, 2 seconds, to ensure sufficient information in the voiceprint feature vector; and the signal-to-noise ratio (SNR) of the speech segment exceeds the minimum SNR threshold to ensure the quality of voiceprint feature extraction. After registering the centroid of the interviewee's role, the initial value of the centroid for subsequent speech segments identified as belonging to the interviewee is not fixed but iteratively updated as speech segments of the interviewee accumulate. The iterative update method is as follows: when a new speech segment is identified as belonging to the interviewee's role, the voiceprint feature vector of that segment is weighted and averaged with the current centroid of the interviewee's role according to preset update weights. For example, the update weights are 0.8 (the original centroid weight) and 0.2 (the new segment weight). The updated weighted average vector is used as the new centroid of the interviewee's role. As the conversation progresses, the centroid of the interviewee's role gradually converges to the more accurate voiceprint feature center of the speaker. The update weights are set to 0.8 for the original centroid and 0.2 for the new segment. This ratio assigns a lower weight to the contribution of the new segment to the centroid, smoothing out voiceprint feature shifts caused by temporary changes in voice or environmental fluctuations in a single speech segment. When the signal-to-noise ratio (SNR) of the current speech segment is detected to be lower than the preset SNR lower limit, or when the similarity between the current speech segment and the character's current centroid drops abnormally by more than the preset drop range, it indicates that the speech segment may be subject to strong noise interference or a drastic change in the speaker's voice. In this case, the centroid update for the segment is skipped to avoid misleading drift of the character's centroid caused by the abnormal segment.

[0041] Then, a similarity comparison is performed on each subsequent speech segment. The purpose of the similarity comparison is to assign the speech segment to the most matching registered role or mark it as a new role. The cosine similarity between the speech segment's voiceprint feature vector and the centroid of each registered role is calculated. The cosine similarity is calculated by taking the inner product of the voiceprint feature vector and the role's centroid vector, dividing it by the product of the magnitudes of the two vectors, and the quotient is the cosine similarity. The cosine similarity value ranges from -1 to 1; a larger value indicates that the two vectors are closer in direction in the voiceprint feature space, meaning that the voiceprint features of the speech segment are more similar to the corresponding role's voiceprint baseline.

[0042] The calculated cosine similarity is compared with a first preset threshold. The first preset threshold is a threshold value for determining whether a speech segment can be assigned to a registered role, determined statistically based on the similarity distribution of the voiceprint model among different speakers and the similarity distribution of the same speaker. When the cosine similarity corresponding to a registered role is greater than the first preset threshold, the speech segment is assigned to the corresponding role with the highest similarity, and the role centroid of that role is updated. When the similarity of a speech segment with the centroids of all registered roles is lower than the first preset threshold, it indicates that the speech segment may come from a new role that has not yet been registered, such as a third party newly added during the inquiry process. In this case, a new role is automatically created, the voiceprint feature vector of the speech segment is set as the temporary role centroid of the new role, and the new role is marked as pending manual confirmation. The operator is then prompted to confirm in the subsequent manual review interface.

[0043] After assigning roles to audio segments, the speaker role labels obtained will be associated with those segments. Speaker role labels include three types: questioner role labels, questioned person role labels, and new role labels awaiting manual confirmation.

[0044] The following technical effects were achieved through this step: First, this step extracts voiceprint feature vectors through a speaker embedding network based on residual networks. It utilizes the direct connection channels in the residual block stacking structure to alleviate the gradient vanishing problem, enabling the network to extract deeper voiceprint features.

[0045] Second, this step involves performing role registration at the beginning of the inquiry process. The interrogator's role centroid is determined according to the priority order of the pre-registered interrogator's voiceprint database, the sound source localization azimuth, and the analysis of the first temporary benchmark of the first speech segment. The first long speech segment of the respondent that meets the criteria is set as the respondent's role centroid, providing a basis for role determination in subsequent speech segments. The iterative update mechanism of the role centroid allows the voiceprint representation of the same speaker to gradually converge and become more accurate during the conversation, while providing a persistent voiceprint benchmark for the reuse of role information across conversations.

[0046] S300: Obtain the stable transcript segment detected during the transcript generation process, bind the speech segment corresponding to the stable transcript segment to the current target question identifier and the speaker role label, and generate a quadruple containing a session identifier, question identifier, role label and speech segment reference; In interrogation recording scenarios, the interrogator asks questions according to a pre-set list, and the interrogated person answers each question. The recording system uses speech recognition to transcribe the answers into text in real time, and generates answer patches incrementally to fill in the recording template after the text stabilizes. However, in streaming speech recognition, the recognition results change continuously as the speech progresses, with frequent changes in intermediate results. If intermediate results are used as the answer text for filling in, the filled content may be overwritten by subsequent recognition results, leading to recording errors. Therefore, the recording system needs to accurately detect the moment when the interrogated person has finished answering and the recognition result no longer changes, i.e., the stable transcription segment at the end of the sentence.

[0047] Step S300 provided in this embodiment of the invention includes: Speech activity detection is continuously performed on the real-time speech stream. When a speech end point is detected, the current transcription result is marked as a candidate stable segment. The system receives a stable state identifier returned by the streaming speech recognition module. When the stable state identifier indicates that the current recognition result no longer changes, the current transcription result is marked as a candidate stable segment. Determine whether the silence duration of the real-time speech stream reaches a preset silence threshold, and at the same time determine whether the current transcription result constitutes a complete semantic unit. When both conditions are met, mark the current transcription result as a candidate stable segment. The system receives the final confirmation flag from the dual-channel recognition results. When the final confirmation flag is active, the current transcription result is marked as a candidate stable segment. The dual-channel recognition results are obtained by comparing the same speech stream after being recognized by two parallel speech recognition engines. When candidate stable segment markers are generated by at least one of the above methods, the corresponding candidate stable segment is confirmed as the sentence-end stable transcription segment, and subsequent steps are triggered.

[0048] The specific implementation method is as follows: This step provides four methods for detecting stable transcripts at the end of sentences. The four methods run in parallel, and any method that is triggered first can be used as the basis for generating candidate stable segment markers.

[0049] Speech activity detection is based on continuous performance monitoring of the real-time speech stream. Speech activity detection distinguishes speech segments from silence segments by analyzing the energy, zero-crossing rate, or spectral characteristics of the audio signal. When a speech activity transitions from a speaking state to a silent state, it is marked as a speech end point. At the speech end point, the latest transcription result output by the streaming speech recognition module before the current moment is marked as a candidate stable segment. This method relies on the high sensitivity of speech activity detection to silence boundaries and is suitable for scenarios where there is a clear pause at the end of a speaker's turn. It is the fastest response method among the four approaches.

[0050] The detection of stable state indicators based on the streaming speech recognition module. During continuous audio input reception, the streaming speech recognition module outputs two types of recognition results: intermediate results and final results. Intermediate results change continuously as the audio input progresses, representing recognition hypotheses that are not yet finalized; the final result represents the recognized text that the system has confirmed will no longer change. The streaming speech recognition module returns a stable state indicator along with the final result, indicating that the text segment has stabilized. Upon receiving the stable state indicator returned by the streaming speech recognition module, when the stable state indicator indicates that the current recognition result is no longer changing, the corresponding final recognition result is marked as a candidate stable segment. This method directly obtains stability confirmation signals from the internal state of the speech recognition engine, without waiting for additional silence time, and has advantages in scenarios where the speaker speaks continuously with short pauses between sentences.

[0051] This method employs a joint detection approach based on silence duration and semantic integrity. It determines whether the silence duration of the real-time speech stream reaches a preset silence threshold, which is set based on the typical duration of pauses between sentences in natural dialogue. When the silence duration reaches the preset threshold, it further determines whether the current transcription result constitutes a complete semantic unit. The determination of a complete semantic unit involves inputting the current transcription result into a semantic integrity analysis model, which outputs a confidence score indicating whether the current transcription result is a complete sentence. The semantic integrity analysis model takes the word sequence of the current transcription result as input and the sentence integrity confidence score as output. It extracts contextual semantic features through a multi-layer self-attention encoder and maps them to the integrity confidence score through a fully connected layer. During training, the model uses complete sentences as positive samples and truncated half-sentences as negative samples, employing a binary classification cross-entropy loss function. This allows it to identify whether the transcription result is complete in terms of grammatical structure and semantic expression, such as whether it ends with a period, question mark, or exclamation mark, and whether there are obvious half-sentence truncation features. When both conditions are met—the silence duration reaching a preset silence threshold and the current transcription result constituting a complete semantic unit—the current transcription result is marked as a candidate stable segment. This method combines pause cues at the sound level with semantic integrity judgment at the language level. Compared with schemes that rely solely on silence duration, it can effectively distinguish between thought pauses within a sentence and expression-ending pauses at the end of a sentence, reducing the possibility of misjudging pauses within a sentence as stable segments at the end of the sentence.

[0052] The detection is based on dual-channel recognition and comparison. The dual-channel recognition result is obtained by comparing the recognition results of the same speech stream by two parallel speech recognition engines. The two speech recognition engines are internally independent and can employ different acoustic models or the same model structure but trained with different random initialization parameters. Each engine independently performs streaming recognition on the same speech stream, outputting intermediate and final results. When both engines output final results for the same speech segment, the two final results are compared to calculate the text consistency rate. The final results output by the two engines may differ in character length. The alignment method uses a dynamic time warping algorithm to construct alignment paths between two text sequences character by character, establishing character-level alignment relationships under the constraint of minimizing edit distance. When the alignment results of the two engines conflict (i.e., the two engines output different characters at the same alignment position), the arbitration algorithm prioritizes the recognition result of the engine with the higher confidence score. The confidence score is output by each engine based on its internal acoustic model's confidence in recognizing the speech segment. If the text consistency rate exceeds a preset consistency rate threshold, a final confirmation flag is returned, indicating that the recognition results of the two independent engines are highly consistent, and the recognized text segment has extremely high reliability. The system receives the final confirmation flag from the dual-channel recognition results. When the final confirmation flag is active, the current transcription result corresponding to the final confirmation flag is marked as a candidate stable segment. This method utilizes cross-validation of two independent recognition engines to improve the confidence of stable transcription segments at the end of sentences, and is particularly suitable for low signal-to-noise ratio scenarios with high environmental noise or heavy accents.

[0053] After a stable transcript segment at the end of a sentence is identified, the corresponding speech segment is bound to the current target question identifier and the speaker role label determined in step S200. The current target question identifier is a unique identifier for the currently executing item in the preset question list in the query system; for example, "Q5" represents the fifth question. The storage reference of the speech segment corresponding to the stable transcript segment, the current session identifier, the current question identifier, and the speaker role label are combined into a quadruple containing the session identifier, question identifier, role label, and speech segment reference. The purpose of the quadruple is to establish a structured association between each answer text in the transcript and its corresponding original speech segment, the session in which it is located, the question it belongs to, and the speaker role, providing complete association data for the subsequent establishment of the speech backtracking index. The generated quadruple is written into the association data structure of the current session, waiting for the subsequent generation of answer patches to trigger the establishment of the speech backtracking index.

[0054] The following technical effects were achieved through this step: First, this step provides four parallel methods for detecting stable sentence-end transcripts. These methods independently determine whether the recognition results are stable from four different dimensions: speech activity boundary, internal stable signal recognition by streaming, silence duration and semantic integrity combined, and dual-channel recognition cross-validation. The four methods complement each other and have appropriate detection methods under different speaking styles and environmental noise conditions, thus improving the reliability and real-time performance of stable sentence-end transcript detection.

[0055] Second, this step triggers quadruple binding immediately after the stable transcribed segment at the end of the sentence is confirmed, associating and storing four types of information—session identifier, question identifier, role label, and speech segment reference—simultaneously. Compared to existing technologies where unstructured role speech intervals cannot be associated with preset question slots, this establishes precise multi-dimensional index anchors for tracing the source of subsequent answer text and original speech.

[0056] S400: When generating an answer patch based on the stable transcript segment at the end of the sentence during the transcript generation process, a speech backtracking index item is established based on the question identifier and answer text of the answer patch, combined with the speech segment reference in the quadruple and the transcript conversation temporal confidence parameter. During the transcript generation process, after a stable transcribed segment at the end of a sentence is identified, the transcript system extracts the answer semantics from this segment and uses it as an incremental answer patch to fill the corresponding question slot in the transcript template, generating an answer record. This answer record is presented as a text in the transcript display interface, but the original audio evidence behind it has not yet been traceably linked to the answer text. If the accuracy of the transcription of a certain answer text needs to be verified in subsequent transcript review, evidence examination, or procedural legality verification, the corresponding original audio segment cannot be directly located based solely on the answer text. Reviewers need to manually search for the corresponding position in the entire recording, which is inefficient and prone to omissions.

[0057] Step S400 provided in this embodiment of the invention includes: The voice backtracking index is stored in a distributed cache, and the voice segments are persistently stored as independent files after lossy compression encoding; or, the voice backtracking index is stored in a relational database, and the voice segments are persistently stored as session-level large files with time offset indexes after lossy compression encoding. The voice backtracking index includes at least the unique answer identifier, session identifier, question identifier, answer text hash, voice segment storage path, start and end timestamps, and speaker role label.

[0058] The specific implementation method is as follows: When the transcription system generates an answer patch based on the sentence-end stable transcript, the answer patch contains the following information: a question identifier, indicating which question in the transcription template the answer belongs to; the answer text, which is the answer text extracted from the sentence-end stable transcript or generated after semantic processing; and a session identifier, indicating a unique identifier for the current query session. The transcription system searches for a matching quadruple in the quadruple association data generated in step S300 based on the question identifier in the answer patch, and extracts the speech segment reference from the quadruple. The speech segment reference is a pointer or reference identifier pointing to the storage location of the corresponding speech segment of the sentence-end stable transcript.

[0059] Based on the information in the voice segment reference and answer patch, a voice backtracking index is constructed, and the corresponding temporal confidence value is extracted from the session temporal confidence parameter configured for the voice segment in step S100 and included in the index. The voice backtracking index item includes at least the following fields: unique answer identifier, a globally unique identifier generated by the answer patch in the transcription system, used to correspond one-to-one with the answer text in the transcription; session identifier, recording the unique identifier of the current query session; question identifier, recording the preset question number corresponding to the answer; answer text hash, a fixed-length summary value obtained by hashing the answer text, used to quickly match the answer text in subsequent searches; voice segment storage path, recording the physical storage path or logical reference address of the voice segment in the persistent storage system; start and end timestamps, recording the start time offset and end time offset of the voice segment in the original voice stream, for accurate positioning during playback; speaker role label, recording the speaker role corresponding to the voice segment, determined by step S200; and transcription session temporal confidence, recording the question-and-answer temporal compliance parameters corresponding to the voice segment, configured by step S100 when segmenting the voice segment.

[0060] The storage location of the voice retrospective index and the persistent storage method of the voice segment files adopt one of the following two schemes or a combination of the two schemes.

[0061] The first approach involves a distributed cache combined with independent file storage. Voice recall index entries are stored as hot data in the distributed cache, which employs a key-value pair storage structure. The key is the unique identifier of the answer, and the value is the complete voice recall index entry, supporting millisecond-level fast query responses. Voice segments are persistently stored as independent files in object storage or a network file system after lossy compression encoding. The purpose of lossy compression encoding is to reduce storage space usage. Voice encoding algorithms are used to compress the original PCM format voice segments, such as AAC or Opus encoding, reducing the original voice file size to about one-tenth of its original size. The compressed voice files are named using the unique identifier of the answer or a voice segment reference, facilitating direct location from the voice recall index. By storing each voice segment as an independent file, when it's necessary to play back the voice corresponding to a specific answer, only the corresponding single voice file needs to be read directly from the voice segment storage path in the voice recall index, without loading other voice data from the session.

[0062] The second approach uses a relational database with session-level large files and a time-offset index for storage. Voice playback index entries are stored as structured records in the relational database, with each field corresponding to a column in the database table. A database index is built using the unique answer identifier as the primary key, allowing for quick retrieval of the corresponding voice playback index entry by answer identifier during subsequent review. Voice segments, after lossy compression encoding, are not split into independent files but are merged and stored as a single session-level large file, organized by session. Within the large file, the lossy compressed encoded data of all voice segments in that session are stored sequentially by time, with each voice segment having a unique start and end offset. The voice segment storage path field in the voice playback index entry records the storage path of the session-level large file, and the start and end timestamp fields record the time offset position of the voice segment within the large file. These two fields allow for precise location and extraction of the corresponding voice segment data for playback within the large file. Compared to the independent file storage approach, the session-level large file approach reduces the number of small files, facilitating the overall archiving and migration of session data. When multiple sessions' voice data need to be backed up in batches or transferred to long-term archive storage, the session-based packaging method is more efficient. For example, Figure 3 Example data table diagram of voice traceability index items provided in embodiments of the present invention.

[0063] The following technical effects were achieved through this step: First, this step automatically creates a voice traceability index when the answer patch is generated, precisely binding each answer text in the transcript with traceability information such as a unique answer identifier, session identifier, question identifier, voice segment storage path, start and end timestamps, and speaker role tags. Compared with existing technologies where unstructured role-based speech intervals cannot be associated with preset question slots and do not support answer traceability, this establishes a precise traceability link from the answer text to the original voice segment, providing a traceable data foundation for transcript review, evidence examination, and procedural legality verification.

[0064] Second, this step provides two optional voice storage solutions, suitable for different deployment scales and retrieval needs. The distributed caching plus independent file storage solution is suitable for high-frequency review scenarios, with voice tracing and retrieval responses completed in milliseconds; the relational database plus session-level large file solution is suitable for scenarios requiring session-level archiving and historical data management.

[0065] Third, this step incorporates the transcript conversation temporal confidence parameter into the voice playback index, storing the question-and-answer temporal compliance information of the voice segments together with the source information. Compared with existing technologies that only store basic source information and lack temporal compliance verification, this provides a parameter basis for determining whether the question-and-answer order conforms to the transcript specifications during subsequent voice playback.

[0066] S500: In the transcript display interface, a voice playback control is generated for each incrementally filled answer field. When the user triggers the voice playback control, the corresponding voice segment is played back according to the voice backtracking index item and the transcript session temporal confidence parameter.

[0067] The specific implementation method is as follows: A voice playback control is generated for each incrementally filled answer field in the transcript display interface. The transcript system retrieves the corresponding index record from the voice playback index storage based on the unique identifier of the answer field. If a matching voice playback index item is found, a voice playback control is generated next to the answer field; for answer fields without an associated voice playback index item, no voice playback control is generated.

[0068] When a user triggers the voice playback control, the transcription system extracts the storage path and start / end timestamps of the corresponding voice segment from the voice playback index, locates and reads the corresponding voice segment file, and plays it back through an audio player after lossy compression decoding. During playback, the corresponding answer field on the transcription display interface is highlighted, and the speaker role label and duration information are displayed simultaneously. At the same time, the voice playback control reads the transcription conversation timing confidence parameter corresponding to the answer field from the voice playback index to determine whether the question-and-answer timing of the voice segment is compliant. If the timing confidence is within the normal range, voice playback is triggered directly; if the timing confidence is below the normal range, it indicates that the voice segment may have interruptions, interruptions, or timing discrepancies, and a timing error message is displayed to the user before playback starts, reminding the reviewer to verify whether the question-and-answer timestamps of the voice segment are consistent with the question-and-answer order recorded in the transcript.

[0069] Preferably, embodiments of the present invention also provide a manual correction step, including: Provides a manual review interface, displaying a list of voice clips that were not successfully assigned a role, as well as a list of voice clips whose similarity comparison scores are lower than the second preset threshold; The system receives manually corrected character types and adds the voiceprint feature vectors of the corrected speech segments to the corresponding character's voiceprint database, then recalculates the character centroid of the corresponding character.

[0070] The specific implementation method is as follows: A manual review interface is provided, which is accessed by the operator after the inquiry ends or during a pause in the inquiry process. The manual review interface displays two lists of voice segments: the first list is a list of voice segments that were not successfully assigned a role, including voice segments whose similarity scores were all below the first preset threshold in step S200, which were created as new roles and marked as awaiting manual confirmation; the second list is a list of voice segments whose similarity comparison scores are below the second preset threshold, including voice segments that were automatically identified as a role but have low similarity scores and limited confidence in their assignment. The second preset threshold is higher than the first preset threshold and is used to mark low-confidence assignment segments located near the automatic determination boundary for operator review and confirmation.

[0071] Each record in the list includes the audio segment number, segment duration, the currently automatically determined role label, similarity score, and the corresponding audio playback control. Operators can click the playback control to listen to the complete audio segment and determine the correct role type based on the speaker's vocal characteristics, content, and context. If the operator believes the automatically determined role label is correct, they click the "Confirm" button to mark the audio segment as reviewed; if the operator believes the automatically determined role label is incorrect, they select the correct role type from the role drop-down list and click the "Correct" button to manually correct it.

[0072] After manual correction is completed, the voiceprint feature vector of the corrected speech segment is added to the corresponding character's voiceprint database. The voiceprint database is a collection of voiceprint feature vectors for all confirmed speech segments of that character. If the character already has a voiceprint database, the new feature vector is appended to it; if the character is newly created, this feature vector is used as the first vector in the voiceprint database. After the voiceprint database is updated, the character centroid is recalculated. The recalculation method is to calculate the arithmetic mean of all voiceprint feature vectors in the character's voiceprint database across all dimensions; the resulting mean vector is the updated character centroid. The updated character centroid takes effect in subsequent queries, gradually improving the accuracy of character determination with the accumulation of manual corrections.

[0073] When the cumulative number of manual corrections reaches a preset number, the similarity threshold for the corresponding character is automatically adjusted to adapt to the actual voiceprint distribution, including: The average similarity distance between all voiceprint feature vectors in the corresponding character's voiceprint database and the centroid of the current character is calculated, as well as the degree of dispersion between each vector. Based on the average similarity distance and dispersion, a new similarity threshold for the corresponding role is calculated according to a preset mapping relationship; Replace the original first preset threshold with the new similarity determination threshold.

[0074] The specific implementation method is as follows: When the number of times a character has been manually corrected in the manual review interface reaches the preset number of corrections, the adaptive adjustment process of the similarity judgment threshold for that character is automatically triggered. The preset number of corrections is a threshold value for accumulating a sufficient sample size in the voiceprint database for that character. Only after a sufficient number of manually confirmed correct voiceprint feature vectors have been accumulated in the voiceprint database can the statistical results be representative and reliable.

[0075] The adaptive adjustment process first calculates the similarity distance between all voiceprint feature vectors in the corresponding character's voiceprint database and the current character's centroid. The similarity distance is calculated as follows: for each voiceprint feature vector in the database, calculate its cosine similarity to the current character's centroid; subtract the cosine similarity from 1 to obtain the similarity distance. The arithmetic mean of all similarity distances is calculated and used as the average similarity distance for the character's voiceprint feature vectors, reflecting the degree of deviation from the center of the character's voiceprint feature vectors. Simultaneously, the standard deviation of each similarity distance is calculated as the dispersion of the character's voiceprint feature vectors, reflecting the consistency fluctuation of the character's voiceprint features across different speech segments.

[0076] Then, a new similarity threshold for the corresponding role is calculated based on the average similarity distance and dispersion. The formula for calculating the new similarity threshold is: Th new=1-(μ d +α×σ d ). Where μ d σ represents the average similarity distance, which is the average cosine distance between the centroid of the current character and all voiceprint feature vectors in the voiceprint database. d The degree of dispersion represents the standard deviation of the cosine distance between each voiceprint feature vector and the character's centroid. α is a preset tolerance coefficient used to adjust the sensitivity of the new threshold to the degree of dispersion. For example, the default value of the tolerance coefficient α is 1.0, at which point the new threshold covers approximately 68% of the similarity range of vectors in the voiceprint database. When α increases to 2.0, the new threshold covers approximately 95% of the similarity range of vectors, resulting in a more lenient attribution determination, suitable for characters with large voiceprint fluctuations. When α decreases to 0.5, the new threshold covers only approximately 38% of the similarity range of vectors, resulting in a more stringent attribution determination, suitable for characters with concentrated voiceprint distributions. The specific value of the tolerance coefficient can be dynamically adjusted during system operation based on the historical accuracy of each character's determination.

[0077] The original first preset threshold is replaced with a new similarity threshold, and the new threshold is used in the attribution determination of subsequent speech segments for this character. If the character reaches the preset number of corrections again in subsequent manual corrections, the adaptive adjustment process is triggered again, and the threshold is updated with the latest voiceprint database statistics, so that the threshold dynamically tracks the long-term changes in the character's voiceprint distribution as the voiceprint database continues to expand.

[0078] The closed-loop combination of manual correction and adaptive adjustment enables the judgment threshold to be dynamically optimized as the character's voiceprint data accumulates. The threshold is automatically tightened for characters with concentrated voiceprint distribution and automatically relaxed for characters with dispersed voiceprint distribution, thus continuously improving the accuracy of character judgment.

[0079] Example 2, as Figure 4 As shown, based on the same inventive concept as the real-time voice interaction and incremental backfilling method for interrogation transcripts provided in Embodiment 1, this embodiment of the invention also provides a real-time voice interaction and incremental backfilling system for interrogation transcripts, the system comprising: The real-time voice segmentation module 11 is used to acquire real-time voice streams and segment them into several voice segments during the transcript generation process based on closed-loop voice interaction and incremental backfilling, and to configure conversation timing confidence parameters for the voice segments to verify the compliance of the transcript question-and-answer timing. The voiceprint role determination module 12 extracts the voiceprint feature vector of each speech segment, compares it with the centroid of at least one registered role, and determines the speaker role label corresponding to the speech segment. The quadruple binding module 13 obtains the stable transcript segment detected during the transcript generation process, binds the speech segment corresponding to the stable transcript segment to the current target question identifier and the speaker role label, and generates a quadruple containing a session identifier, question identifier, role label and speech segment reference. The speech retrospective index module 14, when generating an answer patch based on the stable transcript segment at the end of the sentence during the transcript generation process, establishes a speech retrospective index item based on the question identifier and answer text of the answer patch, combined with the speech segment reference in the quadruple and the transcript conversation temporal confidence parameter; The voice playback control module 15 generates a voice playback control for each incrementally filled answer field in the transcript display interface. When the user triggers the voice playback control, the corresponding voice segment is played back according to the voice backtracking index item and the transcript session timing confidence parameter.

[0080] In one embodiment, the real-time speech segmentation module 11 is further configured to acquire speech streams from different directions through at least two microphone channels, and separate the speech stream from the direction of the questioner and the speech stream from the direction of the questioned person based on the sound source localization.

[0081] Speech segments are segmented based on silence boundaries and semantic boundaries by using a single-channel microphone combined with speech activity detection and speaker change detection.

[0082] In one embodiment, the voiceprint role determination module 12 is further used to preprocess the speech segments: The preprocessing includes resampling, amplitude normalization, and removal of beginning and end silences; The voiceprint feature vector is extracted through a deep voiceprint model, which includes a speaker embedding network based on a residual network.

[0083] At the beginning of the inquiry, role registration is performed: if a pre-registered voiceprint database of the inquirer exists, the matching voice segment is marked as the inquirer role; otherwise, the voiceprint feature vector of the first long voice segment of the inquired object that meets the preset criteria is set as the centroid of the inquired object role. Subsequent audio segments are assigned to corresponding roles by calculating their similarity to the centroids of registered roles. Segments with similarity greater than a first preset threshold are assigned to the corresponding roles, while segments with similarity less than the first preset threshold are created as new roles and marked as awaiting manual confirmation.

[0084] In one embodiment, the quadruple binding module 13 is further configured to continuously perform speech activity detection on the real-time speech stream, and when a speech end point is detected, mark the current transcription result as a candidate stable segment; The system receives a stable state identifier returned by the streaming speech recognition module. When the stable state identifier indicates that the current recognition result no longer changes, the current transcription result is marked as a candidate stable segment. Determine whether the silence duration of the real-time speech stream reaches a preset silence threshold, and at the same time determine whether the current transcription result constitutes a complete semantic unit. When both conditions are met, mark the current transcription result as a candidate stable segment. The system receives the final confirmation flag from the dual-channel recognition results. When the final confirmation flag is active, the current transcription result is marked as a candidate stable segment. The dual-channel recognition results are obtained by comparing the same speech stream after being recognized by two parallel speech recognition engines. When candidate stable segment markers are generated by at least one of the above methods, the corresponding candidate stable segment is confirmed as the sentence-end stable transcription segment, and subsequent steps are triggered.

[0085] In one embodiment, the voice backtracking index module 14 is further configured to store the voice backtracking index item in a distributed cache, and the voice segment is persistently stored as an independent file after lossy compression encoding; or, the voice backtracking index item is stored in a relational database, and the voice segment is persistently stored as a session-level large file with a time offset index after lossy compression encoding. The voice backtracking index includes at least the unique answer identifier, session identifier, question identifier, answer text hash, voice segment storage path, start and end timestamps, and speaker role label.

[0086] In one embodiment, the voice playback control module 15 generates a voice playback control for each incrementally filled answer field in the transcript display interface. When the user triggers the voice playback control, the corresponding voice segment is played back according to the voice backtracking index item and the transcript session timing confidence parameter.

[0087] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A real-time voice interaction and incremental backfilling method for interrogation transcripts, characterized in that, include: In the process of generating transcripts based on closed-loop voice interaction and incremental backfilling, real-time voice streams are acquired and segmented into several voice segments, and conversation timing confidence parameters are configured for the voice segments to verify the compliance of the transcript question-and-answer timing. Extract the voiceprint feature vector of each speech segment, compare it with the centroid of at least one registered role, and determine the speaker role label corresponding to the speech segment; Obtain the stable transcript segment detected during the transcript generation process, bind the speech segment corresponding to the stable transcript segment to the current target question identifier and the speaker role label, and generate a quadruple containing a session identifier, question identifier, role label and speech segment reference; When the transcript generation process generates an answer patch based on the stable transcript segment at the end of the sentence, a speech backtracking index item is established based on the question identifier and answer text of the answer patch, combined with the speech segment reference in the quadruple and the transcript conversation temporal confidence parameter. In the transcript display interface, a voice playback control is generated for each incrementally filled answer field. When the user triggers the voice playback control, the corresponding voice segment is played back according to the voice backtracking index item and the transcript session temporal confidence parameter.

2. The real-time voice interaction and incremental backfilling method for interrogation transcripts according to claim 1, characterized in that, Acquire real-time audio streams and segment them into several audio segments, including: Speech streams from different directions are acquired through at least two microphone channels, and the speech streams from the direction of the questioner and the direction of the questioned person are separated based on sound source localization.

3. The real-time voice interaction and incremental backfilling method for interrogation transcripts according to claim 1, characterized in that, Acquiring real-time audio streams and segmenting them into several audio segments also includes: Speech segments are segmented based on silence boundaries and semantic boundaries by using a single-channel microphone combined with speech activity detection and speaker change detection.

4. The real-time voice interaction and incremental backfilling method for interrogation transcripts according to claim 1, characterized in that, Before extracting the speaker feature vector for each speech segment, preprocessing of the speech segments is also included: The preprocessing includes resampling, amplitude normalization, and removal of beginning and end silences; The voiceprint feature vector is extracted through a deep voiceprint model, which includes a speaker embedding network based on a residual network.

5. A real-time voice interaction and incremental backfilling method for interrogation transcripts according to claim 1, characterized in that, The similarity is compared with at least one registered role centroid to determine the speaker role label corresponding to the speech segment, specifically including: At the beginning of the inquiry, role registration is performed: if a pre-registered voiceprint database of the inquirer exists, the matching voice segment is marked as the inquirer role; otherwise, the voiceprint feature vector of the first long voice segment of the inquired object that meets the preset criteria is set as the centroid of the inquired object role. Subsequent audio segments are assigned to corresponding roles by calculating their similarity to the centroids of registered roles. Segments with similarity greater than a first preset threshold are assigned to the corresponding roles, while segments with similarity less than the first preset threshold are created as new roles and marked as awaiting manual confirmation.

6. The real-time voice interaction and incremental backfilling method for interrogation transcripts according to claim 1, characterized in that, The specific steps for detecting the stable transcribed segments at the end of sentences include: Speech activity detection is continuously performed on the real-time speech stream. When a speech end point is detected, the current transcription result is marked as a candidate stable segment. The system receives a stable state identifier returned by the streaming speech recognition module. When the stable state identifier indicates that the current recognition result no longer changes, the current transcription result is marked as a candidate stable segment. Determine whether the silence duration of the real-time speech stream reaches a preset silence threshold, and at the same time determine whether the current transcription result constitutes a complete semantic unit. When both conditions are met, mark the current transcription result as a candidate stable segment. The system receives the final confirmation flag from the dual-channel recognition results. When the final confirmation flag is active, the current transcription result is marked as a candidate stable segment. The dual-channel recognition results are obtained by comparing the same speech stream after being recognized by two parallel speech recognition engines. When candidate stable segment markers are generated by at least one of the above methods, the corresponding candidate stable segment is confirmed as the sentence-end stable transcription segment, and subsequent steps are triggered.

7. A real-time voice interaction and incremental backfilling method for interrogation transcripts according to claim 1, characterized in that: The voice backtracking index is stored in a distributed cache, and the voice segments are persistently stored as independent files after lossy compression encoding; or, the voice backtracking index is stored in a relational database, and the voice segments are persistently stored as session-level large files with time offset indexes after lossy compression encoding. The voice backtracking index includes at least the unique answer identifier, session identifier, question identifier, answer text hash, voice segment storage path, start and end timestamps, and speaker role label.

8. A real-time voice interaction and incremental backfilling method for interrogation transcripts according to claim 1, characterized in that, It also includes a manual correction step: Provides a manual review interface, displaying a list of voice clips that were not successfully assigned a role, as well as a list of voice clips whose similarity comparison scores are lower than the second preset threshold; The system receives manually corrected character types and adds the voiceprint feature vectors of the corrected speech segments to the corresponding character's voiceprint database, then recalculates the character centroid of the corresponding character.

9. A real-time voice interaction and incremental backfilling method for interrogation transcripts according to claim 8, characterized in that, Once the cumulative number of manual corrections reaches the preset number, the similarity threshold for the corresponding character will be automatically adjusted to match the actual voiceprint distribution, including: The average similarity distance between all voiceprint feature vectors in the corresponding character's voiceprint database and the centroid of the current character is calculated, as well as the degree of dispersion between each vector. Based on the average similarity distance and dispersion, a new similarity threshold for the corresponding role is calculated according to a preset mapping relationship; Replace the original first preset threshold with the new similarity determination threshold.

10. A real-time voice interaction and incremental backfilling system for interrogation transcripts, characterized in that, The system for implementing the method according to any one of claims 1 to 9, the system comprising: The real-time voice segmentation module is used to acquire real-time voice streams and segment them into several voice segments during the transcript generation process based on closed-loop voice interaction and incremental backfilling, and to configure conversation timing confidence parameters for the voice segments to verify the compliance of the transcript question-and-answer timing. The voiceprint role determination module extracts the voiceprint feature vector of each speech segment, compares it with the centroid of at least one registered role, and determines the speaker role label corresponding to the speech segment. The quadruple binding module obtains the stable transcript segment detected during the transcript generation process, binds the speech segment corresponding to the stable transcript segment to the current target question identifier and the speaker role label, and generates a quadruple containing a session identifier, question identifier, role label and speech segment reference. The speech retrospective index module, when generating an answer patch based on the stable transcript segment at the end of the sentence during the transcript generation process, establishes a speech retrospective index item based on the question identifier and answer text of the answer patch, combined with the speech segment reference in the quadruple and the transcript conversation temporal confidence parameter; The voice playback control module generates a voice playback control for each incrementally filled answer field in the transcript display interface. When the user triggers the voice playback control, the corresponding voice segment is played back according to the voice backtracking index item and the transcript session temporal confidence parameter.

Citation Information

Patent Citations

  • Speaker labeling method and device, electronic equipment and storage medium

    CN113421563A

  • Voice interaction method and device, electronic equipment and storage medium

    CN121393423A