Methods and systems for generating submission forms for voiceprint identification, and voiceprint identification methods.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-23
- Publication Date
- 2026-08-14
AI Technical Summary
[0010]本发明实施例通过提供一种声纹鉴定的送检表单生成方法、系统及声纹鉴定方法,声纹鉴定及送检表单生成的全流程依赖人工处理、流程割裂导致办案效率低,无法满足实时业务需求的问题,实现了声纹鉴定的系统化与标准化,大幅提升了办案效率
1、通过将语音质量评估、语音识别转文本、生成引导文本、采集样本语音及声纹特征比对分析集成为自动化的处理流程,解决了声纹鉴定时因人工操作环节多、流程割裂导致的声纹鉴定效率低、标准化程度不足的问题,实现了声纹鉴定流程的系统化、标准化和高效化。
Smart Images

Figure REF-OBJ-1776927908655-000002 
Figure REF-OBJ-1776927908655-000003 
Figure REF-OBJ-1776927908655-000004
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a method for generating a submission form for voiceprint identification, a system for generating a submission form for voiceprint identification, and a voiceprint identification method. Background Technology
[0002] With the rapid development of information technology and the widespread adoption of the internet, social media and instant messaging applications have become fully integrated into social life. While this transformation brings convenience to the public, it has also been exploited by criminals, leading to profound changes in the forms and methods of illegal and criminal activities. Voice recordings on new social media software such as WeChat, QQ, and WhatsApp, as well as traditional mobile phone call recordings, have become important carriers for criminals to transmit information and conduct transactions. This has resulted in a surge in the amount of voice recordings involved in cases. These voice data contain rich information about relationships between individuals and criminal plans, serving as direct or indirect evidence. They play an irreplaceable role in reconstructing criminal facts, identifying suspects, breaking case deadlocks, and conducting in-depth investigations to dismantle criminal gangs. However, faced with massive and complex voice data, investigators generally encounter difficulties in converting it into effective evidence, including low efficiency in acquiring voice information and challenges in maximizing the value of voice evidence. This severely restricts the quality and efficiency of case investigations.
[0003] To address the processing needs of voice data, several auxiliary tools have emerged in the market. Currently, the main types of independent software or tools include: first, voiceprint comparison software, focused on speaker identification of specific voice segments; second, speech-to-text tools, designed to convert speech content into text for easy reading and retrieval; and third, relatively simple audio file management software, providing basic playback, trimming, and annotation functions. Furthermore, in the field of forensic identification, the preparation of submitted materials mainly relies on manual processing, requiring the completion of complex forms containing detailed attribute information of the evidence and samples, according to the different requirements of each identification institution.
[0004] Despite the existence of the aforementioned tools, current technical solutions still fail to meet the needs of systematic and intelligent processing of voice evidence in police operations, mainly due to the following problems: 1. Fragmented tool functions and lack of collaborative workflow: Current voiceprint, transcription, and management tools are all isolated "islands" with single functions. Their data formats are incompatible, and the processing workflow cannot be connected. Investigators need to manually export, convert, and import data between different software programs in a tedious manner, which makes it impossible to form an automated and intelligent workflow from data screening and feature analysis to sample preparation, severely restricting overall work efficiency.
[0005] 2. Difficulty and inefficiency in acquiring massive amounts of voice information: The electronic devices seized with the case have huge storage capacities, containing hundreds of hours of voice data. The traditional method of relying entirely on manual "listening, recording, and filtering one by one" is inadequate when processing such a large amount of data. Not only is it slow, but it also easily leads to the dilemma of "not being able to listen to all the data or remember it all." Furthermore, auditory fatigue and subjective limitations may cause key clues to be missed, resulting in missed opportunities to solve the case.
[0006] 3. Subjective and Blind Screening of Evidence Voice Recordings Leads to Quality Inconsistencies: The quality of the evidence's voice recordings directly determines the feasibility of voiceprint identification. Currently, evidence screening relies entirely on the subjective auditory judgment of investigators, lacking objective and quantifiable voice quality assessment standards and technical support. This blind screening method often results in selected evidence being unusable for identification due to quality defects such as excessive background noise, excessively short voice segments, and unclear content, rendering preliminary work ineffective.
[0007] 4. Repeated and inefficient voice sample collection with a low pass rate: Collecting voice samples that meet the requirements for identification is another major challenge. Suspects often disguise themselves by changing their speech speed, tone, and pitch, while investigators lack professional sampling guidance and real-time feedback tools. This leads to repeated sample collection, yet it is still difficult to obtain valid samples that match the voice samples sufficiently in terms of pronunciation, tone, and rhythm, thus delaying the identification process.
[0008] 5. The preparation of submitted materials is prone to errors and omissions, posing a high risk of procedural compliance issues: The material preparation process before commissioning an appraisal requires strict adherence to procedures, necessitating the accurate completion of a large amount of structured information. Currently, this process heavily relies on manual operation, making it highly susceptible to problems such as errors in information entry, omission of key attributes, and formats that do not meet the requirements of specific appraisal institutions. This not only leads to repeated communication and revisions with the appraisal institution, reducing efficiency, but may also affect the legal validity of evidence due to procedural flaws, dampening the enthusiasm of investigators to apply voiceprint technology.
[0009] Therefore, there is an urgent need for a professional, systematic, and authoritative voice data application and processing solution to provide solid technical support for frontline case handling in using voice data to uncover criminal clues and combat criminal activities. Summary of the Invention
[0010] This invention provides a method, system, and method for generating submission forms for voiceprint identification. It addresses the problem that the entire process of voiceprint identification and submission form generation relies on manual processing, resulting in fragmented workflows, low case-handling efficiency, and inability to meet real-time business needs. This invention achieves the systematization and standardization of voiceprint identification, significantly improving case-handling efficiency.
[0011] This invention provides a method and system for generating a submission form for voiceprint identification. The method for generating the submission form for voiceprint identification includes: Acquire initial voice data, perform multi-dimensional quality assessment on the initial voice data, and obtain voice data of the sample to be processed that meets the preset quality threshold; The speech data of the sample to be processed is subjected to speech recognition to obtain the corresponding text content of the sample; Based on the text content of the sample, sample text material is automatically generated to guide the target object to speak; Obtain sample speech data based on the sample text material; Based on the voice data of the evidence, the voice data of the sample, and the identity information of the target object, a structured voiceprint identification submission form is automatically generated.
[0012] Optionally, the step of performing a multi-dimensional quality assessment on the initial speech data includes: The initial speech data is segmented into multiple consecutive speech analysis segments, wherein there is a temporal overlap between adjacent speech analysis segments; For each of the aforementioned speech analysis segments, quality assessment processing is performed in parallel.
[0013] Optionally, the step of performing quality assessment processing in parallel for each of the speech analysis segments includes: The signal of a predetermined duration at the beginning of the initial speech data is extracted as the background noise benchmark. The speech analysis segment is divided into multiple analysis frames. The average energy ratio of the speech frame to the silence frame is calculated and converted into a decibel value as the signal-to-noise ratio score. The proportion of active speech frames in the total number of analyzed frames is used as the speech activity score. The proportion of silent frames in the total number of analyzed frames in the speech analysis segment is counted and used as the silent proportion score. Calculate the difference between the maximum and minimum amplitudes of the speech analysis segment signal and convert it into a decibel value as a dynamic range score; Based on predefined weighting coefficients, the signal-to-noise ratio score, speech activity score, silence percentage score, and dynamic range score are weighted and summed to obtain the segment quality score of the speech analysis segment.
[0014] Optionally, the step of performing a multi-dimensional quality assessment on the initial speech data to obtain evidence speech data that meets a preset quality threshold includes: Based on the segment quality scores of each speech analysis segment, the overall quality assessment result of the initial speech data is determined, and the speech segments that meet the preset quality threshold are selected or marked according to the overall quality assessment result as the evidence speech data. The methods for determining the overall quality assessment results include at least one of the following: The segment quality scores of each speech analysis segment are weighted and averaged, and the average value is used as the overall quality score. The score of the speech analysis segment with the highest quality score is selected as the representative of the overall quality score; Based on the distribution of quality scores for each segment, the initial speech data is labeled into multiple quality level regions.
[0015] Optionally, before the step of performing a multi-dimensional quality assessment on the initial speech data, the following steps are included: Obtain the context information associated with the initial speech data, and preliminarily associate candidate speaker identifiers with the initial speech data based on the context information; the context information includes acoustic feature information and non-acoustic feature information, and the non-acoustic feature information includes account identifier information and speech generation time information; Based on the candidate speaker identifiers, speaker separation and clustering are performed on the initial speech data; Based on the results of separation and clustering, short-term speech segments with similar acoustic features are grouped to their respective speakers, and speech segments belonging to the same speaker are spliced and labeled to form speaker speech segments with temporal information. Each speaker's speech segment separated is used as the initial speech data for subsequent processing.
[0016] Optionally, the step of automatically generating sample text material for guiding the target object to speak based on the text content of the evidence includes: The text content of the sample was analyzed to obtain key features used to guide vocalization; Obtain textual evidence of verbal evidence associated with the target object; The verbal evidence text material is processed based on the key features to generate the sample text material.
[0017] Optionally, the step of processing the verbal evidence text material based on the key features to generate the sample text material includes: If the key feature is a semantic feature, then extract the sentences that are semantically related to the key feature from the verbal evidence text material and reorganize the sentences to obtain the sample text material; If the key feature is a pronunciation feature, or if there is no sentence in the verbal evidence text material that is directly related to the key feature, then the key feature is inserted into the verbal evidence text material to obtain the sample text material.
[0018] Optionally, the step of embedding the key features as new content into the verbal evidence text material to obtain the sample text material includes: The key features are inserted into the verbal evidence text material to obtain intermediate text material; The sentence order of the intermediate text material is adjusted to obtain the sample text material.
[0019] Optionally, the step of automatically generating a structured voiceprint identification submission form based on the voice data of the evidence, the voice data of the sample, and the identity information of the target object includes: Based on the user's selected submission institution or case type, the system matches and calls the corresponding standard submission form template from a pre-set template library; the template library includes structured templates that conform to the document specifications of different appraisal institutions. The voice data of the evidence, the voice data of the sample, and the identity information of the target object are automatically mapped and filled according to the field definitions of the submission form template to generate the voiceprint identification submission form.
[0020] Furthermore, to achieve the above objectives, embodiments of the present invention also provide a system for generating submission forms for voiceprint identification, the system comprising: The data acquisition and preprocessing module is used to acquire initial voice data, perform multi-dimensional quality assessment on the initial voice data, and obtain sample voice data that meets the preset quality threshold. The speech translation module is used to perform speech recognition on the speech data of the evidence to obtain the corresponding text content of the evidence; The text generation module is used to automatically generate sample text materials to guide the target object to speak based on the text content of the sample. The sample acquisition and verification module is used to obtain sample speech data and the identity information of the target object based on the sample text material. The commissioned submission module is used to automatically generate a structured voiceprint identification submission form based on the voice data of the sample, the voice data of the specimen, and the identity information of the target object. The sample screening module is used to retrieve the sample voice data based on the received keywords, locate and highlight the position of the keywords in the sample voice data, and associate the location with the time point in the original voice data and the context position in the chat sequence.
[0021] Furthermore, to achieve the above objectives, embodiments of the present invention also provide a voiceprint identification method, the method comprising: Receive a voiceprint identification submission form and a voice data packet associated with the voiceprint identification submission form, wherein the voice data packet includes the voice data of the sample and the voice data of the specimen. Extract a first set of voiceprint features from the speech data of the evidence, and extract a second set of voiceprint features from the speech data of the sample; Calculate the similarity score between the first voiceprint feature set and the second voiceprint feature set; Based on the quality assessment information associated with the voiceprint identification submission form, the judgment threshold is dynamically determined. The similarity score is compared with the dynamically determined judgment threshold to generate a preliminary identification opinion; The confidence level of the preliminary identification opinion is calibrated, and the voiceprint identification result containing the identification opinion and the quantified confidence level is output.
[0022] One or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages: 1. By integrating speech quality assessment, speech recognition to text conversion, generation of guiding text, collection of sample speech, and voiceprint feature comparison and analysis into an automated processing flow, the problem of low efficiency and insufficient standardization in voiceprint identification caused by multiple manual operation steps and fragmented processes has been solved, realizing the systematization, standardization, and efficiency of the voiceprint identification process.
[0023] 2. By automatically generating text materials to guide sample collection based on the text identified in the evidence, the problem of lack of correlation between the sample collection content and the evidence speech, which led to the failure of voiceprint comparison, was solved, thereby improving the targeting of sample collection and the success rate of voiceprint comparison.
[0024] 3. By conducting multi-dimensional quality assessment and screening of the audio samples at the front end of the process, the problem of resource waste and inability to conduct identification is solved by allowing low-quality samples to directly enter the subsequent processing stage. This achieves efficient utilization of identification resources and improves the feasibility of identification.
[0025] 4. By segmenting speech data into multiple speech analysis segments with overlapping durations and performing parallel quality assessment, the problem of long single-threaded processing of long audio and low system resource utilization is solved, thereby improving the efficiency of speech data quality assessment. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of the framework of the method for generating a submission form for voiceprint identification according to the present invention. Figure 2 This is a flowchart illustrating the method for generating a submission form for voiceprint identification according to the present invention. Figure 3 This is a schematic diagram illustrating the process of multi-dimensional quality assessment according to an embodiment of the present invention; Figure 4 for Figure 2 A flowchart of step S120 in one embodiment is shown in the corresponding example. Figure 5This is a schematic diagram illustrating the process of separating and clustering initial speech data according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the submission form generation system for voiceprint identification according to the present invention. Figure 7 This is a flowchart illustrating the voiceprint identification method of the present invention; Figure 8 This is a schematic diagram of the terminal structure of the hardware operating environment involved in an embodiment of the present invention. Detailed Implementation
[0027] To address the issues of reliance on experts, cumbersome processes, and low efficiency in voiceprint identification in scenarios such as customs anti-smuggling, this invention proposes a method for generating submission forms for voiceprint identification. By automatically assessing the quality of initial speech and performing text recognition, a guiding text for sample collection is intelligently generated. Then, sample speech obtained based on this guiding text and the identity information of the target object are acquired. Based on the speech data of the evidence, the sample speech data, and the identity information of the target object, a structured voiceprint identification submission form is automatically generated. This automates the core identification process and standardizes the front-end operations, significantly improving case-handling efficiency while ensuring the rigor of judicial procedures.
[0028] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the invention to those skilled in the art.
[0029] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0030] In this embodiment, a method for generating a submission form for voiceprint identification is provided.
[0031] Reference Figure 1 and Figure 2 The method for generating the submission form for voiceprint identification in this embodiment includes the following steps: Step S1: Obtain initial voice data and perform multi-dimensional quality assessment on the initial voice data to obtain voice data of the sample to be processed that meets the preset quality threshold. In this embodiment, initial voice data refers to a set of original voice evidence obtained in a case. In communication software forensics, this typically manifests as a collection of multiple independent voice messages from a single conversation. Multi-dimensional quality assessment refers to performing a standardized set of quantitative evaluation indicators on each independent voice message in the collection in parallel, and outputting a concise, binary quality label. The preset quality threshold refers to the quantitative boundary line for determining whether a single voice message is clear or unclear; this threshold is optimized based on massive amounts of real-world data. Evidence voice data typically refers to the set of all voice messages marked as "clear" selected from the initial voice set; subsequent speech recognition and text analysis are prioritized or performed only on this set.
[0032] As an optional implementation, after receiving the initial voice file, the massive and fragmented voice data is initially screened, automatically distinguishing between "clear" voice data with analytical value and "unclear" voice data that requires manual review or has low value, greatly improving the efficiency of subsequent manual review or automatic processing.
[0033] See Figure 3 In performing multi-dimensional quality assessment, the embodiments of this application include the following steps: Step S110: Divide the initial speech data into multiple consecutive speech analysis segments, wherein there is a time overlap between adjacent speech analysis segments; In this embodiment, since long audio processing takes a long time and cannot fully utilize the parallel processing capabilities of computing devices, and simple non-overlapping cutting may destroy complete speech units, such as words and phrases, at the cutting point, thereby affecting the accuracy of local acoustic feature extraction, the initial speech is cut into speech analysis segments with time overlap, which can improve the efficiency of parallel processing while maximizing the continuity and integrity of the speech content in each analysis segment.
[0034] Optionally, a fixed-duration sliding window approach can be used for segmentation, ensuring that there is a predefined overlap area between adjacent windows. Experiments have shown that setting the duration of each speech analysis segment to 30 to 40 seconds can include sufficient speech content for stable feature analysis, while also serving as a basic task unit for parallel computing, maintaining load balance.
[0035] Optionally, the overlap duration between adjacent analysis segments can be set to 5 to 10 seconds to ensure that no cut point falls exactly in the middle of a key speech event, such as a plosive or formant transition. This provides a continuous context for subsequent feature extraction and comparison, helps to smooth the feature sequence, reduces edge distortion caused by cutting, and the overlapping part can also provide a basis for consistency verification of the analysis results of different analysis segments.
[0036] Alternatively, when performing speech segmentation, preliminary quality screening can be used to avoid setting segmentation points in areas marked as long silences or extremely low quality, so that each analysis segment contains as much effective speech as possible.
[0037] Furthermore, for overlapping segments, if the quality scores calculated in different analysis segments differ significantly, higher weights can be assigned to higher-quality segments during subsequent fusion.
[0038] Step S120: Perform quality assessment processing in parallel for each of the speech analysis segments.
[0039] In this embodiment, the quality assessment tasks of multiple analysis segments are distributed and processed simultaneously, which greatly improves the overall response speed of the system.
[0040] After step S110 is completed, a task queue containing all speech analysis segments is generated. Each task contains the speech data of an independent analysis segment, along with its start time and segment number. Based on the currently available computing resources, a set of worker threads is dynamically created or allocated. Each worker thread retrieves an analysis segment from the task queue and executes the complete quality assessment process independently and without interference.
[0041] like Figure 4 As shown, the steps for quality assessment of speech analysis segments include: Step S121: Extract the signal of a predetermined duration from the beginning part of the initial speech data as the background noise benchmark, divide the speech analysis segment into multiple analysis frames, calculate the average energy ratio of the speech frame and the silence frame and convert it into a decibel value as the signal-to-noise ratio score; In this embodiment, the signal-to-noise ratio (SNR) of the speech analysis segment is detected to quantify the intensity ratio of the speech signal relative to the background noise.
[0042] For example, the initial 100 milliseconds of audio signal from the initial speech data is extracted as the background noise baseline, and the root mean square (RMS) value of this signal is calculated as the noise RMS baseline. The current speech analysis segment is further subdivided into shorter analysis frames, with a frame length of 10 milliseconds and 160 sampling points at a 16kHz sampling rate. The RMS value is calculated for each frame. If the RMS value is greater than twice the RMS value of the noise baseline, the frame is determined to be a speech frame; otherwise, it is determined to be a silence frame. The average energy ratio (SNR) of all speech frames and silence frames is calculated using the formula SNR = 20log10(mean RMS value of speech frames / RMS value of noise). An SNR ≥ 20dB is set as the criterion for speech clarity. This threshold can effectively distinguish between subjectively clear and incomprehensible speech and unclear speech, providing a reliable quality basis for subsequent processing.
[0043] Alternatively, this SNR dB value can be mapped to a standardized score as a signal-to-noise ratio score, for example, setting SNR≥30dB as the full score and SNR≤0dB as 0 points, with a linear mapping within this range.
[0044] Step S122: Calculate the proportion of active speech frames in the speech analysis segment to the total number of analysis frames, and use it as the speech activity score; In this embodiment, speech activity is detected in the speech analysis segment, and the proportion of actual speaking time to the total speech duration is calculated.
[0045] For example, in a 35-second analysis segment, with 10-millisecond frames, there are a total of 3500 frames. If 2400 of these frames are identified as active speech frames, the calculation formula is: Speech activity ratio = Number of active speech frames / Total number of analysis frames. Therefore, the speech activity ratio of this analysis segment is 2400 / 3500 = 0.686, or 68.6%. A speech activity ratio ≥ 60% is set as the criterion for determining speech activity.
[0046] Optionally, this percentage can be directly converted into a voice activity score on a scale of 0-100, for example, 68.6% corresponds to 68.6 points.
[0047] Step S123: Calculate the proportion of silent frames in the total number of analyzed frames in the speech analysis segment, and use it as the silent proportion score; In this embodiment, silence percentage detection is performed on the speech analysis segment to identify excessively long, meaningless pauses or blanks. The silence percentage and speech activity score are theoretically complementary, but because the discrimination thresholds may differ, their sum may not necessarily be 100%.
[0048] For example, the percentage of silent frames = number of silent frames / total number of analyzed frames. A silent frame percentage ≤ 30% is considered a qualified speech analysis segment.
[0049] Optionally, a score of 100 is set for a noise level of ≤30% and a score of 0 is set for a noise level of ≥70%. A reverse linear mapping is performed within this range to obtain the noise level score.
[0050] Step S124: Calculate the difference between the maximum and minimum amplitude of the speech analysis segment signal and convert it into a decibel value as the dynamic range score; In this embodiment, dynamic range detection is performed on the speech analysis segment. Dynamic range reflects the amplitude difference between the loudest and softest sounds in the speech. A moderate dynamic range is beneficial for subsequent processing. An excessively large dynamic range may contain abrupt, impactful noise, while an excessively small dynamic range may indicate insufficient speech gain or severe compression.
[0051] For example, iterate through all sampling points of the current speech analysis segment to find its maximum amplitude value (max) and minimum amplitude value (min). Calculate the dynamic range using the formula: Dynamic Range = 20log10(max - min / max(|min|,|max|)). Set a dynamic range ≤ 10dB as a qualified speech analysis segment.
[0052] Optionally, the dynamic range can be mapped to a dynamic range score of 0-100 points based on a preset ideal dynamic range range, with points deducted for exceeding this range. The preset ideal dynamic range range can be 6-12dB.
[0053] Step S125: Based on predefined weighting coefficients, the signal-to-noise ratio score, speech activity score, silence ratio score, and dynamic range score are weighted and summed to obtain the segment quality score of the speech analysis segment.
[0054] In this embodiment, the scores from the four dimensions are integrated into a comprehensive segment quality score to fully evaluate the overall quality of the speech analysis segment.
[0055] For example, machine learning training and fine-tuning are performed using massive amounts of real-world speech data to obtain the weight values corresponding to each dimension. For instance, the signal-to-noise ratio (SNR) weight is 0.35, speech activity weight is 0.25, silence percentage weight is 0.25, and dynamic range weight is 0.15. The segment quality score is calculated as follows: (SNR score × SNR weight) + (Speech activity score × Speech activity weight) + (Silence percentage score × Silence percentage weight) + (Dynamic range score × Dynamic range weight). This results in a segment quality score between 0 and 100; a higher score indicates better overall clarity and usability of the speech segment.
[0056] As an optional implementation, after obtaining the segment quality score, the overall quality assessment result of the initial speech data is determined based on the segment quality score of each speech analysis segment, and the speech parts that meet the preset quality threshold are selected or marked according to the overall quality assessment result, which are then used as the evidence speech data.
[0057] Optionally, the methods for determining the overall quality assessment results include at least one of the following: The segment quality scores of each speech analysis segment are weighted and averaged, and the average value is used as the overall quality score. The score of the speech analysis segment with the highest quality score is selected as the representative of the overall quality score; Based on the distribution of quality scores for each segment, the audio data of the sample is labeled into multiple quality level regions.
[0058] The overall quality assessment result is compared with a preset quality threshold. If the overall quality score is higher than or equal to the threshold, the entire initial speech data is deemed to meet the quality requirements and is marked as "examined speech data". If the overall quality score is lower than the threshold, but some high-quality regions are identified using the quality distribution annotation method, the speech analysis segments within these high-quality regions are selected, either concatenated or individually marked as "examined speech data", while the low-quality portions are discarded. Simultaneously, the selected speech data is tagged "Passed Quality Screening" and includes relevant quality metadata (such as the overall score or the quality region it belongs to) in the database and interface. The final output is one or more continuous speech segments that meet the quality requirements, i.e., the examined speech data.
[0059] like Figure 5 As shown, before the step of performing a multi-dimensional quality assessment of the evidence's voice data, the following steps are included: Step S010: Obtain the context information associated with the initial speech data, and preliminarily associate candidate speaker identifiers with the initial speech data based on the context information; Step S020: Based on the candidate speaker identifiers, perform speaker separation and clustering on the initial speech data; Step S030: Based on the results of separation and clustering, short-term speech segments with similar acoustic features are grouped to their respective speakers, and speech segments belonging to the same speaker are spliced and labeled to form speaker speech segments with temporal information. Step S040: Use each speaker's speech segment as the initial speech data for subsequent processing.
[0060] In this embodiment, the initial voice data may contain mixed dialogues from multiple speakers. Therefore, at the very beginning of the voiceprint identification process, it is necessary to separate, classify, and reassemble the voices of different speakers in the mixed audio. This provides clean, track-separated input data for subsequent single-speaker voice quality assessment, text transcription, and voiceprint comparison, thereby improving the feasibility, accuracy, and efficiency of the entire identification process. The contextual information associated with the initial voice data includes voice information and non-voice information. The non-voice information includes account identification information and voice generation time information.
[0061] Optionally, account identifiers (such as WeChat IDs) and generation times associated with the voice data are automatically extracted. Based on the principle that voices from the same account are highly likely to come from the same person, a unified candidate speaker identifier is initially assigned to voices from the same account. This provides crucial prior guidance for subsequent voiceprint clustering algorithms, significantly improving the accuracy and efficiency of multi-person voice separation.
[0062] Optionally, based on a voiceprint feature clustering algorithm, short-duration speech segments with similar acoustic features in the initial speech data are automatically grouped into the same speaker, achieving unsupervised initial speaker separation. According to the clustering results, speech segments belonging to the same speaker are spliced and labeled in chronological order to form complete speaker speech segments with temporal information. Finally, each separated speaker speech segment is treated as an independent unit of evidence speech data and sent to subsequent multi-dimensional quality assessment and processing procedures.
[0063] Step S2: Perform speech recognition on the speech data of the sample to be processed to obtain the corresponding text content of the sample; In this embodiment, the voice data of the clean sample after quality assessment and screening is converted into corresponding text content, providing accurate input basis for subsequent intelligent guidance sample text generation based on text content.
[0064] Optionally, a domain-optimized offline speech recognition engine can be used to automatically transcribe the speech data of the evidence to be processed. This engine integrates acoustic and language models optimized for common scenarios and professional terminology in the judicial field to improve recognition accuracy in specific contexts. The recognition process takes high-quality single-speaker speech segments output from previous steps as input and outputs a timestamped verbatim transcript, while recording metadata such as recognition confidence to provide a reliable reference for subsequent processing.
[0065] Step S3: Based on the text content of the sample, automatically generate sample text material to guide the target object to speak; In this embodiment, guiding text is intelligently generated based on the content of the evidence to ensure that the collected sample speech is highly correlated with the evidence in terms of pronunciation features, thereby improving the effectiveness and accuracy of voiceprint comparison. Specifically, this includes the following steps: Step S310: Analyze the text content of the sample to obtain key features used to guide vocalization; In this embodiment, the analysis of the text content of the evidence may include semantic analysis, syntactic analysis, and / or phonetic analysis.
[0066] Optionally, the key features obtained include semantic features and phonetic features. Semantic features are words or phrases with clear lexical meaning, such as specific personal names, place names, verb or noun phrases, etc. These features are key to understanding the content of the discourse. Phonetic features typically include functional words, specific vowel or consonant combinations, and idiomatic expressions or interjections without substantive meaning but with stable pronunciation. Functional words such as demonstrative pronouns like "this" and "that" are of particular interest due to the diversity of their vowel pronunciations, such as / e / and / i / . Specific vowel or consonant combinations are specific phonemes or phoneme combinations that are identified from the sample speech and have diagnostic value.
[0067] Step S320: Obtain verbal evidence text materials associated with the target object; In this embodiment, textual evidence materials related to the target are acquired, serving as the draft for generating sample speech. Using these textual evidence materials as a draft provides a personalized, natural, and case-related language template for subsequent sample collection. By embedding key features extracted from the evidence into the target's own interrogation records, self-reports, and other texts, the target can be guided to speak naturally in a familiar and less guarded language environment, thus efficiently obtaining high-quality samples highly comparable to the evidence's speech in terms of pronunciation habits. This process not only significantly improves the initial success rate of sample collection and overcomes the problem of target individuals disguising themselves or speaking unnaturally due to unfamiliar text in related technologies, but also strengthens the evidentiary relevance between the sample and the evidence, and between the sample and the case, making the entire collection process more in line with judicial practice requirements.
[0068] Optionally, the verbal evidence may come from the audio and written transcripts of the questioning and interrogation of the target; as well as the target's self-statements in the case.
[0069] Step S330: Process the verbal evidence text material based on the key features to automatically generate the polished sample text material.
[0070] In this embodiment, different strategies are used to process the key features based on their type and their relevance to the verbal evidence text material, ultimately generating sample text material for guiding pronunciation.
[0071] As an optional implementation, when the key feature is determined to be a semantic feature, sentences that are semantically related to the key feature are extracted from the verbal evidence text material and the sentences are reorganized to obtain the sample text material.
[0072] Optionally, when extracting statements semantically associated with the key features, a full-text search and semantic similarity calculation are performed on the verbal evidence text material to find all statements that are content-related to the semantic features.
[0073] Optionally, during sentence restructuring, multiple related sentences extracted are spliced, combined, or appropriately grammatically rewritten according to certain logical rules to form a semantically coherent and content-related sample text. This aims to guide the target audience to naturally speak sentences containing that semantic feature.
[0074] As another optional implementation, when the key feature is determined to be a pronunciation feature, or when there is no sentence in the verbal evidence text material that is directly related to the key feature, the key feature is inserted into the verbal evidence text material to obtain the sample text material.
[0075] Optionally, the key feature can be treated as an independent text unit and directly inserted into the verbal evidence text to obtain intermediate text. The insertion point can be at the beginning or end of a sentence, between sentences, or in the intervals of paragraphs. For example, suppose the verbal evidence text is: "I was in the park, and then I went home." After inserting the phonetic feature "here," the intermediate text might become: "I was in this park, and then I went home." Optionally, the intermediate text material can be directly used as the sample text material.
[0076] Optionally, to improve the effectiveness of sample collection and prevent target subjects from changing their pronunciation habits based on memory, the sentence order of the intermediate text material can be adjusted. The adjustment method can be based on a preset algorithm, such as a random shuffling algorithm. The intermediate text material is segmented into sentences, and the order of these sentences is randomly shuffled. After the order adjustment, the final sample text material is obtained.
[0077] Optionally, the sample text material is generated to meet at least one of the following conditions: avoids containing sensitive information or illegal content of the case; covers key pronunciation features appearing in the evidence text; and has a text length and sentence complexity suitable for the target audience to read aloud or repeat.
[0078] In this embodiment, by intelligently analyzing the text content of the audio sample, key features are automatically extracted and combined with the target subject's own verbal evidence to generate highly customized sample text. In particular, for pronunciation features, the "insertion of draft + optional reordering" method can naturally guide the target subject to speak sentences containing specific pronunciation elements, improving the success rate and efficiency of collecting qualified sample audio and solving the problem of blind and inefficient manual writing of sample text.
[0079] Step S4: Obtain sample speech data and the identity information of the target object based on the sample text material; As an optional implementation method, the target object is guided to speak based on the sample text material corresponding to the guiding text, and the obtained initial sample speech is verified through multi-parameter speech quality to obtain effective sample speech data. This ensures that the sample speech meets the professional identification requirements in terms of content relevance, acoustic quality and judicial compliance, providing a reliable data foundation for subsequent accurate voiceprint comparison and reducing identification failures caused by sample quality issues.
[0080] For example, during actual data collection, the target subject can be guided to speak through questioning, repetition, or reading aloud, while the speech is recorded simultaneously. After uploading the initial sample speech, multi-parameter quality verification is performed. Verification criteria include: effective duration ≥ 120 seconds, signal-to-noise ratio ≥ 50dB, average energy ≥ -25dB, clipping ratio ≤ 10%, one speaker, and speech distortion indices: frequency distortion (FD) ≤ 20 and clipping distortion (CILP) ≤ 50. Speech that passes verification is confirmed as valid sample speech data.
[0081] Step S5: Based on the voice data of the evidence, the voice data of the sample, and the identity information of the target object, automatically generate a structured voiceprint identification submission form.
[0082] In this embodiment, the final stage of the voiceprint identification submission process in judicial practice suffers from cumbersome paperwork, high error rates, and insufficient standardization. After completing the preprocessing, quality screening, and standardized collection and verification of the voice samples, the system can automatically integrate key data and information generated throughout the entire process and intelligently generate standardized submission forms that meet the requirements of different identification institutions, thereby improving the work efficiency of frontline case handlers.
[0083] As an optional implementation, based on the user-selected submitting institution or case type, a corresponding standard submission form template is matched and called from a pre-set template library; the template library includes structured templates that conform to the document specifications of different appraisal institutions; the voice data of the evidence, the voice data of the sample, and the identity information are automatically mapped and filled according to the field definitions of the submission form template to obtain a filled template; the filled template is used to generate the voiceprint identification submission form in a standard electronic document format.
[0084] Optionally, the standard submission form templates cover a standardized submission form template library for judicial appraisal institutions at all levels. Each template strictly adheres to the document format specifications, field definitions, and layout requirements of the corresponding institution. When a user selects the target institution or specifies the case type, the corresponding standard template can be precisely matched and retrieved from the template library according to preset mapping rules. The templates adopt a structured design, with each field clearly labeled with its data type and fill rules.
[0085] For example, for the audio data of the evidence, its file metadata and quality assessment summary are automatically extracted and filled into the "Evidence Description" and "Technical Parameters" fields; for the audio data of the samples, the collection verification information is extracted and filled into the "Sample Information" and "Collection Status Description" fields; for the identity information of the target object, it is mapped to relevant fields such as "Information of the Person Being Examined". The entire mapping process automatically handles special format conversions such as timestamp formatting and hash value segmentation, and the filled content is formatted in a standardized manner and meets the readability requirements of judicial documents.
[0086] The technical solutions in the above embodiments of the present invention have at least the following technical effects or advantages: Through automated separation and clustering, and multi-dimensional quality assessment, the chaotic raw speech is transformed into clean, single-speaker, and clearly labeled structured, usable speech fragments for evidence. Guided text materials are then generated based on the text content of these structured speech fragments, ensuring high comparability of samples collected by investigators without specialized skills. Finally, during voiceprint comparison analysis, a multi-model fusion and parallel computing strategy is employed to automatically complete voiceprint feature extraction, similarity calculation, and threshold determination, outputting an expert opinion with quantified confidence levels. This approach significantly lowers the professional threshold while ensuring the rigor of judicial procedures, providing efficient and reliable technical support for real-world scenarios requiring rapid response, such as customs anti-smuggling operations.
[0087] Based on the same inventive concept, embodiments of the present invention also provide a system corresponding to the methods in the above embodiments.
[0088] Reference Figure 6 A system for generating submission forms for voiceprint identification, comprising: The data acquisition and preprocessing module is used to acquire initial voice data, perform multi-dimensional quality assessment on the initial voice data, and obtain sample voice data that meets the preset quality threshold. The speech translation module is used to perform speech recognition on the speech data of the evidence to obtain the corresponding text content of the evidence; The text generation module is used to automatically generate sample text materials to guide the target object to speak based on the text content of the sample. The sample acquisition and verification module is used to acquire sample speech data and the identity information of the target object based on the sample text material. The submission module is used to automatically generate a structured voiceprint identification submission form based on the voice data of the evidence, the voice data of the sample, and the identity information of the target object.
[0089] In this embodiment, the voiceprint identification submission form generation system solves the problems of cumbersome procedures, complex operations for front-line case handlers, and low efficiency in voiceprint identification. Through the collaborative work of the above modules, the complex identification process that requires multiple professionals, multiple sets of software tools, and multiple manual operations is integrated into a standardized automated processing chain.
[0090] Optionally, the data acquisition and preprocessing module is responsible for receiving and initially organizing all input data, greatly simplifying the operational burden on frontline case handlers. This module automatically performs voiceprint feature clustering analysis on the uploaded evidence audio data, automatically separating, classifying, and extracting speech segments from different speakers in mixed audio, laying the foundation for subsequent single-speaker voice processing.
[0091] Optionally, in addition to performing multi-dimensional quality assessment on the initial speech data to obtain sample speech data that meets the preset quality threshold, the data acquisition and preprocessing module can also perform separation and clustering processing on the initial speech data to distinguish and extract speech segments belonging to different speakers.
[0092] Optionally, the data acquisition and preprocessing module also has a retrieval function, that is, to perform exact matching or fuzzy matching based on the words entered by the user to find speech information containing these words or related to these words.
[0093] In addition, the data acquisition and preprocessing module also includes an instruction evaluation function to intelligently control the quality of the separated individual voice recordings. It uses multiple acoustic dimensions, such as signal-to-noise ratio, voice activity, silence ratio, and dynamic range, to quantitatively score the voice recordings and automatically filter out clear and usable sample voice data.
[0094] Optionally, the speech-to-text module automatically converts the selected high-quality speech samples into corresponding text content, i.e., the text content of the samples. It optimizes for common terms used in the judicial field to improve transcription accuracy.
[0095] Optionally, the text generation module can extract key pronunciation features through linguistic analysis based on the translated text content of the evidence, and intelligently reorganize it to generate a sample text material. This text is highly related to the evidence in terms of pronunciation features, but has been semantically anonymized to prevent information leakage that may occur if the person being questioned is not the exact suspect. This sample text material is used to standardize and guide subsequent sample collection work.
[0096] Optionally, the sample collection and verification module is responsible for processing the sample audio. This module can display guiding text output by the text generation module to frontline case handlers, instructing them to complete sample collection using standardized methods such as interrogation, repetition, or reading aloud. In addition, this module can perform a series of preset and rigorous quality checks on the uploaded sample audio, such as effective duration, signal-to-noise ratio, and speaker uniqueness.
[0097] Optionally, the submission module is responsible for intelligently integrating and formatting the technical data generated throughout the voiceprint identification process, such as the voice information of the evidence and the voice information of the sample, with case management information, such as the identity of the target. This module automatically maps and fills standardized fields with the quality parameters of the evidence, the collection and verification records of the sample, and the personnel identity information by calling a pre-built form template library that conforms to the standards of different identification institutions. Finally, it generates a rigorously formatted and complete voiceprint identification submission form or form with one click, and can package related electronic evidence into a complete submission material package that can be directly submitted, solving the problems of low efficiency and error-proneness of manual form filling.
[0098] Optionally, the voiceprint identification submission form generation system also includes a sample screening module, which is used to retrieve the sample voice data based on the received keywords, locate and highlight the position of the keywords in the sample voice data, and associate the location with the time point in the original voice data and the context position in the chat sequence. The evidence screening module features text-based search functionality, allowing users to search for audio recordings based on input keywords. Frontline investigators can quickly locate precise audio segments mentioning specific content within the audio recordings by entering keywords, and can also pinpoint the chat location of that audio segment, significantly improving the efficiency of data analysis and evidence review.
[0099] In addition, to ensure the integrity, authenticity and traceability of data during the voiceprint identification process, a hash value verification mechanism is introduced in the key links of data transmission and storage to ensure the credibility of the entire chain from data entry to the generation of identification conclusions.
[0100] For example, when a voice file containing evidence or a sample is uploaded through the system interface, the system automatically uses the SHA-256 algorithm to calculate the hash value of the file. The calculated hash value, along with the file's metadata, such as filename, size, upload time, and uploader, is stored in the database and permanently bound to subsequent case records.
[0101] Optionally, the data acquisition and preprocessing module can calculate the hash value of each independent speaker speech segment file generated after separation and clustering; the quality assessment and screening module can calculate the hash value of the screened speech data files of the specimen to be processed; the speech translation module can calculate the hash value of the generated text content file of the specimen; and the feature comparison and analysis module can calculate the hash value of the finally generated voiceprint feature vector file and identification report file.
[0102] The series of hash values generated during the above process will be linked in the order of processing to form an immutable data processing hash chain. When a processed data file is accessed in any subsequent stage, the system recalculates its current hash value and compares it with the original hash value stored in the database. If they do not match, an immediate alert is issued, indicating that the file may be corrupted or tampered with. The hash chain can also be exported and attached to the authentication report for third-party organizations to use standard tools to independently verify the integrity of data files at any stage.
[0103] In this embodiment, by constructing a standardized voiceprint identification workflow, the traditional complex identification process, which relies on multiple professionals, multiple sets of tools, and multiple steps of manual operation, is integrated into an automated service that can be triggered by front-line case handlers simply by uploading evidence and sample voice recordings and filling in basic information. The system's embedded intelligent preprocessing (speaker separation, quality screening), content-guided generation, multi-modal feature comparison, and end-to-end hash verification mechanisms greatly reduce the operational threshold and human error while ensuring the standardization of the identification process, the objectivity of the conclusions, and the judicial credibility of the data. Ultimately, this achieves a transformation in voiceprint identification from expert experience-driven to system-intelligent driven, significantly improving the efficiency of voiceprint identification and evidence processing capabilities.
[0104] Since the voiceprint identification submission form generation system described in this embodiment of the invention is the system used to implement the voiceprint identification submission form generation method of this invention, those skilled in the art can understand the specific structure and variations of this system based on the method described in this embodiment of the invention, and therefore will not be repeated here. All systems used in the voiceprint identification submission form generation method of this invention fall within the scope of protection of this invention.
[0105] Based on the same inventive concept, embodiments of the present invention also provide a voiceprint identification method associated with the above embodiments.
[0106] Reference Figure 7 In this embodiment, the voiceprint identification method includes: Step S6: Receive the voiceprint identification submission form and the voice data packet associated with the voiceprint identification submission form, wherein the voice data packet includes the voice data of the sample and the voice data of the specimen. In this embodiment, after receiving the voiceprint identification submission form and its associated voice data packet sent by the voiceprint identification submission form generation system as described in the previous embodiment, voiceprint features are extracted from the voice data of the evidence and the voice data of the sample, respectively, and comparative analysis is performed to generate voiceprint identification results.
[0107] Step S7: Extract the first voiceprint feature set from the evidence speech data and extract the second voiceprint feature set from the sample speech data; Optionally, a multi-model fusion intelligent comparison and dynamic decision-making mechanism can be used to achieve a quantifiable identification of whether the voice samples and the evidence come from the same speaker.
[0108] For example, a pre-trained deep neural network model (such as a model based on TDNN or ResNet architecture) is used to extract high-level, abstract deep features from speech. These features can characterize the speaker's unique pronunciation habits and prosodic patterns. Simultaneously, a pre-defined acoustic model is used to extract stable spectral features from the speech spectrum, which reflect the physical structural characteristics of the sound. Both the first and second voiceprint feature sets contain these two complementary types of features, forming the basis for the comparison.
[0109] Step S8: Calculate the similarity score between the first voiceprint feature set and the second voiceprint feature set; Optionally, a preset similarity measurement algorithm (such as cosine similarity, probabilistic linear discriminant analysis scoring, etc.) is used to calculate the comprehensive similarity score between the first voiceprint feature set and the second voiceprint feature set. This score is a quantitative value used to characterize the probability that the two speech segments come from the same speaker.
[0110] Step S9: Dynamically determine the judgment threshold based on the quality assessment information associated in the voiceprint identification submission form; Optionally, the voiceprint identification submission form generated by the voiceprint identification submission form generation system includes quality assessment information in specific fields. This quality assessment information is generated after an automated multi-dimensional quality assessment of the voice sample. This information includes at least quantitative scores such as signal-to-noise ratio and voice activity.
[0111] Optionally, based on quality assessment information, the corresponding judgment threshold can be dynamically selected or calculated from a preset threshold mapping table. For example, a stricter threshold, such as 0.85, can be used for high-quality samples; while a relatively lenient threshold, such as 0.70, can be used for low-quality samples. This setting allows the identification criteria to adapt to the objective conditions of the samples, improving the feasibility of the identification.
[0112] Step S10: Compare the similarity score with the dynamically determined judgment threshold to generate a preliminary identification opinion; Optionally, if the similarity score is greater than or equal to the judgment threshold, a preliminary identification opinion of "tentatively identifying as identical" is generated. If the similarity score is less than the judgment threshold, a preliminary identification opinion of "tentatively denying identity" is generated.
[0113] Step S11: calibrate the confidence level of the preliminary identification opinion and output the voiceprint identification result containing the identification opinion and the quantified confidence level.
[0114] In this embodiment, the voiceprint identification result includes at least the final identification opinion, such as "the voice in the sample and the voice in the evidence are from the same person", "the voice in the sample and the voice in the evidence are not from the same person", or "cannot be determined", as well as the corresponding quantitative confidence level, and automatically generates a voiceprint identification report that conforms to the standard.
[0115] Optionally, in order to improve the reliability of the identification conclusion, the confidence level of the above preliminary identification opinion can be calibrated.
[0116] Optionally, the calibration process may consider score consistency and speech quality in the speech data packet. Score consistency refers to the similarity scores calculated from different dimensions such as pronunciation patterns and sound spectral structure features. Whether their indices are consistent, a high degree of consistency will result in a higher internal credibility score.
[0117] For example, in a voiceprint comparison, the following intermediate data was obtained: the similarity score calculated from the pronunciation pattern dimension was 0.88, and the score calculated from the sound spectrum structure feature dimension was 0.82. Both scores are higher than the judgment threshold of 0.75 and are close in value. Therefore, the judgment feature conclusions are highly consistent, and a consistency score of 0.9 is assigned. At the same time, the signal-to-noise ratio of the pre-assessed voice sample was 35dB (which is considered good), and a quality reliability score of 0.8 was assigned. Finally, the consistency score and the quality reliability score were combined according to a preset weight ratio, and the overall confidence level of this identification opinion was output as 85%. This process ensures that the voiceprint identification results have objective credibility.
[0118] In this embodiment, by receiving a structured voiceprint identification submission form and voice data packets, the identification process is automatically triggered and data flows seamlessly. Utilizing multimodal features and a dynamic threshold mechanism based on prior quality assessment improves the accuracy, objectivity, and automation level of voiceprint identification under complex conditions.
[0119] In this embodiment of the invention, a device for generating a submission form for voiceprint identification is proposed.
[0120] Reference Figure 8 , Figure 8 This is a schematic diagram of the terminal structure of the hardware operating environment involved in an embodiment of the present invention.
[0121] like Figure 8As shown, the control terminal may include: a processor 1001, such as a CPU, a network interface 1003, a memory 1004, and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The network interface 1003 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1004 may be high-speed RAM or stable non-volatile memory, such as disk storage. Alternatively, the memory 1004 may be a storage device independent of the aforementioned processor 1001.
[0122] Those skilled in the art will understand that Figure 8 The terminal structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0123] like Figure 8 As shown, the memory 1004, which serves as a computer storage medium, may include an operating system, a network communication module, a program for generating voiceprint identification submission forms, and a voiceprint identification program.
[0124] exist Figure 8 In the hardware structure of the device for generating the voiceprint identification submission form shown, the processor 1001 can call the voiceprint identification submission form generation program stored in the memory 1004 and perform the following operations: Acquire initial voice data, perform multi-dimensional quality assessment on the initial voice data, and obtain sample voice data that meets the preset quality threshold; The speech data of the evidence is subjected to speech recognition to obtain the corresponding text content of the evidence; Based on the text content of the sample, sample text material is automatically generated to guide the target object to speak; Obtain sample speech data and the identity information of the target object based on the sample text material; Based on the voice data of the evidence, the voice data of the sample, and the identity information of the target object, a structured voiceprint identification submission form is automatically generated.
[0125] Optionally, the processor 1001 may call the voiceprint identification submission form generation program stored in the memory 1004, and further perform the following operations: The initial speech data is segmented into multiple consecutive speech analysis segments, wherein there is a temporal overlap between adjacent speech analysis segments; For each of the aforementioned speech analysis segments, quality assessment processing is performed in parallel.
[0126] Optionally, the processor 1001 may call the voiceprint identification submission form generation program stored in the memory 1004, and further perform the following operations: The signal of a predetermined duration at the beginning of the initial speech data is extracted as the background noise benchmark. The speech analysis segment is divided into multiple analysis frames. The average energy ratio of the speech frame to the silence frame is calculated and converted into a decibel value as the signal-to-noise ratio score. The proportion of active speech frames in the total number of analyzed frames is used as the speech activity score. The proportion of silent frames in the total number of analyzed frames in the speech analysis segment is counted and used as the silent proportion score. Calculate the difference between the maximum and minimum amplitudes of the speech analysis segment signal and convert it into a decibel value as a dynamic range score; Based on predefined weighting coefficients, the signal-to-noise ratio score, speech activity score, silence percentage score, and dynamic range score are weighted and summed to obtain the segment quality score of the speech analysis segment.
[0127] Optionally, the processor 1001 may call the voiceprint identification submission form generation program stored in the memory 1004, and further perform the following operations: Based on the segment quality scores of each speech analysis segment, the overall quality assessment result of the initial speech data is determined, and the speech segments that meet the preset quality threshold are selected or marked according to the overall quality assessment result as the evidence speech data. The methods for determining the overall quality assessment results include at least one of the following: The segment quality scores of each speech analysis segment are weighted and averaged, and the average value is used as the overall quality score. The score of the speech analysis segment with the highest quality score is selected as the representative of the overall quality score; Based on the distribution of quality scores for each segment, the initial speech data is labeled into multiple quality level regions.
[0128] Optionally, the processor 1001 may call the voiceprint identification submission form generation program stored in the memory 1004, and further perform the following operations: Obtain the context information associated with the initial speech data, and preliminarily associate candidate speaker identifiers with the initial speech data based on the context information; the context information includes acoustic feature information and non-acoustic feature information, and the non-acoustic feature information includes account identifier information and speech generation time information; Based on the candidate speaker identifiers, speaker separation and clustering are performed on the initial speech data; Based on the results of separation and clustering, short-term speech segments with similar acoustic features are grouped to their respective speakers, and speech segments belonging to the same speaker are spliced and labeled to form speaker speech segments with temporal information. Each speaker's speech segment separated is used as the initial speech data for subsequent processing.
[0129] Optionally, the processor 1001 may call the voiceprint identification submission form generation program stored in the memory 1004, and further perform the following operations: The text content of the sample was analyzed to obtain key features used to guide vocalization; Obtain textual evidence of verbal evidence associated with the target object; Based on the key features, the verbal evidence text is processed to automatically generate the polished sample text.
[0130] Optionally, the processor 1001 may call the voiceprint identification submission form generation program stored in the memory 1004, and further perform the following operations: If the key feature is a semantic feature, then extract the sentences that are semantically related to the key feature from the verbal evidence text material and reorganize the sentences to obtain the sample text material; If the key feature is a pronunciation feature, or if there is no sentence in the verbal evidence text material that is directly related to the key feature, then the key feature is inserted into the verbal evidence text material to obtain the sample text material.
[0131] Optionally, the processor 1001 may call the voiceprint identification submission form generation program stored in the memory 1004, and further perform the following operations: The key features are inserted into the verbal evidence text material to obtain intermediate text material; The sentence order of the intermediate text material is adjusted to obtain the sample text material.
[0132] Optionally, the processor 1001 may call the voiceprint identification program stored in the memory 1004 and further perform the following operations: Based on the user's selected submission institution or case type, the system matches and calls the corresponding standard submission form template from a pre-set template library; the template library includes structured templates that conform to the document specifications of different appraisal institutions. The voice data of the evidence, the voice data of the sample, and the identity information are automatically mapped and filled according to the field definitions of the submission form template to generate the voiceprint identification submission form.
[0133] Optionally, the processor 1001 may call the voiceprint identification program stored in the memory 1004 and further perform the following operations: Receive a voiceprint identification submission form and a voice data packet associated with the voiceprint identification submission form, wherein the voice data packet includes the voice data of the sample and the voice data of the specimen. Extract a first set of voiceprint features from the speech data of the evidence, and extract a second set of voiceprint features from the speech data of the sample; Calculate the similarity score between the first voiceprint feature set and the second voiceprint feature set; Based on the quality assessment information associated with the voiceprint identification submission form, the judgment threshold is dynamically determined. The similarity score is compared with the dynamically determined judgment threshold to generate a preliminary identification opinion; The confidence level of the preliminary identification opinion is calibrated, and the voiceprint identification result containing the identification opinion and the quantified confidence level is output.
[0134] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0136] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0137] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0138] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, third, etc., does not indicate any order. These words can be interpreted as names.
[0139] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0140] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for generating a submission form for voiceprint identification, characterized in that, The method includes: Acquire initial voice data, perform multi-dimensional quality assessment on the initial voice data, and obtain sample voice data that meets the preset quality threshold; The speech data of the evidence is subjected to speech recognition to obtain the corresponding text content of the evidence; Based on the text content of the sample, sample text material is automatically generated to guide the target object to speak; Obtain sample speech data and the identity information of the target object based on the sample text material; Based on the voice data of the evidence, the voice data of the sample, and the identity information of the target object, a structured voiceprint identification submission form is automatically generated.
2. The method as described in claim 1, characterized in that, The step of performing a multi-dimensional quality assessment on the initial speech data includes: The initial speech data is segmented into multiple consecutive speech analysis segments, wherein there is a temporal overlap between adjacent speech analysis segments; For each of the aforementioned speech analysis segments, quality assessment processing is performed in parallel.
3. The method as described in claim 2, characterized in that, The step of performing quality assessment processing in parallel for each of the speech analysis segments includes: The signal of a predetermined duration at the beginning of the initial speech data is extracted as the background noise benchmark. The speech analysis segment is divided into multiple analysis frames. The average energy ratio of the speech frame to the silence frame is calculated and converted into a decibel value as the signal-to-noise ratio score. The proportion of active speech frames in the total number of analyzed frames is used as the speech activity score. The proportion of silent frames in the total number of analyzed frames in the speech analysis segment is counted and used as the silent proportion score. The difference between the maximum and minimum amplitudes of the speech analysis segment signal is calculated and converted into a decibel value as a dynamic range score. Based on predefined weighting coefficients, the signal-to-noise ratio score, speech activity score, silence percentage score, and dynamic range score are weighted and summed to obtain the segment quality score of the speech analysis segment.
4. The method as described in claim 2 or 3, characterized in that, The step of performing multi-dimensional quality assessment on the initial speech data to obtain evidence speech data that meets a preset quality threshold includes: Based on the segment quality scores of each speech analysis segment, the overall quality assessment result of the initial speech data is determined, and the speech segments that meet the preset quality threshold are selected or marked according to the overall quality assessment result as the evidence speech data. The methods for determining the overall quality assessment results include at least one of the following: The segment quality scores of each speech analysis segment are weighted and averaged, and the average value is used as the overall quality score. The score of the speech analysis segment with the highest quality score is selected as the representative of the overall quality score; Based on the distribution of quality scores for each segment, the initial speech data is labeled into multiple quality level regions.
5. The method as described in claim 1 or 2, characterized in that, Prior to the step of performing a multi-dimensional quality assessment on the initial speech data, the following steps are included: Obtain the context information associated with the initial speech data, and preliminarily associate candidate speaker identifiers with the initial speech data based on the context information; the context information includes acoustic feature information and non-acoustic feature information, and the non-acoustic feature information includes account identifier information and speech generation time information; Based on the candidate speaker identifiers, speaker separation and clustering are performed on the initial speech data; Based on the results of separation and clustering, short-term speech segments with similar acoustic features are grouped to their respective speakers, and speech segments belonging to the same speaker are spliced and labeled to form speaker speech segments with temporal information. Each speaker's speech segment separated is used as the initial speech data for subsequent processing.
6. The method as described in claim 1, characterized in that, The step of automatically generating sample text material to guide the target object to speak based on the text content of the sample includes: The text content of the sample was analyzed to obtain key features used to guide vocalization; Obtain textual evidence of verbal evidence associated with the target object; Based on the key features, the verbal evidence text is processed to automatically generate the polished sample text.
7. The method as described in claim 6, characterized in that, The step of processing the verbal evidence text material based on the key features to automatically generate the polished sample text material includes: If the key feature is a semantic feature, then extract the sentences that are semantically related to the key feature from the verbal evidence text material and reorganize the sentences to obtain the sample text material; If the key feature is a pronunciation feature, or if there is no sentence in the verbal evidence text material that is directly related to the key feature, then the key feature is inserted into the verbal evidence text material to obtain the sample text material.
8. The method as described in claim 7, characterized in that, The step of embedding the key features as new content into the verbal evidence text material to obtain the sample text material includes: The key features are inserted into the verbal evidence text material to obtain intermediate text material; The sentence order of the intermediate text material is adjusted to obtain the sample text material.
9. The method as described in claim 1 or 6, characterized in that, The step of automatically generating a structured voiceprint identification submission form based on the voice data of the evidence, the voice data of the sample, and the identity information of the target object includes: Based on the user's selected submission institution or case type, the system matches and calls the corresponding standard submission form template from a pre-set template library; the template library includes structured templates that conform to the document specifications of different appraisal institutions. The voice data of the evidence, the voice data of the sample, and the identity information of the target object are automatically mapped and filled according to the field definitions of the submission form template to generate the voiceprint identification submission form.
10. A system for generating submission forms for voiceprint identification, characterized in that, The system includes: The data acquisition and preprocessing module is used to acquire initial voice data, perform multi-dimensional quality assessment on the initial voice data, and obtain sample voice data that meets the preset quality threshold. The speech translation module is used to perform speech recognition on the speech data of the evidence to obtain the corresponding text content of the evidence; The text generation module is used to automatically generate sample text materials to guide the target object to speak based on the text content of the sample. The sample acquisition and verification module is used to acquire sample speech data and the identity information of the target object based on the sample text material. The submission module is used to automatically generate a structured voiceprint identification submission form based on the voice data of the evidence, the voice data of the sample, and the identity information of the target object.
11. A method for voiceprint identification, characterized in that, The method includes: Receive a voiceprint identification submission form and a voice data packet associated with the voiceprint identification submission form, wherein the voice data packet includes the voice data of the sample and the voice data of the specimen. Extract a first set of voiceprint features from the speech data of the evidence, and extract a second set of voiceprint features from the speech data of the sample; Calculate the similarity score between the first voiceprint feature set and the second voiceprint feature set; Based on the quality assessment information associated with the voiceprint identification submission form, the judgment threshold is dynamically determined. The similarity score is compared with the dynamically determined judgment threshold to generate a preliminary identification opinion; The confidence level of the preliminary identification opinion is calibrated, and the voiceprint identification result containing the identification opinion and the quantified confidence level is output.