A voiceprint recognition evaluation data construction method and system based on a large language model double-layer automatic labeling
Patent Information
- Application Number
- CN202611031517.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本发明旨在提供一种基于大语言模型双层自动标注的声纹识别评估数据构建方法及系统,旨在解决现有声纹识别评估数据构建方案中无法兼顾音频端到端理解与双视角主观体验模拟的问题,即如何在消除人工依赖的同时,实现声纹识别结果客观正确性标注与不同身份用户主观体感标注的自动化、精准化协同构建
本发明通过构建双层异构模型协同标注架构,利用多模态大语言模型直接进行端到端的音频对比以获取客观标注结果,并通过格式校验与重试机制保障数据完整性,进而结合客观结果推断说话人身份以动态切换注册与非注册用户的双视角提示词模板,驱动文本大语言模型执行差异化主观体感推理,从而在消除人工依赖的同时,实现了客观正确性与主观体验的双重自动化精准评估,显著提升了声纹识别评估数据的构建效率与评价真实性。
Smart Images

Figure CN122842618A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and system for constructing voiceprint recognition evaluation data based on a large language model with two-layer automatic annotation. Background Technology
[0002] In the field of voiceprint recognition evaluation, constructing a high-quality evaluation dataset is a crucial step in measuring model performance. Existing voiceprint recognition evaluation data construction schemes mainly suffer from the following limitations: First, general-purpose multimodal annotation platforms (such as human-in-the-loop systems based on crowdsourcing, professional, and expert three-tier architectures) heavily rely on manual listening judgment, which is not only costly and inefficient but also difficult to handle the demands of large-scale data annotation. Second, existing dialogue annotation methods based on large language models typically only use a frozen audio base model combined with a large text model to add speaker attribute labels (such as age and gender) to the dialogue, or have the large language model directly compare voiceprint embedding vectors for speaker verification. These schemes lack end-to-end auditory understanding capabilities for audio signals and do not design specific annotation strategies for voiceprint scenarios. Third, existing LLM-assisted subjective annotation and error correction schemes mainly focus on improving the quality of existing text labels and cannot solve the problem of constructing voiceprint evaluation data from scratch that includes both objective correctness and subjective experience dimensions.
[0003] Existing technologies fail to effectively distinguish and process two completely different types of tasks—objective auditory judgment and subjective experiential judgment—within a unified process, making it difficult for evaluation data to accurately reflect the real differences in sensory perception between users with different identities (such as registered users and unregistered users). Summary of the Invention
[0004] This invention aims to provide a method and system for constructing voiceprint recognition evaluation data based on a large language model with two-layer automatic annotation. It aims to solve the problem that existing voiceprint recognition evaluation data construction schemes cannot simultaneously take into account end-to-end audio understanding and dual-perspective subjective experience simulation. That is, how to achieve automated and accurate collaborative construction of objective correctness annotation of voiceprint recognition results and subjective feeling annotation of users with different identities while eliminating human dependence.
[0005] To achieve the above objectives, the technical solution adopted by this invention is: a method for constructing voiceprint recognition evaluation data based on a large language model with dual-layer automatic annotation, comprising: The audio to be labeled, the dialogue text data, and the voiceprint reference sample of the registered user are obtained. The audio to be labeled is processed by voice activity detection to remove silent segments. A multimodal request message is constructed, and the registered user's voiceprint reference sample and the preprocessed audio to be labeled are input into the multimodal large language model. The first layer of objective auditory judgment and reasoning is executed to obtain intermediate data containing objective labeling results. The objective annotation results in the intermediate data are subjected to format validation and line count consistency checks. If the validation fails, a retry mechanism is triggered to re-call the multimodal large language model to generate new objective annotation results until the validation passes. Based on the verified objective annotation results and the original voiceprint determination information, the speaker's true identity is inferred. Based on the inferred true identity, the corresponding perspective prompt word template is selected to construct the second layer of subjective sensory judgment request. The constructed second-layer subjective sensory perception judgment request is input into the text large language model, and the second-layer subjective sensory perception judgment reasoning is executed using the selected perspective cue word template, and the final subjective annotation result is output.
[0006] Preferably, constructing the multimodal request message includes: The registered user's voiceprint reference sample is used as the baseline audio, and the preprocessed audio to be labeled is used as the audio to be tested. The multimodal large language model is also given prompts that require auditory input as the primary factor and confidence as the secondary factor, thus forming a multimodal input vector that includes audio features and semantic logic.
[0007] Preferably, the retry mechanism re-invoking the multimodal large language model includes: A maximum retry threshold is set. When the number of output rows of the objective annotation results in the intermediate data is inconsistent with the number of input audios or the JSON format is invalid, the request parameters are automatically reset and the multimodal large language model is called again for inference. This process is repeated until the objective annotation results that meet the format requirements are obtained or the maximum retry limit is reached.
[0008] Preferably, the inference of the speaker's true identity includes: Comparing the original voiceprint determination result with the objective annotation result, if the original voiceprint is determined to be that of a non-registered user but the objective annotation result is true, or if the original voiceprint is determined to be that of a registered user but the objective annotation result is false, then the current speaker is determined to be a non-registered user; otherwise, the speaker is determined to be a registered user.
[0009] Preferably, the step of selecting the corresponding perspective prompt template includes: If the deduced true identity is a registered user, then select the registered user's perspective prompt template to determine whether the system treats the current user as a registered user; If the inferred true identity is a non-registered user, then select the non-registered user perspective prompt template to determine whether the system is mistakenly treating the current user as a registered user.
[0010] Preferably, the execution of the second-level subjective sensory judgment and reasoning includes: By using a large textual language model to read dialogue records in the historical context window, and combining them with selected perspective cue word templates, the system responses are analyzed to determine whether there are identity misalignment signals, and the final subjective annotation results are determined based on preset conflict resolution rules.
[0011] Preferably, the conflict resolution rules include: When a clear signal of identity mismatch appears, it is preferentially judged as an incorrect state; For ambiguous scenarios, the authenticity of subjective annotation results is determined by whether the system uses exclusive terminology or the degree of context mismatch.
[0012] Preferably, the method further includes S6: assembling the objective annotation results and the subjective annotation results to generate a complete evaluation dataset containing ASR text, voiceprint determination, confidence level, objective dimension labels and subjective dimension labels.
[0013] Preferably, the reasoning process of the multimodal large language model adopts an end-to-end audio comparison method, directly judging the objectivity and correctness of the voiceprint recognition result based on the audio waveform features.
[0014] On the other hand, this invention proposes a voiceprint recognition evaluation data construction system based on a large language model with two-layer automatic annotation, comprising: The input data preparation module is used to obtain the audio to be labeled, dialogue text data, and registered user voiceprint reference samples, and to perform speech activity detection processing on the audio to be labeled to remove silent segments. The first annotation layer module is used to construct a multimodal request message, input the registered user's voiceprint reference sample and the preprocessed audio to be annotated into the multimodal large language model, execute the first layer of objective auditory judgment and reasoning, obtain intermediate data containing objective annotation results, and perform format verification and line number consistency checks on the objective annotation results in the intermediate data. If the verification fails, a retry mechanism is triggered to call the multimodal large language model again to generate new objective annotation results until the verification passes. The identity inference and perspective selection module is used to infer the speaker's true identity based on the verified objective annotation results and the original voiceprint judgment information. Based on the inferred true identity, the module selects the corresponding perspective prompt word template and constructs the second layer of subjective sensory judgment request. The second annotation layer module is used to input the constructed second-layer subjective somatosensory judgment request into the text big language model, perform the second-layer subjective somatosensory judgment reasoning using the selected perspective prompt word template, and output the final subjective annotation result; The annotation output module is used to assemble objective annotation results with subjective annotation results to generate a complete evaluation dataset containing ASR text, voiceprint judgment, confidence score, objective dimension labels and subjective dimension labels.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention constructs a two-layer heterogeneous model collaborative annotation architecture, utilizes a multimodal large language model to directly perform end-to-end audio comparison to obtain objective annotation results, and ensures data integrity through format verification and retry mechanisms. Then, it combines the objective results to infer the speaker's identity and dynamically switches between dual-perspective prompt word templates for registered and unregistered users, driving the text large language model to perform differentiated subjective sensory reasoning. Thus, while eliminating human dependence, it achieves dual automated and accurate evaluation of objective correctness and subjective experience, significantly improving the construction efficiency and evaluation authenticity of voiceprint recognition evaluation data. Attached Figure Description
[0016] Figure 1 This is a flowchart of the speaker recognition evaluation data construction method based on a large language model with two-layer automatic annotation, according to the present invention. Figure 2 This is a block diagram of the voiceprint recognition evaluation data construction system based on a large language model and two-layer automatic annotation, as per the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0018] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a specific posture. If the specific posture changes, the directional indicators will also change accordingly.
[0019] Furthermore, if the embodiments of this invention involve descriptions of "first," "second," etc., such descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the use of "and / or" or "and / or" throughout the text implies three parallel solutions. For example, A and / or B includes solution A, solution B, or a solution that simultaneously satisfies A and B. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0020] Please see Figure 1 This invention provides a method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model, comprising: We acquire audio data to be labeled, dialogue text data, and voiceprint reference samples of registered users. We then perform speech activity detection processing on the audio to be labeled to remove silent segments. By acquiring multi-source data and performing speech activity detection to remove silent segments, we effectively eliminate invalid noise interference, significantly improve the signal-to-noise ratio and inference efficiency of subsequent multimodal model processing of audio signals, and lay a clean data foundation for building a high-quality voiceprint recognition evaluation dataset.
[0021] Construct a multimodal request message, input the registered user's voiceprint reference sample and the preprocessed audio to be labeled into the multimodal large language model, perform the first layer of objective auditory judgment and reasoning, and obtain intermediate data containing objective labeling results; The process of constructing a multimodal request message includes: using a registered user's voiceprint reference sample as the baseline audio, using the preprocessed audio to be labeled as the audio to be tested, and adding a prompt word instruction that requires the multimodal large language model to prioritize auditory perception and supplement it with confidence, thus forming a multimodal input vector that includes audio features and semantic logic.
[0022] By constructing a multimodal input vector containing benchmark and test audio as well as specific auditory priority instructions, the multimodal large language model is driven to perform end-to-end objective auditory reasoning. This effectively solves the problem of the separation between voiceprint features and semantic logic in traditional solutions, significantly improves the accuracy and robustness of speaker identity judgment, and provides highly reliable objective data support for subsequent subjective sensory evaluation.
[0023] The objective annotation results in the intermediate data are formatted and checked for row count consistency. If the verification fails, a retry mechanism is triggered to re-invoke the multimodal large language model to generate new objective annotation results until the verification passes. The retry mechanism involves re-invoking the multimodal large language model. This includes setting a maximum retry threshold. When the number of output rows of the objective annotation results in the intermediate data is inconsistent with the number of input audio tracks, or when the JSON format is invalid, the request parameters are automatically reset and the multimodal large language model is invoked again for inference. This process is repeated until an objective annotation result that meets the format requirements is obtained or the maximum retry limit is reached. The inference process of the multimodal large language model adopts an end-to-end audio comparison method, directly judging the objective correctness of the voiceprint recognition result based on audio waveform features.
[0024] By introducing format validation and row count consistency checks, as well as an adaptive retry strategy, the risks of unparseable large model generation results or mismatched data dimensions are effectively avoided. This significantly improves the robustness and data integrity of the automated annotation process, ensuring the accuracy and reliability of the objective basic data upon which subsequent subjective sensory judgments and inferences rely.
[0025] Based on the verified objective annotation results and the original voiceprint judgment information, the speaker's true identity is inferred. Based on the inferred true identity, the corresponding perspective prompt word template is selected to construct the second layer of subjective sensory judgment request. The process of inferring the speaker's true identity includes comparing the original voiceprint determination result with the objective annotation result. If the original voiceprint is determined to be that of a non-registered user but the objective annotation result is true, or if the original voiceprint is determined to be that of a registered user but the objective annotation result is false, then the current speaker is determined to be a non-registered user; otherwise, the speaker is determined to be a registered user.
[0026] Furthermore, select the corresponding perspective cue template, including: If the deduced true identity is a registered user, then select the registered user's perspective prompt template to determine whether the system treats the current user as a registered user; If the inferred true identity is a non-registered user, then select the non-registered user perspective prompt template to determine whether the system is mistakenly treating the current user as a registered user.
[0027] By comparing objective annotations with the original judgment results, the speaker's true identity is accurately inferred. Based on this, the dual-perspective prompt word templates for registered and unregistered users are dynamically switched. This effectively simulates the differentiated experience perception of different user groups in the voiceprint recognition system, solves the problem that existing technologies cannot distinguish subjective sensory dimensions, and significantly improves the accuracy of the evaluation data in restoring the user's true sensory experience and the relevance of the evaluation.
[0028] The constructed second-layer subjective sensory perception judgment request is input into the text large language model, and the second-layer subjective sensory perception judgment reasoning is executed using the selected perspective cue word template, and the final subjective annotation result is output.
[0029] The second layer of subjective sensory judgment and reasoning includes: using a large textual language model to read the dialogue records in the historical context window, combining the selected perspective prompt word template, analyzing whether there are identity misalignment signals in the system response, and determining the final subjective annotation result based on the preset conflict resolution rules.
[0030] Specifically, the conflict resolution rules include: when there is a clear signal of identity mismatch, it is preferentially judged as an incorrect state; for ambiguous scenarios, the authenticity of subjective labeling results is determined based on whether the system uses exclusive titles or the degree of context mismatch.
[0031] By leveraging a large textual language model combined with historical context and specific perspective templates, we can deeply analyze identity misalignment signals in system responses. Furthermore, by using pre-defined conflict resolution rules, we can effectively solve the judgment problem in ambiguous scenarios. This significantly improves the ability of subjective sensory evaluation to capture user experience details and the rigor of judgment logic, ensuring that the final annotation results can truly reflect the actual perceived quality of users during voiceprint interaction.
[0032] Furthermore, the above method also includes S6: assembling the objective annotation results with the subjective annotation results to generate a complete evaluation dataset containing ASR text, voiceprint determination, confidence level, objective dimension labels, and subjective dimension labels. By constructing a complete evaluation dataset that integrates ASR text, voiceprint determination, confidence level, and both objective and subjective dimension labels, the structured unification and end-to-end correlation of multimodal perception data are achieved. This not only eliminates data silos, facilitating subsequent multidimensional cross-analysis, but also significantly improves the quantitative accuracy and interpretability of the evaluation system for the overall performance of the voiceprint recognition system.
[0033] On the other hand, this invention proposes a voiceprint recognition evaluation data construction system based on a large language model with two-layer automatic annotation, such as... Figure 2 As shown, it includes: The input data preparation module is used to obtain the audio to be labeled, dialogue text data, and registered user voiceprint reference samples, and to perform speech activity detection processing on the audio to be labeled to remove silent segments. Registered user voiceprint reference sample: The voiceprint reference audio of the registered user serves as the benchmark sample for voiceprint comparison, allowing the multimodal large language model to learn its voice features.
[0034] Audio to be labeled: User speech segments in the dialogue to be judged. Silent segments can be removed first using Voice Activity Detection (VAD) to reduce input length.
[0035] Dialogue text data: includes the ASR transcribed text of this round of dialogue, the dialogue system's response text, the identity determination results and confidence scores output by the voiceprint recognition model.
[0036] In intelligent dialogue systems with voiceprint recognition (speaker identification) capabilities, an evaluation dataset needs to be constructed to assess the performance of the voiceprint model. Each evaluation dataset needs to contain annotations in two dimensions: Objective annotation (voiceprint model accuracy): For each round of dialogue, the identity determination output by the voiceprint recognition model is checked against the speaker's actual identity. This requires comparing the audio with a reference sample of registered users' voiceprints to determine if the speaker is a registered user.
[0037] Subjective annotation (actual user experience): Based on the dialogue system's responses, judge from the user's perspective whether they feel the dialogue system correctly identified them. This requires annotators to read the dialogue context and system responses, simulating the user's feelings.
[0038] This invention proposes a two-layer large language model automatic annotation method: the first layer uses a large language model with multimodal (audio + text) understanding capabilities to complete objective auditory judgment annotation; the second layer uses a text large language model to construct judgment prompt words from two identity perspectives to complete subjective sensory annotation.
[0039] The first annotation layer module is used to construct a multimodal request message, input the registered user's voiceprint reference sample and the preprocessed audio to be annotated into the multimodal large language model, execute the first layer of objective auditory judgment and reasoning, obtain intermediate data containing objective annotation results, and perform format verification and line number consistency checks on the objective annotation results in the intermediate data. If the verification fails, a retry mechanism is triggered to call the multimodal large language model again to generate new objective annotation results until the verification passes. The workflow of the first annotation layer module is as follows: (1) Construct a multimodal annotation request. The message contains two audio files (the registered user's voiceprint reference sample and the audio to be annotated) and annotation prompts.
[0040] (2) The original judgment result and confidence score of the voiceprint recognition system are embedded in the prompt words as reference information, but the model is required to rely mainly on auditory judgment and the confidence score is only used as an auxiliary reference.
[0041] (3) Send the constructed multimodal request to the large language model and obtain the model output (JSONLines format, each line corresponds to the True / False / None judgment of an ASR text).
[0042] (4) Perform parsing and verification on the output. If the number of output lines is inconsistent with the number of input data lines or the format is abnormal, a retry will be triggered (up to a maximum of several retry times). Data lines with a result of None (poor audio quality or empty voiceprint result) are marked as invalid.
[0043] The output of the first annotation layer (objective annotation results) is the input to the second annotation layer. The second annotation layer needs to infer the true identity of the current speaker based on the objective annotation results in order to select the corresponding subjective sensory judgment perspective.
[0044] The identity inference and perspective selection module is used to infer the speaker's true identity based on the verified objective annotation results and the original voiceprint judgment information. Based on the inferred true identity, the module selects the corresponding perspective prompt word template and constructs the second layer of subjective sensory judgment request. The second annotation layer module is used to input the constructed second-layer subjective somatosensory judgment request into the text big language model, perform the second-layer subjective somatosensory judgment reasoning using the selected perspective prompt word template, and output the final subjective annotation result; The workflow of the second annotation layer module is as follows: (1) Real identity inference: Based on the objective annotation results of the first annotation layer and the original judgment results output by the voiceprint recognition system, the real identity of the current speaker is inferred. The rule is: if the voiceprint is determined to be an unregistered user (unknown) and the objective annotation is True, or if the voiceprint is not determined to be an unregistered user and the objective annotation is False, then the current speaker is an unregistered user; otherwise, the current speaker is a registered user.
[0045] (2) Perspective selection: Select the corresponding judgment perspective based on the inferred real identity - if it is a registered user, use the registered user perspective prompt words; if it is a non-registered user, use the non-registered user perspective prompt words.
[0046] (3) Constructing judgment prompts: The two types of prompts define different judgment criteria. The judgment from the perspective of a registered user focuses on whether the dialogue system treats the current speaker as a registered user; the judgment from the perspective of a non-registered user focuses on whether the dialogue system mistakenly treats the current speaker as a registered user.
[0047] (4) Call the text big language model: send the constructed prompt words (including registered user identifier, current user input, dialogue system reply, and historical dialogue context) to the text big language model to obtain the judgment result.
[0048] (5) Conflict resolution: The prompt word contains built-in conflict resolution rules - when there is a clear identity mismatch signal in the reply (such as directly calling a stranger by their exclusive nickname), the wrong signal shall prevail; when there is no obvious identity information, the default judgment is that the feeling is correct.
[0049] The second annotation layer relies on the output of the first annotation layer to determine the true identity and the selected perspective. The annotation results of the two layers together constitute the complete annotation of the voiceprint recognition evaluation dataset.
[0050] The annotation output module is used to assemble the objective annotation results and the subjective annotation results to generate a complete evaluation dataset containing ASR text, voiceprint judgment, confidence score, objective dimension labels (True / False / None) and subjective dimension labels (True / False).
[0051] Furthermore, the workflow of this system is as follows: S1. Obtain the judgment records generated by the voiceprint recognition system during dialogue interaction, including voiceprint recognition results (identity determination, confidence score), ASR transcribed text, and dialogue system response text.
[0052] S2. Perform VAD (Voice Activity Detection) preprocessing on the audio to be labeled to remove silence and invalid speech segments and shorten the audio length to reduce the input overhead of the large language model.
[0053] S3. Obtain a voiceprint reference sample of the registered user.
[0054] S4. Construct a multimodal annotation request message list. Each message contains: (a) a reference sample of the registered user's voiceprint and its descriptive text: This is a reference sample of the registered user's voice; please remember its voice characteristics; (b) the audio to be annotated and annotation prompts. The prompts embed the ASR text to be annotated in the current batch, along with its corresponding voiceprint determination result and confidence level data. The model is required to strictly follow the process of first listening to the reference audio to remember the voice characteristics, then listening to the audio to be annotated to determine the speaker's identity, using the reference confidence level for verification, and finally comparing the comprehensive judgment with the voiceprint result.
[0055] S5. Call the multimodal large language model and send the constructed multimodal message to the model inference interface.
[0056] S6. Validate the output text returned by the model: Extract the annotation results in JSONLines format, check whether the number of output lines matches the number of input data lines, and whether the JSON format is valid. If they do not match, proceed to the retry process S7.
[0057] S7. Verification Closed Loop: During retries, the same multimodal input is used to re-request model inference, leveraging the randomness generated by the large language model to obtain different outputs until verification passes or the retry limit is reached. This closed-loop design ensures that the output of the first annotation layer has passed format and integrity verification before flowing into the second annotation layer, preventing abnormal outputs from contaminating the downstream identity inference and perspective selection logic.
[0058] S8. Record the annotation results that have passed the parsing as True (correct voiceprint determination), False (incorrect voiceprint determination) or None (cannot be determined, such as poor audio quality or empty voiceprint result).
[0059] S9. Infer the current speaker's true identity based on the results of the first annotation layer. The inference logic is as follows: if the voiceprint is determined to be from a non-registered user and the objective annotation is True, or if the voiceprint is determined to be from a registered user and the objective annotation is False, then the current speaker is a non-registered user; otherwise, they are a registered user. Key architectural design points: This identity inference step is a crucial "bridge point" connecting the two layers. It transforms the discrete annotation output (True / False / None) of the first annotation layer into the viewpoint selection signal (registered user / non-registered user) required by the second annotation layer. This allows the data dependency between the two layers to be transmitted only through this simple binary signal, avoiding direct parameter coupling or model-level dependency between the two layers.
[0060] S10. Select the judgment perspective based on the real identity and construct the corresponding judgment prompt words.
[0061] S11. The prompt from the perspective of the registered user requires the model to judge whether the system correctly identifies the current speaker as a registered user from the perspective of the registered user. The condition for judging it as True is that the reply logic is consistent with the registered user interacting with the system and there is no obvious identity conflict.
[0062] S12, Non-registered User Perspective Prompts: The model must judge from the perspective of a non-registered user whether the system has mistakenly identified the current speaker as a registered user. It prioritizes three cases: treating the current speaker as a regular non-registered user, mentioning a registered user but not the current speaker, and genuinely treating the current speaker as a registered user—the first two are judged as True, and the third as False. The branching logic for the dual-perspective: The branch selection in S10 completely decouples the judgment tasks of the two perspectives during inference—in the same set of dialogue data, samples inferred as registered users only follow the prompts from the registered user's perspective, and samples inferred as non-registered users only follow the prompts from the non-registered user's perspective. The two types of prompts independently define their judgment criteria and conflict resolution rules, without interfering with each other.
[0063] S13. Send the prompt words to the text big language model for inference. The prompt words contain built-in conflict resolution rules and examples to ensure that the model makes reasonable judgments in ambiguous scenarios.
[0064] S14. Analyze and resolve conflicts in the judgment results returned by the model, and output the final subjective somatosensory annotation results.
[0065] S15. Merge the objective annotation results of the first annotation layer and the subjective annotation results of the second annotation layer into a complete annotation dataset output.
[0066] The above rules for inferring true identity ; Viewpoint selection rules: ; Output verification conditions: ; That is, the total number of labels output by the second labeling layer should be equal to the number of valid labels (not None) in the first labeling layer. Where, The set of all subjective sensory annotation results output by the second annotation layer; i represents the number of dialogue rounds traversed. It is the set of all objective annotation results output by the first annotation layer.
[0067] In the formula, The original judgment result of the Nth round of voiceprint model is True / False / None; The total number of dialogue rounds to be annotated; For the first The subjective sensory perception annotation results (the second annotation layer outputs True / False); This is a prompt template for non-registered users.
[0068] The following will provide further explanation using specific scenarios: Application scenario: Intelligent dialogue systems with voiceprint recognition capabilities require the construction of large-scale labeled datasets for periodic evaluation of voiceprint model performance. High labeling efficiency is required, and the cost of purely manual labeling is unacceptable.
[0069] Data preparation: Prepare judgment records for several dialogue interactions within a certain period. Each data record includes the user's audio segment (after VAD processing), the original voiceprint recognition judgment result (registered user identifier or non-registered user marker) and confidence score, ASR transcribed text, dialogue system response text, and recent historical dialogue context. Simultaneously, prepare voiceprint reference samples for each corresponding registered user.
[0070] The first annotation layer executes as follows: For each dialogue device, the reference audio of its registered user and the audio to be annotated are combined with annotation prompts to form a multimodal request. The prompts require the model to strictly follow the following order: first, listen to the reference audio to remember the sound features; then, use your ear to determine the speaker's identity for each ASR text segment in the audio to be annotated; finally, use reference confidence scores for verification; and then compare the results with the voiceprint determination. The prompts embed the ASR text and voiceprint determination information for the current batch. The request is sent to a multimodal large language model (e.g., a model with audio understanding capabilities) to obtain the annotation output in JSONLines format, parsing each line as True / False / None. If the number of output lines does not match the number of input lines, a retry is triggered. In this embodiment, the maximum number of retries is set to 5.
[0071] The second annotation layer executes the following: Based on the output of the first annotation layer, the true identity of the speaker in each round of dialogue (registered user / non-registered user) is determined according to the real identity inference rules. For registered users, prompts from the registered user's perspective are used to judge whether the response reflects correct identity recognition. Responses that are irrelevant but do not involve identity misalignment are still judged as True. Responses are only judged as False when there is a clear signal of identity misalignment (such as the system treating the current user as a stranger and asking who you are, or treating a registered user as a third person). For non-registered users, prompts from the non-registered user's perspective are used to judge whether the system incorrectly treats the current user as a registered user. Responses that mention the registered user but do not refer to the current user are still judged as True (it would be better if the registered user were present). Responses are only judged as False when the response explicitly places the current user in the registered user's position (such as using a title specific to the registered user, or continuing a context belonging only to the registered user). The prompts are sent to the text-based large language model for inference to obtain the True / False judgment result.
[0072] Annotated output: The final output is a complete dataset containing both objective and subjective annotations, which can be directly used for subsequent evaluation of the voiceprint recognition system (e.g., constructing a two-dimensional confusion matrix, calculating a derived index system, etc.).
[0073] The above description is merely an exemplary embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention specification and drawings under the technical concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.
Claims
1. A method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model, characterized in that, include: The audio to be labeled, the dialogue text data, and the voiceprint reference sample of the registered user are obtained. The audio to be labeled is processed by voice activity detection to remove silent segments. A multimodal request message is constructed, and the registered user's voiceprint reference sample and the preprocessed audio to be labeled are input into the multimodal large language model. The first layer of objective auditory judgment and reasoning is executed to obtain intermediate data containing objective labeling results. The objective annotation results in the intermediate data are subjected to format validation and line count consistency checks. If the validation fails, a retry mechanism is triggered to re-call the multimodal large language model to generate new objective annotation results until the validation passes. Based on the verified objective annotation results and the original voiceprint determination information, the speaker's true identity is inferred. Based on the inferred true identity, the corresponding perspective prompt word template is selected to construct the second layer of subjective sensory judgment request. The constructed second-layer subjective sensory perception judgment request is input into the text large language model, and the second-layer subjective sensory perception judgment reasoning is executed using the selected perspective cue word template, and the final subjective annotation result is output.
2. The method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model according to claim 1, characterized in that, The construction of the multimodal request message includes: The registered user's voiceprint reference sample is used as the baseline audio, and the preprocessed audio to be labeled is used as the audio to be tested. The multimodal large language model is also given prompts that require auditory input as the primary factor and confidence as the secondary factor, thus forming a multimodal input vector that includes audio features and semantic logic.
3. The method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model according to claim 1, characterized in that, The retry mechanism re-invokes the multimodal large language model, including: A maximum retry threshold is set. When the number of output rows of the objective annotation results in the intermediate data is inconsistent with the number of input audios or the JSON format is invalid, the request parameters are automatically reset and the multimodal large language model is called again for inference. This process is repeated until the objective annotation results that meet the format requirements are obtained or the maximum retry limit is reached.
4. The method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model according to claim 1, characterized in that, The inference of the speaker's true identity includes: Comparing the original voiceprint determination result with the objective annotation result, if the original voiceprint is determined to be that of a non-registered user but the objective annotation result is true, or if the original voiceprint is determined to be that of a registered user but the objective annotation result is false, then the current speaker is determined to be a non-registered user; otherwise, the speaker is determined to be a registered user.
5. The method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model according to claim 1, characterized in that, The selection of the corresponding view prompt word template includes: If the deduced true identity is a registered user, then select the registered user's perspective prompt template to determine whether the system treats the current user as a registered user; If the inferred true identity is a non-registered user, then select the non-registered user perspective prompt template to determine whether the system is mistakenly treating the current user as a registered user.
6. The method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model according to claim 1, characterized in that, The execution of the second-level subjective sensory judgment reasoning includes: By using a large textual language model to read dialogue records in the historical context window, and combining them with selected perspective cue word templates, the system responses are analyzed to determine whether there are identity misalignment signals, and the final subjective annotation results are determined based on preset conflict resolution rules.
7. The method for constructing voiceprint recognition evaluation data based on a large language model with dual-layer automatic annotation as described in claim 6, characterized in that, The conflict resolution rules include: When a clear signal of identity mismatch appears, it is preferentially judged as an incorrect state; For ambiguous scenarios, the authenticity of subjective annotation results is determined by whether the system uses exclusive terminology or the degree of context mismatch.
8. The method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model according to claim 1, characterized in that, The method further includes S6: assembling the objective annotation results and the subjective annotation results to generate a complete evaluation dataset containing ASR text, voiceprint determination, confidence level, objective dimension labels and subjective dimension labels.
9. The method for constructing voiceprint recognition evaluation data based on a two-layer automatic annotation of a large language model according to claim 1, characterized in that, The reasoning process of the multimodal large language model adopts an end-to-end audio comparison method, directly judging the objective correctness of the voiceprint recognition result based on the audio waveform features.
10. A system for constructing voiceprint recognition evaluation data based on a large language model with two-layer automatic annotation for implementing the method as described in any one of claims 1-9, characterized in that, include: The input data preparation module is used to obtain the audio to be labeled, dialogue text data, and registered user voiceprint reference samples, and to perform speech activity detection processing on the audio to be labeled to remove silent segments. The first annotation layer module is used to construct a multimodal request message, input the registered user's voiceprint reference sample and the preprocessed audio to be annotated into the multimodal large language model, execute the first layer of objective auditory judgment and reasoning, obtain intermediate data containing objective annotation results, and perform format verification and line number consistency checks on the objective annotation results in the intermediate data. If the verification fails, a retry mechanism is triggered to call the multimodal large language model again to generate new objective annotation results until the verification passes. The identity inference and perspective selection module is used to infer the speaker's true identity based on the verified objective annotation results and the original voiceprint judgment information. Based on the inferred true identity, the module selects the corresponding perspective prompt word template and constructs the second layer of subjective sensory judgment request. The second annotation layer module is used to input the constructed second-layer subjective somatosensory judgment request into the text big language model, perform the second-layer subjective somatosensory judgment reasoning using the selected perspective prompt word template, and output the final subjective annotation result; The annotation output module is used to assemble objective annotation results with subjective annotation results to generate a complete evaluation dataset containing ASR text, voiceprint judgment, confidence score, objective dimension labels and subjective dimension labels.