Speaker log task optimization method based on semantic capability of large language model

By adopting a semantic capability optimization method based on a large language model in the speaker log task, the problem of poor robustness of large-scale annotation data dependence and multi-noise environment is solved, and more accurate and readable speaker log results are achieved.

CN119943055AActive Publication Date: 2025-05-06KUNMING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510110597.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The prior art has high dependence on large-scale annotation data in speaker log tasks, and its robustness in multi-noise environments leads to high error rates, affecting the accuracy and readability of the results.

Method used

The speaker log task optimization method based on the semantic ability of the large language model is adopted to generate time-stamped speech transcript text through speech activity detection and automatic speech recognition. The prompt constructor is used to generate matching prompt words, and the timestamp and text content are parsed through the large language model to generate preliminary speaker log results containing timestamps, sentences and speaker tags, and then post-processing is performed to output accurate speaker log results.

Benefits of technology

It significantly reduces the error rate of speaker log tasks, improves the accuracy and readability of the results, especially in complex interjection scenarios, and improves the stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943055A_ABST
    Figure CN119943055A_ABST
Patent Text Reader

Abstract

The invention relates to a speaker log task optimization method based on large language model semantic capability, and belongs to the technical field of artificial intelligence. The method comprises the following steps: generating a voice transcription text with a timestamp through a voice activity detection and automatic voice recognition module, and integrating the generated timestamp with the transcription text to form a timestamp text stream; analyzing the timestamp text stream by using a prompt constructor, and generating a prompt word matched with the speaker log task; inputting the generated cue word and timestamp text stream into a large language model, analyzing timestamps and text content, and generating a preliminary speaker log result containing the timestamps, sentences and speaker tags; the initial speaker log result is post-processed, an accurate speaker log result is output, and the error rate is obviously reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a speaker log task optimization method based on the semantic capability of a large language model, and belongs to the technical field of artificial intelligence. Background Art

[0002] With the rapid development of speech processing technology, speaker log, as an important part of speech recognition technology, has received more and more attention. The goal of the speaker log task is to identify and annotate the speaker identity of each speech in the audio stream through the time index. In simple terms, it is to solve the problem of "who said what and when". This task has wide application value in media broadcasting, conference records, social media conversations, court records and business meetings. Traditional speaker log methods are mainly divided into two categories: clustering-based modular methods and end-to-end systems. Among them, clustering-based methods mainly rely on voice activity detection, speaker embedding extraction and clustering algorithms, and have good flexibility and scalability. However, these methods are highly dependent on the quality of feature extraction, and the optimization goals between modules are inconsistent, making it difficult to achieve global optimization. The end-to-end method completes the speaker log task through a unified neural network architecture and adopts a permutation-invariant loss function to solve the problem of speaker overlap. This method simplifies the process and improves the performance in complex scenarios, but it has a high dependence on large-scale labeled data, and its robustness in multi-noise environments still needs to be improved. Summary of the invention

[0003] The technical problem to be solved by the present invention is: the present invention proposes a method for optimizing speaker log tasks based on the semantic capabilities of a large language model. The method of the present invention solves the problems of high dependence on large-scale annotated data and poor robustness in a noisy environment. The present invention reduces the error rate and improves the accuracy and readability of speaker log results.

[0004] The technical solution of the present invention is: a speaker log task optimization method based on the semantic ability of a large language model, the method comprising:

[0005] Step 1: Generate a speech transcription text with a timestamp through voice activity detection and automatic speech recognition modules, and integrate the generated timestamp with the transcription text to form a timestamp text stream;

[0006] Step 2: Use the prompt builder to analyze the timestamp text stream and generate prompt words that match the speaker log task;

[0007] Step 3: Input the generated prompt words and timestamp text stream into the large language model, parse the timestamp and text content, and generate a preliminary speaker log result containing timestamps, sentences, and speaker labels;

[0008] Step 4: Post-process the preliminary speaker log results and output accurate speaker log results.

[0009] Further, the Step 1 includes: the voice activity detection module adopts the FSMN monophonic large language model;

[0010] The voice activity detection module first detects valid voice segments in the audio and marks the start and end timestamps of these voice segments; then the automatic speech recognition module transcribes these voice segments to generate corresponding text; finally, the timestamps generated by the voice activity detection module and the text generated by the automatic speech recognition module are integrated into a timestamp text stream.

[0011] Furthermore, the Step 2 includes:

[0012] Firstly, the prompt builder is used to analyze the timestamped text stream to extract the semantic information, contextual association and temporal features in the text, and then the potential speaker transition points are identified based on the characteristics of the dialogue structure.

[0013] Then, the prompt word task is constructed; the prompt word contains semantic content, timestamp and context information.

[0014] Furthermore, in Step 3, by inputting the constructed prompt words into the large language model, the semantic understanding ability of the model and the few-shot prompt strategy are utilized to generate a preliminary speaker log result including timestamps, sentences and speaker labels.

[0015] Furthermore, the Step 3 includes:

[0016] Step 3.1. Combine the large language model of the input prompt word with the input timestamp text stream, analyze the semantic content through the few-shot prompt strategy, identify the speaking areas and transition points of different speakers, and show the model the typical dialogue structure through few-shot, guiding the model to accurately distinguish the speeches of different speakers; at the same time, the model determines the speaker identity based on semantic features and generates preliminary speaker log results containing timestamps, sentences and speaker labels.

[0017] Step 3.2, after that, multiple rounds of dialogue input are performed and preliminary speaker log results are generated multiple times: if some speeches in the generated log fail to mark the speaker identity, they are marked as "Speaker 1" by default; and the content of interjections or semantic transitions is assigned to other speakers based on semantic features.

[0018] Furthermore, in the Step 4, the post-processing of the preliminary speaker log results includes:

[0019] Step 4.1, merge the continuous speeches of the same speaker, adjust the start and end positions of the timestamps, and align the timestamps;

[0020] Step 4.2. Eliminate redundancy and delete repeated statements or invalid speeches.

[0021] The present invention provides a speaker log task optimization system based on the semantic capability of a large language model. The system comprises: a module for executing the speaker log task optimization method based on the semantic capability of a large language model.

[0022] The present invention provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the speaker log task optimization method based on the semantic capability of a large language model when executing the program.

[0023] The present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for optimizing a speaker log task based on the semantic capability of a large language model is implemented.

[0024] The present invention provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for optimizing a speaker log task based on the semantic capability of a large language model is implemented.

[0025] The beneficial effects of the present invention are:

[0026] 1. The method of the present invention constructs special prompts for the speaker log task, and uses a large language model to directly generate the speaker log results, thereby completing the task semantically; the process mainly includes generating a transcript with a timestamp through voice activity detection and an automatic speech recognition system, and then constructing prompt words to generate prompts for specific tasks; these prompt words are processed by the large language model to generate the final speaker log results;

[0027] 2. The method of the present invention uses the clustered speaker log error rate as the evaluation indicator. The experimental results on the MagicData-RAMC dataset show that the error rate of the model is significantly reduced; compared with the baseline system, the performance of the speaker log task is improved by 30% using this method;

[0028] 3. The log results generated by the method of the present invention have accurate timestamp calibration, clear semantic logic and good adaptability to complex interruption scenarios, thereby improving the stability of the system; the present invention can be applied to complex conversations involving interruption scenarios, accurately identifying the speaker's identity and assigning a timestamp through semantic information;

[0029] 4. The present invention quickly captures task characteristics through a small number of examples and generates high-quality speaker log results. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a comparison diagram of the influence of using semantic information and embedded information in the speaker log task proposed by the present invention;

[0031] Figure 2 This is a flowchart of semantic speaker log processing based on large language model driving proposed by the present invention. DETAILED DESCRIPTION

[0032] Example 1: Figure 1-Figure 2 As shown, a speaker log task optimization method based on the semantic capability of a large language model comprises:

[0033] Step 1: Generate speech transcription text with timestamps through voice activity detection (VAD) and automatic speech recognition (ASR) modules, and integrate the generated timestamps with the transcription text to form a timestamp text stream;

[0034] Further, the Step 1 includes: the voice activity detection module uses the FSMN monophonic large language model to efficiently detect the start and end time of the voice segment;

[0035] The voice activity detection module first detects valid voice segments in the audio and marks the start and end timestamps of these voice segments; then the automatic speech recognition module transcribes these voice segments to generate corresponding text; finally, the timestamps generated by the voice activity detection module and the text generated by the automatic speech recognition module are integrated into a timestamp text stream.

[0036] The present invention integrates the text data with timestamps to form a timestamp text stream by integrating the timestamp information generated by the voice activity detection and automatic speech recognition modules with the transcribed text data, and pairs the start and end timestamps of each speech segment with the corresponding transcribed text to form a timestamp text stream, which is then used as input for subsequent steps, and the output of each transcribed text is closely related to the corresponding timestamp information to form a time series text record;

[0037] This process essentially combines the start and end time of each speech segment with the corresponding transcript to form a pairing of timestamps and text.

[0038] Among them, the voice activity detection module is mainly responsible for separating valid voice segments and marking their start and end times, while the automatic speech recognition module is mainly responsible for transcribing voice segments into text data with timestamps. The voice activity detection module first processes the input audio signal, identifies the parts of voice activity and ignores silent or noisy segments. For each valid voice segment, the voice activity detection module will mark the start time and end time, that is, the timestamp of each valid voice segment. After the voice activity detection module is completed, the automatic speech recognition module transcribes each voice segment and converts the voice content into text. The automatic speech recognition module will not only transcribe the text of each voice segment, but also assign timestamps to each transcribed text, that is, mark the start and end time for each text unit, such as a word or sentence.

[0039] After passing through the voice activity detection and automatic speech recognition modules, the final output is a sequence of multiple voice segments and their corresponding transcribed texts and timestamps, namely a timestamp text stream. This stream contains not only the text content of each voice segment, but also the time position of these contents in the conversation.

[0040] In this process, there is no additional integration process, that is, the timestamps and texts generated by the voice activity detection module and the automatic speech recognition module are directly combined to form a timestamp text stream, and their outputs can be directly combined to form a complete timestamp and transcribed text pair.

[0041] Step 2: Use the prompt builder to analyze the timestamp text stream and generate prompt words that match the speaker log task;

[0042] Furthermore, the Step 2 includes:

[0043] The construction of prompt words is the key to realize the generation of speaker logs. First, the prompt builder is used to analyze the timestamp text stream, extract the semantic information, context association and time features in the text, and identify potential speaker transition points based on the characteristics of the dialogue structure.

[0044] Then, a prompt word task is constructed, such as providing an example of a two-person conversation in a standard format, including timestamps, text, and speaker labels, to guide the large language model to understand the task requirements; the prompt words contain semantic content, timestamps, and contextual information.

[0045] In the process of "generating prompt words matching the speaker log task" of the present invention, the prompt constructor first needs to analyze the timestamp text stream. The prompt constructor will extract key information from the timestamp transcribed text generated by the voice activity detection and automatic speech recognition module. For example, it can identify the key content and semantic relationships in the text, such as the theme of each paragraph, the speaker's speech content, etc. In addition, the constructor can understand the relationship between the context, ensure the coherence and logic of the conversation, help the model better capture the changes in the context, and identify the order of speech.

[0046] Secondly, the prompt builder identifies potential speaker switching points based on the semantic and temporal features of the conversation. For example, it identifies whether a certain statement is a response, supplement or rebuttal to the previous statement. If there is a long pause or a change in speech speed, it may also mean a switch in the speaker.

[0047] Finally, after completing the analysis, the prompt builder will generate prompt words that match the speaker log task based on the extracted information. For example, the prompt builder can describe the core information of each speech based on the semantic content, and provide corresponding time information for each speech to ensure the order and time of the speech are consistent. At the same time, the context in the conversation is supplemented based on the context information to help the model understand the background of the speaker's speech more accurately.

[0048] Step 3: Input the generated prompt words and timestamp text stream into the large language model (LLM), parse the timestamp and text content, and generate a preliminary speaker log result containing timestamps, sentences, and speaker labels;

[0049] Furthermore, in Step 3, by inputting the constructed prompt words into the large language model, the semantic understanding ability of the model and the few-shot prompt strategy are utilized to generate a preliminary speaker log result including timestamps, sentences and speaker labels.

[0050] Furthermore, the Step 3 includes:

[0051] Step 3.1. Combine the large language model of the input prompt word with the input timestamp text stream, analyze the semantic content through the few-shot prompt strategy, identify the speaking areas and transition points of different speakers, and show the model the typical dialogue structure through few-shot, guiding the model to accurately distinguish the speeches of different speakers; at the same time, the model determines the speaker identity based on semantic features and generates preliminary speaker log results containing timestamps, sentences and speaker labels.

[0052] Step 3.2, after that, multiple rounds of dialogue input are performed and preliminary speaker log results are generated multiple times: if some speeches in the generated log fail to mark the speaker identity, they are marked as "speaker 1" by default; and the content of the interjection or semantic transition is assigned to other speakers according to the semantic features. For example, "speaker 2". The generated log results have accurate timestamp calibration, clear semantic logic, and good adaptability to complex interjection scenarios, thereby improving the stability of the system.

[0053] Step 4: Post-process the preliminary speaker log results and output accurate speaker log results.

[0054] Furthermore, in the Step 4, the initially generated log results are post-processed to optimize the accuracy of timestamps and speaker assignments, and the post-processing of the initial speaker log results includes:

[0055] Step 4.1, merge the continuous speeches of the same speaker, adjust the start and end positions of the timestamps, and align the timestamps to make the results more accurate;

[0056] Step 4.2. Eliminate redundancy, delete repeated statements or invalid statements, and ensure the conciseness and readability of the log results.

[0057] The present invention provides a speaker log task optimization system based on the semantic capability of a large language model, the system comprising:

[0058] A timestamp text stream acquisition module is used to generate a speech transcription text with a timestamp through a voice activity detection and automatic speech recognition module, and integrate the generated timestamp with the transcription text to form a timestamp text stream;

[0059] A prompt word construction module is used to analyze the timestamped text stream using the prompt builder to generate prompt words that match the speaker log task;

[0060] A preliminary speaker log result generation module is used to input the generated prompt words and timestamp text stream into the large language model, parse the timestamp and text content, and generate a preliminary speaker log result containing timestamps, sentences and speaker labels;

[0061] The post-processing module is used to post-process the preliminary speaker log results and output accurate speaker log results.

[0062] The present invention provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the speaker log task optimization method based on the semantic capability of a large language model when executing the program.

[0063] The present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for optimizing a speaker log task based on the semantic capability of a large language model is implemented.

[0064] The present invention provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for optimizing a speaker log task based on the semantic capability of a large language model is implemented.

[0065] The present invention uses the clustered speaker log error rate (CDER) as an evaluation indicator to ensure the consistency of the timestamp and semantics of the log results. The performance of the model is evaluated by merging the continuous speeches of the same speaker and comparing the consistency of the generated results with the reference results. The experiments of the present invention on the MagicData-RAMC dataset show that the error rate of the model is significantly reduced to 0.244. The performance of the speaker log task is improved by 30% using this method.

[0066] To verify the effectiveness of the method of the present invention, the present invention uses FSMN-Monophone VAD as the baseline large language model of the speaker log generation method of the large language model. The advantage of this large language model is that it takes contextual information into account in the process of structural modeling, provides fast training and reasoning speed, has controllable delay, enhances abstract learning ability, and improves discrimination ability. The FSMN model consists of 24 layers of Transformer encoders and decoders, where the hidden layer dimension is 1024 and the forward propagation layer dimension is 4096, which is used to enhance the nonlinear modeling ability of the model. The length predictor consists of 5 layers of one-dimensional convolutional networks and 2 layers of linear layers, where the hidden layer dimension is 512. Each layer of convolutional network has a ReLU activation function to introduce nonlinearity and enhance the feature learning ability of the model.

[0067] The present invention uses GLM (General Language Model) as the core large language model to perform speaker log tasks. The model is based on a 24-layer Transformer architecture, with a hidden layer dimension of 1024 and a forward propagation layer dimension of 4096. It uses a multi-head self-attention mechanism and GELU activation function, and has powerful semantic modeling capabilities. The model is designed with Few-shot prompts, combining timestamps, semantic information and context to achieve high-precision speaker log generation. At the same time, the large language model introduces layer normalization and Dropout regularization technology, and adopts a multi-round generation strategy to reduce uncertainty.

[0068] Compared with the baseline model, the MagicData-RAMC dataset corpus of the present invention is combined with the speaker log task. The specific experimental results are shown in Table 1 below:

[0069] Table 1 shows the performance comparison of large language models under different prompt formats.

[0070]

[0071] From the comparison results in Table 1, it can be seen that GLM performs best in the Few-shot prompt design, with CDER reaching 0.244, significantly better than the Baichuan model and the Yi model. Under the Chain-of-thought prompt, the performance of GLM drops to 0.498, but it is still better than the other two large language models. Under the Chain-of-thought and CRISPE prompts, although the performance has declined, GLM still outperforms other models. This shows that GLM has the strongest adaptability to Few-shot prompts, and can quickly capture task characteristics and generate high-quality speaker log results through a small number of examples. In contrast, the Chain-of-thought prompt may introduce additional generation biases because it requires the model to have more complex reasoning capabilities. Although the CRISPE format prompt has certain requirements for task parsing capabilities, GLM still shows strong stability and generation capabilities.

[0072] Table 2 shows the comparative experiments of different numbers of examples in Few-shot prompts.

[0073]

[0074] From the comparative experimental results in Table 2, it can be seen that when the Few-shot is 2-shot, the model's clustered speaker log error rate (CDER) is as low as 0.244, and the performance is the best. However, due to insufficient examples, the 1-shot model has limited ability to understand the task, and its clustered speaker log error rate is as high as 0.350. Although 3-shot provides more information, it may interfere with the model performance due to increased complexity, and the clustered speaker log error rate rises to 0.313. The experimental data results show that using 2-shot can bring the best performance for the speaker log task and reduce the clustered speaker log error rate. It can also help GLM balance task parsing and generation capabilities, providing sufficient task guidance while avoiding performance degradation caused by redundant information. Too few or too many examples in the prompt design will affect the model's task understanding and output accuracy, indicating that the Few-shot prompt design needs to consider the rationality of the number of examples to maximize the performance advantage of the model.

[0075] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A speaker log task optimization method based on the semantic capability of a large language model, characterized by: The method comprises: Step 1: Generate a speech transcription text with a timestamp through voice activity detection and automatic speech recognition modules, and integrate the generated timestamp with the transcription text to form a timestamp text stream; Step 2: Use the prompt builder to analyze the timestamp text stream and generate prompt words that match the speaker log task; Step 3: Input the generated prompt words and timestamp text stream into the large language model, parse the timestamp and text content, and generate a preliminary speaker log result containing timestamps, sentences, and speaker labels; Step 4: Post-process the preliminary speaker log results and output accurate speaker log results.

2. The speaker log task optimization method based on the semantic ability of a large language model according to claim 1, characterized in that: Step 1 includes: the voice activity detection module adopts the FSMN monophonic large language model; The voice activity detection module first detects valid voice segments in the audio and marks the start and end timestamps of these voice segments; then the automatic speech recognition module transcribes these voice segments to generate corresponding text; finally, the timestamps generated by the voice activity detection module and the text generated by the automatic speech recognition module are integrated into a timestamp text stream.

3. The speaker log task optimization method based on the semantic ability of a large language model according to claim 1, characterized in that: The Step 2 includes: Firstly, the prompt builder is used to analyze the timestamped text stream to extract the semantic information, contextual association and temporal features in the text, and then the potential speaker transition points are identified based on the characteristics of the dialogue structure. Then, the prompt word task is constructed; the prompt word contains semantic content, timestamp and context information.

4. The speaker log task optimization method based on the semantic ability of a large language model according to claim 1, characterized in that: In the Step 3, by inputting the constructed prompt words into the large language model, the semantic understanding ability of the model and the few-shot prompt strategy are used to generate a preliminary speaker log result containing timestamps, sentences and speaker labels.

5. The speaker log task optimization method based on the semantic ability of a large language model according to claim 1, characterized in that: The Step 3 includes: Step 3.

1. Combine the large language model of the input prompt word with the input timestamp text stream, analyze the semantic content through the few-shot prompt strategy, identify the speaking areas and transition points of different speakers, and show the model the typical dialogue structure through few-shot, guiding the model to accurately distinguish the speeches of different speakers; at the same time, the model determines the speaker identity based on semantic features and generates preliminary speaker log results containing timestamps, sentences and speaker labels. Step 3.2, after that, multiple rounds of dialogue input are performed and preliminary speaker log results are generated multiple times: if some speeches in the generated log fail to mark the speaker identity, they are marked as "speaker 1" by default; and the content of interjections or semantic transitions is assigned to other speakers based on semantic features.

6. The speaker log task optimization method based on the semantic ability of a large language model according to claim 1, characterized in that: In the Step 4, the post-processing of the preliminary speaker log results includes: Step 4.1, merge the continuous speeches of the same speaker, adjust the start and end positions of the timestamps, and align the timestamps; Step 4.

2. Eliminate redundancy and delete repeated statements or invalid speeches.

7. A speaker log task optimization system based on the semantic capability of a large language model, characterized in that: The system comprises: a module for executing the speaker log task optimization method based on the semantic capability of a large language model as claimed in any one of claims 1 to 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the speaker log task optimization method based on the semantic capability of a large language model is implemented as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for optimizing speaker log tasks based on the semantic capability of a large language model is implemented as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for optimizing speaker log tasks based on the semantic capability of a large language model is implemented as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Service quality management method and system based on general big language model

    CN117391515A

  • Conference summary generation method based on large language model

    CN119003759A

Cited By

  • Audio speaker recognition method and system, storage medium and electronic equipment

    CN120220729A

  • Speaker log generation method based on semantic alignment

    CN120673764A

  • Dialogue analysis method for multiple speakers

    CN121393427A