Audio translation annotation and analysis method based on multi-agent cooperation mechanism
Through the multi-agent collaboration mechanism, the efficiency and accuracy problems in the processing of public opinion demands are solved, the audio transcription, information extraction and analysis reports are automated, the transcription accuracy and labeling efficiency are improved, and the audio data processing in multiple scenarios is adapted to the process of audio data in multiple scenarios.
Patent Information
- Application Number
- CN202510714641.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-15
AI Technical Summary
When handling massive public opinion calls, the existing technology is inefficient and insufficient accuracy, and cannot automatically generate analysis reports. It is difficult for traditional methods to process multimodal data and dynamic tag updates.
The multi-agent collaboration mechanism is adopted to translate audio into text through the Whisper audio translation model, combine domain knowledge and language style learning model, extract time expressions and abstract paragraphs, use large language models for annotation and analysis, build multi-level label data sets, and fine-tune the model to generate a document set with labels.
It realizes efficient transcription, information extraction and analysis reports from audio to text, improves transcription accuracy and labeling efficiency, significantly improves label relevance and correctness, and adapts to audio data processing in different scenarios.
Smart Images

Figure CN120492633A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of information technology and artificial intelligence, and relates to a method for audio translation, annotation and analysis based on a multi-agent collaborative mechanism. Background Art
[0002] In recent years, with the continuous development of information technology, the deep learning capabilities of large language models have achieved unprecedented improvements, enabling this technology to be more widely applied in daily life and helping people improve work efficiency. Furthermore, with the advent of digital government, digital technology has been widely applied to government management and services, promoting the digitalization and intelligent operation of government. In the field of public administration, analyzing public opinion is crucial. By collecting, analyzing, and responding to public opinions, suggestions, and demands, public departments can better formulate policies and provide services. Specifically, in the case of public opinion hotlines, hundreds or even thousands of calls are received daily in a single region. Relying on human resources to record, categorize, archive, and analyze these calls is unrealistic. People now need tools to help them translate and categorize the information they need to process, as well as summarize and analyze it, to accurately understand public concerns.
[0003] Faced with the massive amount of public opinion calls that need to be processed, in order to help relevant departments obtain the required information more quickly and accurately, we first need to translate the audio information and convert it into text information that is easier for the model to understand. This will improve the efficiency of recording and organizing the content of the call. Secondly, we need to mark the labels of the call content and use these labels to express the category of the problem. In this way, government workers can understand the content of the article more accurately through the content labels without reading the full text. Finally, this invention makes full use of the advantages of large language models in data analysis and content generation to provide a summary report of the selected content files.
[0004] As a classic model, the LDA topic model generates labels by mining the potential topic distribution of text, solving the problem of traditional TF-IDF methods ignoring semantic associations. Its advantage is that implicit topics can be discovered without labeled data, but its limitation is that it is difficult to handle multimodal data and dynamic label update requirements. At the research level, in 2020, Tian Feng et al. mapped image visual features and text labels to the same space in the "Semantic Label Generation Method Based on Multimodal Subspace Learning" and generated labels through nearest neighbor search. At the practical level, in 2025, Wei Xiaoyan et al.'s patent "A Label Generation Method, Device, Electronic Device and Storage Medium" proposed a job search label generation model based on embedding layers and feature recognition. The label generation model has gradually evolved from traditional topic models to a comprehensive technology system driven by multimodal fusion and deep learning, covering more and more application scenarios.
[0005] Large language models (LLMs) achieve structured parsing and semantic extraction of icon content by integrating visual encoders with language model backbones. The DeepSeek-VL2 model uses a mixture of experts (MoE) architecture, which can extract key information from complex charts and generate logically coherent text summaries. The multimodal pre-training strategy further optimizes the alignment of images and text, enabling the model to combine label semantics to generate accurate analysis conclusions. In terms of specific applications, in 2025, Yu Heke et al. pointed out in "Retrieved generation for 10 large language models and its generalizability in assessing medical fitness" that the GPT4-RAG model generates surgical risk reports by analyzing patient indicators with higher accuracy than human assessors. Summary of the Invention
[0006] In order to solve the above problems, in a first aspect, according to some embodiments of the present application, the method for audio translation labeling based on a multi-agent collaborative mechanism includes: S1. Translate the audio file into a text file; S2. The text file is passed through a model that incorporates domain knowledge and language style learning, and the output is a dialogue set with a well-organized text structure. S3. By incorporating the large language model of the prompt project, we extract and structure time expressions, place names, and summary paragraphs, and associate the event context with the conversation context to obtain an information set. S4. Concatenate the conversation set and the information set to obtain a document set; S5. A series of texts and their corresponding labels at all levels with an inclusion relationship constitute a text-label pair dataset, and the text-label pair dataset is constructed into an instruction dataset that conforms to the instruction fine-tuning format specification; S6. fine-tune the large language model base according to the instruction dataset to obtain a fine-tuned model; S7. Input the document set into the fine-tuned model to annotate labels at all levels to obtain a document set with labels.
[0007] According to the method of audio translation labeling based on a multi-agent collaborative mechanism in some embodiments of the present application, it also includes screening and judging according to the inclusion relationship of labels at all levels. For the labels at all levels determined by the dialogues and key information in the document collection, the labels with incorrect relationships that do not have inclusion relationships in the labels at all levels are eliminated, and the dialogues and key information in the document collection are re-input into the fine-tuned model until the dialogues and key information in the document collection are obtained. It has an inclusion relationship in the labels at all levels.
[0008] According to the method for audio translation and labeling based on a multi-agent collaborative mechanism in some embodiments of the present application, step S1 specifically includes: using the Whisper audio translation model based on the Transformer architecture to process the input audio file Q, and translating the audio file Q into a text file T.
[0009] According to the method of audio translation labeling based on multi-agent collaborative mechanism in some embodiments of the present application, in which the model of domain knowledge and language style learning is introduced in step S2 , where the model The network is represented by matrices and tensors during training. Abstract into a function function Word vector for input text file data To process, The parameters of the model are a set of weights, , , A and B are the training parameters to be updated, A is initialized using a random Gaussian distribution, and B is initialized with a zero matrix; Represents the pre-trained language model parameters, A, B are represented by After matrix transformation, it is obtained as bypass matrix; The update during training is expressed as: Where, , where rank ; , represents the dimension of the vector, Indicates a size of A real matrix of ; The training weight update steps are as follows: Lora formula: (1) Derivative of loss L: Where, represents the loss function, Represents the output of the model, which is composed of the input After weight matrix transformation, Represents the loss function For weight parameters gradient; Weight update amount: For the first step of back propagation, Where, Indicates the next moment The matrix parameter A, Indicates the current time The matrix parameter A, Represents the learning rate, which is a positive hyperparameter. Representation matrix The transpose of Indicates the The loss function value calculated in the iteration, Indicates the next moment The matrix parameter B, Indicates the current time Matrix parameter B; Since the matrix B is initialized to all 0s and the matrix A is initialized to a random Gaussian distribution, for any step t in the training, Where, represents the initial matrix parameter A, Represents the changing function of A with time step t, which is used to characterize the changing trend of A. Represents the changing function of B with time step t, which is used to characterize the changing trend of B; Where, represents the learning increment of B during the training process, Represents the learning increment of A during the training process; Substituting equation (4) into equation (5) we have Use language optimization through text file T to get instruction data set and set rank , for the model Perform training, import the model weights obtained from training, and obtain the model ; Input the text file T into the model Get the conversation set , Indicates the A conversation.
[0010] According to the method of audio translation labeling based on multi-agent collaborative mechanism in some embodiments of the present application, in step S3, a large language model with a prompting project is added. From the dialogue Extract and structure time expressions, place names and summary paragraphs, and automatically associate event background with conversation context to obtain information sets ,in Represents a conversation set The key information in the article, including the summary, the place and time mentioned, Indicates the A key information.
[0011] According to the method for audio translation labeling based on multi-agent collaborative mechanism in some embodiments of the present application, in step S4, the conversation set With information set Splicing to get a document set ; According to the method for audio translation and labeling based on a multi-agent collaborative mechanism in some embodiments of the present application, the fine-tuned model in step S6 uses Qwen2.5-7B-Instruct as the large model base, and is fine-tuned using the LoRA algorithm. The batch size is set to 16, the learning rate is 5e-5, and the temperature is 0.8. The fine-tuned model is Qwen4TG.
[0012] According to the method for audio translation labeling based on multi-agent collaborative mechanism in some embodiments of the present application, step S7 includes For document sets After word segmentation, the word after word segmentation is , Indicates the Word segmentation, each word passes through the encoding layer to generate a word embedding vector , expressed by the following formula: Where, Indicates the participles; The word is positionally encoded, where the position encoding is expressed as follows: in, Represents the dimension index of the position vector, 2 Indicates an even dimension, 2 +1 indicates an odd dimension, pos is the position index, and d is the word vector dimension. Indicates the location ,latitude The positional encoding value on , Indicates the location ,latitude Positional encoding value on ; Word vectors The word vector The weighted aggregation of each word is obtained through the attention mechanism to obtain the The semantic representation after the vector is updated is formally expressed as Where, represents the attention output result, Represents query matrix, key matrix, value matrix, represents the dimension of the key vector; Through linear mapping Obtained in The sequence is mapped to the feature space. Each set of linearly projected vector representations is called a head. ScaledDot-productAttention is then applied to each set of mapped sequences. Each attention head focuses on a certain aspect of semantic similarity. Multiple attention heads allow the model to focus on multiple aspects simultaneously. The formal representation is: Where, Indicates the An attention head, represents the linear transformation matrix that maps the input sequence to Q, used for the i-th attention head, represents the linear transformation matrix that maps the input sequence to K, Represents the linear transformation matrix that maps the input sequence to V; Concatenate multiple attention heads to get the result sequence: Where, Represents the output of the multi-head attention mechanism, Indicates the an attention head; Enter Perform residual connection with the multi-head attention result sequence and normalize it to get the input residual: Where, represents the output result of the multi-head attention sub-layer at the i-th position, is the standard layer normalization operation; Will Input feedforward sublayer, which contains two layers of linear transformation and an activation function: Where, represents the output of the feedforward sublayer, represents the input vector, represents the first-layer linear transformation, represents the second-layer linear transformation, represents the activation function; Will Add to the input residual and then normalize: Where, Represents the final output after being processed by the feedforward neural network sublayer and obtained through residual connection and layer normalization; After linear mapping to the vocabulary space through the decoding layer, the label content at each level is obtained.
[0013] According to the method for audio translation labeling based on multi-agent collaborative mechanism in some embodiments of the present application, the label in step S5 includes a first-level label , secondary tag , third-level label , where the second-level tag is a subtag of a first-level tag, and the third-level tag is a subtag of a second-level tag; where, Indicates the First-level tags; Indicates that it belongs to The first level label Secondary tags, Indicates that it belongs to The first level label The second level label three-level labels; Among them, the document set with labels ,in Indicates labels at all levels, Means 1, 2, 3.
[0014] On the second aspect, the method for analyzing audio translation annotation labels based on the multi-agent collaborative mechanism in some embodiments of the present application includes any of the methods for annotating labels; and further includes The set of labeled documents obtained for the input multiple audio files Q , extract the label content and reassemble it into a label set ,in Indicates the The first audio file Level label, =1,..., , Indicates the number of audios, =1,2,3; The tag set P is modeled and aggregated according to the dimensions of frequency, hierarchy, and co-occurrence relationship, and the visualization results are output to conduct horizontal comparison and vertical trend analysis.
[0015] Beneficial effects: On the first aspect, the present invention mainly addresses the problems existing in traditional audio transcription and analysis methods, such as low efficiency, insufficient accuracy and inability to automatically generate analysis reports, and provides an efficient, accurate and automated solution for the processing and analysis of audio data.
[0016] On the second aspect, the present invention is an audio transcription, annotation and analysis method based on a multi-intelligent collaborative mechanism. By integrating multiple language models, it realizes the full-process automated processing from audio transcription to text optimization, information extraction, annotation and analysis report generation.
[0017] On a third level, this method can efficiently process audio data in various scenarios, achieving high transcription accuracy, powerful text optimization capabilities, and precise information extraction. Furthermore, by building a multi-agent collaborative working mechanism, this method achieves heterogeneous capability collaboration, adaptive task allocation, and inter-module linkage, demonstrating excellent generalization capabilities.
[0018] In the fourth aspect, the present invention introduces a large language model into the text optimization process, which effectively solves the defect of poor quality of speech model translation effect in the face of dialect translation.
[0019] In the fifth aspect, the label association search test method: the present invention proposes a multi-level label classification query test method, which significantly improves the relevance and correctness of text labels by determining the relationship between labels at all levels, matching and testing them one by one. It is particularly difficult to screen and judge according to the inclusion relationship of labels at all levels. For the labels at all levels determined by the dialogues and key information in the document collection, the labels with incorrect relationships that do not have inclusion relationships in the labels at all levels are eliminated, and the dialogues and key information in the document collection are re-input into the fine-tuned model until the dialogues and key information in the document collection are obtained. The labels at all levels have inclusion relationships. Compared with the model that determines each level at the same time, the present invention is lower than that of each level. If there is no inclusion relationship between the levels, the level is eliminated and re-determined, which can significantly improve the accuracy of the hierarchical label determination. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Provide a technical roadmap for multi-agent collaborative audio translation and analysis methods.
[0021] Figure 2 This is a model diagram of the multi-agent collaborative audio translation analysis method.
[0022] Figure 3 Classify matching rules for the annotation model.
[0023] Figure 4 Set up an example for the report partition.
[0024] Figure 5Obtain an instruction dataset for the text file T using language optimization.
[0025] Figure 6 Construct an instruction dataset for the text-label pair dataset. DETAILED DESCRIPTION
[0026] The examples of the present invention are further described in detail below in conjunction with the accompanying drawings and technical solutions.
[0027] The present invention proposes a method for audio translation and annotation analysis based on a multi-agent collaborative mechanism, comprising the following steps: S1. Obtain audio translation results through the whisper audio translation model: Input the original audio file Q, which is preferably less than three minutes, and process it using the Whisper audio translation model based on the Transformer architecture. The model has a self-attention mechanism, supports end-to-end audio recognition tasks, and can automatically segment and restore semantics based on contextual information. It can maintain high transcription accuracy when facing different speaking speeds, noise interference, and heavy-accented corpus, and supports processing the input audio Q in a streaming or segmented manner. After a short period of time, a piece of translated text T is generated. Among them, the original audio file Q of any length is input (three minutes is the most suitable), and the whisper audio translation model generates a piece of translated text T after a short period of time. S2. Optimizing Translated Texts with Large Language Models For the translated text T obtained in S1, the Qwen2.5-1.5B-Instruct large language model trained by fine-tuning Get a collection of conversations with optimized text and formatting ,in Represents each complete conversation, and n represents the number of conversations. The training process has the following basic steps: The model network is represented by matrices and tensors. The large model can be abstracted into a function: function Word vectors for text file data To process, As the parameters of the model, here is a set of weights consisting of 1.5 billion floating point numbers. , , only A and B are training parameters that need to be updated and will be adjusted for the task. A is initialized using a random Gaussian distribution and B is initialized with a zero matrix. Represents the pre-trained language model parameters, A, B are represented by After matrix transformation, it is obtained as the bypass matrix.
[0028] The update during training can be expressed as: Where, , where rank . , represents the dimension of the vector, Indicates a size of A real matrix of .
[0029] The training weight update steps are as follows: LoRA formula: (1) Derivative of loss L: Where, represents the loss function, Represents the output of the model, which is composed of the input After weight matrix transformation, Represents the loss function For weight parameters gradient.
[0030] Weight update amount: For the first step of back propagation, Where, Indicates the next moment The matrix parameter A, Indicates the current time The matrix parameter A, Represents the learning rate, which is a positive hyperparameter. Representation matrix The transpose of Indicates the The loss function value calculated in the iteration, Indicates the next moment The matrix parameter B, Indicates the current time The matrix parameter B.
[0031] Since the matrix B is initialized to all 0s and the matrix A is initialized to a random Gaussian distribution, for any step t in the training, Where, represents the initial matrix parameter A, Represents the changing function of A with time step t, which is used to characterize the changing trend of A. Represents the changing function of B with time step t, which is used to characterize the changing trend of B.
[0032] Where, represents the learning increment of B during the training process, Represents the learning increment of A during the training process.
[0033] Bringing (4) into (5) yields The instruction data set is obtained by language optimization of the text file T, which is constructed as a json data text pair of instruction-input-output. Figure 5 As shown, and set the rank , for the model Perform training, import the model weights obtained from training, and obtain the model .
[0034] T passes the model Finally, a high-quality dialogue set with clear structure and clear roles is formed. ,in Represents each complete conversation, and n represents the number of conversations.
[0035] S3, for the dialogue set obtained in S2 , through Qwen2.5-1.5B-Instruct large language model Perform summary analysis, location and time extraction to obtain information set ,in Represents key information in the conversation set, such as summary, mentioned place, time, etc. Large model This is achieved by adding a prompt project. The basic prompt parameter is set as follows: "You are a text analysis expert and need to extract information from incoming text conversations. Extract the time, location, and summary and output them in JSON format. Do not output other content." This model can accurately extract and structure time expressions (including relative and absolute time), location names (including nested hierarchies), and summary paragraphs from conversation sets, and automatically associate event background with conversation context.
[0036] S4, splicing document set: the conversation set obtained in S2 The information set obtained from S3 Stitching together a complete set of documents .
[0037] S5. Text Conversation Annotation Using Large Language Models S5.1. Constructing label inclusion relationships at all levels Designing a primary label . Secondary tags . Level 3 label The hierarchical structure is shown in Table 1.
[0038] Table 1 Example of label hierarchical relationship S5.2, fine-tuning model: The present invention uses a series of texts and their corresponding labels at all levels to form a 51542-item "text-label pair" data set, and constructs the text-label pair data set into an instruction data set that conforms to the instruction fine-tuning format specification, such as Figure 6 Then, using the instruction dataset and Qwen2.5-7B-Instruct as the large model base, the LoRA algorithm is used for fine-tuning, with the batch size set to 16 and the learning rate set to 5e-5. The fine-tuned model is called Qwen4TG.
[0039] S5.3, three-level label annotation: using the text annotation model Qwen4TG, according to the labeling strategy, the document set obtained in S4 is generated in a generative way Label them one by one to obtain a set of documents with labels ,in, Indicates the levels of labels, there are three levels of labels here. The specific steps are as follows: For the input document set , first undergoes word segmentation, and the word after word segmentation is , each word passes through the encoding layer to generate a word embedding vector , expressed by the following formula: The word is positionally encoded, where the position encoding is represented as , expressed by the following formula: in, Represents the dimension index of the position vector, 2 Indicates an even dimension, 2 +1 indicates an odd dimension, pos is the position index, and d is the word vector dimension. Indicates the location ,latitude The positional encoding value on , Indicates the location ,latitude The positional encoding value on .
[0040] This constitutes the word vector , the word vector The weighted aggregation of each word is obtained through the attention mechanism to obtain the updated semantic representation of the i-th vector, which is formally expressed as Where, represents the attention output result, Represents query matrix, key matrix, value matrix, Indicates the dimension of the key vector.
[0041] Through linear mapping Obtained in The sequence is mapped to the feature space. Each set of linearly projected vector representations is called a head. ScaledDot-productAttention is then applied to each set of mapped sequences. Each attention head is responsible for focusing on a certain aspect of semantic similarity. Multiple heads can allow the model to focus on multiple aspects at the same time. The formal representation is: Where, Indicates the An attention head, represents the linear transformation matrix that maps the input sequence to Q, used for the i-th attention head, represents the linear transformation matrix that maps the input sequence to K, Represents the linear transformation matrix that maps the input sequence to V.
[0042] Where, Represents the output of the multi-head attention mechanism, Indicates the attention heads. Among them, , , is the mapping matrix, are the query matrix, key matrix, and value matrix, respectively, and h is the number of attention heads. Finally, the results of multiple heads are concatenated to obtain the final result sequence.
[0043] Then enter Make residual connection with multi-head attention and normalize: Where, represents the output result of the multi-head attention sub-layer at the i-th position, is the standard layer normalization operation.
[0044] It is a standard layer normalization operation (mean normalization and variance scaling in the feature dimension) The feedforward sublayer consists of two layers of linear transformation and an activation function: Where, represents the output of the feedforward sublayer, represents the input vector, represents the first-layer linear transformation, represents the second-layer linear transformation, Represents the activation function.
[0045] Add the FFN output to the input residual, and then : Where, represents the final output after being processed by the feedforward neural network sublayer and normalized by residual connection and layer. final, After linear mapping to the vocabulary space through the decoding layer, the label content at each level is obtained.
[0046] According to the inclusion relationship of each level of labels, we filter and judge, remove the labels with wrong relationships and re-input them into the model until we get the labels of each level with correct relationships. Finally, we get the document set with labels. ,in Indicates labels at all levels. There are three levels of labels here.
[0047] S6. Count the number of tags and draw the corresponding chart The set of labeled documents obtained for the input multiple audio files Q , extract the label content and reassemble it into a label set ,in, Indicates the The first audio file Then, based on the data processing and visualization module of the programming language, Perform data statistics separately and generate visual charts based on the overall calculation ratio.
[0048] Customize report templates and generate analysis reports According to our own requirements, we set up report filling partitions, inserted the data information and charts obtained in S6 into the predetermined positions, and then used the large language model to fill in each partition, conduct analysis and summary, and finally complete the report.
[0049] The method of the present invention sets two groups of experiments for audio translation effect and classification labeling effect as follows: Experiment 1: 187 audio samples totaling 263.2 minutes were manually sampled and verified. The experimental results are shown in Table 2.
[0050] Table 2 Audio translation efficiency Experiment 2: 13,060 texts that have been categorized and annotated were annotated by multiple large language models. The experimental results are shown in Table 3.
[0051] Table 3 Classification and labeling results Through Experiments 1 and 2, it can be concluded that the method of the present invention not only improves the audio translation speed but also effectively ensures the translation quality. Moreover, from Experiment 2, it can be seen that in the existing large language model, the method of the present invention greatly improves the ability of the large language model to perform classification tasks.
[0052] Some examples of this implementation are as follows: Audio translation optimization results: A: Hello, is this student XXX? I saw that you submitted feedback about the course.
[0053] B: Yes, I found that I had finished watching the video during class yesterday, but the system did not record my learning progress.
[0054] A: I see. Do you learn through the web or the app? B: I learned it through the mobile app, and it kept showing 0%, but I had actually finished it.
[0055] A: OK, in this case we recommend that you log out of your account first and then log in again, the system will automatically refresh the progress.
[0056] B: I tried it, but it still didn’t update.
[0057] A: I will record the problem for you first. Our technicians will handle it and synchronize the progress later. We will notify you by text message after the problem is solved.
[0058] B: Okay, thank you.
[0059] A: You’re welcome. If you have any other questions, please feel free to contact us.
[0060] B: Okay, bye.
[0061] A: Goodbye.
[0062] First-level label: online learning management Secondary label: Course playback issues Level 3 label: Progress out of sync Summary Analysis: After learning the video on the APP, the system did not synchronize the learning progress, requiring technical processing The method of the present invention realizes audio translation, summary analysis, text annotation and report generation. The multi-agent collaborative mechanism fully utilizes the capabilities of each agent model to improve translation efficiency and quality. The reasonable label relationship setting strategy effectively improves the correctness of text annotation, and has great application value in speech recognition, text annotation, analysis and summary, etc.
[0063] Finally, it should be noted that the purpose of publishing some examples is to facilitate a further understanding of the present invention. However, those skilled in the art will appreciate that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the contents disclosed in the examples; the scope of protection claimed by the present invention shall be determined by the scope defined in the claims.
Claims
1. A method for audio translation and labeling based on a multi-agent collaborative mechanism, characterized in that: include S1. Translate the audio file into a text file; S2. The text file is passed through a model that incorporates domain knowledge and language style learning, and the output is a dialogue set with a well-organized text structure. S3. By incorporating the large language model of the prompt project, we extract and structure time expressions, place names, and summary paragraphs, and associate the event context with the conversation context to obtain an information set. S4. Concatenate the conversation set and the information set to obtain a document set; S5. A series of texts and their corresponding labels at all levels with an inclusion relationship constitute a text-label pair dataset, and the text-label pair dataset is constructed into an instruction dataset that conforms to the instruction fine-tuning format specification; S6. fine-tune the large language model base according to the instruction dataset to obtain a fine-tuned model; S7. Input the document set into the fine-tuned model to annotate labels at all levels to obtain a document set with labels.
2. The method for audio translation and labeling based on a multi-agent collaborative mechanism according to claim 1 is characterized in that: It also includes screening and judging according to the inclusion relationship of labels at all levels. For the labels at all levels determined by the conversations and key information in the document concentration, the labels with incorrect relationships that do not have inclusion relationships are removed, and the conversations and key information in the document concentration are re-input into the fine-tuned model until the conversations and key information in the document concentration have inclusion relationships in the labels at all levels.
3. The method for audio translation and labeling based on multi-agent collaborative mechanism according to claim 1 is characterized in that: in, Step S1 specifically includes: using the Whisper audio translation model based on the Transformer architecture to process the input audio file Q and translate the audio file Q into a text file T.
4. The method for audio translation and labeling based on a multi-agent collaborative mechanism according to claim 1, characterized in that: in, Step S2 introduces a model for domain knowledge and language style learning , where the model The network is represented by matrices and tensors during training. Abstract into a function function Word vector for input text file data To process, The parameters of the model are a set of weights, , , A and B are the training parameters to be updated, A is initialized using a random Gaussian distribution, and B is initialized with a zero matrix; Represents the pre-trained language model parameters, A, B are represented by After matrix transformation, it is obtained as bypass matrix; The update during training is expressed as: The training weight update steps are as follows: Lora formula: (1) Derivative of loss L: Where, represents the loss function, Represents the output of the model, which is composed of the input After weight matrix transformation, Represents the loss function For weight parameters gradient; Weight update amount: For the first step of back propagation, Where, Indicates the next moment The matrix parameter A, Indicates the current time The matrix parameter A, Represents the learning rate, which is a positive hyperparameter. Representation matrix The transpose of Indicates the The loss function value calculated in the iteration, Indicates the next moment The matrix parameter B, Indicates the current time Matrix parameter B; Since the matrix B is initialized to all 0s and the matrix A is initialized to a random Gaussian distribution, for any step t in the training, Where, represents the initial matrix parameter A, Represents the changing function of A with time step t, which is used to characterize the changing trend of A. Represents the changing function of B with time step t, which is used to characterize the changing trend of B; Where, represents the learning increment of B during the training process, Represents the learning increment of A during the training process; Substituting equation (4) into equation (5) we have Use language optimization through text file T to get instruction data set and set rank , for the model Perform training, import the model weights obtained from training, and obtain the model ; Input the text file T into the model Get the conversation set , Indicates the A conversation.
5. The method for audio translation and labeling based on multi-agent collaborative mechanism according to claim 1 is characterized in that: in, In step S3, a large language model is added by adding the prompt project From the dialogue Extract and structure time expressions, place names and summary paragraphs, and automatically associate event background with conversation context to obtain information sets ,in Represents a conversation set The key information in the article, including the summary, the place and time mentioned, Indicates the A key information.
6. The method for audio translation and labeling based on multi-agent collaborative mechanism according to claim 2 is characterized in that: in, In step S4, the conversation set With information set Splicing to get a document set .
7. The method for audio translation and labeling based on multi-agent collaboration mechanism according to claim 1 is characterized in that: in, The fine-tuned model in step S6 uses Qwen2.5-7B-Instruct as the large model base and is fine-tuned using the LoRA algorithm. The batch size is set to 16, the learning rate is 5e-5, and the temperature is 0.
8. The fine-tuned model is Qwen4TG.
8. The method for audio translation and labeling based on multi-agent collaboration mechanism according to claim 1 is characterized in that: in, Step S7 includes For document sets After word segmentation, the word after word segmentation is , Indicates the Word segmentation, each word passes through the encoding layer to generate a word embedding vector , expressed by the following formula: Where, Indicates the participles; The word is positionally encoded, where the position encoding is expressed as follows: in, Represents the dimension index of the position vector, 2 Indicates an even dimension, 2 +1 indicates an odd dimension, pos is the position index, and d is the word vector dimension; Indicates the location ,latitude The positional encoding value on , Indicates the location ,latitude Positional encoding value on ; Word vectors The word vector The weighted aggregation of each word is obtained through the attention mechanism to obtain the The semantic representation after the vector is updated is formally expressed as Where, represents the attention output result, Represents query matrix, key matrix, value matrix, represents the dimension of the key vector; Through linear mapping Obtained in The sequence is mapped to the feature space. Each set of linearly projected vector representations is called a head. ScaledDot-productAttention is then applied to each set of mapped sequences. Each attention head focuses on a certain aspect of semantic similarity. Multiple attention heads allow the model to focus on multiple aspects simultaneously. The formal representation is: Where, Indicates the An attention head, represents the linear transformation matrix that maps the input sequence to Q, used for the i-th attention head, represents the linear transformation matrix that maps the input sequence to K, Represents the linear transformation matrix that maps the input sequence to V; Concatenate multiple attention heads to get the result sequence: Where, Represents the output of the multi-head attention mechanism, Indicates the an attention head; Enter Perform residual connection with the multi-head attention result sequence and normalize it to get the input residual: Where, represents the output result of the multi-head attention sub-layer at the i-th position, is the standard layer normalization operation; Will Input feedforward sublayer, which contains two layers of linear transformation and an activation function: Where, represents the output of the feedforward sublayer, represents the input vector, represents the first-layer linear transformation, represents the second-layer linear transformation, represents the activation function; Will Add to the input residual and then normalize: Where, represents the final output after being processed by the feedforward neural network sublayer and normalized by residual connection and layer; After linear mapping to the vocabulary space through the decoding layer, the label content at each level is obtained.
9. The method for audio translation and labeling based on multi-agent collaboration mechanism according to claim 2 is characterized in that: in, The tags in step S5 include the first-level tags , secondary tag , third-level label , where the second-level tag is a subtag of a first-level tag, and the third-level tag is a subtag of a second-level tag; where, Indicates the First-level tags; Indicates that it belongs to The first level label Secondary tags, Indicates that it belongs to The first level label The second level label three-level labels; Among them, the document set with labels ,in Indicates labels at all levels, Means 1, 2, 3.
10. An analysis method for audio translation annotation tags based on a multi-agent collaborative mechanism, characterized in that: A labeling method comprising any one of claims 1 to 9; further comprising The set of labeled documents obtained for the input multiple audio files Q , extract the label content and reassemble it into a label set ,in Indicates the The first audio file Level label, =1,..., , Indicates the number of audios, =1,2,3; The tag set P is modeled and aggregated according to the dimensions of frequency, hierarchy, and co-occurrence relationship, and the visualization results are output to conduct horizontal comparison and vertical trend analysis.
Citation Information
Cited By
Multi-agent decision execution method and device, electronic equipment and storage medium
CN121052342A
Large language model dialogue type intelligent report generation method based on LoRA fine tuning
CN121072500A
Large language model dialogue intelligent report generation method based on LoRA fine-tuning
CN121072500B