Intelligent voice data defect detection method based on multi-modal large model

Through the intelligent speech data defect detection method based on multimodal large model, the answers to descriptive and verified questions are generated and evaluated, and the problem that the existing technology can only detect a single type of data defect is solved, and simultaneous identification and accurate detection of multiple types of defects are achieved.

CN120199232APending Publication Date: 2025-06-24BEIJING INST OF COMP TECH & APPL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237272.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-02
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art can only detect a single type of data defect, cannot identify multiple types of defects at the same time, and lack effective external information introduction, resulting in a false recall when the data defect is not obvious.

Method used

The intelligent speech data defect detection method based on multimodal large model is adopted. By designing a problem template, descriptive and verified questions are generated, and the questions and speech samples are input into the multimodal large model. After obtaining the answer, the answer is evaluated using the large language model or string matching, and the error score of the sample is calculated to judge the data defect.

Benefits of technology

The simultaneous identification of various types of data defects is achieved, the accuracy of data defect detection is improved, the situation of misrecalls is reduced, and the effectiveness of detection is enhanced by introducing external information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199232A_ABST
    Figure CN120199232A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent voice data defect detection method based on a multi-modal large model, and belongs to the field of artificial intelligence. The method comprises the steps that firstly, a problem template is designed, and related problems are selected according to tasks; then, inputting the question and the corresponding voice sample into a multi-modal large model to obtain a question answer about the voice sample; and finally, evaluating different types of problems by using large language models such as ChatGPT or character string matching, calculating an error score of the sample under a specific threshold value, and finally judging whether the tag is wrong or not by using the score. According to the method, a multi-modal large model is used for identifying data defects, and more accurate data defect detection is carried out in combination with a large amount of potential external knowledge in a large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to an intelligent voice data defect detection method based on a multimodal large model. Background Art

[0002] With the development of artificial intelligence technology, high-quality training data is becoming increasingly crucial for model development. Intelligent voice datasets are an essential resource in the field of artificial intelligence, especially in the development of speech recognition and synthesis technologies. Through high-quality intelligent voice datasets, researchers and developers can continuously optimize algorithms to enable intelligent devices to better understand human language and provide more intelligent and personalized interaction experiences. However, due to problems such as manual annotation errors, existing datasets usually contain data defects. Data defects mainly include label noise and poisoned data, etc. Among them, label noise refers to data with incorrect labels in the dataset, and poisoned data refers to a data poisoning method for backdoor attacks by adding triggers. Existing technologies distinguish normal samples from defective samples by comparing information such as the distribution characteristics of samples or the loss function of the model.

[0003] Current data defect detection technologies distinguish normal samples from defective samples by comparing information such as the distribution characteristics of samples or the loss function of the model. Existing methods usually can only detect a single type of data defect and cannot identify multiple types of defects simultaneously. For example, the annotation noise detection method can only effectively detect annotation noise and cannot detect poisoned data. At the same time, existing technologies compare the information of defective data with that of normal data, lacking the introduction of effective external information, which is prone to false recalls when data defects are not obvious. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] The technical problem to be solved by the present invention is how to provide an intelligent voice data defect detection method based on a multimodal large model to solve the problems that current data defect detection technologies can only detect a single type of data defect, cannot identify multiple types of defects simultaneously, lack the introduction of effective external information, and are prone to false recalls when data defects are not obvious.

[0006] (2) Technical Solutions

[0007] To solve the above technical problems, the present invention proposes an intelligent voice data defect detection method based on a multimodal large model, which includes the following steps: question generation, question answering, and answer scoring;

[0008] Step 1: Question Generation

[0009] In the question generation section, the speech tasks are divided into three categories, namely: audio text classification task, audio source classification task, and emotion classification task, and each category has corresponding question templates; these templates include descriptive questions and verification questions; the descriptive questions are generated manually and are related to the audio commonalities of the tasks; the verification questions are related to specific tags and tasks; this module focuses on generating descriptive questions and verification questions according to specific tasks and tags.

[0010] Step 2: Question Answering

[0011] In the question answering section, the generated descriptive questions and verification questions are input into the multi-modal large model to generate answers; by inputting the audio samples and the corresponding questions into the multi-modal large model, the answers of the samples are obtained.

[0012] Step 3: Answer Scoring

[0013] In the answer scoring section, the samples are scored by judging whether the answers are consistent with the corresponding tags; a large language model is used to judge whether the answers to the descriptive questions conform to the corresponding tags, and string matching is used to determine whether the answers to the verification questions are consistent with the tags; the samples are scored according to the evaluation results of the answers, and data errors are detected.

[0014] (III) Beneficial Effects

[0015] The present invention proposes an intelligent speech data defect detection method based on a multi-modal large model. First, question templates are designed, and relevant questions are selected according to the tasks; then, the questions and the corresponding speech samples are input into the multi-modal large model to obtain the question answers about the speech samples; finally, large language models such as ChatGPT or string matching are used to evaluate different types of questions, calculate the error scores of the samples under specific thresholds, and finally use this score to judge whether the tag is incorrect. Through training on a large amount of data, the large model can learn a large amount of data knowledge and better generalize to unseen data or scenarios. The multi-modal large model is designed to process multiple types of inputs (such as text, images, and sounds), which enables them to understand and process richer data types. The present invention uses a multi-modal large model to identify data defects and combines the potential large amount of external knowledge in the large language model to perform more accurate data defect detection. Brief Description of the Drawings

[0016] Figure 1 It is a flowchart of the technical solution of the present invention. Detailed Embodiments

[0017] To make the objectives, content, and advantages of the present invention clearer, the following further describes the detailed embodiments of the present invention in conjunction with the drawings and embodiments.

[0018] The overall technical solution of the present invention is as follows Figure 1 shown, including: problem generation, question answering, and answer scoring.

[0019] In the problem generation part, the speech tasks are divided into three categories, namely: audio text classification task, audio source classification task, and emotion classification task, and each category has a corresponding question template. These templates include descriptive questions and verification questions. The descriptive questions are generated manually and are related to the audio commonalities of the tasks. The verification questions are related to specific tags and tasks. This module focuses on generating descriptive questions and verification questions according to specific tasks and tags.

[0020] In the question answering part, the generated descriptive questions and verification questions are input into a multimodal large model to generate answers. By inputting the audio samples and the corresponding questions into the multimodal large model, the answers of the samples can be obtained.

[0021] In the answer scoring part, the samples are scored by judging whether the answers are consistent with the corresponding tags. A large language model is used to judge whether the answers to the descriptive questions conform to the corresponding tags, and string matching is used to determine whether the answers to the verification questions are consistent with the tags. The samples are scored according to the evaluation results of the answers, and data errors are detected.

[0022] Step 1: Problem generation

[0023] The present invention proposes to obtain audio information by querying a multimodal large model and compare the consistency of the information and the tags. Therefore, the first step is to design questions according to the given dataset x i is the input audio, y i is the label corresponding to the audio x i For task T, generate its corresponding descriptive questions and verification questions where Q i is the question set of the i-th sample, is the j-th question of the i-th sample, f desc (·) is a function for generating descriptive questions based on the task, f ver (·) is a function for generating verification questions, N q is the total number of generated questions, N desc is the number of descriptive questions. The descriptive questions are used to obtain general information about the audio samples, while the verification questions are used to confirm whether the audio samples conform to specific tags. Different types of questions need to be set for different tasks. For example, if the dataset consists of spoken digits, the question settings may focus on identifying spoken digits or differentiating between words with similar pronunciations.

[0024] The descriptive question aims to obtain the characteristics or features or background of the sample related to the main content, without relying on specific tags. For example, the descriptive question is "Describe the background of the audio you heard", and the answer to this question is related to the audio background content. The present invention designs different questions for different tasks. In the audio source classification dataset, the question can be "Describe the characteristics of the audio you heard". In the emotion dataset, the question can be "Describe the emotion of the voice you heard". These questions help to understand the overall information of the sample, so as to detect whether the overall information of the sample is consistent with the tag.

[0025] The verification question is generated according to the tag of the dataset. Its purpose is to extract more semantics from the audio, such as audio features, and confirm whether the tag is accurate, usually by directly asking for detailed information related to the tag. For example, given the tag "dog", the question can be "Can you hear the sound of a dog?". Since the audio types in the same dataset are the same, the present invention designs a template for this dataset. The template is manually created for different tags, such as "Can you hear the sound of {tag}?". The questions for each tag will be automatically generated according to the template. These questions directly check the matching degree between the specific tag of the sample and the audio, and verify the accuracy of the tag.

[0026] The template is manually created, and the tag is the data tag contained in the dataset. Suppose there is a dataset containing two types of data, "cat" and "dog", and the template is "Can you hear the sound of {tag}?", the following questions will be generated: "Can you hear the sound of a cat?" "Can you hear the sound of a dog?"

[0027] Step Two: Question Answering

[0028] The next step is to obtain the answers to the generated questions. Input the audio of each sample and the corresponding question into the multimodal large model and obtain the answers. In this step, the input question and the corresponding audio are converted into the answers of the multimodal large model, that is where M(·) is the multimodal large model, the corresponding answers, A i is the set of answers corresponding to the sample (x i ,y i ). The answer to the descriptive question is a description of the audio content. The answer to the verification question is "yes" or "no", which is used to describe whether the audio is consistent with the tag. For example, the descriptive question is "Describe the audio context", the label of the input audio is "one", and the answer of the large language model to the descriptive question is "I can hear someone saying 'one'.

[0029] Step Three: Answer Scoring

[0030] After obtaining the answer, the present invention needs to determine whether the answer is consistent with the label. Since the verification question is associated with a specific label in the dataset and the answer directly contains "yes" or "no", string matching can be used to obtain the result. If the answer is "yes", the result is "true", and "no" corresponds to "false". However, the answer to the descriptive question is a detailed description of the features of the audio or the content contained therein, and the answer does not directly contain "yes" or "no", so string matching cannot be used to obtain the result. Therefore, a large language model such as ChatGPT (this model is different from the multi-modal large model in step two) is selected to determine whether the label is consistent with the answer to the descriptive question. The present invention inputs the answer to the descriptive question and the corresponding label into the LLM, and the large language model determines whether the label matches the answer to the descriptive question, giving "true" or "false". where h(·) is a consistency function used to determine whether the answer is consistent with the label, indicating whether the answer to the j-th question of the i-th sample of the sample is consistent with the label, being "true" or "false", C i is the evaluation result of the i-th sample.

[0031] After evaluating the answers to the descriptive questions and verification questions, for each sample, the present invention can obtain an array containing "true" or "false". Subsequently, the present invention calculates a consistency score S for each sample based on this data i , The present invention sets a threshold θ. If the score S of the sample i is less than θ, the sample is considered a defective sample with data defects.

[0032] The present invention proposes an intelligent voice data defect detection method for a multi-modal large model. First, a question template is designed, and relevant questions are selected according to the task; then, the questions and the corresponding voice samples are input into the multi-modal large model to obtain the answers to the questions about the voice samples; finally, large language models such as ChatGPT or string matching are used to evaluate different types of questions, calculate the error score of the sample under a specific threshold, and finally use this score to determine whether the label is incorrect. Through training on a large amount of data, the large model can learn a large amount of data knowledge and better generalize to unseen data or scenarios. The multi-modal large model is designed to process multiple types of inputs (such as text, images, and sounds), which enables them to understand and process richer data types. The present invention uses the multi-modal large model to identify data defects and combines the potential large amount of external knowledge in the large language model for more accurate data defect detection.

[0033] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. An intelligent speech data defect detection method based on a multimodal large model, characterized in that: The method comprises the following steps: question generation, question answering and answer scoring; Step 1: Question generation In the question generation part, speech tasks are divided into three categories: audio text classification tasks, audio source classification tasks, and emotion classification tasks. Each category has a corresponding question template; these templates include descriptive questions and verification questions; descriptive questions are manually generated and related to the audio commonalities of the task; verification questions are related to specific labels and tasks; This module focuses on generating descriptive and verification questions based on specific tasks and labels; Step 2: Answer the questions In the question answering part, the generated description questions and verification questions are input into the multimodal large model to generate answers; By inputting the audio sample and the corresponding question into the multimodal large model, the answer to the sample is obtained; Step 3: Scoring answers In the answer scoring part, the samples are scored by judging whether the answers are consistent with the corresponding labels; Use large language models to determine whether answers to descriptive questions match the corresponding labels, and use string matching to determine whether answers to verification questions are consistent with the labels; score samples based on the evaluation results of the answers and detect data errors.

2. The intelligent voice data defect detection method based on multimodal large model according to claim 1, characterized in that: In the step 1, according to the given data set Design problem, x i is the input audio, y i It's audio x i Corresponding labels, N is the number of samples; for task T, generate its corresponding description question and verification question Where Q i is the problem set of the i-th sample, is the jth question of the i-th sample, f dese (·) is a function for generating descriptive questions based on the task, f ver (·) is the function for generating verification questions, N q is the total number of generated questions, N desc is the number of description questions.

3. The intelligent voice data defect detection method based on multimodal large model as claimed in claim 2, characterized in that: The description question aims to obtain the characteristics of the sample or the characteristics or background of the audio related to the main content without relying on specific labels. The description question helps to understand the overall information of the sample and thus detect whether the overall information of the sample is consistent with the label.

4. The intelligent voice data defect detection method based on multimodal large model as claimed in claim 2, characterized in that: Verification questions are generated based on the labels of the dataset, with the goal of extracting more semantics from the audio and confirming whether the labels are accurate by directly asking for detailed information related to the labels. Since the audio types in the same dataset are the same, a template is designed for the dataset, which is manually created for different labels, and questions for each label will be automatically generated based on the template. These questions directly check the match between the specific label and the audio of the sample, verifying the accuracy of the label.

5. The intelligent voice data defect detection method based on multimodal large model as claimed in claim 2, characterized in that: In step 2, the input question and the corresponding audio are converted into the answer of the multimodal large model, that is, Where M(·) is a multimodal large model, yes The corresponding answer is, A i is a sample (x i ,y i ) The answer to the description question is a description of the audio content, and the answer to the verification question is "yes" or "no", which is used to describe whether the audio is consistent with the label.

6. The intelligent voice data defect detection method based on multimodal large model according to claim 5, characterized in that: In step three, after obtaining the answer, determine whether the answer and the label are consistent; since the verification question is associated with a specific label in the data set, the answer directly contains "yes" or "no", allowing string matching to be used to obtain the result. If the answer is "yes", the result is "true", and "no" corresponds to "false".

7. The intelligent voice data defect detection method based on multimodal large model according to claim 5, characterized in that: In the step three, the answer to the description question is a detailed description of the characteristics of the audio or the content contained therein. The answer does not directly contain "yes" or "no", and string matching cannot be used to obtain the result; therefore, a large language model is selected to determine whether the label is consistent with the answer to the description question; the answer to the description question and the corresponding label are input into the LLM, and the large language model determines whether the label and the answer to the description question match, and gives "true" or "false".

8. The intelligent voice data defect detection method based on multimodal large model according to claim 7, characterized in that: The large language model is ChatGPT.

9. The intelligent voice data defect detection method based on multimodal large model according to claim 6 or 7, characterized in that: In the step three, Among them, h(·) is the consistency function, which is used to determine whether the answer and the label are consistent. Is the answer to the jth question of the i-th sample consistent with the label, "true" or "false", C i is the evaluation result of the i-th sample.

10. The intelligent voice data defect detection method based on multimodal large model according to claim 9, characterized in that: In step 3, the samples are scored according to the evaluation results of the answers, and data errors are detected, including: after evaluating the answers to the description questions and the verification questions, for each sample, an array containing "true" or "false" is obtained; then, based on this data, a consistency score S is calculated for each sample. i , Set the threshold θ, if the score of the sample S i If it is less than θ, the sample is considered to be a defective sample and there is a data defect.