Multi-modal large model-based surface access volume quality control method and system

By combining multimodal large-scale models with interview voice and questionnaire data, the low efficiency of face-to-face questionnaire data quality control and the problem of consistency verification were solved, realizing automated quality control and improving the authenticity and credibility of the data.

CN121808052AActive Publication Date: 2026-04-07INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing questionnaire surveys, face-to-face questionnaires have low data quality control efficiency, cannot effectively verify the consistency between the interviewer's recorded answers and the interviewee's actual verbal expression, and suffer from high labor costs, insufficient coverage and content depth.

Method used

A quality control method based on a multimodal large model is adopted. By simultaneously acquiring interview voice data and structured electronic questionnaire data, multimodal feature extraction and fusion analysis are performed to locate standardized voice answers, and consistency comparison is conducted to generate quality control output information.

Benefits of technology

It has automated and intelligentized the face-to-face survey quality control process, significantly improving efficiency, reducing labor costs, enabling full coverage of survey samples, identifying data distortion issues, and improving the authenticity and accuracy of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808052A_ABST
    Figure CN121808052A_ABST
Patent Text Reader

Abstract

The invention provides a surface access volume quality control method and system based on a multi-modal large model, and relates to the technical field of data quality management.The method comprises the steps that interview voice data synchronously collected for the same survey session and structured electronic questionnaire data associated with the interview voice data are obtained, the structured electronic questionnaire data comprises question information and a logical relationship; performing multi-modal feature extraction on the interview voice data based on the question information and the logic relationship, and performing fusion analysis by using a multi-modal large model so as to position and extract standardized voice answers corresponding to the questions in the questionnaire from continuous dialogue contents; consistency comparison is carried out on the extracted standardized voice answers and recorded answers of the corresponding questions in the structured electronic questionnaire data; and generating quality control output information based on a consistency comparison result. According to the method, full-sample automatic quality control can be realized, the limitation of manual sampling inspection is eliminated, and the surface access data quality and auditing efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data quality management technology, specifically to a method and system for surface access volume quality control based on a multimodal large model. Background Technology

[0002] In questionnaire surveys, especially face-to-face questionnaires, data quality control is crucial. Currently, the industry mainly relies on manual post-survey checks and simple rule checks built into electronic questionnaire platforms (such as logical checks and setting quality control questions). While these methods can detect some logical errors or obvious perfunctory answers, they have two fundamental limitations: first, they are inefficient and costly; second, they cannot access the "black box" of the interview process, making it difficult to verify the consistency between the answers recorded by the investigator and the respondents' actual verbal expressions.

[0003] This contradiction is particularly pronounced in open-ended interview scenarios involving subjective evaluations and descriptions of feelings. For example, when a questionnaire asks respondents to rate a service on a scale of 1-5, they often don't directly state their score but instead express their feelings in descriptive language. The interviewer needs to interpret this complex verbal description and translate it into a score for the answer choices. In this process, the interviewer's subjective interpretation, oversights, or even unintentional mishearing can easily lead to discrepancies between the final recorded answer and the respondent's true intentions. Furthermore, in extreme cases, there may be instances of interviewers failing to strictly follow procedures, answering on behalf of others, or even filling out questionnaires themselves—all violations of regulations.

[0004] To address these issues, existing technologies include solutions that support simultaneous recording of interviews, aiming to ensure data authenticity through post-interview audio verification. For example, the Third National Border Region Survey utilized an automatic recording function on a platform to allow auditors to reconstruct interview scenarios. However, this method faces significant challenges in practice: auditors must listen to lengthy recordings one by one and manually compare their content with the written answers to each question, a massive and extremely tedious task. This results in long review cycles and high labor costs, often allowing only sampling of a small number of samples in practice, failing to achieve comprehensive coverage. Therefore, the risks of numerous potential data inconsistencies and low-quality reporting remain largely uncontrollable.

[0005] In summary, existing quality control methods have significant shortcomings in terms of efficiency, coverage, and content depth. How to automate and intelligently utilize interview audio content to achieve efficient and accurate verification of questionnaire consistency and process authenticity has become a pressing technical challenge for improving the quality of survey data. Summary of the Invention

[0006] To overcome at least some of the problems existing in related technologies, this application proposes a facet access questionnaire quality control method and system based on a multimodal large model. The system uses a multimodal large model to achieve automated quality control processing, which is conducive to effectively improving the overall quality of survey data.

[0007] First aspect This application provides a surface-access volume quality control method based on a multimodal large model, which includes: Acquire interview audio data and associated structured electronic questionnaire data that are collected synchronously for the same survey session. The structured electronic questionnaire data includes question information and logical relationships. Based on the question information and logical relationships, multimodal features are extracted from the interview voice data, and a multimodal large model is used for fusion analysis to locate and extract standardized voice answers corresponding to the questions in the questionnaire from the continuous dialogue content. The extracted standardized voice answers are compared with the recorded answers to the corresponding questions in the structured electronic questionnaire data for consistency. Based on the consistency comparison results, quality control output information is generated.

[0008] In some possible implementations, the multimodal feature extraction includes: Extract the transcribed text sequence and audio feature vector from the interview speech data to form text channel feature representation and audio feature channel representation; The transcribed text sequence includes timestamps and speaker role information.

[0009] Among some possible implementations, the use of a multimodal large model for fusion analysis includes: The questionnaire's question information and logical relationships are modeled as graph structure constraints; When the multimodal large model extracts answers, a joint loss function is constructed for global optimization. The joint loss function includes a semantic loss term for constraining the consistency between the answer and the speech content, and a logical loss term for penalizing violations of the graph structure constraints.

[0010] Among some possible implementations, the use of a multimodal large model for fusion analysis also includes: For the question to be answered, the first answer probability distribution and the second answer probability distribution are obtained based on the text channel feature representation and the audio feature channel representation, respectively. The probability distributions of the first and second answers are weighted and fused to obtain a comprehensive answer distribution. The standardized voice answer is determined based on the comprehensive answer distribution.

[0011] In some possible implementations, the method further includes: After obtaining the comprehensive answer distribution through weighted fusion, a self-consistent reasoning step is further performed to verify and optimize the standardized speech answer; The self-consistent reasoning step includes: taking the text channel feature representation and the audio feature channel representation as input, performing multiple independent reasonings to obtain multiple intermediate answer distributions; aggregating the intermediate answer distributions to obtain an aggregated answer distribution; and generating a verification answer based on the aggregated answer distribution.

[0012] In some possible implementations, the method further includes: Based on the comprehensive answer distribution obtained by weighted fusion, the uncertainty index of the standardized speech answer is calculated, and the step of generating quality control output information refers to this uncertainty index.

[0013] In some possible implementations, the consistency comparison includes: Calculate the semantic consistency score and behavioral consistency score at the question level, and combine them to obtain the consistency index at the question level and questionnaire level.

[0014] In some possible implementations, the calculation of the question-level semantic consistency score includes: The standardized voice answer and the recorded answer are input into a trained consistency determination model to obtain classification labels, wherein the classification labels are selected from at least one of the label set including consistent, inconsistent, semantically contradictory and information missing.

[0015] In some possible implementations, the generation of quality control output information includes: Based on the results of the consistency comparison, multi-dimensional quality control data is generated. The multi-dimensional quality control data includes at least questionnaire-level quality control scores, question-level anomaly indicators, surveyor-level quality indicators, and early warning information. The multi-dimensional quality control data is displayed through a visualization interface, which supports data filtering, sorting, and drill-down analysis by project, time period, region, or survey personnel.

[0016] Second aspect This application provides a questionnaire quality control system based on a multimodal large model, the system comprising: The data acquisition module is used to acquire interview voice data and associated structured electronic questionnaire data collected synchronously for the same survey session. The structured electronic questionnaire data includes question information and logical relationships. The multimodal analysis module is used to extract multimodal features from the interview voice data based on the question information and logical relationships, and to perform fusion analysis using a multimodal large model to locate and extract standardized voice answers corresponding to the questions in the questionnaire from the continuous dialogue content. The consistency comparison module is used to compare the extracted standardized voice answers with the recorded answers of the corresponding questions in the structured electronic questionnaire data. The quality control output module is used to generate quality control output information based on the results of the consistency comparison.

[0017] The multimodal large-scale model-based quality control method for face-to-face interviews provided in this application automates and automates the quality control process by simultaneously fusing interview audio and structured questionnaire data, significantly improving efficiency and reducing labor costs. The multimodal large-scale model, relying on question logic and semantic understanding capabilities, accurately locates the colloquial expressions corresponding to each question in the dialogue, transforming respondents' subjective descriptive language into standardized audio answers, thus achieving accurate verification of recorded answers and true meaning. Figure 1 Consistency verification: This method can fully cover all survey samples, completely eliminating the limitations of manual sampling and shortening the review cycle. At the same time, it can effectively identify data distortion problems caused by investigators' misunderstandings, omissions in summarization, mishearing, etc. The generated structured quality control output information not only provides a clear basis for data correction, but also provides data support for investigator behavior supervision and process optimization, thereby systematically improving the authenticity, accuracy and credibility of face-to-face questionnaire data. It has built an efficient, reliable and traceable new quality control paradigm for scenarios such as large-scale social surveys and market research surveys that involve face-to-face questionnaire surveys. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a surface access volume quality control method based on a multimodal large model, provided in one embodiment of this application. Figure 2 This is a schematic diagram illustrating an application scenario of the multimodal large model-based surface access volume quality control method in one embodiment of this application; Figure 3 for Figure 2 The diagram illustrates the overall process steps corresponding to the application scenario shown. Detailed Implementation

[0019] To make the purpose, technical solution and advantages of this application clearer, the technical solution of this application will be described in detail below.

[0020] As described in the background section, existing technologies include solutions that support simultaneous recording of interviews, aiming to ensure data authenticity through post-interview audio verification. For example, the Third National Survey of Border Regions utilized an automatic recording function to allow auditors to reconstruct interview scenarios. However, this method faces significant challenges in practice: auditors must listen to lengthy recordings one by one and manually compare their content with the written answers to each question, resulting in a massive and extremely tedious workload. This leads to long review cycles and high labor costs, often allowing only random checks on a small sample size, failing to achieve comprehensive coverage. Therefore, the risks of numerous potential data inconsistencies and low-quality reporting remain largely uncontrollable.

[0021] Based on this, this application proposes a method and system for surface access questionnaire quality control based on a multimodal large model. The system uses a multimodal large model to achieve automated quality control processing, thereby effectively improving the overall quality of survey data.

[0022] First, let me introduce the overall concept of this application. This application aims to provide an efficient and automated questionnaire quality control solution. It utilizes a multimodal artificial intelligence model to perform speech recognition and semantic understanding on recorded interviews, automatically comparing the interview audio with the electronic questionnaire answers to identify data inconsistencies, false reporting, or low-quality records. Compared to traditional manual sampling methods, this technical solution can comprehensively cover every interview, significantly improving the timeliness and accuracy of quality control while reducing labor costs.

[0023] Specifically, such as Figure 2 and Figure 3 As shown, in the application scenario of this application's technical solution, when the researcher uses the App to conduct a questionnaire interview, the relevant devices record and save the interview dialogue in real time; after the researcher completes the electronic questionnaire, the audio and questionnaire answers are uploaded to the server together. Figure 2 The system utilizes a cloud-based quality control server. On the server side, a multimodal large-scale model is used to transcribe and semantically analyze the audio content, automatically extracting key points from the respondents' answers. The extracted audio answers are then compared question-by-question with the respondents' completed questionnaires to determine semantic consistency. Finally, the system generates a quality control score and anomaly alerts for each interview. If the audio content does not match the recorded answers, the system will mark the inconsistency. If anomalies are detected (e.g., poor consistency across multiple questions or unusually short recording duration), it indicates potential false reporting or perfunctory interviewing. Furthermore, project teams / quality control personnel can view the relevant quality control results through the platform and provide personalized quality control feedback reports to respondents to help them improve their interviewing skills and recording habits, thereby improving overall data quality.

[0024] It should also be noted that, based on the development of AI technology, existing multimodal large language models already possess qualified speech-text-instruction understanding capabilities, such as Qwen-Omni and Qwen-VL. This provides a feasible basis for adopting AI automatic quality control in the technical solution of this invention. This application does not involve the improvement of the multimodal large language model itself, but rather the application solution based on such models in specific scenarios.

[0025] The following is a specific embodiment of the surface access volume quality control method based on a multimodal large model proposed in this application.

[0026] like Figure 1 As shown, in one embodiment, the surface access volume quality control method based on a multimodal large model includes: Step S110: Obtain interview voice data and associated structured electronic questionnaire data that are collected synchronously for the same survey session. The specific structured electronic questionnaire data includes question information and logical relationships. Specifically, in practice, combined with Figure 2 As shown, this process includes recording and filling out the questionnaire simultaneously via a mobile terminal (running the relevant app), and then sending the voice data and questionnaire data to the server after associating them with the same session identifier; In this process, the terminal equipped with the app is primarily responsible for presenting the questionnaire, completing electronic answers, and recording interview audio, ensuring a one-to-one correspondence between the recordings and the questionnaire data. In practice, the app first retrieves the questionnaire definition from the server, including the question number, stem, question type (single choice, multiple choice, scale, numerical, open-ended, etc.), candidate options, question importance level, question openness, and logical transition rules for each question. Based on this, the app displays the questions sequentially on the interface, and researchers fill in the answers one by one based on the respondents' verbal responses. For example, when the researcher clicks "Start Interview," the app automatically starts recording, capturing the audio signal of the entire interview. During the questionnaire interaction, the app can also record behavioral data such as the time spent on each question and the timestamp of completion. After the interview, the app stops recording and associates the generated audio file with the current questionnaire response using a unique session ID.

[0027] Furthermore, the mobile terminal (based on App functionality) also possesses offline working capabilities. When in an environment with unstable or interrupted network connectivity, the App can securely cache the interview audio data and structured electronic questionnaire data generated during the current session locally. Once the network connection is restored, the associated data packets are automatically re-transmitted to the server, thereby ensuring data integrity and the reliability of the methodology in various real-world scenarios. In other words, for scenarios with unstable network connectivity, the App can securely cache audio recordings and questionnaire data locally first, and then upload them all at once when the network is restored.

[0028] For data uploads from mobile terminals, the server receives the data, verifies and separates it for storage, and schedules subsequent processing through an asynchronous task queue.

[0029] Specifically, for example, a data receiving interface can be set up on the server side to receive upload requests from the app. Each upload should include at least the following: session ID, corresponding audio file, electronic questionnaire answers (e.g., using structured JSON), and session metadata. After verifying the integrity of the uploaded data, the server stores the audio file in object storage or a file system, stores the questionnaire answers and metadata in a database, and maintains the mapping relationship between session IDs, audio paths, and questionnaire records. To improve processing throughput, the server can place each session ID in a task queue for asynchronous processing by subsequent voice processing and quality control pipelines. This supports large-scale concurrent questionnaire survey scenarios.

[0030] Based on this, such as Figure 1 As shown, in step S120, based on the question information and logical relationship, multimodal features are extracted from the interview voice data, and a multimodal large model is used for fusion analysis to locate and extract standardized voice answers corresponding to the questions in the questionnaire from the continuous dialogue content.

[0031] Specifically, as one implementation method, the above-mentioned multimodal feature extraction includes: extracting the transcribed text sequence and audio feature vector of the interview speech data to form text channel feature representation and audio feature channel representation, wherein the transcribed text sequence includes timestamps and speaker role information.

[0032] In practical implementation, the multimodal feature extraction process first involves speech processing. During speech processing, necessary preprocessing is performed on the audio, such as sampling rate resampling, silence detection, and noise suppression, to improve the robustness of subsequent recognition. Then, two processing channels are established: First, there's the speech recognition channel. This involves calling the speech understanding module within a multimodal large-scale model or a standalone automatic speech recognition model. This automatic speech recognition model can be implemented based on mature cloud service APIs (such as those from Alibaba Cloud and iFlytek), transcribing the audio stream into a text sequence. The transcription result includes a timestamp and speaker tags, and can be represented as several records. Furthermore, as a preferred approach, the speech of the researcher and the respondent can be separately labeled using a speaker separation algorithm (e.g., a clustering method based on x-vectors of voiceprint features, or a dedicated speaker logging tool such as PyAnnote) to distinguish between questions and answers during subsequent answer extraction.

[0033] Second, the audio feature channel uses a multimodal large-scale speech encoder or a dedicated speech feature network to map audio segments into high-dimensional vector representations. This information is used for subsequent answer prediction directly from the audio feature space. By simultaneously preserving both transcribed text and speech features, the technical solution of this application constructs a multi-channel feature input of "text channel + audio channel," providing a foundation for subsequent semantic understanding and fusion.

[0034] Specifically, as one implementation method, fusion analysis is performed using a multimodal large model, including: The questionnaire's question information and logical relationships are modeled as graph structure constraints. When extracting answers in a multimodal large model, a joint loss function is constructed for global optimization. The joint loss function includes a semantic loss term to constrain the consistency between the answer and the speech content, and a logical loss term to penalize violations of the graph structure constraints.

[0035] In this application, the entire questionnaire can be modeled as a graph structure with logical constraints, where each question... Let be a node in the graph. Each node has attributes such as question type, set of options, importance, and openness. Nodes are connected by edges. This represents logical or constraint relationships, including jump relationships, mutual exclusion relationships, and dependencies. When extracting answers, this application does not process each question independently, but instead constructs a joint loss function to globally optimize the answers to all questions. For example, the expression for the joint loss function here is: (1) In expression (1), Used to constrain the semantic consistency between the extracted answers and the audio content. Used to penalize answer combinations that violate the internal logical constraints of the questionnaire. This represents a set of problem pairs that have logical constraints. These are the weighting coefficients. Indicates the topic The standardized audio answers are extracted from interview audio content aligned with the question. These are structured data consistent with the question type: for single-choice / scale questions, they are option numbers or rating values; for multiple-choice questions, they are a set of options; for numerical questions, they are numerical values ​​(and / or units); and for open-ended questions, they are standardized text. Through joint optimization, the accumulation of single-question extraction errors can be avoided, improving the overall consistency of the answer set.

[0036] In this embodiment, the fusion analysis using a multimodal large model further includes: for the question to be answered, obtaining a first answer probability distribution and a second answer probability distribution based on text channel feature representation and audio feature channel representation, respectively; weighted fusion of the first answer probability distribution and the second answer probability distribution to obtain a comprehensive answer distribution; and determining standardized speech answers based on the comprehensive answer distribution. In other words, this application employs a multi-channel fusion mechanism in the semantic understanding and answer extraction part, under the constraint of questionnaire structure. Specifically: Regarding multi-channel fusion, this application utilizes both transcribed text and audio features for answer prediction. For each question... The multimodal large model provides the probability distribution of answers based on text channels. And the probability distribution of answers based on audio feature channels. The two are then combined according to their weights to obtain a comprehensive answer distribution, which can be represented by the following expression: (2) Expression (2), This represents the probability distribution of the combined answer after fusing the text and audio channels. It can be a preset constant, or it can be adaptively adjusted according to the current audio quality or transcription confidence.

[0037] Based on expression (2), the final extracted audio answer Through the The method that yields the highest probability is used. By combining text and audio perspectives, the robustness of answer prediction can be enhanced even when speech recognition errors are large, accents are heavy, or environmental noise is strong.

[0038] As a preferred implementation, a self-consistent reasoning mechanism can be introduced for the answer output of each question to improve the stability of answer extraction. Specifically, the method of this application further includes: after obtaining a comprehensive answer distribution through weighted fusion, a self-consistent reasoning step is further performed to verify and optimize the standardized speech answer; the self-consistent reasoning step includes: taking text channel feature representation and audio feature channel representation as input, performing multiple independent reasonings to obtain multiple intermediate answer distributions; aggregating the intermediate answer distributions to obtain an aggregated answer distribution; and generating a verification answer based on the aggregated answer distribution.

[0039] In other words, by independently calling the multimodal large model multiple times on the same audio and questionnaire (by changing the sampling seed or fine-tuning the prompts), we can obtain K sets of answer distributions. Then average them as shown in the following expression: (3) In expression (3), by The final output answer is determined, and the variance between each inference is used as a stability reference indicator. If the variance value is high, it indicates that the model's judgment of the answer to the question fluctuates greatly and has low stability. , For the title The set of candidate answers, and the output of each reasoning step is a correct answer. The probability distribution is calculated. Based on the aggregated answer distribution, the candidate answer with the highest probability value is selected as the verification answer.

[0040] Regarding the verification or optimization of standardized speech answers, the following approach can be adopted: when the verification answer differs from the answer determined based on the comprehensive answer distribution, the verification answer shall be used as the final standardized speech answer.

[0041] In addition, this embodiment also includes: calculating the uncertainty index of standardized voice answers based on the comprehensive answer distribution obtained by weighted fusion, and the subsequent steps of generating quality control output information can refer to this uncertainty index.

[0042] Specifically, regarding the question Based on the distribution of comprehensive answers The uncertainty index, such as when using entropy as a measure, is calculated using the following expression: (4) in, Indicates the topic The uncertainty index; i is the question index; a is the candidate answer to the question (it takes all candidate answers); For the question obtained based on multi-channel fusion The overall probability of the candidate answer a (satisfying that the summation of all normalized a's is 1); Indicates the topic The summation of all candidate answers is performed; log is a logarithmic operation (the base can be the natural logarithm or base 2, which does not affect the relative size, but only changes the scaling factor).

[0043] It's easy to understand that the higher the uncertainty, the lower the model's confidence in the answer to the question, and thus the more likely it is to succeed. As a weighted reference in subsequent quality control, questions with high uncertainty are specially marked to suggest priority for manual review.

[0044] In this way, in the semantic understanding and answer extraction stages, this application can significantly enhance the reliability and interpretability of the algorithm by implementing joint questionnaire structure constraints, multi-channel fusion, uncertainty estimation and self-consistent reasoning.

[0045] It should be noted that, in this application, "standardized voice answers" specifically refers to structured answer data extracted from interview voice data using a multimodal large-scale model, which matches the requirements of structured electronic questionnaire questions. Its format may vary depending on the question type, including option numbers, numerical values, or standardized text, and is used for consistency comparison with recorded answers in the electronic questionnaire.

[0046] Continue back Figure 1 Based on the standardized voice answers extracted in step S120, step S130 is performed to compare the extracted standardized voice answers with the recorded answers of the corresponding questions in the structured electronic questionnaire data.

[0047] For clarity, the answers extracted from the audio and the answers recorded in the electronic questionnaire (also known as form answers) constitute two sets, with the elements in each set represented by [missing information]. and Specifically, in this application's technical solution, consistency modeling and comparison are performed at three levels: question level, questionnaire level, and researcher level. Correspondingly, consistency comparison is performed in step S130, including: calculating the semantic consistency score and behavioral consistency score at the question level, and comprehensively obtaining the consistency index at the question level and questionnaire level.

[0048] For example, at the question level, semantic similarity metrics can be used for evaluation. and Semantic consistency between them, such as encoding them separately as vector representations. and The cosine similarity index is calculated based on the following expression: (5) In expression (5), As an evaluation metric, the closer it is to 1, the more similar the content is.

[0049] Furthermore, this application also introduces a behavioral consistency factor. Using behavioral characteristics such as answering time to aid in judgment. For example, suppose... This represents the approximate duration of the response to question i in the recording. The time researchers spend on a given question within the app can be defined as: (6) In expression (6), This indicates that the parameters are adjusted. If the researcher barely pauses to answer the question before filling in the answer, but the recording shows a longer response, or vice versa, then the behavioral consistency is low.

[0050] Furthermore, by combining the aforementioned semantic consistency and behavioral consistency indicators, a question-level comprehensive consistency score can be defined. For example, the overall consistency score at the question level can be expressed as follows: (7) In expression (7), It is used to balance the influence of both semantics and behavior.

[0051] Furthermore, as a specific implementation method, this application employs a dynamic threshold strategy to set consistency criteria for different question types and attributes, such as setting a threshold for each question. Depending on the openness of the question and importance level Adaptive adjustment, for example, threshold adjustment determination based on the following expression: (8) In expression (8), The threshold represents the basic threshold, and α and β are the weights. In practice, a higher threshold can be set for factual and high-importance questions, while a relatively lenient threshold can be set for open-ended and low-importance questions, in order to better reflect the actual survey scenario.

[0052] In some embodiments, as another specific implementation, the determination of the question-level semantic consistency score can also be achieved directly through a consistency determination model. That is, in these embodiments, the process of calculating the question-level semantic consistency score includes: Standardized voice answers and recorded answers are input into a trained consistency determination model to obtain classification labels, wherein the classification labels are selected from at least one of the following sets of labels: consistent, inconsistent, semantically contradictory, and information missing.

[0053] Specifically, in practical implementation, rule-based judgment based on similarity and behavior can be pre-trained using a specialized consistency judgment model to analyze the question text. Audio-based answer extraction and form answers As a joint input, the output consists of multiple category labels, such as "consistent," "inconsistent," "semantic contradiction," and "missing information." Formally, this can be represented as: (9) In expression (9), This refers to a classification model or a finely tuned large language model obtained through sample learning. This model can capture complex semantics in natural language, such as negation and conditional relations, making consistency determination more precise and reliable.

[0054] Regarding the questionnaire-level consistency in step S130, this application uses a comprehensive consistency score based on all items. The questionnaire-level consistency index is calculated. For example, the weighted average calculation method shown in the following expression can be used to obtain the questionnaire-level consistency index. : (10) Expression (10), weight The parameters can be set based on factors such as the importance and uncertainty of the questions. In practice, questionnaire-level indicators can be converted into quality control scores to determine the overall consistency between the questionnaire record and the audio recording.

[0055] Regarding the researcher-level consistency in step S130, this application constructs an individual quality profile by aggregating the questionnaire-level consistency indicators of all questionnaires handled by the corresponding researcher. For example, the consistency indicator for researcher E can be defined as: (11) In expression (11), This indicates the number of questionnaires completed by the researcher. For the first The consistency score of the questionnaires can be used. Furthermore, statistical characteristics such as the consistency distribution of the survey participants and the proportion of abnormal questionnaires can be analyzed to identify whether there is systematic false reporting or data entry bias.

[0056] Based on step S130, step S140 is performed to generate quality control output information based on the consistency comparison results.

[0057] Specifically, in this step, quality control output information is generated, including: generating multi-dimensional quality control data based on the results of consistency comparison. The multi-dimensional quality control data includes at least questionnaire-level quality control scores, question-level anomaly indicators, surveyor-level quality indicators, and early warning information; and displaying the multi-dimensional quality control data through a visualization interface, which supports data filtering, sorting, and drill-down analysis by project, time period, region, or surveyor.

[0058] In other words, based on the consistency indicators at the question level, questionnaire level, and researcher level obtained in step S130, and based on specific software configurations, this application can generate multi-dimensional quality control results, including quality control scores for individual questionnaires, question-level anomaly lists, researcher quality rankings, and warning lists.

[0059] In implementation, such as Figure 2 , Figure 3 As shown, a quality control platform can be built for project managers based on actual needs. The platform supports filtering and sorting questionnaires by time period, region, survey personnel, etc., as well as focusing on low consistency questionnaires and their corresponding audio segments.

[0060] Furthermore, for researchers, based on specific service configurations, this application can provide personalized quality control feedback reports, indicating which question types or respondents are prone to recording bias, helping them improve their interviewing skills and recording habits. On the other hand, through subsequent human-machine collaboration, the results of manual review can be fed back for further fine-tuning of the multimodal large model and consistency judgment model, enabling continuous optimization of quality control capabilities.

[0061] The multimodal large-scale model-based quality control method for face-to-face questionnaires provided in this application automates and automates the quality control process by simultaneously fusing interview audio and structured questionnaire data, significantly improving efficiency and reducing labor costs. The multimodal large-scale model, relying on question logic and semantic understanding capabilities, accurately locates the colloquial expressions corresponding to each question in the dialogue, transforming respondents' subjective descriptive language into standardized audio answers, thus achieving accurate verification of recorded answers and true meaning. Figure 1 Consistency verification: This method can fully cover all survey samples, completely eliminating the limitations of manual sampling and shortening the review cycle. At the same time, it can effectively identify data distortion problems caused by investigators' misunderstandings, omissions in summarization, mishearing, etc. The generated structured quality control output information not only provides a clear basis for data correction, but also provides data support for investigator behavior supervision and process optimization, thereby systematically improving the authenticity, accuracy and credibility of face-to-face questionnaire data. It has built an efficient, reliable and traceable new quality control paradigm for scenarios such as large-scale social surveys and market research surveys that involve face-to-face questionnaire surveys.

[0062] In one embodiment, this application also proposes a questionnaire quality control system based on a multimodal large model, the system comprising: The data acquisition module is used to acquire interview voice data and associated structured electronic questionnaire data collected synchronously for the same survey session. The structured electronic questionnaire data includes question information and logical relationships. The multimodal analysis module is used to extract multimodal features from interview voice data based on question information and logical relationships, and to perform fusion analysis using a multimodal large model to locate and extract standardized voice answers corresponding to the questions in the questionnaire from continuous dialogue content. The consistency comparison module is used to compare the extracted standardized voice answers with the recorded answers of the corresponding questions in the structured electronic questionnaire data. The quality control output module is used to generate quality control output information based on the results of consistency comparison.

[0063] Regarding the questionnaire quality control system in the above-mentioned embodiments, the specific methods of each module's operation have been described in detail in the embodiments of the relevant method, and will not be elaborated here.

[0064] Based on the above embodiments, it can be seen that this application constructs an end-to-end automated quality control process around "face-to-face interview recording + electronic questionnaire + multimodal large-scale model quality control". Compared with existing solutions that mostly rely on manual sampling or simple rule verification, it has many substantial innovations, as follows: First, at the semantic understanding and answer extraction level, this application does not treat each question as an independent task. Instead, it models the entire questionnaire as a logically constrained graph structure, introduces a joint loss function, and performs global joint reasoning on the audio answers of all questions. By explicitly introducing logical constraints between questions into the loss function, penalties are imposed on answer combinations that violate the internal jump logic and mutual exclusion relationships of the questionnaire. This allows the model to consider the structural consistency of the entire questionnaire while extracting answers to individual questions, thereby reducing the accumulation of local errors caused by independent prediction of each question. This design of "questionnaire structural constraints + joint reasoning" is significantly different from the conventional "ASR + single question extraction" pipeline.

[0065] Secondly, regarding the utilization of multimodal features, this application proposes an answer prediction mechanism that integrates audio and text. Simultaneously, it utilizes feature vectors extracted from the transcribed text obtained through speech recognition and the original audio through a speech encoder. These feature vectors are then used to output answer probability distributions in a large multimodal model, and subsequently fused using adjustable weights to obtain a more robust comprehensive answer distribution. Compared to semantic extraction relying solely on ASR text, this dual-view fusion mechanism maintains high extraction accuracy even with complex accents, significant background noise, or transcription errors, significantly enhancing its adaptability to real-world interview scenarios.

[0066] Furthermore, regarding the control of answer uncertainty and stability, this application introduces uncertainty estimation and self-consistent reasoning mechanisms during the answer extraction process. On the one hand, by calculating indicators such as the entropy value of the answer probability distribution, the confidence level of each question extraction result is quantified. Questions with low confidence levels are marked as "fuzzy questions," and their weight is reduced in subsequent quality control, with manual review recommended as a priority. On the other hand, by performing multiple independent inferences on the same session and voting or averaging the multiple results, the stability of the final answer is improved, and the variance between the inference results is used as a stability reference. This "built-in confidence interval" design effectively avoids misjudging model uncertainty as researcher error, improving the reliability of quality control decisions.

[0067] In terms of consistency comparison and quality control indicator design, this application establishes for the first time a multi-level consistency indicator system at the question level, questionnaire level, and researcher level in questionnaire quality control. At the question level, the system not only calculates the semantic similarity between voice answers and form answers, but also introduces a behavioral consistency factor, comparing the answer duration of a question in the recording with the dwell time of that question in the app to assess whether the question actually occurred during the face-to-face interview from a time behavior perspective; and constructs a comprehensive consistency score through geometric or weighted combination. At the questionnaire level, the quality control score of the entire questionnaire is obtained by weighted aggregation of the consistency scores of each question; at the researcher level, a quality profile of the researcher is formed by statistically analyzing the consistency distribution and abnormality rate of all their questionnaires, which is used to identify systemic false reporting or data entry deviations. This "semantic + behavioral" multi-view Figure 1 Consistency modeling, combined with a multi-layered indicator system, enables this invention to conduct comprehensive analysis from single issues to individuals, and from local anomalies to systemic deviations.

[0068] Furthermore, in terms of the judgment strategy, this application adopts a combination of dynamic thresholds and a specialized consistency judgment model. For different question types and attributes (such as factuality, openness, and importance level), the system adaptively adjusts the semantic consistency judgment threshold based on the question's metadata, achieving a refined strategy that is strict for key factual questions and moderately lenient for open-ended questions. Simultaneously, by training a specialized question-answer consistency classification model, which uses question text, audio-extracted answers, and form answers as unified input, it outputs multi-category labels such as "consistent," "inconsistent," "semantic contradiction," and "information missing," enabling the identification of complex linguistic phenomena such as negation expressions and conditional relationships. Compared to simply using vector similarity with a fixed threshold, this invention significantly improves both the accuracy and interpretability of consistency judgment.

[0069] In summary, this application achieves breakthroughs in "questionnaire structure modeling + multimodal fusion semantic understanding + uncertainty management + multi-view..." Figure 1 The system combines consistency comparison and multi-layered quality control indicators to form a holistic innovative solution that effectively solves the problems of traditional questionnaire quality control relying on manual sampling, inability to utilize the semantic content of audio recordings, and inability to form quantitative profiles of survey personnel, thus significantly improving the authenticity and usability of face-to-face questionnaire data.

[0070] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims. Without changing the core idea of ​​this invention, various optional technical paths and variations can be adopted in the specific implementation process to adapt to different technical conditions and application needs.

[0071] For example, in selecting a multimodal large model, besides prioritizing multimodal models with speech-text-instruction understanding capabilities such as Qwen-Omni and Qwen-VL, a combination of "dedicated ASR + large text model" can also be used. For instance, open-source multilingual speech recognition models such as Whisper can be used for high-precision transcription, and then large language models such as GPT-4, Claude, and Gemini can be used for question-answer extraction and consistency assessment. As long as the overall process still includes "audio transcription or encoding – questionnaire-based answer extraction – consistency comparison between voice and form answers," it falls within the scope of this invention.

[0072] For example, regarding feature channel configuration, the audio-text dual-channel fusion mechanism proposed in this application is not the only implementation. For scenarios with limited hardware computing power or high real-time requirements, a simplified version can be deployed first, using only the speech recognition text channel for answer extraction and comparison, and then gradually introducing the audio feature channel to enhance robustness. For projects requiring extremely high accuracy, more complex fusion strategies can be adopted, such as introducing an attention mechanism to dynamically allocate weights between the two channels, or adaptively adjusting the fusion coefficients based on the audio signal-to-noise ratio to achieve adaptive processing under different sound quality conditions.

[0073] Regarding consistency determination methods, besides using cosine similarity of embedding vectors combined with dynamic thresholds, end-to-end trained consistency classifiers can also replace some rules. For example, a lightweight text matching network can be built, taking "question + audio answer + form answer" as input and directly outputting consistency labels, suitable for local deployment or resource-constrained environments. Alternatively, a large model can be fine-tuned with a small number of samples to better suit the expression characteristics of questionnaires in specific industries (such as healthcare, education, and finance). For sensitive projects, an expert knowledge rule layer can be added to perform secondary correction on the model's determination results, constructing a hybrid consistency framework of "rules + model".

[0074] Regarding the indicator system and quality control strategies, consistency indicators at the question level, questionnaire level, and researcher level can be tailored and expanded according to the actual project. For example, for short-term, small-sample pilot projects, only questionnaire-level quality control scores can be output to quickly screen out obviously abnormal questionnaires; for large-scale, long-term follow-up surveys, researcher-level profiles can be retained, and a time dimension can be introduced to analyze researcher quality trends, automatically issuing warnings for those with persistent abnormalities. Item weights can also be dynamically adjusted according to the focus at different stages; for example, increasing the weight of key screening questions in the early stages of the survey, and focusing more on the stability and consistency of follow-up questions in later stages.

[0075] In terms of system deployment and integration, this invention can be deployed as a standalone cloud-based quality control service (such as...). Figure 2(As shown), it can also be embedded in existing survey platforms or CATI / CAPI systems as a backend module. In addition to a native app, mobile devices can also utilize mini-programs or web apps, as long as they can record audio, complete questionnaires, and upload data. This invention can also be used in conjunction with existing sampling verification systems. For example, questionnaires marked as "high-risk" by the system can be prioritized for manual follow-up, thereby significantly improving the coverage and hit rate of quality control without increasing overall manpower.

[0076] Regarding privacy and security, the modules of the system of this invention can be adjusted according to the data compliance requirements of specific scenarios. For example, for projects with extremely high privacy requirements, the speech recognition and semantic extraction modules can be deployed on local servers or private clouds to prevent audio data from leaving the network; alternatively, local encrypted storage and hierarchical access control strategies can be designed to restrict access to the original recordings and quality control results to only authorized quality control personnel. For collaborative projects across regions or institutions, anonymization or identifier substitution strategies can be adopted to decouple the recordings from sensitive identity information, thereby ensuring the effectiveness of quality control while meeting relevant regulatory requirements.

[0077] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0078] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A surface-access volume quality control method based on a multimodal large model, characterized in that, include: Acquire interview audio data and associated structured electronic questionnaire data that are collected synchronously for the same survey session. The structured electronic questionnaire data includes question information and logical relationships. Based on the question information and logical relationships, multimodal features are extracted from the interview voice data, and a multimodal large model is used for fusion analysis to locate and extract standardized voice answers corresponding to the questions in the questionnaire from the continuous dialogue content. The extracted standardized voice answers are compared with the recorded answers to the corresponding questions in the structured electronic questionnaire data for consistency. Based on the consistency comparison results, quality control output information is generated.

2. The surface access volume quality control method based on a multimodal large model according to claim 1, wherein, The multimodal feature extraction includes: Extract the transcribed text sequence and audio feature vector from the interview speech data to form text channel feature representation and audio feature channel representation; The transcribed text sequence includes timestamps and speaker role information.

3. The surface access volume quality control method based on a multimodal large model according to claim 1 or 2, wherein, The fusion analysis using a multimodal large model includes: The questionnaire's question information and logical relationships are modeled as graph structure constraints; When the multimodal large model extracts answers, a joint loss function is constructed for global optimization. The joint loss function includes a semantic loss term for constraining the consistency between the answer and the speech content, and a logical loss term for penalizing violations of the graph structure constraints.

4. The surface access volume quality control method based on a multimodal large model according to claim 3, wherein, The method of using a multimodal large model for fusion analysis also includes: For the question to be answered, the first answer probability distribution and the second answer probability distribution are obtained based on the text channel feature representation and the audio feature channel representation, respectively. The probability distributions of the first and second answers are weighted and fused to obtain a comprehensive answer distribution. The standardized voice answer is determined based on the comprehensive answer distribution.

5. The surface access volume quality control method based on a multimodal large model according to claim 4, wherein, The method further includes: After obtaining the comprehensive answer distribution through weighted fusion, a self-consistent reasoning step is further performed to verify and optimize the standardized speech answer; The self-consistent reasoning step includes: taking the text channel feature representation and the audio feature channel representation as input, performing multiple independent reasonings to obtain multiple intermediate answer distributions; aggregating the intermediate answer distributions to obtain an aggregated answer distribution; and generating a verification answer based on the aggregated answer distribution.

6. The surface access volume quality control method based on a multimodal large model according to claim 4, wherein, The method further includes: Based on the comprehensive answer distribution obtained by weighted fusion, the uncertainty index of the standardized speech answer is calculated, and the step of generating quality control output information refers to this uncertainty index.

7. The surface access volume quality control method based on a multimodal large model according to claim 1, wherein, The consistency comparison includes: Calculate the semantic consistency score and behavioral consistency score at the question level, and combine them to obtain the consistency index at the question level and questionnaire level.

8. The surface access volume quality control method based on a multimodal large model according to claim 7, wherein, The calculation of the question-level semantic consistency score includes: The standardized voice answer and the recorded answer are input into a trained consistency determination model to obtain classification labels, wherein the classification labels are selected from at least one of the label set including consistent, inconsistent, semantically contradictory and information missing.

9. The surface access volume quality control method based on a multimodal large model according to claim 1, wherein, The generation of quality control output information includes: Based on the results of the consistency comparison, multi-dimensional quality control data is generated. The multi-dimensional quality control data includes at least questionnaire-level quality control scores, question-level anomaly indicators, surveyor-level quality indicators, and early warning information. The multi-dimensional quality control data is displayed through a visualization interface, which supports data filtering, sorting, and drill-down analysis by project, time period, region, or survey personnel.

10. A questionnaire quality control system based on a multimodal large model, characterized in that, include: The data acquisition module is used to acquire interview voice data and associated structured electronic questionnaire data collected synchronously for the same survey session. The structured electronic questionnaire data includes question information and logical relationships. The multimodal analysis module is used to extract multimodal features from the interview voice data based on the question information and logical relationships, and to perform fusion analysis using a multimodal large model to locate and extract standardized voice answers corresponding to the questions in the questionnaire from the continuous dialogue content. The consistency comparison module is used to compare the extracted standardized voice answers with the recorded answers of the corresponding questions in the structured electronic questionnaire data. The quality control output module is used to generate quality control output information based on the results of the consistency comparison.

Citation Information

Patent Citations

  • Voice recognition method and device, computer readable medium and electronic equipment

    CN113421551A

  • Service logic verification method and device based on questionnaire questions

    CN113868369A

  • Voice question answering method and device, electronic equipment and storage medium

    CN115525749A

  • Knowledge graph question and answer reasoning method and system based on electric power and fusion field

    CN116303918A

  • Intelligent voice customer service quality inspection method and device based on multi-modal large model

    CN118631939A

Cited By

  • A collaborative reasoning method and system based on cross-family debate of heterogeneous large language models and a storage medium

    CN122242775A