A method for automatic coding and attribution analysis of qualitative interview data

By constructing a multi-stage collaborative analysis support network for qualitative interview data, the problems of low coding efficiency and poor consistency in qualitative research were solved, enabling efficient, traceable, and professional qualitative interview analysis, thereby improving the credibility and efficiency of scientific research.

CN122174961APending Publication Date: 2026-06-09EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EAST CHINA NORMAL UNIV
Filing Date
2026-03-06
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing qualitative research suffers from problems such as low coding efficiency, poor coding consistency, lack of structured mapping and continuous learning, generation of pseudo-concepts by general large models, and lack of methodological adherence and evidence alignment mechanisms, resulting in long analysis cycles and low credibility.

Method used

A multi-stage collaborative framework of coding and attribution is adopted to construct an analysis support network system, which includes a relation extraction module and an attribution analysis module. Through semi-supervised coding and precision prompting engineering, the system realizes automatic coding and attribution analysis of qualitative interview data, ensuring that the module's decisions are verifiable and the content is accurate.

Benefits of technology

It significantly improves the efficiency and consistency of qualitative interview analysis, achieves traceability and professionalism of coding results, reduces the generation of pseudo-concepts, and ensures the rigor and auditability of scientific research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122174961A_ABST
    Figure CN122174961A_ABST
Patent Text Reader

Abstract

This invention discloses an automatic coding and attribution analysis method for qualitative interview data. Its key feature is the adoption of a multi-stage collaborative framework driven by both coding and attribution, integrating code generation, relation extraction, and attribution analysis into a dynamically collaborative, data-driven closed-loop system. This system achieves automatic coding and attribution analysis of qualitative interview data, specifically including steps such as data acquisition and segmentation, candidate code generation, semi-supervised coding, dual-track output, category aggregation, relation extraction and network construction, core category identification, attribution analysis illusion suppression, and human-machine feedback loop. Compared with existing technologies, this invention improves analysis efficiency, coding consistency, and research credibility, effectively solving problems such as low efficiency of manual coding and the inability to backtrack manual corrections. It can complement or replace existing computer-aided qualitative data analysis software, providing a new generation of intelligent analysis solutions for the field of qualitative research and showing promising application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, specifically to a method for automatic coding and attribution analysis of qualitative interview data based on natural language processing. Background Technology

[0002] Grounded theory, a classic methodology in qualitative research, requires researchers to extract theories from interview texts from the bottom up through three stages: open coding, axial coding, and selective coding. However, this process is labor-intensive, and researchers are prone to "coding fatigue" when faced with large-scale educational interview data, leading to decreased coding consistency and excessively long analysis cycles. With interview transcripts often exceeding hundreds of thousands of words, researchers are frequently forced to reduce sample sizes due to excessive time costs, thus affecting the saturation and universality of the theory. To address these challenges, researchers have begun exploring the application of artificial intelligence technologies, particularly large language models, to support qualitative interview data analysis. While some existing AI educational applications, such as general chatbots or simple question-and-answer systems, can provide some immediate information, they have significant limitations in deep scientific research scenarios, especially in the analysis of complex and lengthy qualitative interview data.

[0003] Computer-aided qualitative data analysis (CAQDAS) software such as NVivo, MAXQDA, and ATLAS.ti, while making progress in data management and visualization, still largely rely on keyword retrieval and manual classification, lacking structured mapping and continuous learning capabilities for grounded theory's three-level coding. Although these tools can assist researchers in data management, the coding process still requires researchers to analyze text word by word, failing to achieve true automated coding. They are essentially digital containers, unable to understand the semantic depth of text, automatically identify causal chains between concepts, or make accurate attributions in ambiguous contexts like human experts. Researchers still need to manually drag and drop each citation for classification; the technology has not addressed the core of cognitive load.

[0004] General-purpose language models often exhibit insufficient adaptability in specific educational and research scenarios. Their powerful versatility does not equate to professional research and analytical capabilities, as they lack a deep understanding of specific disciplinary knowledge systems, research methodologies, and qualitative analysis theories. Educational terminology is highly specialized and context-dependent. For example, "scaffolding" refers to entirely different concepts in the fields of architecture and educational psychology. If general-purpose models lack domain constraints, they are prone to generating seemingly professional but ultimately misleading codes (confabulations), or over-interpreting respondents' colloquial expressions as pathological characteristics. Such unfounded generated content is unacceptable for rigorous research. Furthermore, general-purpose models often lack adherence to the rigorous logic of grounded theory, generating fragmented and disjointed codes that fail to form structured conceptual networks.

[0005] Existing systems are typically functionally singular and fragmented, failing to form a cohesive and comprehensive analytical support system. For example, an AI assistant for conversational analysis, a tool for diagnosing coding quality, and an application for assisting in report generation are usually independent of each other. The system lacks a collaborative mechanism that organically integrates multiple key stages such as code generation, relation extraction, attribution analysis, and continuous learning. A code generation module may excel at extracting keywords, but it may not deeply understand the logical relationships between concepts; a relation extraction module may be powerful, but its output is often template-based, failing to deeply integrate with coding results and evidence fragments.

[0006] General-purpose large models pose a serious "illusion" risk in vertical fields such as education and scientific research—generating seemingly professional but actually unfounded codes, lacking methodological structure and evidence alignment mechanisms, and failing to meet the auditability requirements of rigorous scientific research. Manual corrections using existing tools only take effect in the current session, unable to dynamically update model parameters or codebook rules, thus failing to achieve "increasing accuracy with use." Research and practice both demonstrate that an efficient intelligent analysis system must not only possess powerful content generation and logical reasoning capabilities, but also demonstrate "scientific rigor," namely, the ability to understand domain knowledge, adhere to methodological norms, and provide traceable evidence.

[0007] In summary, existing qualitative research on technologies has the following problems:

[0008] 1) Failure to construct a multi-stage collaborative framework driven by encoding: Existing technologies lack a system architecture that can organically integrate multiple stages such as code generation, relation extraction, attribution analysis, and continuous learning. Although existing CAQDAS tools have automatic / semi-automatic encoding capabilities, they mostly remain at the level of "coarse statistics / basic assistance," lacking structured mapping with grounded theory's three-level encoding, semi-supervised continuous learning, interpretable attribution chains, and rigorous manual correction feedback loops.

[0009] 2) Inaccurate role positioning of AI: Existing AI applications are mostly "generators," which does not match the "facilitator" role required by qualitative research, and is not conducive to ensuring the rigor and traceability of research. Existing systems often directly output encoded results, lacking adherence to methodological structures and evidence alignment mechanisms.

[0010] 3) Insufficient professional support: Due to the lack of precise modeling and application of domain ontology, methodological structure, and evidence alignment, the coding and analysis of existing systems are often generic and template-based, failing to achieve true "professional reliability." General large models lack domain constraints, are prone to generating pseudo-concepts, and are difficult to form structured concept networks.

[0011] Therefore, there is an urgent need for a technical solution that can efficiently automate three-level coding while ensuring professionalism and traceability. Summary of the Invention

[0012] The purpose of this invention is to provide an automatic coding and attribution analysis method for qualitative interview data, addressing the shortcomings of existing technologies. It employs a multi-stage collaborative framework driven by both coding and attribution, constructing an analysis support network composed of multiple specialized, collaborative modules. This organically integrates code generation, relation extraction, and attribution analysis into a dynamically collaborative, data-driven closed-loop system. It dynamically generates and utilizes multi-dimensional coding results to achieve automatic coding and attribution analysis of qualitative interview data. This provides researchers with accurate, professional, and traceable support throughout the entire qualitative interview analysis process, effectively improving analysis efficiency, coding consistency, and research credibility. It effectively solves the problems of low efficiency in traditional qualitative interview analysis due to manual coding, and the lack of structured mapping and continuous learning in existing CAQDAS tools, the high risk of generalized large-scale model illusions, and the inability to revert manual corrections. This invention can be implemented as a software system, a cloud service platform, or an embedded analysis module, complementing or replacing existing computer-aided qualitative data analysis software (CAQDAS), providing a new generation of intelligent analysis solutions for the field of qualitative research.

[0013] The specific technical solution to achieve the purpose of this invention is: an automatic coding and attribution analysis method for qualitative interview data. Its characteristic is the use of a multi-stage collaborative framework for coding and attribution, constructing an analysis support network system composed of multiple specialized, collaborative modules to achieve automatic coding and attribution analysis of qualitative interview data. The modules include: a relation extraction module and an attribution analysis module. The analysis support network system is a dual-core analysis engine consisting of a candidate code generation module and a semi-supervised coding module, which is automatically invoked at each key analysis node to continuously perform in-depth analysis of multi-source heterogeneous data in the interview text, thereby constructing and dynamically updating a multi-dimensional coding model containing open codes, axial categories, and core categories for each segment. The dynamic coding results and concept networks generated by the dual-core analysis are used as the core basis for decision-making by downstream modules. When constructing the concept network, the relation extraction module retrieves and integrates the coding results in real time as its "decision context," thereby determining the choice of relation type (causality, strategy, context, outcome, etc.) and the calculation of edge weights. Similarly, when generating evidence packages, the attribution analysis module also uses this encoded information as a "generation basis," thereby highly customizing the structure, depth, and even content suggestions of the evidence. Furthermore, this invention, through sophisticated prompt engineering technology and a rule engine, clarifies the role and behavioral patterns of each module in the analysis interaction. The semi-supervised coding module is strictly defined as an "assistant" rather than a "substitute," and its core task is to improve coding quality through self-training using domain ontology anchors and pseudo-labels; the attribution analysis module is defined as a "traceability guarantee," aiming to provide evidentiary support for each coding conclusion. To ensure that module decisions are verifiable and the content is accurate, the network system constructed in this invention also integrates Retrieval Enhanced Generation (RAG) technology and Natural Language Inference (NLI) consistency detection, enabling the system to accurately retrieve relevant domain terminology, methodological specifications, and other information from its internal knowledge base before generating codes. This invention significantly improves the specialization and traceability of the analysis process by decomposing complex analytical tasks into different specialized modules and establishing a data flow and decision-making closed loop centered on coding. It provides a new and feasible technical solution for automated and specialized analytical methods in the field of qualitative research in the future.

[0014] This invention employs a multi-module collaborative approach, using a multi-stage collaborative framework driven by both encoding and attribution. It organically integrates the analysis stage, encoding generation, relation extraction, and attribution analysis into a dynamic, collaborative, data-driven closed-loop system, enabling automatic encoding and attribution analysis of qualitative interview data. Specifically, it includes the following steps:

[0015] S1: Data Acquisition and Segmentation

[0016] The acquired educational interview texts are processed by speaker separation and privacy desensitization based on named entity recognition. The interview texts are then segmented into structures at the paragraph or sentence level, and a unique identifier, speaker role label, timestamp, and context window information are generated for each paragraph.

[0017] S2: Candidate Code Generation

[0018] Candidate open codes are generated for the set of segments using the following steps:

[0019] S2-1: Use the TextRank or RAKE algorithm to extract key phrases and obtain candidate open code generation packets;

[0020] S2-2: Generate segment vectors using a pre-trained language model, and extract representative phrases after dimensionality reduction and hierarchical clustering;

[0021] S2-3: Output candidate open code library. Each candidate code includes: code name, draft definition, typical example sentences and negative example hints.

[0022] S3: Semi-supervised coding

[0023] Based on a small number of manually labeled segments and a large number of unlabeled segments, a semi-supervised multi-label open coding method is implemented using a pseudo-label self-training branch and a label propagation branch. The coding results of the two branches are weighted and fused with a weight ratio of 0.6:0.4. A confidence score and uncertainty are calculated for each coding result. The pseudo-label self-training branch generates prediction probabilities for unlabeled segments and selects predictions with a confidence score of not less than 0.85 as pseudo-labels to be added to the training set for iteration, thus obtaining the coding for each segment. The label propagation branch constructs a segment similarity graph and uses nodes as segments and edge weights as the cosine similarity of sentence vectors with a threshold of not less than 0.7. The label propagation algorithm is then applied to the segment similarity graph to obtain the coding for each segment.

[0024] S4: Dual-rail output

[0025] The system performs dual-track output, consisting of native code and academic code. The dual-track output includes: using a key phrase extraction algorithm to extract phrases from the original speech of the respondents as native code; calculating the cosine similarity between the phrase vector of the native code and the term vector of the domain ontology, and matching terms with a similarity of not less than 0.75 as academic code.

[0026] S5: Category Aggregation

[0027] Category aggregation is performed on the open code set, including co-occurrence association aggregation, semantic similarity aggregation, and constraint aggregation based on context-action-outcome paradigm slots.

[0028] S6: Relation Extraction and Network Construction

[0029] The process involves relation extraction and concept network construction. The relation extraction includes: identifying grammatical relations in a segment using dependency parsing and semantic role labeling, and mapping the grammatical relations to grounded theory coding categories using a coding paradigm mapping rule engine. The coding categories include causal conditions, strategies, contexts, and results. The concept network is constructed with nodes as open codes or core categories, edges as relation type labels, and edge weights calculated by fusing co-occurrence strength and semantic similarity.

[0030] S7: Core Category Identification

[0031] Selective coding is performed based on the concept network, the betweenness centrality and eigenvector centrality of nodes are calculated, and multi-objective ranking is performed by combining coverage and relevance to the research question to determine the core categories and key relationship chains.

[0032] S8: Attribution Analysis

[0033] A three-tiered attribution analysis—Token or phrase-level attribution, segment-level attribution, and structural attribution—is performed on the encoded conclusions and their associated segments and concept networks. Segments that contradict the encoded conclusions are identified, a list of counterexample segments is output, and a conclusion-evidence alignment matrix is ​​generated. The Token or phrase-level attribution analysis outputs a list of highlighted trigger phrases based on attention weights or gradient methods. The segment-level attribution analysis outputs a segment-level evidence package based on the link weights from open codes to segments and segment credibility scores. The structural attribution analysis calculates the contribution weights of the core category connection paths in the concept network and outputs a core category evidence package.

[0034] S9: Hallucination Suppression

[0035] The system performs retrieval enhancement generation, similarity comparison, consistency detection, decision branching, and blacklist filtering for illusion suppression verification on the encoded results and concept network node labels. The retrieval enhancement generation retrieves terms that are semantically nearest neighbors to the encoded labels from an authoritative educational knowledge base. The similarity comparison calculates the cosine similarity between the generated word and the search term. The consistency detection uses a natural language reasoning model to determine the implication relationship between the encoded conclusion and the evidence passage. The decision branching process judges a similarity greater than 0.8 as implication and passes; a similarity between 0.5 and 0.8 is judged as neutral and a correction suggestion is output; a similarity less than 0.5 is judged as contradictory and a downgrade output or rejection is performed. The blacklist filtering outputs a warning and suggests alternative terms when a preset list of pseudo-concepts is triggered.

[0036] S10: Human-Machine Feedback Loop

[0037] The system executes a human-machine feedback loop and online updates. The human-machine feedback loop includes: active learning sampling, receiving manual correction instructions, pairwise constraint transformation, incremental learning to update model weights, and generating a new codebook. The active learning sampling selects segments to be manually reviewed based on uncertainty, concept network centrality, and representative coverage strategies. The manual correction instructions include: merging, splitting, renaming, deleting, and adding. The pairwise constraint transformation converts merging operations into mandatory connection constraints and splitting operations into non-connectable constraints, and adds these constraints as penalty terms to the clustering loss function. The new codebook includes code names, definitions, synonyms, and negative examples, and is bound to the model version number and review records.

[0038] Compared with the prior art, the present invention has the following beneficial technical effects and significant technical progress:

[0039] 1) Significantly improved coding efficiency: Through semi-supervised learning and automated processes, the traditional manual coding cycle of several months is shortened to several days.

[0040] 2) Enhanced coding consistency: Domain ontology seed word injection and pairwise constraint mechanism reduce coding drift and subjective bias, and improve coding consistency among multiple researchers.

[0041] 3) Auditable and traceable: The three-layer attribution structure and conclusion-evidence alignment matrix enable each coded conclusion to be traced back to the original passage and trigger phrase, meeting the needs of scientific research review.

[0042] 4) Hallucination suppression: RAG + NLI detection + downgrade / refusal strategy effectively eliminates unfounded or controversial terminology output, ensuring professionalism.

[0043] 5) Continuous learning and evolution: Manual correction updates the model in real time through pairwise constraints and incremental learning. Versioning the codebook facilitates review and rollback, achieving "the more you use it, the more accurate it becomes".

[0044] 6) Effectively solves the problems of low efficiency of manual coding in traditional qualitative interview analysis, lack of structured mapping and continuous learning in existing CAQDAS tools, high risk of illusion in general large models, and inability to backflow manual correction.

[0045] 7) This invention can complement or replace computer-aided qualitative data analysis software (CAQDAS) in function, providing a new generation of intelligent analysis solutions for the field of qualitative research. Attached Figure Description

[0046] Figure 1 This is a flowchart of the present invention;

[0047] Figure 2 This is a schematic diagram of the architecture of a semi-supervised multi-label coding module;

[0048] Figure 3 Example diagram of a conceptual network;

[0049] Figure 4 A flowchart for hallucination suppression verification;

[0050] Figure 5 This is a schematic diagram of the human-machine feedback loop interaction interface. Detailed Implementation

[0051] The present invention will be further described below with reference to the accompanying drawings and embodiments. Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code, including but not limited to disk storage, CD-ROM, optical storage, etc. In the description of the present invention, it should be noted that, unless otherwise expressly specified and limited, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0052] Example 1

[0053] In this embodiment, the input is the educational interview text, and the output is a three-level coding result + concept network + auditable attribution report. An example structure of the interview text is as follows:

[0054] { "interview_id": "INT_2024_001", "metadata": { "school_type": "Public Junior High School", "grade": "Eighth Grade", "subject": "Mathematics", "date": "2024-03-15"}, "segments": [ { "segment_id": "SEG_001", "speaker_role": "Respondent", "timestamp": "00:01:23", "text": "I often can't keep up with the teacher's pace in class, and my mind goes blank.", "context_window": { "prev": "Interviewer: Can you describe your feelings during class?", "next": "Interviewer: When did this start?"}} ]}

[0055] See Figure 1 An automatic coding and attribution analysis method for qualitative interview data based on natural language processing, specifically including steps S1 to S10:

[0056] Step S1: Data Preprocessing

[0057] The system acquires educational interview texts, performs speaker separation and privacy desensitization processing based on named entity recognition, segments the interview texts into segments by paragraph or sentence level, and generates a unique identifier, speaker role label, timestamp and context window information for each segment.

[0058] The speaker separation employs a dual strategy of audio stream processing and text feature analysis to ensure accurate identification. The audio stream processing strategy involves using the pyannote.audio library for speaker separation when the input is an audio file, extracting speaker features using a pre-trained speaker embedding model (such as speechbrain / spkrec-ecapa-voxceleb), and clustering speakers using a clustering algorithm (such as K-means or HDBSCAN) with a cluster size of 2 (interviewers and respondents). Each audio segment is then assigned a speaker label.

[0059] The text feature strategy involves using text features for speaker recognition when the input is plain text. The extracted features include:

[0060] 1) Question mark frequency (number of question marks / total length of the passage; the interviewer's question mark frequency is usually higher than 0.3).

[0061] 2) Segment length (interviewer segments are usually shorter, with an average length of less than 50 words, while respondent segments are longer, with an average length of more than 100 words).

[0062] 3) Turn-taking mode (interviewer usually speaks at the beginning of a segment, and interviewee usually speaks in the middle or at the end of a segment)

[0063] 4) Frequency of first-person pronouns ("I", "we", etc., the frequency of respondents is usually higher than 0.1)

[0064] 5) Frequency of interrogative words ("what", "how", "why", etc., the frequency of which is usually higher than 0.2 by the interviewee)

[0065] The above features are trained using a machine learning classifier (such as random forest or SVM), with a classification accuracy of over 90%. By default, only the respondent's speech is encoded and analyzed, while the interviewer's speech is preserved as context.

[0066] The privacy-de-identifying process employs a BERT-based named entity recognition model (such as bert-base-chinese-NER) to identify sensitive information such as personal names, place names, school names, and organization names in the text. The accuracy rate is required to reach over 95%, and a placeholder replacement strategy of the same type is used, specifically:

[0067] 1) All personal names will be uniformly replaced with anonymous identifiers such as "Respondent A" and "Respondent B";

[0068] 2) Replace place names with "a certain city", "a certain district", or "a certain county";

[0069] 3) Replace the school name with "a certain school", "a certain middle school", or "a certain primary school";

[0070] 4) Replace the organization name with "a certain organization".

[0071] Ensure that the semantic coherence of the text is not affected after replacement, and save the replacement record in a privacy-de-identified log to support subsequent data recovery (only for authorized personnel), in compliance with GDPR and educational ethics guidelines.

[0072] The segmentation employs a multi-level segmentation strategy combining coarse and fine-grained segmentation. The first level, coarse segmentation, divides the text by natural paragraphs or line breaks, identifying paragraph boundaries. The second level, fine-grained segmentation, utilizes a sentence boundary detection algorithm, which judges segmentation based on punctuation (periods, question marks, exclamation marks), syntactic structure (subject-verb-object integrity), and semantic integrity (semantic unit integrity). The segmented text forms structured data, facilitating subsequent processing and traceability.

[0073] Each segment is assigned a globally unique segment_id (in the format "SEG_XXX", where XXX is an incrementing sequence number), which is reserved:

[0074] 1) speaker_role (Speaker role: interviewer / interviewee);

[0075] 2) timestamp (Timestamp, in the format of "HH:MM:SS");

[0076] 3) context_window (Context window, including the text content of the three paragraphs before and after, used to maintain context coherence);

[0077] 4) Traceable pointer (Location information pointing to the original document, including file name, line number, and character offset).

[0078] Step S2: Generation of candidate open codes

[0079] Generate candidate open codes by executing a three-way fusion strategy on the paragraph set. The first-way fusion strategy uses the TextRank or RAKE algorithm for key phrase extraction based on graph ranking. The specific implementation process includes:

[0080] 1) Text preprocessing: Perform word segmentation and词性标注 on the paragraph set, using the jieba or LTP word segmentation tool, filter stop words (using a stop word list specific to the education field, including common stop words such as "的", "了", "在", etc., and high-frequency but non-discriminatory words such as "学生", "老师", "学校", etc.) and low-frequency words (words with a frequency less than 3), and retain content words such as nouns, verbs, and adjectives.

[0081] 2) Construction of word co-occurrence graph: Construct an undirected weighted graph, where the nodes are words and the edge weights are co-occurrence frequencies. The co-occurrence window size is set to 5 words (i.e., two words are considered co-occurring if they appear within a distance of 5 words). The co-occurrence frequency is calculated by the following formula:

[0082]

[0083] Where, is the co-occurrence count of word and word , is the total occurrence count of word , is the total occurrence count of word .

[0084] 3) Graph ranking algorithm: Use the PageRank algorithm to calculate the node weights. The PageRank iteration formula is:

[0085] .

[0086] Where, is the damping factor (default 0.85), is the node It should be noted that the term "词性标注" in the original text seems a bit unclear. It might be a specific term in Chinese for a certain language processing concept, but it's not directly translatable in a standard way. Here I've just left it as is for the purpose of accurately reflecting the original text. If it's a well-defined term in the relevant context, it may need to be adjusted according to its specific meaning.The out-degree is 20-50 iterations, and the convergence threshold is 0.001.

[0087] 4) Phrase Extraction: Extract the top N words with the highest weights as candidates, where N ranges from 5 to 20, with a default value of 10. Then, process the extracted words... Extend This forms a list of candidate phrases.

[0088] The second fusion strategy employs a domain ontology-based guided clustering method, specifically semantic clustering based on the domain ontology. The implementation process includes:

[0089] 1) Vectorization: 768-dimensional sentence vectors are generated from the text segment using a pre-trained language model (SciBERT or EduBERT). The SciBERT model is pre-trained on scientific literature and has a better understanding of educational terminology. The EduBERT model is specifically trained for the education field. The mean pooling strategy is used during vectorization to average the sentence vectors of all token vectors in the text segment.

[0090] 2) Dimensionality Reduction: The UMAP (Uniform Manifold Approximation and Projection) algorithm is used to reduce the 768-dimensional vector to 2-50 dimensions. The UMAP parameters are set as follows:

[0091] (Number of neighbors);

[0092] (Minimum distance);

[0093] (Dimensionality after dimensionality reduction);

[0094] (The distance is measured using cosine distance.)

[0095] During the dimensionality reduction process, seed word vectors are injected from the domain ontology (ERIC thesaurus) to modify the distance metric function, bringing text vectors containing seed words closer together in the vector space with a convergence coefficient of 0.3. Specifically, for text vectors containing seed words, a weighted average of the text vectors and the seed word vectors is calculated with a weight of 0.3 to form guided clustering.

[0096] 3) Clustering: Hierarchical density clustering is performed using the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm. The HDBSCAN parameters are set as follows:

[0097] (Minimum cluster size);

[0098] (Minimum sample size);

[0099] (Cluster selection method).

[0100] For each cluster, a representative phrase is extracted as a candidate code. The selection criteria for the representative phrase are: the highest frequency of occurrence within the cluster, the highest similarity to the cluster center vector, and a moderate length (2-10 characters).

[0101] 4) Denoising: Filter out noise points (points marked as -1 by HDBSCAN) and excessively small clusters (clusters with a size less than 3) to ensure the quality of candidate codes.

[0102] The third-path fusion strategy employs a strictly constrained hint engineering approach, and the specific implementation process includes:

[0103] 1) Prompt word design: Design a structured prompt word template, including task description ("Please extract candidate codes from the following educational interview passages"), output format requirements ("Output JSON format, including code_name, definition, typical_examples, negative_examples"), constraints ("Codes must be based on the passage content and cannot be fabricated. Each code must provide at least 3 typical examples and 2 negative examples"), and example demonstrations (few-shot learning, providing 2-3 examples).

[0104] 2) Batch processing: Input the segments into a large model (such as GPT-4, Claude, etc.) in batches, with each batch consisting of 10-20 segments, to avoid performance degradation caused by excessively long contexts.

[0105] 3) Result verification: Verify the candidate codes output by the model, check the correctness of the format, the rationality of the content, and the relevance to the paragraph, and filter out the outputs that do not meet the requirements.

[0106] 4) Evidence binding: Each candidate code is bound to the corresponding evidence segment. The evidence segment is a segment containing the code keyword or semantically related. The binding relationship is stored in the candidate code data structure for easy tracing and verification later.

[0107] The output data structure is as follows:

[0108] { "code_id": "OC_001", "code_name": "Cognitive Overload", "definition_draft": "Learner's state of exceeding working memory capacity during information processing", "typical_examples": ["Brain blank", "Can't keep up with the teacher's pace"], "negative_examples": ["Deliberately daydreaming", "Playing on the phone"]}.

[0109] The three results are deduplicated and merged to output a candidate open code library. Each candidate code includes a code name, a draft definition (50-200 words), typical example sentences (3-5), negative example hints (2-3), and a source tag (marking which strategy it comes from).

[0110] Step S3: Semi-supervised multi-label open coding

[0111] See Figure 2 The semi-supervised multi-label open coding system adopts a semi-supervised multi-label coding module architecture, which includes: seed labeled samples, pseudo-label self-training branch, label propagation branch, confidence / uncertainty calculation and fusion output.

[0112] The semi-supervised multi-label open coding is based on a small number of manually labeled segments and a large number of unlabeled segments. The semi-supervised learning includes a two-branch fusion mechanism.

[0113] The first branch: Employing an iterative self-training method, the specific implementation process of pseudo-label self-training includes:

[0114] 1) Seed Annotation: 5% to 10% of the total number of segments are manually annotated with multiple labels using stratified sampling to ensure sufficient labeled samples for each major topic category and avoid class imbalance. The annotation interface provides candidate code lists, segment highlighting, one-click annotation, and other functions to improve annotation efficiency. An inter-annotator agreement is used; when multiple annotators annotate the same segment, Krippendorff's Alpha coefficient is calculated, requiring an Alpha coefficient of no less than 0.7. Inconsistent annotations are discussed and unified.

[0115] 2) Initial model training: Train the initial multi-label classifier using seed samples, using BERT-base as the encoder (or specific models in fields such as SciBERT, EduBERT, etc.), with the output layer being a sigmoid-activated multi-label classification head, the loss function being the binary cross-entropy loss, the optimizer being Adam, the learning rate being set to 2e-5, the number of training rounds being 3-5, and the early stopping mechanism being set to stop if the F1 score on the validation set does not improve for 2 consecutive times.

[0116] 3) Pseudo-label generation: Using the trained initial classifier, a predicted probability distribution is generated for the unlabeled segments. For each segment-code pair, the predicted probability is calculated, and predictions with a confidence score (i.e., predicted probability) of not less than 0.85 are selected as pseudo-labels. Pseudo-labels must meet the following conditions:

[0117] a) Predicted probability ≥ 0.85

[0118] b. Semantic similarity with labeled samples ≥ 0.7

[0119] c. Not belonging to the noise category

[0120] 4) Iterative training: Add pseudo-labels to the training set and retrain the classifier. The pseudo-labels are self-trained for 3 to 10 iterations. After each iteration, the credibility score is recalculated and the pseudo-label set is updated to gradually improve the model performance. The learning rate decay coefficient for each iteration is 0.9. The early stopping mechanism is set to stop if the F1 score on the validation set does not improve after 3 consecutive iterations to avoid overfitting.

[0121] 5) Dynamic threshold adjustment: The pseudo-label threshold is dynamically adjusted according to the model performance. The initial threshold is 0.85. If the model performance improves, the threshold can be appropriately reduced (minimum 0.75) to increase the number of pseudo-labels. If the model performance decreases, the threshold can be increased (maximum 0.95) to improve the quality of pseudo-labels.

[0122] The second branch employs a graph-based semi-supervised learning method, the specific implementation process of which includes:

[0123] 1) Graph Construction: Construct a "segment similarity graph", where nodes are segments and edge weights are the cosine similarity of sentence vectors. The construction strategy using k-nearest neighbor graphs (k-NN graphs) or ε-nearest neighbor graphs (ε-NN graphs) is as follows:

[0124] a. In a k-NN graph, each node is connected to its k nearest neighbors (k=5-10).

[0125] b. In the ε-NN graph, the similarity threshold is set to 0.7. Edges with similarity below 0.7 are pruned to form a sparse graph structure, reducing computational complexity.

[0126] 2) Label Propagation Algorithm: An iterative label propagation algorithm is adopted. During initialization, the labels of labeled nodes remain fixed, while the labels of unlabeled nodes are initialized to a uniform distribution. The iterative formula is expressed by the following formula:

[0127] .

[0128] in, The propagation coefficient (range 0.1 to 0.5) The normalized similarity matrix, For the first The label matrix of the next iteration This is the initial label matrix.

[0129] 3) Boundary smoothing: By using a graph structure, the labels of labeled nodes are propagated to adjacent unlabeled nodes to achieve semi-supervised learning. During the propagation process, smoothing constraints are used to ensure that the label distribution of adjacent nodes is similar and to avoid the label distribution being too sharp.

[0130] 4) Multi-label processing: For multi-label scenarios, each label category is propagated independently, and then the results are merged to form a multi-label encoding result.

[0131] The fusion method employs a weighted fusion approach, combining the advantages of two semi-supervised learning branches. The specific implementation of the fusion strategy includes:

[0132] 1) Probability normalization: Normalize the output probabilities of the pseudo-label branch and the label propagation branch to ensure that the probability distribution of the two branches is on the same scale.

[0133] 2) Weighted Fusion: The pseudo-label branch and the label propagation branch are weighted and fused, with a default weight ratio of 0.6:0.4. The fusion strategy uses a weighted average. For each segment-code pair, the final encoding probability is expressed by the following formula:

[0134] .

[0135] in, For the final encoded probability, The probability of a pseudo-label branch. This represents the probability of a branch in the tag propagation process. The weight can be adjusted based on the actual effect: if the pseudo-tag branch performs better, its weight can be increased (up to 0.8); if the tag propagation branch performs better, its weight can be increased (up to 0.6).

[0136] 3) Thresholding: The probabilities after fusion are thresholded, and the encoded labels with a probability ≥ 0.5 are retained to form the final multi-label encoding result.

[0137] 4) Consistency check: Check the consistency of the results of the two branches. If the predictions of the same segment-code pair by the two branches differ too much (difference > 0.3), they are marked as high uncertainty samples and require manual review.

[0138] The credibility of the coding results is quantified using information theory methods. Specific implementation includes:

[0139] 1) Credibility Score Calculation: Calculate a credibility score (probability value, range 0-1) for each coded record. The credibility score directly uses the model's predicted probability, i.e. The credibility score is used to evaluate the reliability of the coding result; the higher the score, the more reliable the coding result.

[0140] 2) Uncertainty Calculation: Calculate the uncertainty (entropy value) of the coding result. The uncertainty is calculated using the following formula:

[0141] .

[0142] in, Let Σ be the probability of the i-th encoded label. Σ is the sum of all possible encoded labels. The entropy value ranges from 0 to log(N), where N is the total number of encoded labels. The higher the entropy value, the greater the uncertainty.

[0143] 3) Sample labeling: Samples with uncertainty higher than 0.6 are labeled as high uncertainty samples and require manual review; samples with uncertainty lower than 0.3 are labeled as high certainty samples and can be used directly; samples with uncertainty between 0.3 and 0.6 are labeled as medium uncertainty samples and can be reviewed in batches.

[0144] 4) Active learning sampling: Reliability scores and uncertainties are used for subsequent active learning sampling. Samples with high uncertainty and low reliability are selected first for manual annotation to improve annotation efficiency.

[0145] Step S4: Dual-track output of native code and academic code

[0146] The system executes both native code and academic code outputs, the dual-track output including:

[0147] 1) Track A (Native Code Generation): Utilize KeyBERT or YAKE algorithms to extract high-frequency n-gram phrases from the original speech segment as native code. KeyBERT uses the BERT model to extract keywords, with the top_n parameter set to 5-10 and the n-gram range set to 1-3. The extracted phrases must meet the following conditions:

[0148] a) The word frequency is no less than 3 times;

[0149] b. Length between 2 and 10 characters;

[0150] c. Does not contain stop words;

[0151] d. Prioritize retaining phrases containing emotional or action words.

[0152] 2) Track B (Academic Code Mapping): Calculate the cosine similarity between the vector of the native code phrase and the term vector of the domain ontology (ERIC Education Resource Information Center Thesaurus). The ERIC Thesaurus contains thousands of standard educational terms and their hierarchical relationships. Each term has a pre-calculated 768-dimensional vector representation. Using batch similarity calculation, for each native code phrase, retrieve the top K academic terms (K=3-5) with the highest similarity. Match terms with a similarity of not less than 0.75 as academic codes. If the highest similarity is less than 0.75, it is marked as "no matching academic term" and requires manual review.

[0153] Finally, the native code-academic code mapping table is output, which includes the native code, the corresponding academic code, the similarity score, and a list of supporting phrase IDs, forming a bidirectional and traceable mapping relationship.

[0154] Step S5: Category Aggregation

[0155] Category aggregation is performed on the open code set, and the category aggregation adopts a three-strategy parallel aggregation mechanism:

[0156] Strategy 1: Co-occurrence Association Aggregation

[0157] The co-occurrence frequency of open codes within the same or adjacent segments (with a window size of two segments before and after) is statistically analyzed, and a code-code co-occurrence matrix is ​​constructed. The matrix dimension is N×N (N is the total number of open codes). Matrix elements... Let i represent the number of times code i and code j co-occur. The co-occurrence frequency is normalized using the following formula:

[0158] .

[0159] in, For open code With open codes The number of times they co-occur within the same or adjacent paragraphs; This is the maximum value of all elements in the code-code co-occurrence matrix, used for normalization; For normalized code With code The co-occurrence intensity ranges from 0 to 1.

[0160] Code pairs with a co-occurrence frequency higher than 0.3 are considered strongly correlated and included in the same core category of candidates.

[0161] Strategy 2: Semantic Similarity Aggregation

[0162] Vector representations (using the SciBERT model) are generated from the description text and typical example sentences of the open codes. The cosine similarity between codes is calculated. Hierarchical clustering algorithms (such as Ward's connection method or average connection method) are used to cluster the similarity matrix. The clustering threshold is set to 0.6-0.8. Each cluster forms an axial category, and the number of codes within a cluster is controlled to be 3-10.

[0163] Strategy 3: Constraint Aggregation Based on Slots in the Context-Action-Outcome Paradigm

[0164] The default Strauss & Corbin coding paradigm template includes the following four slots:

[0165] a. Context;

[0166] b. Phenomenon / Conditions;

[0167] c. Actions / Strategies;

[0168] d. Consequences / Outcomes;

[0169] Through rule engines and semantic analysis, open codes are forced to be assigned to corresponding paradigm slots. Each core category must contain at least one context code and one result code, forming an interpretable structured category.

[0170] Step S6: Relation Extraction and Concept Network Construction

[0171] The process involves relation extraction and concept network construction, with the relation extraction comprising a multi-stage processing flow as follows:

[0172] 1) Syntax Analysis

[0173] Dependency parsing is performed using SpaCy or Stanza to parse the sentence dependency tree, identify subject-verb-object structures, modification relations, and clause relations, and semantic role labeling (SRL) is used to identify semantic roles such as subject, verb, object, time adverbial, place adverbial, reason adverbial, and result adverbial.

[0174] 2) Relationship identification

[0175] The coding paradigm mapping rule engine maps syntactic relations to grounded theory coding categories, which include: causal conditions, strategies, context, and consequences.

[0176] 3) Mapping rules

[0177] When causal trigger words (because, due to, leading to, causing, causing, etc.) are detected, the relevant concepts are marked as causal conditions; when action verbs (adopt, try, use, implement, execute, etc.) are detected, the relevant concepts are marked as strategies; when time or place adverbs (at a certain time, during a certain period, under certain circumstances, etc.) are detected, the relevant concepts are marked as contexts. Combined with the sentiment analysis results (using a pre-trained sentiment analysis model, outputting positive / negative / neutral), the result type is determined to be positive or negative.

[0178] 3) Conceptual Network Construction

[0179] See Figure 3 The nodes of the conceptual network are open codes / core categories / different shapes / colors, and the edges are relation type labels plus edge weights. The nodes are open codes or core categories (different levels are distinguished by different shapes / colors), and the edges are relation type labels (causality, strategy, context, result, inclusion, comparison, etc.). The edge weights are calculated by fusing co-occurrence strength and semantic similarity, and are calculated using the following formula:

[0180] .

[0181] in, This is the weighting coefficient (value range 0.3 to 0.7, default 0.5). For normalized code With code co-occurrence intensity, For normalized code With code Semantic similarity is used. Edges with weights higher than 0.6 are retained, while those with weights lower than 0.6 are pruned, forming a sparse concept network.

[0182] Step S7: Selective coding

[0183] Selective coding is performed based on a concept network. This selective coding employs a core category identification method that integrates multiple indicators, specifically including:

[0184] 1) Centrality calculation

[0185] The betweenness centrality and eigenvector centrality of nodes are calculated, and the betweenness centrality is used to measure the importance of a node as a "bridge". Calculated by the following formula:

[0186] .

[0187] Where σst is the number of shortest paths from node s to node t. This is the number of shortest paths passing through node v.

[0188] The eigenvector centrality measures the association between a node and other nodes with high centrality, and is obtained by solving the principal eigenvectors of the network adjacency matrix.

[0189] 2) Coverage calculation

[0190] The coverage is calculated by counting the number of segments and open codes that each node (category) can interpret, and then using the following formula:

[0191] .

[0192] in, For nodes Number of related segments This represents the total number of paragraphs. For nodes Number of open codes included This represents the total number of open codes.

[0193] 3) Research Question Relevance Score

[0194] A relevance score is calculated based on the semantic similarity between the research question keywords and node labels, with the relevance score ranging from 0 to 1.

[0195] 4) Multi-objective ranking

[0196] The three indicators are combined using a weighted summation method, and the overall score is calculated by the following formula:

[0197] .

[0198] in, For nodes The normalized value of the betweenness centrality indicates the importance of the node as a "bridge" in the conceptual network; For nodes The normalized value of the eigenvector centrality represents the degree of association between the node and nodes with high centrality; For nodes The normalized value of the coverage represents the ratio of the segments that the node can interpret to the open code; For nodes The normalized value for relevance to the research question is calculated from the semantic similarity between the keywords of the research question and the node labels, with weights of 0.4, 0.3, 0.2, and 0.1. These weights can be adjusted according to actual analytical needs to emphasize one of the following: centrality, coverage, or relevance. All indicators are normalized and sorted in descending order of their overall scores. The top 1-3 nodes with the highest scores are selected as the core categories.

[0199] 5) Extraction of key relationship chains

[0200] Starting from the core category, traverse the network along the path with the highest edge weight to extract the complete relational chain of core category → main axis category → open code → evidence segment. Each relational chain contains at least 3 nodes, forming the theoretical narrative skeleton.

[0201] Step S8: Attribution Analysis

[0202] A three-tiered attribution analysis, along with counterexample identification and conclusion-evidence alignment matrix generation, is performed. The attribution analysis employs the following multi-tiered traceability mechanism:

[0203] First level: Token or phrase-level attribution

[0204] 1) Perform interpretability analysis on the encoding classifier, and use the attention weight method, gradient method and masking perturbation method to identify key phrases that trigger encoding. The attention weight method extracts the attention score of each token through the self-attention mechanism of the BERT model; the gradient method identifies important tokens by calculating the gradient value of the encoding probability on the input token; the masking perturbation method identifies key tokens by observing the changes in encoding probability of each masked token.

[0205] Combining the results of the three methods above, tokens with attention weights, gradient values, and perturbation effects all exceeding the threshold are extracted as trigger phrases. A list of highlighted trigger phrases is output, with each trigger phrase including its position in the segment, importance score, and corresponding code.

[0206] Second level: Segment-level attribution

[0207] The segment-level evidence weight is calculated based on three dimensions: the link weight from the open code to the segment (encoding probability value), the segment credibility score (the reliability of the segment itself, calculated based on segment length, completeness, and contextual consistency), and the relational chain position (the position of the segment in the conceptual network relational chain, with segments on the core path having higher weights). The evidence weight is calculated by the following formula:

[0208] .

[0209] in, This is the segment-level evidence weight, used to measure the degree to which the segment supports the encoded conclusion; This represents the link weight from the open code to the segment, i.e., the encoding probability value given by the encoding model; The credibility score of the passage is calculated based on its length, completeness, and contextual consistency, representing the reliability of the passage itself. The score represents the position of the segment within the conceptual network's relational chain; segments located on core paths receive higher scores. Weights of 0.5, 0.3, and 0.2 can be adjusted based on the emphasis placed on link strength, segment credibility, and path importance. The output is a segment-level evidence package containing the segment ID, text content, evidence weight, and a list of associated codes.

[0210] Third level: Structural attribution

[0211] In the conceptual network, the contribution weight of the core category connection path is calculated. A path traversal algorithm is used to start from the core category and traverse along the path with the highest edge weight to the underlying open code and evidence segment. The contribution weight of each path is calculated. The path contribution weight is the product of the edge weights on the path. The core category evidence package is output, which includes the core category, the list of supporting paths, the contribution weight of each path, and all nodes and edges on the path.

[0212] For each encoded conclusion, the counterexample identification process retrieves segments that exhibit low consistency or semantic conflict. The determination of low consistency employs a Natural Language Inference (NLI) model. The encoded conclusion and the segment are input into the NLI model (e.g., RoBERTa-large-MNLI). Segments that show a contradiction or have a semantic similarity below 0.3 are considered counterexamples. The identified counterexample segments and the reasons for the conflict (NLI judgment result, similarity score, and specific content of the conflict) are recorded in the counterexample segment list.

[0213] The generation of the conclusion-evidence alignment matrix specifically includes: constructing a two-dimensional matrix, where rows represent coded conclusions (including open codes, axial categories, and core categories), columns represent evidence fragments (including trigger phrases, segments, and relational paths), and matrix elements are evidence weights. Evidence with a weight higher than 0.5 is considered strong supporting evidence, and evidence with a weight lower than 0.3 is considered weak supporting evidence. This matrix is ​​used to establish a complete mapping relationship between coded conclusions and supporting evidence, ensuring that each conclusion has clear evidence support and supporting bidirectional queries (finding evidence from conclusions and conclusions from evidence).

[0214] Step S9: Hallucination Suppression Verification

[0215] See Figure 4 The hallucination suppression verification includes: a RAG retrieval module, a similarity comparison module, an NLI consistency detection module, and decision branches (pass / correct / downgrade / reject). Hallucination suppression verification is performed on the encoding results and concept network node labels, employing a five-stage pipeline mechanism.

[0216] Phase 1: RAG Search

[0217] We constructed an authoritative knowledge base for education, including: ERIC database abstracts (tens of thousands of educational research abstracts), classic educational textbooks (such as "Educational Psychology", "Curriculum and Teaching Theory", etc., totaling 50+ books), and high-impact factor journal articles (AERA, Educational Researcher, etc., papers from the last 5 years, totaling 100,000+ articles). We used a vector database (Milvus or Pinecone) to store the text block embeddings of these documents (chunk size=512 tokens, overlap=50 tokens). When the encoding engine generates a new encoding tag, the system retrieves the K terms (K=5-10) that are semantically closest to the encoding tag in the vector database. The retrieval uses Approximate Nearest Neighbor Search (ANN), and the similarity calculation uses cosine similarity.

[0218] Phase Two: Similarity Comparison

[0219] Calculate the cosine similarity between the generated word and the search term, with the similarity score ranging from 0 to 1. For the K retrieved terms, calculate the average similarity and the highest similarity.

[0220] Phase 3: NLI Conformance Testing

[0221] Input the encoded conclusion and the corresponding evidence segment into the NLI model (RoBERTa-large-MNLI or DeBERTa-large-MNLI) to determine whether it is entailment, contradiction, or neutral. The NLI model outputs the probability distribution of the three categories, and the category with the highest probability is taken as the judgment result.

[0222] Phase Four: Decision Branches

[0223] The following multi-condition judgment mechanism is adopted:

[0224] Condition 1: If the average similarity is not less than 0.8 and the highest similarity is not less than 0.85, and NLI determines it to be implication, then it passes the verification and is marked as "verified".

[0225] Condition 2: If the similarity is between 0.5 and 0.8 or the NLI judgment is neutral, then a correction suggestion will be output. The system will automatically generate a correction suggestion text, suggest alternative terms (select the term with the highest similarity from the search results as the suggestion), and mark it as "needs correction".

[0226] Condition 3: If the similarity is below 0.5 or NLI determines it to be contradictory, then perform a downgraded output (marked as "to be verified", reducing the credibility score of the code to below 0.5) or reject the answer (transfer to manual review, marked as "requires manual confirmation").

[0227] Phase 5: Blacklist Filtering

[0228] It includes a built-in list of falsified educational pseudo-concepts, such as "learning style theory," "left-brain and right-brain learning," "multiple intelligences theory (over-interpreted version)," and "VARK learning style." The blacklist can be customized and expanded by the user. Each time it is triggered, it outputs a warning message (highlighted in red), recommends alternative standard terms (selected from the ERIC thesaurus), and records the trigger log for subsequent analysis.

[0229] The verification results of the above five stages are summarized into a verification report, which includes verification status, similarity score, NLI judgment result, correction suggestions, blacklist triggering situation, etc. All verification results are bound to the corresponding coded tags to form a traceable verification chain.

[0230] Step S10: Human-Machine Feedback Loop and Online Update

[0231] See Figure 5 The human-machine feedback loop interface includes: a segment highlighting area, a candidate code list, merge / split / add buttons, a confidence indicator, and a historical correction record panel. The human-machine feedback loop executes and updates online, employing the following closed-loop iterative mechanism:

[0232] 1) Proactive learning and sample selection

[0233] A multi-strategy fusion sample selection method is adopted. Strategy 1: uncertainty priority, selecting segments with uncertainty higher than 0.6; Strategy 2: high impact priority, selecting segments located near critical paths or high centrality nodes in the concept network; Strategy 3: representative coverage, ensuring that the sample selection covers all major open code categories and avoiding category imbalance.

[0234] The three strategies described above are sorted according to a preset weight ratio of 0.4:0.3:0.3. The overall score is calculated using the following formula:

[0235]

[0236] in, The overall sampling score is used to determine the priority of samples to enter the manual review queue. This is the normalized value of the uncertainty of the segment. According to the "uncertainty priority" strategy, segments with an uncertainty higher than 0.6 are more likely to be selected. This is the normalized value of the influence of the segment. According to the "high influence priority" strategy, segments located near the critical path of the concept network or high centrality nodes score higher. This is the normalized value of the segment in terms of representative coverage, corresponding to the "representative coverage" strategy, used to ensure that the selected samples cover all major open code categories and avoid category imbalance. The weight ratio of the three strategies is 0.4:0.3:0.3, which can be adjusted according to annotation cost and quality requirements.

[0237] Sort the calculated scores in descending order and select the top 10%-20% of the passages for manual review.

[0238] Step 2: Manually correct the interface interaction

[0239] It provides a visual interactive interface including: segment highlighting (highlighting trigger phrases and associated codes), candidate code list (displaying system-recommended codes and confidence scores), one-click operation buttons (accept, merge, split, rename, delete, add), confidence indicator (displaying the confidence of the current code, coded with color: green > 0.8, yellow 0.5-0.8, red < 0.5), and historical correction record panel (displaying the historical correction records of the segment, supporting undo operations).

[0240] Step 3: Receive manual correction instructions

[0241] The system executes manual correction instructions, including merging, splitting, renaming, deleting, and adding operations. Each correction instruction contains the operation type, the involved code ID, the corrected content, the reason for correction, and the correction timestamp. The merging operation combines two or more codes into a new code; the splitting operation splits a code into two or more sub-codes; the renaming operation modifies the code name or definition; the deleting operation deletes erroneous codes; and the adding operation adds new codes.

[0242] Step 4: Pair constraint transformation

[0243] Transform manual correction instructions into mathematical constraints:

[0244] 1) The merge operation is converted into a Must-Link constraint, which means that the two codes must be in the same cluster;

[0246] 2) The splitting operation is transformed into a Cannot-Link constraint, indicating that the two codes must be in different clusters.

[0248] 3) Transform the above two constraints into penalty terms of the loss function:

[0249] The penalty for Must-Link constraints is:

[0250] .

[0251] in, For the Must-Link constraint set, This is the penalty coefficient (default 0.5). , They are respectively the codes , Feature representation, The squared Euclidean distance between the two eigenvectors is given.

[0252] The penalty for a Cannot-Link constraint is:

[0253] .

[0254] in, For the set of Cannot-Link constraints, This is the penalty coefficient (default 0.3). This is the boundary value (default 1.0). , The encoded corresponding segments or samples involved in pairwise constraints.

[0255] 4) The total loss function is expressed by the following formula:

[0256] .

[0257] in, For the total loss, Clustering loss (or unsupervised / classification loss) is used. , The above two constraints and penalties apply.

[0258] Step 5: Incremental learning and updating

[0259] Online learning technology is used to fine-tune the model weights without training from scratch. Mini-batch SGD is used with a batch size of 32 and a learning rate of 0.0001. The Adam optimizer is used to update only the weights of the last few layers (classification head). The underlying parameters of the BERT encoder are frozen. The number of training rounds is 1-3. The early stopping mechanism is set to stop when the validation set loss no longer decreases.

[0260] Step 6: Codebook version iteration

[0261] 1) Each revision generates a new codebook including: code name, definition (50-200 characters), synonym list (3-5), negative example list (2-3), typical example sentences (3-5), associated segment ID list, creation time, last modification time, and modification history. The codebook version number adopts semantic version control (e.g., v1.2.3), is bound to the model version number (e.g., educode-ai-2024.03.15) and review records (including reviewer, review time, and review comments), and supports version rollback function, which can restore to any historical version, facilitating auditing and reproduction.

[0262] 2) Output data structure

[0263] The following format is used to represent the three-level coding result plus the concept network:

[0264] { "coding_result": { "open_codes": [ { "code_id": "OC_001", "in_vivo": "Mind blank", "academic": "Cognitive overload", "segments": ["SEG_001", "SEG_015"], "confidence": 0.92} ], "axial_categories": [ { "category_id": "AC_001", "name": "Classroom cognitive difficulties", "open_codes": ["OC_001", "OC_003"], "paradigm_slot": "Phenomenon"} ], "core_category": { "category_id": "CC_001", "name": "Learning adjustment disorder", "centrality": {"betweenness": 0.45, "eigenvector": 0.78}, "related_axial": ["AC_001", "AC_002"]}}, "concept_network": { "nodes": [ {"id": "OC_001", "type": "open_code", "label": "cognitive overload"}, {"id": "AC_001", "type": "axial_category", "label": "classroom cognitive difficulties"}, {"id": "CC_001", "type": "core_category", "label": "learning adjustment disorder"} ], "edges": [ {"source": "OC_001","target": "AC_001", "relation": "belongs to", "weight": 0.88}, {"source": "AC_001","target": "CC_001", "relation": "belongs to", "weight": 0.76} ]}, "attribution":{ "conclusion_evidence_matrix": [ {"conclusion": "CC_001", "evidence_segments": ["SEG_001", "SEG_015", "SEG_023"], "total_weight": 2.56} ], "counter_evidence": [ {"conclusion": "CC_001", "conflicting_segment": "SEG_042", "reason": "The respondents indicated that their classroom performance was good"} ]}, "codebook_version": "v1.2.0","model_version": "educode-ai-2024.03", "audit_log_ref": "AUDIT_2024_001"}.

[0265] The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for automatic coding and attribution analysis of qualitative interview data, characterized in that, Employing a multi-stage collaborative framework driven by both coding and attribution, this system organically integrates code generation, relation extraction, and attribution analysis into a dynamic, collaborative, data-driven closed-loop system. This enables automatic coding and attribution analysis of qualitative interview data, specifically including the following steps: S1: Data Acquisition and Segmentation S1-1: Perform speaker separation and privacy desensitization processing based on named entity recognition on the acquired educational interview text; S1-2: Structurally segment the interview text according to paragraphs or sentences; S1-3: Generate a unique identifier, speaker role label, timestamp, and context window information for each segment; S2: Candidate Code Generation S2-1: Use the TextRank or RAKE algorithm to extract key phrases and obtain candidate open code generation packets; S2-2: Generate segment vectors using a pre-trained language model, and extract representative phrases after dimensionality reduction and hierarchical clustering; S2-3: Output the candidate open code library. Each candidate code includes: code name, draft definition, typical example sentences, and negative example hints. S3: Semi-supervised coding Based on a small number of manually labeled segments and a large number of unlabeled segments, a semi-supervised multi-label open coding method is implemented using a pseudo-label self-training branch and a label propagation branch. The coding results of the two branches are then weighted and fused to calculate the confidence score and uncertainty. The pseudo-label self-training branch generates prediction probabilities for unlabeled segments and selects predictions with a confidence score of not less than 0.85 as pseudo-labels to be added to the training set for iteration, thus obtaining the code for each segment. The label propagation branch constructs a segment similarity graph and uses nodes as segments and edge weights as the cosine similarity of sentence vectors with a threshold of not less than 0.

7. The label propagation algorithm is then applied to the segment similarity graph to obtain the code for each segment. The weight ratio of the weighted fusion is 0.6:0.

4. S4: Dual-rail output S4-1: Using a key phrase extraction algorithm, extract phrases from the original speech of the respondents as raw code from the original speech segment; S4-2: Calculate the cosine similarity between the native code phrase vector and the domain ontology term vector, and use terms with a matching similarity of not less than 0.75 as academic code; S5: Category Aggregation Perform category aggregation on the open code set, the category aggregation including: co-occurrence association aggregation, semantic similarity aggregation, and constraint aggregation based on context-action-outcome paradigm slots; S6: Relation Extraction and Network Construction S6-1: Use dependency parsing and semantic role labeling to identify grammatical relations in a segment, and use an encoding paradigm mapping rule engine to map grammatical relations to grounded theory encoding categories, which include: causal conditions, strategies, contexts and results; S6-2: The construction of the conceptual network, with nodes as open codes or core categories and edges as relation type labels, and the edge weights are calculated by fusing co-occurrence strength and semantic similarity; S7: Core Category Identification Selective coding is performed based on the constructed conceptual network, and the betweenness centrality and eigenvector centrality of nodes are calculated. Multi-objective ranking is performed by combining coverage and relevance to the research question to determine the core categories and key relationship chains. S8: Attribution Analysis A three-tiered attribution analysis—Token or phrase-level attribution, segment-level attribution, and structural attribution—is performed on the encoded conclusions and their associated segments and concept networks. Segments that contradict the encoded conclusions are identified, a list of counterexample segments is output, and a conclusion-evidence alignment matrix is ​​generated. The Token or phrase-level attribution analysis outputs a list of highlighted trigger phrases based on attention weights or gradient methods. The segment-level attribution analysis outputs a segment-level evidence package based on the link weights from open codes to segments and segment credibility scores. The structural attribution analysis calculates the contribution weights of core category connection paths in the concept network and outputs a core category evidence package. S9: Hallucination Suppression The encoded results and concept network node labels are subjected to retrieval enhancement generation, similarity comparison, consistency detection, decision branching, and blacklist filtering for illusion suppression verification. The retrieval enhancement generation retrieves terms that are semantically nearest neighbors to the encoded labels in an authoritative educational knowledge base. The similarity comparison calculates the cosine similarity between the generated word and the search term; the consistency detection uses a natural language reasoning model to determine the implication relationship between the encoded conclusion and the evidence segment; the decision branch judges a similarity greater than 0.8 as implication and passes, a similarity between 0.5 and 0.8 is judged as neutral and a correction suggestion is output, and a similarity less than 0.5 is judged as contradictory and a downgrade output or rejection is performed; when the blacklist filtering triggers the preset pseudo-concept list, a warning is output and alternative terms are suggested; S10: Human-Machine Feedback Loop The system executes a human-machine feedback loop and online updates. The human-machine feedback loop includes: active learning sampling, receiving manual correction instructions, pairwise constraint transformation, incremental learning to update model weights, and generating a new codebook. The active learning sampling selects segments for manual review based on uncertainty, concept network centrality, and representative coverage strategies. The manual correction instructions include: Merging, splitting, renaming, deleting, and adding; the pairwise constraint transformation converts merging operations into mandatory connection constraints and splitting operations into non-connectable constraints, and adds mandatory connection constraints and non-connectable constraints as penalty terms to the clustering loss function; the new codebook includes code names, definitions, synonyms, and negative examples, and is bound to the model version number and review records.