A universal qualitative analysis coding system based on LLM

By combining large language models with traditional qualitative research methods, an end-to-end qualitative analysis system is built, which addresses the shortcomings of existing qualitative analysis technologies in efficiency and cross-domain adaptability, and realizes efficient and reliable qualitative analysis.

CN120449823BActive Publication Date: 2025-09-16INST OF MOUNTAIN HAZARDS & ENVIRONMENT CHINESE ACADEMY OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510955977.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-16
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing qualitative analysis techniques are inefficient and lack semantic understanding when processing unstructured text data. They are difficult to adapt to large-scale data and have limited cross-domain applications, which affects the reliability and reproducibility of research results.

Method used

By combining the LLM-based large language model with traditional qualitative research methods, an end-to-end qualitative analysis system is built through automatic generation of intelligent codebooks, dynamic prompt engineering, batch document automatic coding, and consistency assurance modules to improve analysis efficiency and consistency.

Benefits of technology

It improves the efficiency and depth of qualitative analysis, adapts to ultra-large-scale data, enhances the flexibility and scientificity of cross-domain applications, and ensures the reliability and transparency of the results.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention relates to the field of information analysis and processing technology, and specifically, to a general qualitative analysis coding system based on LLM. The present invention breaks through the efficiency bottleneck of traditional qualitative analysis and realizes the automation, intelligence and efficiency of qualitative data analysis. The system can significantly improve the efficiency of qualitative analysis, the depth of semantic understanding, and the consistency of analysis results, especially when processing ultra-large-scale data sets. It shows outstanding advantages, enabling researchers in more disciplines to quickly master and apply qualitative analysis methods, and also has stronger adaptability and flexibility in cross-domain applications. Compared with traditional qualitative analysis methods, the innovation of the present invention is that it can realize the analysis and processing of ultra-large-scale research samples, providing efficient, scientific and innovative research methods and technical support for qualitative research, with high academic value, and promoting the automation and intelligence process of qualitative research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information analysis and processing, and in particular to a universal qualitative analysis coding system based on LLM. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Qualitative analysis, a core method for understanding complex phenomena in fields such as social science, psychology, education, and medicine, aims to deeply explore the meaning and patterns inherent in unstructured textual data. Qualitative analysis encompasses methods such as grounded theory, content analysis, and Husserl's phenomenological analysis. It helps researchers uncover deeper meanings that are difficult to capture using qualitative research methods, thereby forming a systematic understanding and theoretical construction of complex social phenomena. It can effectively identify and explore important scientific issues that traditional qualitative research methods cannot address. Its core value lies in transforming unstructured text into an organized conceptual system, thereby crystallizing scientific questions and providing theoretical guidance and a basis for action.

[0004] With the rapid development of information technology, computer-assisted qualitative data analysis software such as NVivo, ATLAS.ti, MAXQDA, and DiVoMiners have been used in traditional qualitative research to provide basic support functions such as text management, coding, and visualization. At the same time, rule-based text analysis and rudimentary machine learning methods are also beginning to be introduced into the field of qualitative analysis to improve data processing efficiency.

[0005] However, existing qualitative analysis techniques and tools primarily serve as a supplement to digital pen and paper, with the core analytical work still relying on manual labor and facing severe efficiency bottlenecks. In particular, current machine learning-based qualitative analysis tools generally rely on pre-set lexicons, often breaking text samples into single lexical units for mechanical comparison. These tools lack a deep understanding of the text's context and linguistic context, making them unable to capture the deeper meaning and subtle contextual information within the text. In an era of rapidly growing information data, traditional qualitative analysis methods are clearly unable to adapt to this explosive growth. Some of these methods rely on pre-set AI capabilities and rely heavily on a priori, pre-set frameworks, making them ill-suited to the more open and diverse research needs of the postmodern era. Furthermore, existing qualitative analysis techniques exhibit significant limitations in addressing coding consistency, analytical reliability, cognitive burden, depth of semantic understanding, and cross-disciplinary adaptability, impacting the scientific nature of qualitative analysis and the reliability of its results. Furthermore, these techniques exhibit significant limitations in cross-coder consistency and longitudinal consistency, impacting the reliability and reproducibility of research findings. Existing qualitative research methods and approaches are relatively traditional, requiring high professional knowledge and cognitive abilities of researchers, and are more prone to cognitive limitations, which restricts the widespread application of qualitative research methods.

[0006] Based on the above reasons, the present invention has designed a universal qualitative analysis coding system based on LLM. By organically combining artificial intelligence large language model (LLM) technology with traditional qualitative research methods, it can effectively improve the efficiency of qualitative analysis, the depth of semantic understanding, and the consistency of quality analysis. It can also handle very large-scale data, lowering the threshold for research and learning, and has greater adaptability and flexibility in cross-disciplinary applications. Its inherent transparency and verifiability are also guaranteed, enhancing scientificity and credibility. This invention provides innovative, efficient, and scientific research methods and support for identifying and exploring important scientific issues that traditional qualitative research methods cannot address. Summary of the Invention

[0007] The present invention aims to overcome the shortcomings of existing qualitative analysis technology and proposes a universal qualitative analysis coding system based on LLM to improve analysis efficiency, depth and consistency of semantic understanding, and enhance its ability to adapt to ultra-large-scale data sets. By organically combining artificial intelligence large language model technology with traditional qualitative research methods, it is possible to effectively improve the efficiency of qualitative analysis, the depth of semantic understanding and the consistency of quality analysis, effectively process ultra-large-scale data sets, lower the threshold for qualitative research, and enable researchers in more fields to conduct efficient analysis through intuitive operations. In addition, it also has better adaptability and flexibility in cross-domain applications. While its own transparency and verifiability are guaranteed, it also enhances the scientificity and credibility of the research. The present invention provides innovative, efficient and scientific research methods and tool support for identifying and exploring deep-seated social science issues that traditional qualitative research methods cannot reveal.

[0008] In order to achieve the above-mentioned object, the present invention provides a universal qualitative analysis coding system based on LLM, which includes intelligent codebook automatic generation technology, dynamic prompt engineering and document preprocessing module, batch document automatic coding module, LLM response parsing and consistency assurance module, and structured result integration and export module;

[0009] Intelligent codebook automatic generation technology: This technology introduces a large language model to assist in codebook generation and combines iterative optimization processes with human-machine collaboration to ensure the final form and content of the codebook. Specifically, it includes:

[0010] S1, automatic construction and import of the initial codebook: This process supports the generation of a preliminary codebook framework for a large language model, or optimizes it by loading the user's existing codebook.

[0011] S2, human-computer collaborative iteration and continuous optimization of the codebook: continuous optimization and iterative improvement of the codebook; through multiple human-computer collaborative feedback, the structure and content of the codebook are gradually improved to ensure that it meets the needs of the research task.

[0012] The dynamic prompt engineering and document preprocessing modules include:

[0013] S3, document content reading and preprocessing;

[0014] S4, dynamic construction of targeted analysis prompt words: A dedicated processing unit receives the currently confirmed and effective codebook, the list of questions included in the specific analysis batch in the codebook, and the pre-processed content of the current document to be analyzed as input information to generate complete prompt words;

[0015] Batch document automatic coding module: This module supports batch processing of large-scale documents to ensure the efficiency and accuracy of the analysis process.

[0016] S5, single document processing flow optimization; logical grouping, problems and batch processing of each document in the encoding process to improve processing efficiency.

[0017] S6, Language Model Interaction and Fault Tolerance: All communications between the system and the large language model service are performed through a dedicated interaction interface;

[0018] S7, master batch processing: responsible for coordinating and managing the automated coding analysis of multiple documents in the entire qualitative analysis workflow;

[0019] The LLM response parsing and consistency assurance module includes:

[0020] S8, structured parsing of LLM responses: converting the responses returned by the large language model for a complete batch of questions, which are usually in free text format, into a structured data format that is easy to process and store later;

[0021] S9, Coding Consistency Guarantee Mechanism: Through the system-designed consistency guarantee mechanism, we ensure coding consistency across documents and problem batches, maximizing the reliability of analysis results.

[0022] The structured results integration and export module includes:

[0023] S10, Coding result integration and formatting: Once the coding process of the current document is completed, the system will call the result saving and export program to organize the analysis results into a standardized format to ensure the standardization and consistency of the data;

[0024] S11, result export: After all analysis results are prepared, the system exports the structured data into Excel or CSV files through a special auxiliary function to facilitate subsequent data processing and analysis. The auxiliary function receives the sorted header information and a list of data rows containing all document encoding data as input, and writes these contents in a standardized format to the file pointed to by the output file path previously specified by the user.

[0025] S1 is specifically to confirm the initial version of the codebook, which includes three initialization methods:

[0026] S1-1, Large language model assisted generation of initial codebook framework: by inputting research project information, the large language model constructs detailed, context-aware high-level instructions;

[0027] The instructions guide the large language model to play the role of a qualitative research expert and require it to output a structured codebook draft based on the research information provided; the draft includes a descriptive name for the codebook, an overall description of the purpose and scope of the codebook, and a structured list of questions;

[0028] Each independent coded question or entry in the list contains: a unique internal identifier ID, a specific textual description of the question (a logical field name used for subsequent data storage and management), supplementary instructions or operational instructions for the question, and the analytical batch number to which the question logically belongs;

[0029] The instruction requires the large language model to provide brief descriptive text for each set analysis batch;

[0030] The encoding system sends this high-level instruction to the model through its interactive interface with the large language model, obtains the draft codebook text generated by it, and the encoding system automatically parses it into its internal operational codebook data structure;

[0031] S1-2, load the user's existing codebook:

[0032] The coding system allows users to load their existing complete codebooks stored in JSON or other standardized structured formats through designated import functions as a starting point for subsequent iterative optimization;

[0033] S1-3, template-based or interactively created from scratch:

[0034] The coding system prompts the user and provides a basic, empty codebook template, or the user can build it item by item from scratch through the interactive interface provided by the coding system.

[0035] S2 specifically involves combining the auxiliary analysis capabilities of the large language model with the researcher's professional judgment and final decision-making power, through multiple cycles of feedback and modification, until the quality and applicability of the codebook reach a state that satisfies the user. Specifically, it includes the following steps:

[0036] S2-1 presents the current version of the codebook and the optimization suggestions provided by the large language model in the previous round through a visual and editable user interface;

[0037] S2-2, the interface clearly displays all the components of the codebook and supports intuitive and convenient editing of the codebook contents, such as the specific wording of questions, the selection of answer formats, the addition and deletion of predefined options, and the allocation of questions between different analysis batches;

[0038] S2-3, editing operations include "save current edits", "request AI optimization suggestions", or "finally confirm and solidify the codebook".

[0039] The specific steps for document content reading and preprocessing are:

[0040] S3-1, extracts text information from the user-specified file path by reading the character encoding standard to avoid garbled characters and improve compatibility with documents from different sources;

[0041] S3-2, document truncation and content summarization mechanisms are used to avoid the text length limitations of single inputs in large language models;

[0042] S3-3, the mechanism receives the original text content and the preset maximum allowed number of characters or the equivalent number of tokens as parameters. If it detects that the text length exceeds this limit, the mechanism does not truncate it from the end, but instead prioritizes retaining the beginning and end of the document containing important introductions, conclusions or contextual information, and inserts clear omission or truncation prompt information between the two.

[0043] The specific steps for dynamically constructing targeted analysis prompt words are:

[0044] S4-1, based on the input information of S4, constructs an introductory opening text to set the specific role of the large language model in the current task, and clearly indicates the name of the research project to which the current analysis belongs and the theme or focus of the current batch of questions;

[0045] S4-2, emphasizes the answer format specifications to the large language model;

[0046] S4-3, traversing each specific coding question in the current question batch, and for each coding question, integrating the unique identification ID and complete text content of the question into the prompt word being constructed;

[0047] S4-4, dynamically generates highly targeted answer instructions based on the answer format type preset in the codebook for the question and the predefined options it may contain;

[0048] S4-5, if the question object in the codebook also contains additional explanations, precautions, or specific contextual prompts for the question, they are also integrated into the prompt words as additional guidance for the large language model to accurately understand and answer;

[0049] S4-6, append the pre-processed document content currently to be analyzed to the end of the above series of instructions and question lists.

[0050] The specific steps of S5 are:

[0051] S5-1, the key input information received includes the storage path of the document to be analyzed, the previously constructed and finalized codebook, the initialized language model service client instance, and the specific language model name selected by the user;

[0052] S5-2, calling S3 to complete the loading, encoding conversion and intelligent truncation of the specified document content based on the length limit;

[0053] S5-3, the auxiliary functional unit reorganizes and groups all the coding questions in the codebook according to the "batch" attribute attached to each coding question, forming a series of logically related question batch units suitable for step-by-step processing, which is used to manage the complexity of the analysis and optimize the interaction efficiency with the language model;

[0054] S5-4, for the current document, process each batch of questions in turn, and perform the following core operation sequence for all questions in each batch of questions:

[0055] S5-4-1, mobilizing S4 to generate highly customized dynamic analysis prompt words containing detailed answer instructions for the current question batch and the current document or its pre-processed content;

[0056] S5-4-2, through S6, the generated analysis prompt words together with the necessary model selection and temperature parameter accompanying information are sent to the large language model for processing. The deep semantic understanding ability of the large language model is used to interpret the complex instructions in the prompt and the subtle information in the document;

[0057] S5-4-3: Receive and temporarily store the original text responses returned by the large language model for all questions in the current batch;

[0058] S5-4-4 calls S8 to obtain the free text response from the large language model, parses it and converts it into structured data according to the definition of the question in the codebook and the expected answer format;

[0059] S5-4-5, accumulate and merge the structured coding results obtained from parsing the current batch of questions into the total result set maintained for the document;

[0060] S5-5, S5-1~S5-4 are executed in a loop until all the problem batches defined in the coding book have been processed for the current document. The process returns a structured data object containing the complete analysis results of all coding problems for the single document.

[0061] S6 specifically:

[0062] S6-1, the interface is responsible for receiving the prompt text to be sent to the large language model, the initialized language model client instance, the specified model name, and optionally, specific system prompt text that may be used to override or enhance the default system-level instructions;

[0063] S6-2, the interface includes assembling a request body, which includes:

[0064] S6-2-1, system-level role or behavior instructions, i.e., general guiding instructions, used to set the basic behavior pattern of the language model;

[0065] S6-2-2, user-level, task-specific instructions, i.e., dynamic analysis prompts containing the text to be analyzed and coding questions;

[0066] S6-3, the interface has a built-in fault tolerance and retry mechanism: within the preset maximum number of attempts, if an API call fails due to network timeout or temporary server error, the interface will automatically wait for a preset delay time and then re-initiate the same request or use streaming data transmission to process the request and response;

[0067] S6-4, the interface needs to capture the final answer generated by the model, synchronously receive and record the intermediate steps or reasoning chains in the process of generating the answer, and organically combine this information into a more complete and transparent response result; if all preset retries have been exhausted and the API call is still unsuccessful, the interface will no longer try and will return an identifier that clearly contains the error information for the upper-level caller to know and take corresponding measures.

[0068] The specific steps of S7 are:

[0069] S7-1, obtains the path list of all documents to be analyzed from the input directory specified by the user through the file system access mechanism;

[0070] S7-2, the main control process will traverse this document path list, and for each document path in the list, it will call S5 in turn to complete the comprehensive coding analysis of the document;

[0071] S7-3, whenever the processing of a document is completed and its encoding results are returned, these results are added to the global total result list used to store the analysis results of all processed documents;

[0072] S7-4: The system selectively invokes the result saving mechanism after processing each document or after a certain number of documents, periodically saving the currently accumulated analysis results to a designated intermediate output file. This is used to enhance the robustness of the system and prevent the loss of processed data due to unexpected interruptions such as system failures or power problems during long-term operation.

[0073] S7-5, the main control process is responsible for updating the processing status and recording the list of successfully processed files so that it can be restored from the breakpoint after the task is unexpectedly interrupted to avoid repeated processing;

[0074] S7-6: The main control process introduces a short configurable delay between processing one document and preparing to process the next one. This is used to avoid triggering the service provider's rate limit policy due to too frequent API requests, or to reduce the instantaneous pressure on the language model server.

[0075] The specific steps of S8 are:

[0076] S8-1 receives the original response text of the large language model and the original codebook question batch data structure that defines each question in the batch. The data structure contains the ID of each question, the expected answer format, and the field name key metadata;

[0077] S8-2, traverse each coded question object in the current question batch, and try to apply a series of pre-designed parsing rules to each question. The parsing rules are based on the pattern matching ability of regular expressions or supplemented by precise string search and positioning methods to accurately find and extract the answer part corresponding to the ID or specific identifier of the current question.

[0078] S8-3, to improve the robustness and adaptability of the parsing process, design and try out multiple alternative matching modes in sequence. If the original answer text of a question is successfully extracted, after a series of necessary cleaning and normalization steps, carry out targeted subsequent processing according to the expected answer format preset in the codebook for that question;

[0079] S8-4, the parsing result of each question will be stored in a mapping object or dictionary structure with its preset logical field name in the code book as the key and the processed answer content as the value, and returned as the result of the batch parsing; if the answer to a specific question cannot be successfully parsed for any reason during the entire parsing process, the corresponding field of the question will be given a clear error mark or default value to ensure the integrity of the data structure and the clarity of subsequent processing.

[0080] The specific steps for S9 are:

[0081] S9-1, Unified coding instruction generation mechanism: Through standardized templates, the processing instruction format for each problem and batch is guaranteed to be consistent, thereby improving the efficiency of analysis and the consistency of results;

[0082] S9-2, Standardized Response Format Requirements: When constructing dynamic prompt words, indicate the preset format in which the large language model should organize and output its analysis results;

[0083] S9-3, structured result parsing logic: S8 uses a precise parsing strategy based on pattern matching to forcibly map the relatively free text output returned by the large language model back to the pre-defined coding fields in the codebook, ensuring that the data extracted from different responses is structurally unified and standardized.

[0084] S9-4, Centralized and Ordered Result Storage and Export: In the structured result integration and export module, analytical data are organized and presented according to the field order and names defined in the codebook. Whether generating intermediate summary tables or ultimately exporting Excel or CSV files, their column structures are directly derived from the codebook definitions. This ensures that the output data has the same structure across different analytical tasks, facilitating subsequent cross-case comparisons, data consolidation, and statistical analysis.

[0085] Compared with the prior art, the present invention has the following beneficial effects:

[0086] This paper combines large language model technology with traditional qualitative research methods to build an end-to-end intelligent coding framework, breaking through the efficiency and depth bottlenecks of traditional qualitative analysis. Unlike traditional lexicon-based analysis methods, this paper understands texts in their full context, maintaining the integrity and coherence of the analysis, and achieving more accurate qualitative analysis.

[0087] The present invention innovatively designs an intelligent codebook automatic generation technology based on a large language model to solve the subjectivity, inconsistency and low efficiency problems in traditional codebook construction.

[0088] This invention uses customized prompt words to enable large language models to understand and perform professional qualitative analysis tasks, greatly improving the accuracy and depth of analysis.

[0089] The present invention proposes a batch document automatic coding module to overcome the limitation of traditional qualitative analysis methods that are difficult to apply to ultra-large-scale samples.

[0090] The present invention designs a complete set of consistency assurance modules, which can effectively ensure the coding consistency in large-scale sample analysis and improve the scientificity and reliability of qualitative analysis. DETAILED DESCRIPTION

[0091] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art to which the present application belongs.

[0092] The present invention provides a universal qualitative analysis coding system based on LLM, which specifically includes the following contents:

[0093] 1. System initialization and API connection module:

[0094] (1) System configuration and environment preparation:

[0095] This step is responsible for initializing the basic configuration required for system operation, including the user-specified output directory path, logging level, and so on. It also establishes a connection to the Large Language Model service based on the user-provided API key, the selected Large Language Model service provider and its corresponding API base URL, and specific model names such as gpt-4o and deepseek-chat. This section also loads or sets global configuration parameters such as the maximum number of API retries, retry delay time, and default LLM system-level instructions. Subsequently, it receives the API key (apiKey), API base URL, and vendor information as parameters, and attempts to instantiate and return an LLM client object. If initialization fails, an error is logged and the user is prompted to check the configuration.

[0096] (2) Acquisition of research information:

[0097] After the system's core services and connection to the large language model are in place, the system must accurately capture the specific information of the user-defined research project. Through a specific human-computer interaction process, the system will collect the key elements that form the core of the research task, which will serve as the basic input for subsequent intelligent codebook generation and in-depth analysis tasks. The acquired information is stored in a structured manner, including a clear title for the research project; a detailed description of the research background, objectives, and scope, such as in-depth interview transcripts, policy and regulatory texts, and open-ended questionnaire responses; a list of core questions that form the research problem framework; and a list of coding dimensions or categories initially conceived or suggested by the user.

[0098] 2. Intelligent codebook automatic generation technology:

[0099] In traditional qualitative analysis, codebook development is a time-consuming task that relies heavily on the researcher's expertise and experience. Its quality directly impacts the depth and accuracy of subsequent analysis. Furthermore, manually developed codebooks can be subject to subjective limitations, blind spots, and inconsistency issues.

[0100] To address these pain points, this intelligent codebook automatic generation technology introduces a large language model (LLM) to assist in the generation of codebooks, and designs a set of human-computer collaborative iterative optimization processes, allowing users to deeply participate in and dominate the final form and content of the codebook.

[0101] (1) Automatic construction or import of the initial codebook:

[0102] To start the codebook construction process, the system first requires an initial version of the codebook. Considering that researchers may start from different starting points, this module supports three initialization methods.

[0103] First, the large language model assists in generating an initial codebook framework. Given that building a codebook from scratch is not only time-consuming and labor-intensive, but also susceptible to the potential influence of the researcher's personal academic background or theoretical preferences, this method leverages the large language model's extensive knowledge base and powerful logical reasoning capabilities to quickly generate a preliminary, structured codebook draft.

[0104] This process is achieved through a specific mechanism. Specifically, based on the research project information entered by the user during system initialization (covering research topics, core questions, material types, and preliminary coding categories, etc.), the system dynamically constructs detailed, context-aware high-level instructions.

[0105] The instruction explicitly directs the large language model to play the role of a qualitative research expert and requires it to output a structured draft codebook based on the research information provided, usually in a format such as JSON that is easy for machines to parse.

[0106] This draft should include a descriptive name for the codebook, an overall description of the codebook's purpose and scope, and a structured list of questions. Each individual coding question (or item) in the list should include: a unique internal identifier (ID); a specific textual description of the question, including predefined response formats such as "yes or no," "selective classification," "key information extraction," "underlying meaning summary," and "degree or attitude evaluation"; a list of predefined options for selective classification or evaluation questions (if applicable); logical field names for subsequent data storage and management; supplementary explanations or operational instructions for the question; and the analysis batch number to which the question logically belongs. Furthermore, this instruction requests the large language model to provide a brief descriptive text for each specified analysis batch. The system then sends this high-level instruction to the large language model through its interface, obtaining the generated codebook draft text and automatically parsing it into an internally operable codebook data structure.

[0107] Second, loading the user's existing codebook. If the user already has a partially formed or complete codebook accumulated from previous research, such as stored in JSON or other standardized structured formats, the system allows it to be loaded through a designated import function as a starting point for subsequent iterative optimization.

[0108] Third, template-based or interactive creation from scratch. If neither of the above two methods can obtain a suitable initial codebook, or the user prefers to build it completely independently, the system can also prompt the user and may provide a basic, empty codebook template, or allow the user to build it from scratch one by one through the interactive interface provided by the system.

[0109] (2) Human-machine collaborative iteration and continuous optimization of the codebook:

[0110] After obtaining the initial codebook (whether generated by AI, imported by users, or created from templates), given that the understanding of the large language model may deviate from the researcher's specific domain knowledge or specific research needs, or the user wishes to make more refined adjustments and improvements to the codebook,

[0111] The core iterative optimization process of this module then began. The core design of this process is to organically combine the auxiliary analysis capabilities of the large language model with the professional judgment and final decision-making power of the researcher. Through potentially multiple cycles of feedback and revisions, the quality and applicability of the codebook reach a state of user satisfaction.

[0112] Throughout the iterative optimization cycle, the system first presents the current version of the codebook, along with (if any) the optimization suggestions provided by the large language model in the previous round, to the researcher through a visual, editable user interface. This interface should clearly display all the components of the codebook and support intuitive and convenient editing of various codebook contents, such as the specific wording of questions, the choice of answer format, the addition and deletion of predefined options, and the allocation of questions between different analysis batches. The system will then wait for the user to clarify their next operation intention through interface interaction. The main operation options usually include "Save current edits", "Request AI optimization suggestions", or "Finalize and solidify the codebook".

[0113] (1) If the user selects "Save Current Edit", the system will capture all modifications made by the user to the codebook on the interface. To ensure the logical integrity and subsequent usability of the codebook, the internal structure verification mechanism will perform validity verification on the modified codebook, such as checking whether it meets the expected structural format requirements, whether the key required fields exist and are valid. Only codebooks that pass the structure verification will be accepted and updated to the current working version. At the same time, since the user's active editing may have made the optimization suggestions previously provided by the large language model partially or completely invalid, the old suggestions will usually be cleared.

[0114] (2) If the user selects "Request optimization suggestions from the large language model", it means that the user wants to use the power of the large language model again to conduct diagnostic analysis and constructive improvements on the current version of the codebook. At this time, the system will construct a new high-level instruction specifically for requesting optimization suggestions based on the content of the current codebook and the research project information initially entered. This instruction will require the large language model to conduct a detailed evaluation of the overall structure of the existing codebook, the clarity and unambiguity of the question statements, the comprehensiveness of the content coverage, and the logical correlation between the questions, and propose specific and actionable optimization suggestions on how to improve it, such as which questions to modify the wording to make it more precise, which new coding dimensions or questions to add to cover research blind spots, and which questions to adjust the answer format to better meet analysis needs. This instruction is sent through the interactive interface with the large language model. The feedback returned by the large language model is usually a text-based analysis report and a list of improvement suggestions. This information will be stored by the system and presented to the user together with the current codebook at the beginning of the next iteration cycle for the user to refer to, adopt, and decide how to further manually modify the codebook.

[0115] (3) If the user selects "Finalize and solidify the codebook", the system will first call the internal structure verification mechanism again to perform a final validity verification on the codebook currently to be confirmed. Only when the codebook structure is confirmed to be valid and meets the requirements will this version of the codebook be officially adopted as the final version, thus ending the entire iterative optimization cycle. If the structure verification fails, the system will prompt the user to make necessary corrections.

[0116] This iterative loop design provides researchers with tremendous flexibility, allowing them to independently and meticulously edit every detail of the codebook while also conveniently accessing the intelligent analysis and suggestions of the large language model as needed. They can then use this feedback to further refine and improve the codebook. This cycle continues until the codebook reaches optimal quality, depth, breadth, and suitability for the specific research task. The finalized codebook serves as the core basis and operational guide for all subsequent document analysis and coding modules.

[0117] 3. Dynamic prompt engineering and document preprocessing module:

[0118] (1) Document content reading and preprocessing:

[0119] Before encoding the specific document to be analyzed, the first step is to read its content accurately and completely.

[0120] The system's built-in document content reading function is responsible for extracting text information from user-specified file paths. To improve compatibility with documents from different sources, this function typically attempts to decode and read using various common character encoding standards such as UTF-8 and GBK to minimize garbled characters.

[0121] Since large language models usually have certain limitations on the length of text when processing a single input, namely the so-called "context window size", document truncation and content summarization mechanisms are used to properly handle documents that are too long.

[0122] This mechanism receives the original text content and a preset maximum allowed number of characters or equivalent number of tokens as parameters. If the text length exceeds this limit, the mechanism does not simply truncate it. Instead, it prioritizes retaining the beginning and end of the document, which contain important introductions, conclusions, or contextual information, and inserts explicit omission or truncation warnings between them. This strategy preserves the most valuable key information fragments in the document while ensuring that the processed text length meets the input requirements of large language models.

[0123] (2) Dynamic construction of targeted analysis prompt words:

[0124] Dynamically constructing targeted analysis prompts is one of the core technologies that enables intelligent, context-aware coding in this system. Its core mechanism is a dedicated processing unit that receives as input the currently validated and effective codebook, a list of questions for a specific analysis batch within the codebook, and the preprocessed content of the current document to be analyzed.

[0125] Based on the input information, the processing unit first constructs an introductory opening text to set the specific role of the large language model in the current task, and clearly indicates the name of the research project to which the current analysis belongs and the theme or focus of the current batch of questions.

[0126] The large language model is then clearly and explicitly reminded of the strict answer format specifications that must be followed. For example, the answer to each question must begin with a specific prefix and the actual answer content must be placed within specific tags.

[0127] The processing unit then iterates through each specific coded question in the current batch. For each question, it integrates its unique ID and full text into the prompt being constructed. It then dynamically generates highly targeted response instructions based on the question's pre-defined answer format and possible predefined options in the codebook.

[0128] If the question object in the codebook also contains additional explanations, precautions, or specific contextual prompts for the question, they will also be integrated into the prompt words as additional guidance for the large language model to accurately understand and answer.

[0129] Finally, the pre-processed document content to be analyzed is appended to the end of the aforementioned list of instructions and questions. The resulting complete prompt is highly task-specific and context-relevant. Furthermore, its structured instruction design effectively guides the large language model to complete specialized coding tasks in accordance with the researcher's intent and the norms of qualitative research.

[0130] 4. Batch document automatic encoding module:

[0131] (1) Single document processing process arrangement:

[0132] This module designs an automated processing flow that is specifically responsible for performing complete, codebook-based coding analysis on a single independent document file.

[0133] When the process starts, the key input information received includes the storage path of the document to be analyzed, the previously built and finalized code book, the initialized language model service client instance, and the specific language model name selected by the user.

[0134] When the process starts, the aforementioned document content reading and preprocessing mechanism is first called to complete the loading, encoding conversion, and intelligent truncation based on length restrictions of the specified document content. Subsequently, the auxiliary functional unit will reorganize and group all questions according to the "batch" attributes attached to each encoding question in the coding book, forming a series of logically related question batch units suitable for step-by-step processing. The batch design helps to manage the complexity of the analysis and may also be used to optimize the efficiency of interaction with the language model. Next, the system will process each question batch in turn for the current document. For all questions in each question batch, the following core operation sequence is performed:

[0135] (1) Calling the dynamic construction of targeted analysis prompt words to generate highly customized dynamic analysis prompt words containing detailed answer instructions for the current question batch and the current document or its pre-processed content.

[0136] (2) Through interaction and fault tolerance with the language model, the generated analysis prompt words, along with necessary accompanying information such as model selection and temperature parameters, are sent to the large language model for processing. In this step, the language model's deep semantic understanding capabilities are fully utilized to interpret the complex instructions in the prompts and the subtle information in the document.

[0137] (3) Receive and temporarily store the original text responses returned by the large language model for all questions in the current batch.

[0138] (4) Call the structured parsing of the LLM response, parse and convert the free text response obtained from the language model into structured data (usually a set of key-value pairs, where the key is the field name of the question and the value is the parsed answer) based on the definition of the question in the codebook and the expected answer format.

[0139] (5) Accumulate and merge the structured coding results obtained from parsing the current batch of questions into the total result set maintained for the document.

[0140] The entire process will be executed repeatedly until all the question batches defined in the codebook have been processed for the current document. Finally, the process returns a structured data object containing the complete analysis results of all coding questions for the single document.

[0141] (2) Language Model Interaction and Fault Tolerance Mechanism

[0142] All communication between the system and the Large Language Model service is performed through a specially designed interaction interface. This interface is responsible for receiving the prompt text to be sent to the language model, the initialized language model client instance, the specified model name, and optionally, specific system prompt text that may be used to override or enhance the default system-level instructions.

[0143] Within its internal logic, this interface assembles a request body based on the technical specifications of the selected language model service API. This request body typically consists of two main parts: system-level role or behavior instructions, such as the aforementioned general guidance instructions, which set the basic behavior of the language model; and user-level, task-specific instructions, such as dynamic analysis prompts containing the text to be analyzed and encoding issues.

[0144] Taking into account the complexity of actual network environments and the intermittent instability that may occur in API services, the interactive interface has built-in fault tolerance and retry mechanisms. Within the preset maximum number of attempts, if an API call fails due to network timeout, temporary server error, etc., the interface will automatically wait for a preset delay period before re-initiating the same request. For large language models with special designs or specific functions, such as displaying their internal "thinking process" or performing complex reasoning, the interactive interface may need to use streaming data transmission for request and response processing.

[0145] In this mode, the API captures the final answer generated by the model, while also receiving and recording the intermediate steps or reasoning chains involved in generating the answer. This information is then combined into a more complete and transparent response. If all pre-set retries are exhausted and the API call remains unsuccessful, the API will not attempt further and will return an error message clearly indicating the error, allowing the caller to identify and address the issue.

[0146] (3) Master control batch processing flow

[0147] In the entire qualitative analysis workflow, there is a master batch processing process that is responsible for coordinating and managing the automated coding analysis of multiple documents.

[0148] After the process is started, it will first obtain the path list of all documents to be analyzed from the input directory specified by the user through the file system access mechanism.

[0149] The master control process then iterates through this list of document paths. For each document path in the list, it sequentially invokes the aforementioned single-document processing workflow to complete a comprehensive encoding analysis of the document. Once a document is processed and its encoding results are returned, they are added to a global result list that stores the analysis results of all processed documents. To enhance system robustness and prevent the loss of processed data due to unexpected interruptions such as system failures or power outages during extended operations, the system optionally invokes a result saving mechanism after processing each document or after a certain number of documents, periodically saving the accumulated analysis results to designated intermediate output files. The master control process also updates the processing status, recording the list of successfully processed files, allowing for recovery from interruptions and avoiding duplicate processing if the task is unexpectedly interrupted. Furthermore, to prevent excessive API requests from triggering the service provider's rate limiting policy or to alleviate transient pressure on the language model server, the master control process introduces a short, configurable delay between completing one document and preparing to process the next.

[0150] 5. LLM response analysis and consistency assurance module:

[0151] (1) Structural analysis of LLM response:

[0152] The core function of this module is to convert the response content returned by the large language model for a complete batch of questions, which is usually in free text format, into a structured data format that is easy to process and store later.

[0153] When the parsing process starts, the main input received is the original response text of the large language model and the original codebook question batch data structure that defines each question in the batch. The latter contains key metadata such as the ID of each question, the expected answer format, and the field name.

[0154] The parser will traverse each coded question object in the current question batch. For each question, it will try to apply a series of pre-designed parsing rules. The rules are based on the pattern matching ability of regular expressions, or supplemented by precise string search and positioning methods, to accurately find and extract the answer part corresponding to the ID or specific identifier of the current question. To improve the robustness and adaptability of the parsing process, the system designs and tries multiple alternative matching modes in sequence. Once the original answer text of a question is successfully extracted, it will also go through a series of necessary cleaning and normalization steps. At the same time, targeted subsequent processing is carried out according to the expected answer format preset in the coding book for that question.

[0155] After the aforementioned extraction, cleaning, and formatting, the parsed results for each question are stored in a mapping object or dictionary structure, using the logical field name preset in the codebook as the key and the processed answer content as the value. This is then returned as the result of the batch parsing. If the answer to a particular question cannot be successfully parsed for any reason during the parsing process, the system will assign a clear error flag or default value to the corresponding field in the question to ensure the integrity of the data structure and the clarity of subsequent processing.

[0156] (2) Mechanisms to ensure coding consistency:

[0157] Coding consistency is one of the core elements to ensure the quality and reliability of qualitative research results. This system maximizes the consistency of the automated coding process through a collaborative mechanism designed across multiple core modules.

[0158] (1) Unified encoding instruction generation: In the dynamic construction of targeted analysis prompts, the prompts generated for each question or batch of questions to be analyzed strictly follow a set of standardized templates and instruction structures. Regardless of the document or question being processed, the core task requirements, role settings, and basic expectations for the output received by the large language model are highly consistent.

[0159] (2) Standardized response format requirements: When constructing dynamic prompts, the system will clearly and specifically instruct the large language model in which preset format to organize and output its analysis results. The pre-agreed and mandatory requirements for the output format greatly improve the accuracy and efficiency of subsequent automated analysis and promote consistency in response content from the source.

[0160] (3) Structured result parsing logic: The dynamic construction of targeted analysis prompt words adopts a precise parsing strategy based on pattern matching, which forcibly maps the relatively free text output returned by the large language model back to the various pre-defined coding fields in the coding book, ensuring that the data extracted from different responses are structurally unified and standardized.

[0161] (4) Centralized and orderly result storage and export: In the final structured result integration and export module, the system will strictly organize and present the analysis data according to the field order and name defined in the codebook. Whether it is the generated intermediate summary table or the final exported Excel or CSV file, its column structure is directly derived from the definition of the codebook. Ensuring that the output data has the same structure across different analysis tasks also facilitates subsequent cross-case comparisons, data merging, and statistical analysis, thereby further ensuring the consistency and comparability of the analysis results at the macro level.

[0162] 6. Structured result integration and export module:

[0163] (1) Integration and formatting of coding results:

[0164] Once the coding process for the current document is complete, the system invokes a results saving and exporting program to organize the analysis results into a standardized format, ensuring data integrity and consistency. Key inputs to this program include a cumulative list of all processed documents, each containing its own analysis results; the full path to the target output file specified by the user; and the original codebook, which serves as the baseline for the entire analysis process.

[0165] The result saving and exporting program first determines the header and order of the final output table based on the structure of the coding book. Usually, the "file name" or "document identifier" will be set as the first column of the table to clearly identify the original document corresponding to each row of data. The remaining columns strictly follow the logical field names corresponding to each coding problem in the coding book, or to enhance readability, the text description of the coding problem itself is directly used as the column name, and arranged according to the order or importance preset in the coding book. If specific error information is recorded during the processing of a single document, the field used to store this error information will also be dynamically added to the header so that users can quickly locate and troubleshoot the problem.

[0166] The program then iterates through each document-level analysis result object in the result list. For each analysis result object, the values ​​corresponding to each field are extracted according to the previously determined header order. If a specific coded field does not exist in the analysis results for a document, the cell is filled with an empty string or a predefined placeholder to maintain the table structure. In this way, a complete data row is constructed for each processed document. All the data rows of these documents are ultimately aggregated into a two-dimensional data row list.

[0167] (2) Result export:

[0168] Once all analysis results are ready, the system uses a dedicated helper function to export the structured data to an Excel or CSV file for subsequent processing and analysis. This helper function takes as input the organized header information and a list of rows containing all document encoding data, and writes this content to a file pointed to by the user's previously specified output file path in a standardized format.

[0169] The system typically supports exporting results to widely used structured file formats, such as Microsoft Excel spreadsheets or comma-separated value text files. These formats offer excellent versatility and compatibility, enabling researchers to use various office software, statistical analysis packages, database management systems, or specialized qualitative data analysis software for subsequent in-depth data mining, statistical analysis, chart visualization, content review, and cross-validation. The success or failure of file write operations, as well as any problems encountered, are recorded in detail in the run log for user auditing and troubleshooting.

[0170] In addition, the present invention is also equipped with the following technical components:

[0171] 1) API Interaction Layer: Through functions such as init_openai_client and send_to_ai, it achieves stable connection and interaction with the large language model API, supports multiple model options, and meets analysis needs of different complexity and depth.

[0172] 2) Codebook generation engine: Through the generate_codebook_prompt and generate_codebook_with_ai functions, intelligent codebook generation based on research information is realized automatically to ensure the generation of a structured and expected coding framework.

[0173] 3) Dynamic prompt builder: Through the create_dynamic_prompt function, prompts for different types of coding questions are dynamically constructed, and large language models can perform qualitative analysis tasks.

[0174] 4) Batch document processor: Through the process_files_background and process_document functions, a batch strategy is adopted to efficiently process a large number of documents, and multi-threading technology is used to improve processing efficiency.

[0175] 5) Response parsing engine: Through the parse_dynamic_response function, the response of the large language model is structured and parsed to ensure the correct extraction of analysis results and conversion into a standardized research data structure.

[0176] 6) Result integration and export system: Through the save_results_to_excel function, the analysis results are integrated into structured data and support Excel format export, so that the result format that meets the researcher's expectations can be obtained.

[0177] The above are merely preferred embodiments of the present invention and are intended to help understand the method and core concept of this application. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the concept of the present invention fall within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

[0178] The present invention comprehensively solves the core technical problems of efficiency bottlenecks, uncontrollable consistency, cognitive limitations, and insufficient intelligence in the current field of qualitative analysis technology in the existing technology. By organically integrating cutting-edge artificial intelligence technology with traditional qualitative research methods, an end-to-end intelligent coding framework is constructed, which can efficiently, consistently and deeply carry out ultra-large-scale qualitative analysis, significantly improving the efficiency, depth and accuracy of qualitative analysis. The present invention adopts a modular design. The core advantage of the system lies in utilizing the deep semantic understanding and reasoning capabilities of large language models, organically combining the professional knowledge of qualitative analysis with advanced artificial intelligence technology, and being able to deeply analyze ultra-large-scale sample data in batches, fundamentally breaking through the bottleneck of traditional methods.

Claims

1. A general qualitative analysis coding system based on LLM, characterized by: Including intelligent codebook automatic generation technology, dynamic prompt engineering and document preprocessing module, batch document automatic coding module, LLM response analysis and consistency assurance module, structured result integration and export module; The intelligent codebook automatic generation technology introduces a large language model to assist in the generation of the codebook, and combines it with an iterative optimization process of human-computer collaboration to participate in and dominate the final form and content of the codebook. Specifically, it includes: S1, automatic construction and import of the initial codebook: supports the generation of a preliminary codebook framework for the large language model, or optimizes it by loading the user's existing codebook; S2, Human-computer collaborative iteration and continuous optimization of the codebook: Continuous optimization and iterative improvement of the codebook. Through multiple human-computer collaborative feedback, the structure and content of the codebook are gradually improved to ensure that it meets the needs of the research task; The dynamic prompt engineering and document preprocessing module includes: S3, document content reading and preprocessing; S4, dynamic construction of targeted analysis prompt words: a dedicated processing unit receives the currently confirmed and effective codebook and the list of questions included in the analysis batch in the codebook, as well as the pre-processed content of the current document to be analyzed as input information, and generates complete prompt words; The batch document automatic encoding module supports batch processing of large-scale documents to ensure the efficiency and accuracy of the analysis process, including: S5, single document processing optimization: logical grouping, problems and batch processing of each document in the coding process to improve efficiency; S6, language model interaction and fault tolerance mechanism: all communications between the system and the large language model service are performed through a specially designed interaction interface; S7, master batch processing: responsible for coordinating and managing the automated coding analysis of multiple documents in the entire qualitative analysis workflow; The LLM response parsing and consistency assurance module includes: S8, structured parsing of LLM responses: converting the responses returned by the large language model for the complete batch of questions, which are usually in free text format, into a structured data format that is easy to process and store later; S9, Coding Consistency Guarantee Mechanism: This mechanism is used to ensure coding consistency across documents and problem batches, maximizing the reliability of analysis results. The structured result integration and export module includes: S10, Coding result integration and formatting: Once the coding process of all documents is completed, the system will call the result saving and export program to organize the analysis results into a standardized format to ensure the standardization and consistency of the data; S11, result export: After all analysis results are prepared, the system exports the structured data into Excel or CSV files through a special auxiliary function to facilitate subsequent data processing and analysis. The auxiliary function receives the sorted header information and a list of data rows containing all document encoding data as input, and writes these contents into the file pointed to by the output file path previously specified by the user in a standardized format.

2. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: S1 specifically includes: confirming the initial version of the codebook, which includes three initialization methods: S1-1, Large language model assisted generation of initial codebook framework: by inputting research project information, the large language model constructs detailed, context-aware high-level instructions; The instructions guide the large language model to play the role of a qualitative research expert and require it to output a structured codebook draft based on the provided research information; the draft includes a descriptive name of the codebook, an overall description of the purpose and scope of the codebook, and a structured list of questions; Each independent coded question or entry in the list includes: a unique internal identifier ID, a specific textual description of the question, a logical field name for subsequent data storage and management, supplementary instructions or operational instructions for the question, and the analytical batch number to which the question logically belongs; The instructions require the large language model to provide brief descriptive text for each set analysis batch; The encoding system sends the high-level instructions to the model through its interactive interface with the large language model, obtains the codebook draft text generated by the model, and automatically parses it into its internal operational codebook data structure by the encoding system; S1-2, load the user's existing codebook: The coding system allows users to load their existing complete codebooks stored in JSON or other standardized structured formats through a designated import function as a starting point for subsequent iterative optimization; S1-3, template-based or interactively created from scratch: The coding system prompts the user and provides a basic, empty codebook template, or the user constructs the codebook item by item from scratch through an interactive interface provided by the coding system.

3. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: S2 specifically involves combining the auxiliary analysis capabilities of the large language model with the researcher's professional judgment and final decision-making power, through multiple cycles of feedback and modification, until the quality and applicability of the codebook reach a state that satisfies the user. Specifically, it includes the following steps: S2-1, presenting the current version of the codebook and the optimization suggestions provided by the large language model in the previous round through a visual and editable user interface; S2-2, the interface clearly displays all components of the codebook and supports intuitive and convenient editing of various codebook contents, such as the specific wording of questions, the selection of answer formats, the addition and deletion of predefined options, and the allocation of questions between different analysis batches; S2-3, the editing operations include "save current edit", "request AI optimization suggestions" or "finally confirm and solidify the codebook".

4. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: The specific steps of document content reading and preprocessing are: S3-1, extracts text information from the user-specified file path by reading the character encoding standard to avoid garbled characters and improve compatibility with documents from different sources; S3-2, using document truncation and content summarization mechanisms to avoid the text length limitation of the large language model when inputting a single input; S3-3, the mechanism receives the original text content and a preset maximum allowed number of characters or an equivalent number of tokens as parameters. If it is detected that the text length exceeds this limit, the mechanism does not truncate from the end, but instead prioritizes retaining the beginning and end of the document containing important introductions, conclusions, or contextual information, and inserts clear omission or truncation prompt information between the two.

5. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: The specific steps of dynamically constructing the targeted analysis prompt words are: S4-1, based on the input information of S4, constructs an introductory opening text to set the role of the large language model in the current task and clearly indicates the name of the research project to which the current analysis belongs and the theme or focus of the current batch of questions; S4-2, emphasizing the answer format specification to be followed to the large language model; S4-3, traversing each specific coding question in the current question batch, and for each coding question, integrating the unique identification ID and complete text content of the question into the prompt word being constructed; S4-4, dynamically generates highly targeted answer instructions based on the answer format type preset in the codebook for the question and the predefined options it may contain; S4-5, if the question object in the codebook also contains supplementary explanations, notes, or contextual prompts for the question, these are also integrated into the prompt words as additional guidance for accurately understanding and answering the large language model; S4-6, appending the pre-processed document content currently to be analyzed completely to the end of S4-1 to S4-5 and the question list.

6. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: The specific steps of S5 are: S5-1, the key input information received includes the storage path of the document to be analyzed, the previously constructed and finalized codebook, the initialized language model service client instance, and the specific language model name selected by the user; S5-2, calling S3 to complete the loading, encoding conversion and intelligent truncation of the specified document content based on the length limit; S5-3, the auxiliary functional unit reorganizes and groups all coding questions in the codebook according to the "batch" attribute attached to each coding question, forming a series of logically related question batch units suitable for step-by-step processing, which is used to manage the complexity of analysis and optimize the interaction efficiency with the language model; S5-4, for the current document, process each batch of questions in turn, and perform the following core operation sequence for all questions in each batch of questions: S5-4-1, mobilizing S4 to generate highly customized dynamic analysis prompt words containing detailed answer instructions for the current question batch and the current document or its pre-processed content; S5-4-2, sending the generated analysis prompt word along with the model selection and temperature parameter accompanying information to the large language model through S6 to request it to process, and the deep semantic understanding ability of the large language model is used to interpret the complex instructions in the prompt and the subtle information in the document; S5-4-3, receiving and temporarily storing the original text responses returned by the large language model for all questions in the current batch; S5-4-4, calling the free text response obtained from the large language model in S8, parsing and converting it into structured data according to the definition of the question in the codebook and the expected answer format; S5-4-5, accumulate and merge the structured coding results obtained from parsing the current batch of questions into the total result set maintained for the document; S5-5, S5-1 to S5-4 are executed in a loop until all problem batches defined in the coding book have been processed for the current document. The process returns a structured data object containing the complete analysis results of all coding problems of the single document.

7. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: The S6 is specifically: S6-1, the interface is responsible for receiving prompt text to be sent to the large language model, an initialized language model client instance, a specified model name, and optional system prompt text that may be used to override or enhance default system-level instructions; S6-2, the interface includes assembling a request body, the request body including: S6-2-1, system-level role or behavior instructions, i.e., the general default system-level instructions, used to set the basic behavior mode of the language model; S6-2-2, user-level, task-specific instructions, i.e., dynamic analysis prompts containing the text to be analyzed and coding questions; S6-3, the interface has a built-in fault tolerance and retry mechanism: within the preset maximum number of attempts, if an API call fails due to network timeout or temporary server error, the interface will automatically wait for a preset delay time and then re-initiate the same request or use streaming data transmission to process the request and response; S6-4, the interface needs to capture the answer content finally generated by the model, synchronously receive and record the intermediate steps or reasoning chain in the process of generating the answer, and organically combine this information into a more complete and transparent response result; if all preset retry times have been exhausted and the API call is still unsuccessful, the interface will no longer try again, and will return an identifier that clearly contains the error information for the upper-level caller to know and take corresponding measures.

8. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: The specific steps of S7 are: S7-1, obtains the path list of all documents to be analyzed from the input directory specified by the user through the file system access mechanism; S7-2, the main control process will traverse this document path list, and for each document path in the list, it will call S5 in sequence to complete a comprehensive encoding analysis of the document; S7-3, whenever the processing of a document is completed and its encoding results are returned, these results are added to the global total result list used to store the analysis results of all processed documents; S7-4, the system selectively invokes the result saving mechanism after processing each document or after a certain number of documents, and periodically saves the currently accumulated analysis results to a designated intermediate output file. This is used to enhance the robustness of the system and prevent the loss of processed data due to unexpected interruptions such as system failures or power problems during long-term operation. S7-5, the main control process is responsible for updating the processing status and recording the list of successfully processed files so that the task can be restored from the breakpoint after an unexpected interruption to avoid repeated processing; S7-6, the main control process introduces a short configurable delay between processing one document and preparing to process the next document, to avoid triggering the service provider's rate limit policy due to too frequent API requests, or to reduce the instantaneous pressure on the language model server.

9. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: The specific steps of S8 are: S8-1, receiving the main input as the original response text of the large language model and the original codebook question batch data structure that defines each question in the batch, wherein the data structure includes the ID of each question, the expected answer format, and the field name key metadata; S8-2, iterate through each coded question object in the current question batch and attempt to apply a series of pre-designed parsing rules to each question. The parsing rules are based on the pattern matching capabilities of regular expressions or supplemented by precise string search and location methods to accurately find and extract the answer portion corresponding to the ID or identifier of the current question; S8-3, to improve the robustness and adaptability of the parsing process, multiple alternative matching modes are designed and tried in sequence. If the original answer text of a question is successfully extracted, after a series of cleaning and normalization steps, targeted subsequent processing is performed according to the expected answer format preset in the codebook for the question; S8-4, the parsing result of each question will be stored in a mapping object or dictionary structure using its preset logical field name in the codebook as the key and the processed answer content as the value, and will be returned as the result of the batch parsing; If the answer to a question cannot be successfully parsed for any reason during the entire parsing process, a clear error identifier or default value will be assigned to the corresponding field of the question to ensure the integrity of the data structure and the clarity of subsequent processing.

10. The universal qualitative analysis coding system based on LLM according to claim 1 is characterized in that: The specific steps of S9 are as follows: S9-1, unified coding instruction generation mechanism: using standardized templates to ensure that the processing instruction format of each question and batch is consistent, thereby improving analysis efficiency and consistency of results; S9-2, standardized response format requirements: when constructing dynamic prompt words, instructing the large language model in which preset format to organize and output its analysis results; S9-3, structured result parsing logic: The S8 adopts a precise parsing strategy based on pattern matching to forcibly map the relatively free text output returned by the large language model back to the various pre-defined coding fields in the coding book, so that the data extracted from different responses are structurally unified and standardized; S9-4, centralized and orderly result storage and export: In the structured result integration and export module, the analysis data is organized and presented according to the field order and name defined in the coding book; whether it is the generation of the intermediate summary table or the final exported Excel or CSV file, its column structure is directly derived from the definition of the coding book; ensuring that the output data has the same structure between different analysis tasks, which is used for subsequent cross-case comparison, data merging and statistical analysis.

Citation Information

Patent Citations

  • Intelligent knowledge assistant construction method and device based on large language model

    CN118194999A

  • Data mining and analysis method based on LLM large model

    CN118210914A