Stream tuning question and answer real-time analysis method based on code unified knowledge representation
Through multi-model collaborative processing and semantic-driven architecture, it achieves second-level feedback from voice input to structured output, solving the problems of low efficiency and poor adaptability of traditional epidemic investigation. It is suitable for voice transcription and structured information extraction in epidemiological surveys and telephone interview scenarios.
Patent Information
- Application Number
- CN202510780629.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional epidemic investigation work relies on manual operations, which is inefficient and has serious information errors and omissions. The existing semi-automatic system has a high error rate in noisy environments and cannot achieve real-time voice analysis and dynamic adaptation, resulting in insufficient timeliness in epidemic prevention and control.
A real-time analysis method for epidemic investigation and question answering adopts multi-model collaborative processing, dynamic pattern conversion and semantic-driven architecture. Through the collaboration of speech recognition, role labeling, coded pattern representation and large language model, it achieves feedback from speech input to structured output in seconds.
Significantly reduce labor costs and error rates, support streaming data processing, achieve feedback in seconds, and meet the needs of efficient and accurate flow investigation in epidemic prevention and control.
Smart Images

Figure CN120706438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and information extraction, and specifically to a real-time parsing method for epidemic investigation and question answering based on unified knowledge representation of code. Background Art
[0002] Epidemiological investigation (hereinafter referred to as "epidemiological investigation") is a core link in epidemic prevention and control. It provides key evidence for blocking the chain of transmission of the epidemic by quickly collecting information such as the activity trajectory of infected persons and the persons they have contacted. However, traditional epidemiological investigation is highly dependent on manual operations. Investigators need to ask respondents one by one through telephone or on-site interviews, and manually enter the answers into structured forms. This model faces severe challenges when responding to sudden and large-scale epidemics. On the one hand, manual form filling is inefficient, and a single epidemiological investigation takes tens of minutes or even hours. When an epidemic breaks out, hundreds or even thousands of investigations often need to be handled simultaneously, and labor costs rise sharply. On the other hand, manual record keeping can easily lead to information errors and omissions, such as incorrectly filling in fields, omitting key time nodes, or misinterpreting ambiguous expressions, which seriously affect data quality and the accuracy of prevention and control decisions.
[0003] To reduce manual reliance, some technologies attempt to achieve semi-automated processing through speech recognition or rule-based engines, but these limitations remain significant. Speech recognition models exhibit high error rates in the presence of noise, dialects, or multi-person conversations, and transcripts still require manual correction. Questionnaire parsing systems often rely on fixed templates or manually coded rules, making it difficult to dynamically adapt to frequent adjustments to epidemiological survey fields (such as the addition of fields like "exposure location type" and "vaccination batch"), resulting in high system maintenance costs. Furthermore, non-standardized statements from respondents (such as "lives in Jiading District" and "ordinary cadre and staff") often cannot be automatically mapped to standardized fields, requiring secondary intervention by epidemiologists. The sudden and rapid spread of the epidemic places even higher demands on the timeliness of epidemiological surveys. Existing systems often utilize offline batch processing, which cannot support real-time parsing and rolling output of streaming speech. This results in delays in critical information and hinders the speed of prevention and control responses. For example, a lag of several hours in the movement trajectory of an infected person could further spread the epidemic. While academics have explored coded pattern representation methods (such as the Know Coder framework) to enhance structured information extraction, these methods focus on static text processing and fail to address real-time speech parsing and dynamic pattern adaptation. While multi-model collaborative research has improved the performance of single tasks such as speech recognition and punctuation recovery, it has yet to achieve a deep integration of role labeling, semantic mapping, and automated value assignment.
[0004] To sum up, although the existing epidemic investigation technology has been optimized in single-task performance (such as improved voice recognition accuracy), it lacks end-to-end collaborative design and cannot achieve full-process automation from voice input to structured output. The core links still rely on manual operation and have not yet broken through the core bottleneck of "manual filling in forms", making it difficult to meet the comprehensive requirements of epidemic investigation work for efficiency, cost and real-time performance. Summary of the Invention
[0005] The purpose of the present invention is to provide a real-time parsing method for epidemic investigation and question-answering based on unified knowledge representation of code in response to the deficiencies of the existing technology. It adopts a method of converting streaming voice input into high-precision structured data in real time, and through a structured knowledge extraction framework of multi-model collaborative processing, dynamic mode conversion and semantic-driven architecture, it systematically solves the core problems of low efficiency of manual form filling, poor adaptability of semi-automatic systems, and insufficient ability to understand complex semantics in traditional epidemic investigation. The method supports dynamically adjusted questionnaire schema and realizes feedback in seconds, significantly reducing labor costs and error rates, meeting the demand for efficient and accurate epidemic investigation in epidemic prevention and control, supporting streaming data processing mechanism, and can generate and feedback parsing results in real time when voice recognition results are scrolled in. The method is simple and effective, and is particularly suitable for speech transcription and structured information extraction in scenarios such as epidemiological surveys and telephone interviews, and has good application prospects.
[0006] The specific technical solution for achieving the objectives of the present invention is: a real-time parsing method for epidemic investigation and question answering based on unified knowledge representation of code. The method is characterized by achieving noise-robust speech transcription and accurate role labeling through a multi-model collaborative mechanism, dynamically adapting complex schema constraints using a code-based schema representation method (CSR), and generating executable structured assignment statements based on the semantic understanding capabilities of a large language model (LLM). The method specifically includes:
[0007] 1) Streaming of interview recordings
[0008] The interview recordings are streamed through a speech recognition module with a multi-model collaborative architecture, converting the audio of the interview conversation into a text stream, i.e., structured text with role annotations. The speech recognition module includes: a Chinese speech recognition model, a voice activity detection model, a Chinese punctuation model, and a speaker recognition model;
[0009] 2) Questionnaire mode conversion
[0010] Extract the questionnaire table schema through a string processing algorithm and convert it into a Python class definition containing field annotations. The Python class definition includes: field identifiers, field descriptions, and data types, and is embedded in the code in the form of annotations.
[0011] 3) Extraction of questionnaire field information
[0012] Input Python class definitions and text streams into the large language model, and generate assignment statements through structured prompt word templates to extract questionnaire field information;
[0013] 4) Execute post-processing
[0014] 3. Convert the extracted information into the target output format of JSON, SQL, or XML, submit it to the downstream service interface for processing operations including field filtering and format conversion, and output the parsed results.
[0015] The voice activity detection model is used to segment the start and end time of the voice stream, and the speaker recognition model distinguishes the roles of the questioner and the answerer; the large language model includes: number standardization and address completion.
[0016] The string processing algorithm supports extracting field information from the DDL, TSV or JSON format questionnaire description of the SQL data table, thereby automatically generating Python class definitions and retaining the annotation information used to describe the fields. The field annotations meet the LLM semantic understanding requirements.
[0017] The structured prompt word template implements the following field assignment rules:
[0018] 1) Assign values based on the real name, gender, address hierarchy, and time information explicitly mentioned in the conversation;
[0019] 2) Assignment of fuzzy information is prohibited, and the output only contains executable Python assignment statements.
[0020] The target output format is generated by concatenating Python class definitions and assignment statements, and dynamically converted after execution by the Python interpreter.
[0021] The multi-model collaboration includes: a multi-model collaborative speech-to-text unit, a pattern representation unit, a knowledge extraction unit and a streaming data processing interface. The multi-model collaborative speech-to-text unit integrates the paraformer-zh, fsmm-vad, ct-punc-c and cam++ models; the pattern representation unit converts the questionnaire schema into a Python class definition that can be parsed by LLM; the knowledge extraction unit generates structured assignment statements based on LLM; the streaming data processing interface receives the audio stream in real time and scrolls to output the parsing results.
[0022] The real-time analysis supports a streaming data processing mechanism, and can generate and feed back analysis results in real time when the speech recognition results are scrolled into the input.
[0023] Compared with the prior art, the present invention has the following beneficial technical effects and significant technical progress:
[0024] 1) Combining speech processing, pattern representation, and LLM parsing effectively solves the problems of low accuracy and poor real-time performance of traditional methods in question-answering parsing tasks in disease control and epidemic investigation scenarios.
[0025] 2) Supports dynamically adjusted questionnaire schemas and provides feedback within seconds, significantly reducing labor costs and error rates, and meeting the demand for efficient and accurate epidemic investigation in epidemic prevention and control.
[0026] 3) Supports streaming data processing mechanism, which can generate analysis results and feedback in real time when the speech recognition results are scrolled into input.
[0027] 4) The method is simple and effective, and is particularly suitable for speech transcription and structured information extraction in scenarios such as epidemiological surveys and telephone interviews, and has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 Flowchart of the present invention. DETAILED DESCRIPTION
[0029] See Figure 1 The present invention achieves noise-robust speech transcription and accurate character labeling through a multi-model collaborative mechanism. It uses Code-based Schema Representation (CSR) to dynamically adapt to complex schema constraints and generates executable structured assignment statements based on the semantic understanding capabilities of the Large Language Model (LLM). The method specifically includes:
[0030] Step 1: Multi-model collaborative parsing of streaming speech
[0031] 1-1: Voice Activity Detection and Segmentation
[0032] The voice activity detection module is used to distinguish valid speech from background noise, and the threshold is adjusted dynamically. , the system can adapt to the noise level in different environments. For example, in a noisy epidemiological survey site, the energy standard deviation If the value is higher, the threshold will automatically increase to reduce false detection; conversely, in a quiet environment, the threshold will be lowered to retain low-volume speech. It reflects the model's assessment of the quality of the speech segment. Subsequent modules can prioritize high-confidence segments to improve overall efficiency.
[0033] The input streaming audio signal The FSMN-VAD (Feedforward Sequential Memory Network Voice Activity Detection) model is used to segment valid speech segments. The model analyzes the audio energy spectral density. and short-time zero-crossing rate , dynamically adjust the detection threshold, its dynamic threshold Calculated by the following formula:
[0034] .
[0035] in, 、 is an empirical coefficient, which has been verified through experiments to effectively balance noise suppression and speech segment integrity.
[0036] Output segmentation results It is expressed by the following formula:
[0037] .
[0038] in, For the The starting time of the speech segment; For the The end time of the speech segment; Represents the confidence of the speech segment, which is used for weighted fusion of subsequent modules.
[0039] 1-2: Multi-task cascade speech-to-text
[0040] A cascaded speech-to-text module improves transcription quality layer by layer. Paraformer-zh's non-autoregressive architecture significantly reduces inference latency, making it suitable for real-time scenarios. The punctuation recovery module not only adds punctuation marks but also captures contextual dependencies through a conditional random field model, such as automatically adding a question mark at the end of an interrogative sentence. The speaker recognition module uses voiceprint features to distinguish between the questioner and the answerer, ensuring that subsequent fields are correctly associated. The post-processing stage combines domain knowledge (such as administrative division libraries) for semantic completion. For example, the ambiguous "Jiading District" is parsed into the complete "Shanghai-Jiading District," avoiding manual secondary processing.
[0041] For each speech , execute the following processing flow in sequence:
[0042] 1) Speech recognition: Use the Paraformer-zh model (non-autoregressive structure, based on the self-attention mechanism) to transform the audio frame ( is the number of frames, is the Mel filter bank dimension) converted to original text .
[0043] 2) Punctuation restoration: Insert punctuation through the CT-PUNC-C model (conditional random field sequence labeling) and output , whose punctuation accuracy satisfies the following formula:
[0044]
[0045] Where .
[0046] 3) Speaker recognition: Extract voiceprint features based on the CAM++ model (multi-scale convolutional attention network) , calculate the speaker similarity matrix , and output the role label shown in the following formula :
[0047] .
[0048] Where represents the speaker identity label of the th voice segment.
[0049] 4) Post-processing: Use the LLM to perform semantic error correction and standardization:
[0050] a) Digital standardization: Map "twenty-seven" in "The age is twenty-seven years old" to a numerical value ;
[0051] b) Address completion: Based on the geographical knowledge graph , expand "Jiading District" to a hierarchical address ;
[0052] c) Semantic error correction: Correct recognition errors (such as "dry cloth clerk" → "cadre clerk").
[0053] Finally, generate the structured text represented by the following formula :
[0054] .
[0055] Where is the statement confidence; The time information of the jth voice segment.
[0056] Step 2: Dynamic mode conversion and coding representation
[0057] 2-1: Mode parsing
[0058] Extract field names , data types and comments For example, for the SQL statement: CREATE TABLE Questionnaire (xm VARCHAR(50) COMMENT 'patient name', nl INT COMMENT'age'), after parsing, the triple set represented by the following formula is generated: :
[0059] .
[0060] 2-2: Python class code generation
[0061] The code generation module converts schema descriptions into code structures understandable by the LLM. Using type mapping rules, the system automatically handles compatibility between different data types, for example, mapping a database VARCHAR to a Python str. Field annotations are embedded to help the LLM accurately understand field semantics. Furthermore, if the schema contains complex relationships (such as a hierarchical association between "address" and "administrative division"), the algorithm generates inheritance classes to capture constraints and ensure that subsequent assignments conform to business logic. The code generation strategy is as follows:
[0062] 1) Type mapping: convert SQL / JSON type Convert to Python type ,For example:
[0063] VARCHAR(50) → str
[0064] INT → int
[0065] Use the typing syntax supported by Python 3.6 and higher to annotate the field type together with the field.
[0066] 2) Annotation embedding: field description Inserting code in the form of comments helps LLM understand the meaning of the fields and accurately extract information. The following is an example of the generated Python class definition code:
[0067] class Questionnaire:
[0068] # Patient Name
[0069] xm: str
[0070] # Age (unit: years)
[0071] nl: int
[0072] # nationality
[0073] mz: str
[0074] Step 3: Semantic-driven structured knowledge extraction
[0075] 3-1: Prompt instruction design
[0076] Deeply integrate task instructions (Instruction), context fragments (Context) and pattern definitions (Class) to guide the large language model (LLM) to accurately generate structured assignment statements. Specifically, Prompt consists of three parts and is expressed as follows:
[0077] .
[0078] The instructions explicitly require the LLM to generate assignment statements based solely on the field information in the Context and Class fields within the current window, prohibiting assumptions about unmentioned and unreasonably inferred information. For example, if the Context contains the question "What's your age?" and the answer is "27 years old," and the class definition contains a field nl annotated with "Age (unit: years)," the LLM must generate obj.nl = 27. If the conversation only mentions "Ms. Wang" without providing a real name, the name field xm is skipped.
[0079] In order to achieve efficient processing in streaming scenarios, the present invention introduces a sliding window mechanism. For example, for a window size of Turn dialogue, step length round, then only the most recent The turn-based dialogue is used as context input to the LLM. This design ensures that the context length is always limited to In the round, we can avoid errors caused by LLM being distracted by long text, and the overlap of context can also ensure that key information can be processed repeatedly in multiple windows to prevent omissions. For the above example, the window overlap ratio Calculated by the following formula:
[0080] .
[0081] Generating Python assignment statements has multiple advantages. First, Python allows multiple assignments to the same field, with subsequent values automatically overwriting previous values. For example, if the first window generates obj.nl = 27, and the second window generates obj.nl = 28 due to a more complete context, the final value after execution is 28. This feature is perfectly suited for overlapping sliding windows, eliminating the need for complex deduplication logic. Second, Python code can be executed directly, dynamically updating object instances through the interpreter. For example, when step 3 generates a Python assignment statement:
[0082] obj.zy = 'cadres and staff' # from window 1;
[0083] obj.zy = 'Civil servant' # From window 2 (correction);
[0084] After executing the above assignment statement, the final value of the zy field is "civil servant".
[0085] Through the above design, the present invention reduces the computational burden of LLM while ensuring high accuracy and real-time performance in streaming scenarios, providing reliable technical support for rapid decision-making in epidemic prevention and control.
[0086] 3-2: LLM parsing and code generation
[0087] The LLM module strictly controls the generation parameters (such as low temperature value ) ensures output stability. Lower temperature values increase the probability that the model will select the highest-probability word, reducing random errors. A duplicate penalty mechanism suppresses redundant output, for example, by preventing multiple assignments to the same field. Generated assignment statements must strictly conform to Python syntax and field type constraints. For example, numeric fields must only accept integers or floating-point numbers. Illegal assignments (such as those that violate the anti-speculation rule) are filtered to ensure compliance of the output data.
[0088] Transcribe the text With class code Enter LLM and set the generation parameter temperature (reduce randomness), repeated punishment , the maximum generated length .
[0089] Step 4: Execution and post-processing
[0090] After obtaining the assignment statement, the necessary post-processing code segments are spliced before and after the code segment, and then the spliced Python program script is executed using the Python interpreter. The output obtained in JSON or other formats is the result of information extraction.
[0091] The present invention will be further described in detail below with reference to specific embodiments.
[0092] Example 1
[0093] Step 1: Multi-model collaborative parsing of streaming speech
[0094] The epidemiologists spoke with the subjects of the investigation and obtained a streaming call recording. Through multi-model collaborative analysis of the streaming voice, the transcript of part of the epidemiological investigation conversation was obtained as follows:
[0095] Epidemiologist: Hello, Ms. Wang. We are from a disease control center. We have detected that you are positive for a certain infectious disease. We have some questions for you. Is it convenient for you?
[0096] Subject of epidemiological investigation: Yeah, that’s convenient. It’s positive, right? Now?
[0097] Epidemiologist: Yes, it is positive. So, how old are you?
[0098] Epidemiological investigation subject: currently 27 years old.
[0099] Epidemic investigator: Okay, then what is your ethnicity?
[0100] Epidemiological investigation subject: I am Tibetan.
[0101] Epidemic investigator: What is your occupation?
[0102] Epidemiological investigation subjects: ordinary cadres and staff.
[0103] Some fields of the SQL table are as follows:
[0104] xm patient name varchar(100)
[0105] jzxm Parent name varchar(100)
[0106] lxfs contact phone varchar(20)
[0107] zjlx certificate type varchar(5)
[0108] Step 2: Dynamic mode conversion and coding representation
[0109] Through dynamic mode conversion, the Python class definition is as follows:
[0110] from datetime import datetime
[0111] class Questionnaire:
[0112] # Patient Name
[0113] xm: str
[0114] # Parent's name
[0115] jzxm: str
[0116] # Contact Number
[0117] lxfs: str
[0118] # Document type
[0119] zjlx: str
[0120] Step 3: Semantic-driven structured knowledge extraction
[0121] Through semantic-driven structured knowledge extraction, some of the assignment statements output by LLM are as follows:
[0122] obj = Questionnaire()
[0123] obj.nl = 27
[0124] obj.mz = 'Tibetan'
[0125] obj.zy = 'Ordinary cadres and staff'
[0126] Step 4: Execution and post-processing
[0127] Executed through the Python interpreter, some of the output is as follows:
[0128] {"nl": 27; "mz": 'Tibetan'; "zy": 'Ordinary cadre and staff'}.
[0129] The above is only a further explanation of the present invention and is not intended to limit the patent of the present invention. Any equivalent implementation of the present invention should be included in the scope of the claims of the patent of the present invention.
Claims
1. A real-time parsing method for epidemic investigation and question answering based on unified knowledge representation of code, characterized by: A multi-model collaborative approach is used to achieve speech transcription and role labeling. A coded pattern representation method is used to dynamically adapt to complex schema constraints. An executable structured assignment statement is generated based on a large language model. The method specifically includes the following steps: 1) Streaming of interview recordings The interview recordings are streamed through a speech recognition module with a multi-model collaborative architecture, converting the audio of the interview conversation into a text stream, i.e., structured text with role annotations. The speech recognition module includes: a Chinese speech recognition model, a voice activity detection model, a Chinese punctuation model, and a speaker recognition model; 2) Questionnaire mode conversion Extract the questionnaire table schema through a string processing algorithm and convert it into a Python class definition containing field annotations. The Python class definition includes: field identifiers, field descriptions, and data types, and is embedded in the code in the form of annotations. 3) Extraction of questionnaire field information Input Python class definitions and text streams into the large language model, and generate assignment statements through structured prompt word templates to extract questionnaire field information; 4) Execute post-processing Convert the extracted information into the target output format of JSON, SQL or XML, submit it to the downstream service interface for processing operations including field filtering and format conversion, and output the parsing results.
2. The method for real-time analysis of epidemic investigation and question answering based on unified code knowledge representation according to claim 1 is characterized in that: The voice activity detection model is used to segment the start and end time of the voice stream, and the speaker recognition model is used to distinguish the questioner and the answerer roles; The large language model includes: number standardization and address completion.
3. The real-time parsing method for epidemic investigation and question answering based on unified code knowledge representation according to claim 1 is characterized in that: The string processing algorithm supports extracting field information from the DDL, TSV or JSON format questionnaire description of the SQL data table, thereby automatically generating Python class definitions and retaining the annotation information used to describe the fields. The field annotations meet the LLM semantic understanding requirements.
4. The method for real-time analysis of epidemic investigation and question answering based on unified code knowledge representation according to claim 1 is characterized in that: The structured prompt word template implements the following field assignment rules: 1) Assign values based on the real name, gender, address hierarchy, and time information explicitly mentioned in the conversation; 2) Assignment of fuzzy information is prohibited, and the output only contains executable Python assignment statements.
5. The method for real-time analysis of epidemic investigation and question answering based on unified code knowledge representation according to claim 1 is characterized in that: The target output format is generated by concatenating Python class definitions and assignment statements, and dynamically converted after execution by the Python interpreter.
6. The method for real-time analysis of epidemic investigation and question answering based on unified code knowledge representation according to claim 1 is characterized in that: The multi-model collaboration includes: a multi-model collaborative speech-to-text unit, a pattern representation unit, a knowledge extraction unit and a streaming data processing interface. The multi-model collaborative speech-to-text unit integrates the paraformer-zh, fsmm-vad, ct-punc-c and cam++ models; the pattern representation unit converts the questionnaire schema into a Python class definition that can be parsed by LLM; the knowledge extraction unit generates structured assignment statements based on LLM; the streaming data processing interface receives the audio stream in real time and scrolls to output the parsing results.