Medical record automatic generation method and system based on multi-modal data and active perception
Patent Information
- Application Number
- CN202610517158.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-20
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-04-20
AI Technical Summary
[0003]当前病历生成主要有以下几类方式:人工录入与模板化、语音识别与NLP抽取、图像与视频识别和设备数据接入与自动填充;这些方式在一定程度上提高了病历生成效率,但仍存在不足:被动采集、多模态融合不足、字段间缺乏一致性等
[0060]本发明提出基于多模态数据与主动感知的病历自动生成方法及系统,通过采集多模态原始数据,构建字段-证据图,并结合临床知识库与时序规则形成约束集合,再依据信息增益策略主动触发补充采集动作,直至关键字段满足阈值条件,最后对候选值进行约束解码并输出带有证据指纹链的病历;该方法不仅能全面覆盖病历所需的多模态信息,还能在采集过程中动态发现信息缺口并主动补充,保证字段间的逻辑一致性与时序合理性,同时通过证据链条提供可审计的溯源能力,提升了病历生成效率,同时显著增强结果的完整性、可靠性与合规性。
Smart Images

Figure CN122050674B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, specifically a method and system for automatically generating medical records based on multimodal data and proactive perception. Background Technology
[0002] With the rapid development of smart healthcare and artificial intelligence technologies, the application of electronic medical records has gradually become widespread. How to generate medical records efficiently, accurately, and compliantly has become a key issue in the construction of medical information systems.
[0003] Currently, medical record generation mainly takes the following forms: manual input and templates, speech recognition and NLP extraction, image and video recognition, and device data access and automatic filling. These methods have improved the efficiency of medical record generation to some extent, but still have shortcomings: passive collection, insufficient multimodal fusion, and lack of consistency between fields.
[0004] To address the aforementioned issues, researchers began exploring the application of multimodal data fusion and proactive sensing strategies in medical record generation. This approach breaks through the traditional static recognition model of medical record generation, possessing the capability of closed-loop data collection, fusion reasoning, field filling, and result auditing. Although multimodal and proactive sensing have brought new directions to automatic medical record generation, the following issues remain: How to proactively sense and generate medical records based on multimodal data. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention proposes an automatic medical record generation method and system based on multimodal data and proactive perception. This method can fully utilize multimodal data and introduce a proactive perception mechanism to form a closed loop in the stages of data collection, fusion, reasoning, filling, and tracing, thereby meeting the multiple clinical needs for efficiency, accuracy, and compliance.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] Automatic medical record generation methods based on multimodal data and proactive sensing include:
[0008] Collect multimodal raw data from patients and doctors to construct field evidence graphs;
[0009] A constraint set is generated based on the clinical ontology and temporal rules, and priority and dynamic thresholds are configured for each medical record field.
[0010] Based on information gain, additional data collection actions are generated, and the field evidence graph is iteratively updated. The iterative update controls the additional data collection actions of the target field until the target field reaches the dynamic threshold or satisfies the constraint set. The target field is a medical record field that does not satisfy the constraint context or is below the dynamic threshold.
[0011] Constraint decoding is performed on the medical record fields in the iteratively updated field evidence graph to obtain the values of each target field, and the medical record with an attached evidence fingerprint chain is output.
[0012] Specifically, the collection of multimodal raw data from patients and doctors, and the construction of a field evidence graph, includes:
[0013] Collect multimodal raw data from patients and doctors, including audio streams, video streams, and medical device data from consultations; preprocess the multimodal raw data, including time alignment and channel identification marking.
[0014] A pre-defined set of medical record fields is used to generate a set of corresponding candidate values for the fields based on the pre-processed multimodal raw data. The units, names, and time annotations are standardized to establish a one-to-one reference relationship between each candidate value of the field and its source fragment.
[0015] For each field candidate value and its source fragment, an evidence fingerprint is generated. The evidence fingerprint includes time interval, channel and device identifier, spatial pose parameters, sampling configuration and fragment summary information.
[0016] Using medical record fields, field candidate values, and evidence fingerprints as graph nodes, and the support relationships, rebuttal relationships, chronological relationships, and source relationships between graph nodes as graph edges, a field evidence graph is constructed, and the edges of the field evidence graph are labeled with source type and channel label.
[0017] Specifically, the process of generating a constraint set based on clinical ontology and temporal rules, and configuring priorities and dynamic thresholds for each medical record field, includes:
[0018] The terms, entities and relations in the clinical ontology are parsed and normalized to extract the time-series anchor point set and scene label set related to the medical record field, and a constraint candidate library for the medical record field is generated.
[0019] The constraint candidate library is mapped to a preset medical record field set to form a field-level rule template. The field-level rule template includes mutual exclusion relationship, implication relationship, value range, legal unit of measurement, chronological relationship and necessary field set, and the source, version, applicable department and applicable population conditions are marked for each rule.
[0020] Based on the medical visit scenario and patient stage information, a related subset is selected from the field-level rule template to establish cross-visit vertical association constraints and stage-based time window constraints, and key fields in historical medical records are synchronized to the current constraint context to generate a constraint set.
[0021] A hierarchical priority queue is assigned to each medical record field, and a dynamic threshold strategy is set for the field according to the type of evidence source, channel status, patient interaction status and scenario complexity to obtain an adjustable mapping from field to threshold.
[0022] The associated subset is subjected to field-level rule conflict adjudication and effective orchestration. The adjudication results are determined according to the source trust level, version timestamp and department domain priority order, and the session-level hard constraint set and soft constraint set are output.
[0023] Specifically, the process of adjudicating and orchestrating field-level rule conflicts in the associated subset, determining the adjudication results according to source trust level, version timestamp, and departmental domain priority, and outputting a session-level set of hard constraints and a set of soft constraints includes:
[0024] The associated subsets are aggregated and indexed to generate a conflict candidate set containing mutual exclusion, overriding, overlapping and out-of-domain references, and the source trust level, version timestamp and effective conditions are recorded for each field-level rule in the associated subsets;
[0025] A conflict resolution sequence is generated based on the source trust level, version timestamp, and effective conditions. The conflict resolution sequence is then sorted again using the patient's multimodal raw data to obtain the execution order for each conflict tuple.
[0026] The conflicting tuples are adjudicated one by one in the order described, including the processes of retention, downgrading, freezing, and replacement;
[0027] The rules at the field level in the associated subset after the adjudication are arranged to take effect, and the triggering stage, triggering conditions, interlocking relationships and coverage boundaries are marked. The session-level hard constraint set and soft constraint set are output.
[0028] Specifically, the step of generating additional data collection actions based on information gain and iteratively updating the field evidence graph includes:
[0029] Retrieve medical record fields that do not meet the constraint set or are below the dynamic threshold from the field evidence graph to form a target field set, and group the target fields according to the priority of medical record fields and the stage of medical treatment;
[0030] For each target field, a corresponding set of additional collection actions is generated from the preset action library, and the expected requirements are labeled for each action.
[0031] The action candidate set is evaluated and sorted according to the information gain strategy to generate an additional collection action execution sequence, and actions with dependencies are split and actions that can be parallelized are merged.
[0032] Additional data collection actions are issued according to the execution order to obtain new evidence. The new evidence is then aligned with the time base, verified for channel consistency, and deduplicated for fragments to generate evidence identifiers and evidence fingerprints.
[0033] The newly added evidence is attached to the field evidence graph, and the supporting relationships, rebuttal relationships, sequential relationships and source relationships are updated. The target field set and action candidate set are calculated based on the updated field evidence graph until the constraint set is satisfied or the dynamic threshold is exceeded.
[0034] Specifically, the step of issuing additional data collection actions according to the execution order to acquire new evidence, performing time base alignment, channel consistency verification, and fragment deduplication on the new evidence, and generating evidence identifiers and evidence fingerprints includes:
[0035] According to the execution order, additional acquisition actions are issued to each acquisition channel. The additional acquisition action information is encapsulated into control instructions containing source type, channel label and sampling window, and a tracking number is registered for each control instruction.
[0036] Receive additional collected multimodal raw data, segment it according to the sampling window, generate corresponding evidence fragments, and attach an initial timestamp and channel label to each evidence fragment;
[0037] The evidence fragments are time-aligned and channel-consistent. Fragments that fail the verification are moved to an isolation list, while fragments that pass the verification are moved to an acceptance list.
[0038] The fragments in the acceptance list are deduplicated to generate a unique evidence identifier, and an evidence fingerprint is generated based on the initial timestamp, channel label and evidence identifier.
[0039] Specifically, the process of performing constraint decoding on the medical record fields in the iteratively updated field evidence graph to obtain the values of each target field and outputting the medical record with an attached evidence fingerprint chain includes:
[0040] Based on the iteratively updated field evidence graph, the hard and soft constraint versions of the current session are locked, and a set of medical record fields to be decoded and a list of constraint references are generated according to the medical record fields and time series anchors.
[0041] The candidate values of the medical record fields to be decoded are preprocessed to achieve consistency, forming a set of decodable candidates;
[0042] Based on the constraint reference list and the decodable candidate set, a decoding priority sequence is constructed, and a conflict handling order is set.
[0043] Based on the decoding priority sequence, the decodeable candidate set is constrained and decoded. Under the premise of satisfying hard constraints, a target candidate value is selected for each medical record field, and a freeze mark and re-entry condition are set for fields that cannot be directly parsed.
[0044] A medical record field mapping table is generated for the target candidate value, a pointer chain between field values and corresponding evidence fragments is established, and the medical record with the evidence fingerprint chain is output by combining the medical record field mapping table and the evidence fingerprint chain.
[0045] Specifically, a medical record field mapping table is generated for the target candidate value, a pointer chain between field values and corresponding evidence fragments is established, and the medical record with the evidence fingerprint chain is output by combining the medical record field mapping table and the evidence fingerprint chain, including:
[0046] Establish a medical record field mapping table according to the field identifier of the target candidate value, register each medical record field with the selected target candidate value, candidate source and constraint reference number, and assign a unique mapping sequence number to each entry;
[0047] Retrieve evidence fragments associated with the target candidate value, and concatenate them according to the mapping sequence number of the medical record field mapping table to form an evidence chain structure arranged in the order of medical record fields;
[0048] Add evidence fingerprints to the evidence fragments in the evidence chain structure. The evidence fingerprints include the initial timestamp and channel parameters of the evidence fragments. Merge the evidence fingerprints with the evidence fragments to generate an evidence fingerprint chain.
[0049] The medical record field mapping table is combined with the evidence fingerprint chain to form a medical record with an attached evidence fingerprint chain.
[0050] The system for automatically generating medical records based on multimodal data and proactive perception is used to implement an automatic medical record generation method based on multimodal data and proactive perception. It includes: a graph construction module, a constraint configuration module, an iterative update module, and a medical record generation module.
[0051] The graph construction module is used to collect multimodal raw data from patients and doctors and construct field evidence graphs.
[0052] The constraint configuration module is used to generate a constraint set based on the clinical ontology and time sequence rules, and to configure priority and dynamic threshold for each medical record field.
[0053] The iterative update module generates additional acquisition actions based on information gain and iteratively updates the field evidence graph.
[0054] The medical record generation module is used to perform constraint decoding on the medical record fields in the iteratively updated field evidence graph to obtain the values of each target field and output the medical record with an attached evidence fingerprint chain.
[0055] Specifically, the medical record generation module includes: a decodable candidate set generation unit, a constraint decoding unit, and a medical record generation unit;
[0056] The decodable candidate set generation unit generates a set of medical record fields to be decoded and a list of constraint references based on the iteratively updated field evidence graph, according to the medical record fields and time-series anchors, and performs consistency preprocessing on the candidate values of the medical record fields to be decoded to form a decodable candidate set.
[0057] The constraint decoding unit constructs a decoding priority sequence based on the constraint reference list and the decodable candidate set, performs constraint decoding on the decodable candidate set based on the decoding priority sequence, selects target candidate values for each medical record field under the premise of satisfying hard constraints, and sets freeze flags and re-entry conditions for fields that cannot be directly parsed.
[0058] The medical record generation unit is used to generate a medical record field mapping table for the target candidate value, establish a pointer chain between field values and corresponding evidence fragments, and output a medical record with an attached evidence fingerprint chain by combining the medical record field mapping table and the evidence fingerprint chain.
[0059] Compared with the prior art, the beneficial effects of the present invention are:
[0060] This invention proposes an automatic medical record generation method and system based on multimodal data and proactive perception. It collects raw multimodal data, constructs a field-evidence graph, and combines it with a clinical knowledge base and temporal rules to form a constraint set. Then, based on an information gain strategy, it proactively triggers supplementary data collection actions until key fields meet threshold conditions. Finally, it performs constraint decoding on candidate values and outputs a medical record with an evidence fingerprint chain. This method not only comprehensively covers the multimodal information required for medical records but also dynamically identifies and proactively fills information gaps during the collection process, ensuring logical consistency and temporal rationality between fields. Simultaneously, it provides auditable traceability through the evidence chain, improving medical record generation efficiency and significantly enhancing the completeness, reliability, and compliance of the results. Attached Figure Description
[0061] Figure 1 The flowchart of the automatic medical record generation method based on multimodal data and active perception provided by the present invention;
[0062] Figure 2 A schematic diagram of the field evidence provided by this invention;
[0063] Figure 3 This is a schematic diagram of a medical record provided by the present invention;
[0064] Figure 4 This is a diagram illustrating the architecture of the automatic medical record generation system based on multimodal data and proactive perception provided by this invention. Detailed Implementation
[0065] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.
[0066] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0067] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. In addition, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.
[0068] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.
[0069] Example 1
[0070] Please see Figure 1 The present invention provides an embodiment of an automatic medical record generation method based on multimodal data and active perception, comprising the following specific steps:
[0071] Step S1: Collect multimodal raw data from patients and doctors, and construct a field evidence graph.
[0072] The specific steps of step S1 are as follows:
[0073] Step S101: Collect multimodal raw data from patients and doctors, including audio streams, video streams, and medical device data from consultations. Preprocess the multimodal raw data, including time alignment and channel identification marking.
[0074] In this embodiment, audio streams, video streams, and medical device data are generated during the doctor-patient consultation process. Specifically, during the acquisition of the audio stream, valid speech segments are extracted through speech detection, generating continuous speech intervals and writing timestamps. The video stream is segmented into keyframes and frame segments during the acquisition process, with each frame segment having an additional capture time and channel sequence number. The medical device data is based on sampling points as the smallest unit, with each data point accompanied by a device ID, measurement channel, and time label. Subsequently, time alignment is performed on segments from different sources, that is, each segment is mapped to a unified session timeline through a global clock or synchronization signal. After alignment, the segments are then marked according to the channel identifier.
[0075] Step S102: Preset the medical record field set, generate the corresponding field candidate value set based on the preprocessed multimodal raw data, and standardize the unit, naming and time labeling to establish a one-to-one reference relationship between each field candidate value and its source fragment.
[0076] In this embodiment, data fragments of different modalities contain potential descriptions of specific medical record fields. For example, audio fragments can reveal a patient's chief complaint or past medical history, video fragments can present abnormal movements or facial expressions, and device data can directly correspond to vital sign values. By matching these fragments with a set of medical record fields, a preliminary set of candidate values is constructed for each field. After transcription and semantic analysis, audio fragments yield candidate values for the chief complaint field. After action recognition, video fragments yield candidate values for the gait feature field. Device data generates numerical candidate values such as body temperature and heart rate. The units, names, and time annotations of the candidate values are standardized. After standardization, each candidate value establishes a one-to-one reference relationship with its source fragment. This relationship includes not only the text or numerical content of the candidate value but also the source fragment identifier, acquisition channel, and time interval information.
[0077] Step S103: Generate evidence fingerprints for each candidate value of a field and its source fragment. The evidence fingerprints include time interval, channel and device identifier, spatial pose parameters, sampling configuration and fragment summary information.
[0078] In this embodiment, the candidate value only reflects the possible values of a certain field, but cannot explain the specific source and acquisition conditions of the value. By constructing an evidence fingerprint, multi-dimensional source information can be attached to each candidate value. The time interval is used to accurately identify the start and end positions of the corresponding segment on a unified time axis; the channel and device identifiers are used to distinguish different acquisition sources; the spatial pose parameters are used for video and action segments; the sampling configuration records the parameter settings during data acquisition; and the segment summary information is formed by compressing and hashing the segment data to form a verifiable segment summary. It should be noted that the generation process of the evidence fingerprint is not a simple label attachment, but rather a unified binding of multi-dimensional meta-information such as time, channel, space, configuration, and summary to the candidate value object.
[0079] Step S104: Construct a field evidence graph using medical record fields, field candidate values, and evidence fingerprints as graph nodes, and support relationships, rebuttal relationships, chronological relationships, and source relationships between graph nodes as graph edges, and label the edges of the field evidence graph with source type and channel label.
[0080] In this embodiment, medical record field nodes are used as first-level nodes, candidate values corresponding to the fields are used as second-level nodes, and evidence fingerprints associated with the candidate values are used as third-level nodes, forming a basic three-layer graph architecture. In this structure, there are belonging edges between field nodes and candidate value nodes, and there are one-to-one correspondence edges between candidate value nodes and evidence fingerprint nodes. Between candidate value nodes, edges for support and rebuttal relationships are established based on homology, semantic conflict, or data complementarity. At the same time, the temporal relationship between candidate values and evidence fingerprints is also explicitly represented in the graph.
[0081] like Figure 2 As shown, in Figure 2 Solid lines represent chronological relationships, solid arrows represent support relationships, and dashed lines represent rebuttal relationships. First, candidate values A1, B1, and C1 are generated based on fields A, B, and C. Then, candidate value C2 is obtained based on the chronological relationship between the candidate values (represented by solid lines). Subsequently, each candidate value provides support information to evidence fingerprints E1 to E3 via solid arrows, while some candidate values provide rebuttal information to the corresponding evidence fingerprints via dashed lines. This creates an evidence structure where support and rebuttal coexist under the same evidence fingerprint, enabling multi-source candidate value fusion reasoning and evidence determination based on medical record information.
[0082] Step S2: Generate a constraint set based on the clinical ontology and temporal rules, and configure priority and dynamic threshold for each medical record field.
[0083] The specific steps of step S2 are as follows:
[0084] Step S201: Parse and normalize the terms, entities and relations in the clinical ontology library, extract the time-series anchor point set and scene label set related to the medical record field, and generate a constraint candidate library for the medical record field.
[0085] In this embodiment, the clinical ontology consists of a large number of medical terms, disease classifications, symptom characteristics, examination methods, treatment pathways, and their logical relationships. First, the terms, entities, and relationships in the clinical ontology are parsed and normalized to unify different expressions into a standardized naming system and clean up redundant or duplicate entity descriptions. The parsing mechanism identifies entity types in the ontology, such as symptoms, signs, examinations, diagnoses, and treatment measures, and extracts causal relationships, temporal dependencies, and mutual exclusion relationships between these entities. Subsequently, through normalization operations, synonyms, abbreviations, and cross-language expressions are unified into standard medical terms, and temporal anchors are extracted for relationships involving time. At the same time, scene tags are configured for fields based on scene context information. After the above parsing and normalization processes, a constraint candidate library directly related to the medical record fields is generated.
[0086] Step S202: Map the constraint candidate library to the preset medical record field set to form a field-level rule template. The field-level rule template includes mutual exclusion relationship, implication relationship, value range, legal unit of measurement, chronological relationship and necessary field set, and marks the source, version, applicable department and applicable population conditions for each rule.
[0087] In this embodiment, for each preset medical record field, relevant knowledge items are selected from the constraint candidate library and converted into field-level rule templates. First, the medical record field set is traversed item by item, such as chief complaint, past medical history, physical examination, auxiliary examinations, diagnosis, prescription, etc. For each field, the corresponding semantic entities and relations are retrieved from the constraint candidate library, and a rule template is generated. The rule template includes: mutual exclusion relations, such as no history of allergy and penicillin allergy cannot be true at the same time; implication relations, such as the gender should be implied as female when the pregnant woman field is true; value range, such as the body temperature field should be within the medically reasonable range; legal units of measurement, such as blood pressure is expressed in mmHg; sequential relations, such as the examination precedes the diagnosis; and a necessary set of fields, such as the diagnosis field must exist before the prescription is generated.
[0088] Step S203: Based on the medical visit scenario and patient stage information, select the associated subset from the field-level rule template, establish cross-visit vertical association constraints and stage time window constraints, and synchronize the key fields in the previous medical records to the current constraint context to generate a constraint set.
[0089] In this embodiment, the current patient's medical visit scenario and stage label are identified, such as preoperative preparation for emergency visits. Subsequently, relevant constraints are retrieved from the rule template based on these labels. For example, in the emergency scenario, the rule that vital signs must be filled in and the diagnosis field must be before the prescription is retained, while in the follow-up scenario, the continuation of previous diagnoses and medication adherence records are emphasized. At the same time, vertical association constraints across medical visits are established, and key fields such as basic medical history, allergy history, and primary diagnosis are extracted from previous medical records and aligned with the current constraint subset to avoid duplicate collection or logical conflicts.
[0090] It should be noted that during the synchronization of past medical record fields, a constraint context is generated, which includes both rules specific to the current scenario and mandatory constraints that continue vertically.
[0091] Step S204: Assign a hierarchical priority queue to each medical record field, and set a dynamic threshold strategy for the field according to the evidence source type, channel status, patient interaction status and scenario complexity to obtain an adjustable mapping from field to threshold.
[0092] In this embodiment, medical record fields are prioritized according to their clinical sensitivity and importance in the treatment process. Simultaneously, considering the quality differences of different evidence sources and the complexity of patient interaction scenarios, dynamic thresholds are set for each field to achieve adjustable field validation conditions. Specifically, firstly, fields are assigned to priority queues based on clinical risk and key nodes in the treatment pathway. Then, thresholds are initially set according to the type of evidence source; for example, thresholds for numerical fields directly obtained from equipment can be set lower, while fields derived from spoken descriptions can have higher confidence thresholds. Considering channel status and patient interaction, such as when the video channel is obstructed or there is insufficient lighting, the thresholds for corresponding action recognition fields are increased. If the patient demonstrates high cooperation during follow-up, the thresholds for some descriptive fields are lowered. Finally, dynamic adjustments are made based on the complexity of the treatment scenario.
[0093] Step S205: Perform field-level rule conflict adjudication and effective orchestration on the associated subset, determine the adjudication result according to the source trust level, version timestamp and department domain priority order, and output the session-level hard constraint set and soft constraint set.
[0094] The specific steps of step S205 are as follows:
[0095] Step S2051: Aggregate and index the associated subset to generate a conflict candidate set containing mutual exclusion, overriding, overlapping and cross-domain references, and record the source trust level, version timestamp and effective conditions for each field-level rule in the associated subset.
[0096] In this embodiment, associated subsets are merged according to medical record field keys, applicable scenarios, and time-series anchors, aggregating similar or nearly identical rules under a unified index. Subsequently, the aggregated rules are compared item by item to identify the types of relationships that exist between them, such as mutual exclusion (two rules cannot satisfy the same condition simultaneously), coverage (one rule's scope completely includes another rule), overlap (two rules are applicable under certain conditions but have differences), and cross-domain reference (a rule imposes constraints on other fields beyond its own scope). In this way, a conflict candidate set covering multiple conflict modes is generated.
[0097] Step S2052: Generate a conflict resolution sequence based on the source trust level, version timestamp, and effective conditions, and sort the conflict resolution sequence a second time by combining the patient's multimodal raw data to obtain the execution order for each conflict tuple.
[0098] In this embodiment, preliminary priority is assigned based on the rule's metadata: rules with higher source credibility are prioritized over rules from empirical or secondary sources; rules with updated version timestamps are prioritized over previous versions; rules whose effectiveness conditions perfectly match the current medical scenario and patient characteristics are prioritized over rules with ambiguous conditions or broad applicability. Then, a conflict resolution sequence is generated. After generating the preliminary resolution sequence, the patient's multimodal raw data is sorted a second time. The multimodal data includes symptom descriptions extracted from speech, vital signs identified in video, and vital sign parameters recorded by the device. This information reflects the patient's real-time status. By comparing the applicable conditions of the rules with the patient's real-time data, the sorting is dynamically adjusted. For example, if a rule involves abnormal heart rate conditions, and the device data clearly shows that the heart rate is outside the normal range, the rule will be given higher priority in the sorting. Conversely, if a rule requires the patient to be in the postoperative recovery stage, but the multimodal data has not yet detected relevant markers, the rule will be postponed or frozen.
[0099] Step S2053: Adjudicate the conflicting tuples one by one according to the execution order, including the processing of retention, downgrading, freezing and replacement.
[0100] In this embodiment, the conflicts between rules are not single-dimensional, but are manifested as semantic mutual exclusion, scope overlap or temporal contradiction. A layered processing strategy is adopted, which introduces four types of processing methods: retention, downgrading, freezing and replacement. Under the premise of ensuring that the constraint set is executable, the knowledge value is preserved to the maximum extent. The above-mentioned adjudication actions are all completed one by one under the guidance of the execution order to avoid circular adjudication or parallel conflicts.
[0101] Step S2054: Arrange the field-level rules in the associated subset after the adjudication to take effect, mark the triggering stage, triggering conditions, interlocking relationships and coverage boundaries, and output the session-level hard constraint set and soft constraint set.
[0102] In this embodiment, the field-level rules after adjudication are transformed from a static set into executable orchestration logic to ensure that the triggering and effectiveness of rules in the session are sequential and context-dependent. Since different rules may only be applicable at different stages or need to be interlocked with other rules, the rules are hierarchically labeled and executed in an orchestration manner, outputting a set of hard constraints and a set of soft constraints.
[0103] Step S3: Generate additional collection actions based on information gain, and iteratively update the field evidence graph. The iterative update controls the target field to add collection actions until the target field reaches the dynamic threshold or satisfies the constraint set. The target field is a medical record field that does not satisfy the constraint context or is below the dynamic threshold.
[0104] The specific steps of step S3 are as follows:
[0105] Step S301: Retrieve medical record fields that do not meet the constraint set or are below the dynamic threshold from the field evidence graph to form a target field set, and group the target fields according to the priority of the medical record fields and the stage of medical treatment.
[0106] In this embodiment, through the interaction between the field-evidence graph and the constraint set, fields that have not yet met the logical constraints or whose confidence level is lower than the dynamic threshold are automatically identified, and these fields are organized into a target set. Specifically, each medical record field is first retrieved from the field-evidence graph one by one to determine whether its candidate value meets the established hard constraints, soft constraints, and time sequence logic. Fields that do not meet the constraints or lack necessary evidence are marked as incomplete. Subsequently, it is further detected whether the confidence level of the candidate value corresponding to the field is lower than the dynamic threshold set for it. For example, if the noise of the device data is too high or the confidence level of speech recognition is insufficient, such fields are also included in the target set.
[0107] It should be noted that logical constraints include hard constraints, soft constraints, and temporal constraints. Hard constraints refer to deterministic rules that must be strictly met; once violated, the field is deemed invalid. For example, test values must fall within medically permissible physical ranges, such as blood pressure, which cannot be negative. Soft constraints are probabilistic constraints based on statistical laws or empirical models, used to evaluate the rationality of a field. For example, the co-occurrence probability between a combination of symptoms and a specific disease. Temporal constraints describe the rationality of a field's evolution over time. For example, admission time must be earlier than discharge time. Dynamic thresholds are not fixed constants but are dynamically updated based on the following factors: data source quality (increasing the threshold when noise is high to enhance reliability); field importance level (critical fields such as diagnosis results correspond to higher thresholds); thresholds can be appropriately lowered during the consultation and initial diagnosis stages to improve recall, while thresholds are increased during the review stage to ensure accuracy; historical statistical distribution, etc.
[0108] Step S302: For each target field, generate a corresponding set of additional collection action candidates from the preset action library, and label each action with the expected requirements.
[0109] In this embodiment, each field corresponds to several potential data sources. When the existing evidence is insufficient to meet the threshold requirements, it is necessary to actively design new collection methods. Through the mapping relationship between the action library and the field, several optional collection paths can be quickly generated for the target field. At the same time, the expected input or conditions required for the action are marked during the generation. The action library usually has multiple types of operation preset, such as guided questioning, body position or action instructions, equipment retesting, channel switching or external data calling. During the generation of the action candidate set, each action needs to be marked with the corresponding expected requirement information.
[0110] Step S303: Evaluate and sort the action candidate set according to the information gain strategy, generate an additional collection action execution sequence, split actions with dependencies, and merge actions that can be parallelized.
[0111] In this embodiment, the utility evaluation of the candidate set of additional acquisition actions generated for the target field is performed, and priority ranking is performed based on the information gain strategy to form an ordered execution sequence. For each candidate action, the information gain is evaluated, and the relative utility is calculated based on the evidence gaps it can cover, the modal resources required, and the possible confidence improvement. The actions are ranked according to the evaluation results to generate a preliminary execution sequence. The dependencies and parallelism between actions are further processed. If an action must depend on the result of another action, the dependent action is split to ensure sequential execution. If multiple actions can be performed simultaneously on different channels, they are merged into parallel execution units.
[0112] Step S304: Issue additional acquisition actions according to the execution order to obtain new evidence, perform time base alignment, channel consistency verification and fragment deduplication on the new evidence, and generate evidence identifier and evidence fingerprint.
[0113] The specific steps of step S304 are as follows:
[0114] Step S3041: Issue additional acquisition actions to each acquisition channel according to the execution order, encapsulate the additional acquisition action information into control instructions containing source type, channel label and sampling window, and register a tracking number for each control instruction.
[0115] In this embodiment, actions are parsed one by one according to the execution order, extracting their source type (e.g., voice, video, medical device), channel labels (e.g., microphone, camera, monitor number), and sampling window (e.g., sampling duration, frame rate, or measurement interval). These parameters together constitute the core content of the instruction. Subsequently, the instructions are uniformly encapsulated into a format that can be directly recognized by each acquisition channel, and a tracking number is registered for each instruction. It should be noted that the tracking number and channel label work together to ensure that in a scenario of multi-channel parallel acquisition, the source and execution order of different actions can be accurately distinguished.
[0116] Step S3042: Receive additional collected multimodal raw data, segment it according to the sampling window, generate corresponding evidence fragments, and attach an initial timestamp and channel label to each evidence fragment.
[0117] In this embodiment, the received audio stream, video stream, and device data stream are first segmented according to the sampling window parameters set in the issued instruction. The audio stream is sliced into audio segments according to a fixed duration, the video stream is extracted into key frames according to the frame rate and time window, and the device data is formed into discrete measurement segments according to the numerical sampling interval. Each segment is generated with an initial timestamp attached to identify the starting point of the segment in the unified session timeline. In addition to the timestamp, a channel label is attached to each segment to distinguish its source, such as voice channel - microphone a, video channel - camera b, and device channel - monitor c. The combination of timestamp and channel label enables the evidence segment to achieve unique positioning and multimodal cross-indexing when updating the field - evidence graph in the subsequent process.
[0118] Step S3043: Perform time alignment and channel consistency verification on the evidence fragments, transfer fragments that fail the verification to the isolation list, and transfer fragments that pass the verification to the acceptance list.
[0119] In this embodiment, the initial timestamps attached to each evidence fragment are globally calibrated. By comparing the sampling window boundary with the session timeline, fragments with drift or overlap are corrected. After time alignment is completed, channel consistency verification is performed, which checks whether the fragment source matches the channel label of the issued command, whether the acquisition device is within the preset state range, and whether the logical correspondence between different modalities is satisfied. The output results of the verification process are divided into two categories: fragments that pass the verification are included in the acceptance list and become valid inputs for subsequent evidence graph updates; fragments that fail the verification are transferred to the isolation list, and the reason for the anomaly and the corresponding command number are recorded in the list.
[0120] Step S3044: Deduplicate the fragments in the acceptance list to generate a unique evidence identifier, and generate an evidence fingerprint based on the initial timestamp, channel label and evidence identifier.
[0121] In this embodiment, the fragments in the acceptance list are deduplicated. This process includes comparing whether the time intervals of the fragments highly overlap, whether the channel labels are consistent, and whether the fragment summary information is similar, thereby identifying and eliminating duplicate or highly similar fragments. For the fragments that are confirmed to be retained, a unique evidence identifier is assigned to them. This identifier is independent of the tracking number. After the unique evidence identifier is generated, the initial timestamp, channel label, and evidence identifier of the fragment need to be further combined to form an evidence fingerprint. The evidence fingerprint not only preserves the time location and channel source of the fragment, but also ensures reusability and verifiability through the uniqueness of the identifier.
[0122] Step S305: Connect the newly added evidence to the field evidence graph, update the supporting relationship, rebuttal relationship, sequential relationship and source relationship, and calculate the target field set and action candidate set based on the updated field evidence graph until the constraint set is satisfied or the dynamic threshold is exceeded.
[0123] In this embodiment, newly added evidence is first attached to the corresponding field candidate value node, and evidence identifiers and evidence fingerprints are generated as node-attached information. Subsequently, based on the logical relationship between the newly added evidence and existing candidate values, the supporting relationships, rebuttal relationships, chronological relationships, and source relationships are updated in the graph. After the update is completed, the target field set and action candidate set need to be recalculated based on the new field evidence graph. This process includes re-searching for fields that still do not meet the constraint set or are below the dynamic threshold, and adjusting the priority and action requirements according to the updated evidence information. If the target field set is empty or the confidence of all field candidate values exceeds the dynamic threshold, the iteration process terminates. If there are still gaps, additional collection actions are generated, and a new round of iteration begins. Through this cyclical update mechanism, the medical record generation process is ensured to continuously approach the requirements of completeness and consistency.
[0124] Step S4: Perform constraint decoding on the medical record fields in the iteratively updated field evidence graph to obtain the values of each target field, and output the medical record with the attached evidence fingerprint chain.
[0125] The specific steps of step S4 are as follows:
[0126] Step S401: Based on the iteratively updated field evidence graph, lock the hard and soft constraint versions of the current session, and generate a set of medical record fields to be decoded and a list of constraint references based on the medical record fields and time series anchors.
[0127] In this embodiment, it is first necessary to lock the versions of hard and soft constraints applicable in the current session. When multiple constraint libraries coexist, filtering is required based on session identifier, version number, and applicable scenario to ensure that the constraints called in subsequent decoding are the currently valid versions. Subsequently, based on the latest topology of the field-evidence graph, medical record fields with undetermined target values are filtered, and combined with time-series anchor information, these fields are organized into a set to be decoded. While forming the set to be decoded, a constraint reference list is generated for each field in the set. This list contains hard constraint entries, soft constraint entries, time-series anchors, and cross-field dependencies associated with the field.
[0128] Step S402: Perform consistency preprocessing on the candidate values of the medical record fields to be decoded to form a decodeable candidate set.
[0129] In this embodiment, the preprocessing includes semantic, numerical, and temporal unification, and the candidate values after unification are organized into a decodable candidate set.
[0130] Step S403: Construct a decoding priority sequence based on the constraint reference list and the decodable candidate set, and set the conflict handling order.
[0131] In this embodiment, based on the temporal anchors and interlocking relationships recorded in the constraint reference list, the fields to be decoded are initially sorted. For example, if a diagnostic field can only be decoded after a test field, the test field is prioritized in the sequence. Secondly, considering the confidence level of candidate values, the fields within the same priority are further refined, prioritizing the decoding of high-confidence candidate values to reduce unnecessary conflicts. While generating the decoding priority sequence, a conflict handling order must also be set. When multiple candidate values form a mutually exclusive relationship within the same field, the decision is made first based on the principle that hard constraints take precedence over soft constraints. When constraints are of the same level but from different sources, they are sorted according to the source's credibility level and version timestamp. If they still cannot be distinguished, the strength of evidence and temporal consistency of the candidate values are used as the final criterion.
[0132] Step S404: Perform constrained decoding on the decodable candidate set based on the decoding priority sequence, select target candidate values for each medical record field under the premise of satisfying hard constraints, and set freeze flags and re-entry conditions for fields that cannot be directly parsed.
[0133] In this embodiment, starting from the beginning of the decoding priority sequence, the constraint reference list is called for each field to be decoded, and the degree of conformity between the candidate value and the hard constraint is compared item by item. If the candidate value satisfies all hard constraint conditions, it is marked as the target candidate value, and the field is marked as decoded. If the candidate value can only satisfy the soft constraint, or there is a partial conflict, the optimal solution is selected by confidence ranking, and the correlation with the conflict rule is recorded. For fields where no candidate value can be found that satisfies the hard constraint, or where the evidence is insufficient to complete the parsing, a freeze mark is set, and re-entry conditions are registered, such as the need to supplement device readings or wait for multimodal data at subsequent time points. The re-entry conditions of frozen fields not only record the gaps at the data level, but also include a trigger mechanism, such as specifying that the field will re-enter the decoding sequence after the next round of active collection. In this way, constraint decoding can ensure the logical correctness of the completed fields and provide an interface for subsequent correction of the incomplete fields. The final result is a set of medical record fields that are partially decoded and partially frozen for processing.
[0134] Step S405: Generate a medical record field mapping table for the target candidate value, establish a pointer chain between field values and corresponding evidence fragments, and output a medical record with an attached evidence fingerprint chain by combining the medical record field mapping table and the evidence fingerprint chain.
[0135] The specific steps of step S405 are as follows:
[0136] Step S4051: Establish a medical record field mapping table according to the field identifier of the target candidate value, register each medical record field corresponding to the selected target candidate value, candidate source and constraint reference number, and assign a unique mapping sequence number to each entry.
[0137] In this embodiment, an index column of the mapping table is established based on the unique identifier of the medical record field, such as the field code or field name. The confirmed target candidate value is extracted from the decoding result and registered in correspondence with the field identifier. At the same time, the source information of the candidate value is recorded, including the corresponding evidence fingerprint number and acquisition channel type, to ensure that the original fragment can be traced when the result is called. Furthermore, the constraint reference number called by the candidate value during the decoding process is synchronously written into the table entry to reflect the correspondence between the field result and the rule condition. After registration, a unique mapping sequence number is assigned to each entry in the mapping table. This sequence number is independent of the field identifier and evidence number and is used to provide an index in subsequent medical record output and evidence chain generation.
[0138] Step S4052: Retrieve evidence fragments associated with the target candidate value, and concatenate them according to the mapping sequence number of the medical record field mapping table to form an evidence chain structure arranged in the order of medical record fields.
[0139] In this embodiment, the assigned mapping numbers are first read one by one according to the medical record field mapping table, using the natural order of the fields as the framework for concatenation. Then, evidence fragments associated with the target candidate value are retrieved according to the mapping numbers. These fragments come from audio dialogues, video clips, or device data. During the retrieval, it is necessary not only to confirm the uniqueness of the evidence identifier, but also to call its attached timestamp and channel tag to ensure the accuracy of the fragment order. After the retrieval is completed, the fragments are connected sequentially according to the mapping number order to generate an evidence chain covering all fields. It should be noted that the generated evidence chain is not just a simple splicing of fragments, but an ordered multimodal index sequence. Each link in the chain is composed of a ternary relationship of field identifier-target candidate value-evidence fragment, so that any field value in the medical record can be traced back to the corresponding original evidence fragment, and its channel source and constraint reference number can be further tracked.
[0140] Step S4053: Add evidence fingerprints to the evidence fragments in the evidence chain structure. The evidence fingerprints include the initial timestamp and channel parameters of the evidence fragments. Merge the evidence fingerprints with the evidence fragments to generate an evidence fingerprint chain.
[0141] In this embodiment, the initial timestamp and channel parameters are extracted from each evidence fragment in the evidence chain. The initial timestamp is used to determine the position of the fragment in the global timeline of the session, and the channel parameters are used to identify the physical channel and acquisition device information from which it originates. Subsequently, these metadata and fragment contents are merged to generate an evidence fingerprint, so that each fragment in the chain is no longer just a data body, but an evidence unit with an identity mark. After the evidence fingerprint is generated, all fragments with fingerprints are reconnected to form a complete evidence fingerprint chain. This chain not only maintains the sequential relationship of medical record fields, but also provides a traceable identifier for each node.
[0142] Step S4054: Combine the medical record field mapping table with the evidence fingerprint chain to form a medical record with an attached evidence fingerprint chain.
[0143] In this embodiment, the medical record field mapping table is first used as a framework to traverse the field entries one by one. For each entry, its unique mapping number is read and the corresponding evidence node is retrieved in the evidence fingerprint chain. Then, the field value, source information and evidence fingerprint are integrated to form a multi-binding of field-value-evidence fingerprint in the medical record text or structured record.
[0144] Figure 3The diagram shown illustrates the hierarchical structure of a medical record generated automatically using a multimodal data and proactive sensing method. This structure comprises a medical record field layer, a candidate value layer, evidence fragments, and evidence fingerprints, illustrating the mapping relationship between medical record results and evidence. The final output medical record not only contains complete treatment records but also carries a chain of evidence, ensuring the content can be traced back to the original data and improving the authenticity and compliance of the medical record. It should be noted that... Figure 3 A simplified version of the medical record is provided.
[0145] Example 2
[0146] Please see Figure 4 Another embodiment of the present invention provides: an automatic medical record generation system based on multimodal data and active perception, comprising: a graph construction module, a constraint configuration module, an iterative update module, and a medical record generation module;
[0147] The graph construction module is used to collect multimodal raw data from patients and doctors and construct field evidence graphs.
[0148] The constraint configuration module is used to generate a constraint set based on the clinical ontology and time sequence rules, and to configure priority and dynamic threshold for each medical record field.
[0149] The iterative update module generates additional acquisition actions based on information gain and iteratively updates the field evidence graph.
[0150] The medical record generation module is used to perform constraint decoding on the medical record fields in the iteratively updated field evidence graph to obtain the values of each target field and output the medical record with an attached evidence fingerprint chain.
[0151] The medical record generation module includes: a decodable candidate set generation unit, a constraint decoding unit, and a medical record generation unit;
[0152] The decodable candidate set generation unit generates a set of medical record fields to be decoded and a list of constraint references based on the iteratively updated field evidence graph, according to the medical record fields and time-series anchors, and performs consistency preprocessing on the candidate values of the medical record fields to be decoded to form a decodable candidate set.
[0153] The constraint decoding unit constructs a decoding priority sequence based on the constraint reference list and the decodable candidate set, performs constraint decoding on the decodable candidate set based on the decoding priority sequence, selects target candidate values for each medical record field under the premise of satisfying hard constraints, and sets freeze flags and re-entry conditions for fields that cannot be directly parsed.
[0154] The medical record generation unit is used to generate a medical record field mapping table for the target candidate value, establish a pointer chain between field values and corresponding evidence fragments, and output a medical record with an attached evidence fingerprint chain by combining the medical record field mapping table and the evidence fingerprint chain.
[0155] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.
[0156] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for automatically generating medical records based on multimodal data and active perception, characterized in that: include: Collect multimodal raw data from patients and doctors to construct field evidence graphs; A constraint set is generated based on the clinical ontology and temporal rules, and priority and dynamic thresholds are configured for each medical record field. Additional data collection actions are generated based on information gain, and the field evidence graph is iteratively updated. The iterative update controls the target field to perform additional data collection actions until the target field satisfies the constraint set or the confidence level reaches the dynamic threshold. The target field is a medical record field that does not satisfy the constraint set or has a confidence level lower than the dynamic threshold. Constraint decoding is performed on the medical record fields in the iteratively updated field evidence graph to obtain the values of each target field, and the medical record with an attached evidence fingerprint chain is output. The process of collecting multimodal raw data from patients and doctors to construct a field evidence graph includes: Collect multimodal raw data from patients and doctors, including audio streams, video streams, and medical device data from consultations; preprocess the multimodal raw data, including time alignment and channel identification marking. A pre-defined set of medical record fields is used to generate a set of corresponding candidate values for the fields based on the pre-processed multimodal raw data. The units, names, and time annotations of the candidate values are standardized so that each candidate value for a field establishes a one-to-one reference relationship with its source fragment. For each field candidate value and its source fragment, an evidence fingerprint is generated. The evidence fingerprint includes time interval, channel and device identifier, spatial pose parameters, sampling configuration and fragment summary information. Using medical record fields, field candidate values, and evidence fingerprints as graph nodes, and the support relationships, rebuttal relationships, chronological relationships, and source relationships between graph nodes as graph edges, a field evidence graph is constructed, and the edges of the field evidence graph are labeled with source type and channel label. The process of generating a constraint set based on clinical ontology and temporal rules, and configuring priorities and dynamic thresholds for each medical record field, includes: The terms, entities, and relationships in the clinical ontology are parsed and normalized to extract a set of time-series anchor points related to medical record fields. At the same time, a set of scene labels is configured for medical record fields to generate a candidate library of constraints for medical record fields. In particular, through normalization operations, synonyms, abbreviations, and cross-language expressions are unified into standard medical terms, and time-series anchor points are extracted from relationships involving time. For each preset medical record field in the preset medical record field set, the corresponding semantic entities and relations are retrieved from the constraint candidate library to form a field-level rule template. The field-level rule template includes mutual exclusion relations, implication relations, value range, legal units of measurement, chronological relations and necessary field sets, and the source, version, applicable department and applicable population conditions are marked for each rule. Based on the medical visit scenario and patient stage information, a related subset is selected from the field-level rule template to establish cross-visit vertical association constraints and stage-based time window constraints, and key fields in historical medical records are synchronized to the current constraint context to generate a constraint set. A hierarchical priority queue is assigned to each medical record field, and a dynamic threshold strategy is set for the medical record field according to the evidence source type, channel status, patient interaction status and scenario complexity to obtain an adjustable mapping from medical record field to threshold; wherein, the dynamic threshold is used to realize the adjustability of field validation conditions. The associated subset is subject to field-level rule conflict adjudication and effective arrangement. The adjudication results are determined according to the source trust level, version timestamp, and department domain priority, and the session-level hard constraint set and soft constraint set are output. The process involves adjudicating and orchestrating field-level rule conflicts in the associated subset, determining the adjudication results based on source trust level, version timestamp, and departmental domain priority, and outputting a session-level set of hard constraints and a set of soft constraints, including: The associated subsets are aggregated, and similar or nearly identical rules are grouped under a unified index. Then, the aggregated rules are compared item by item to identify the types of relationships between the rules, generating a conflict candidate set that includes mutual exclusion, overriding, overlapping, and cross-domain reference relationships. For each field-level rule in the associated subset, the source trust level, version timestamp, and effective conditions are recorded. Among them, cross-domain reference relationship refers to a rule imposing constraints on other fields that are outside the scope of this field. A conflict resolution sequence is generated based on the source trust level, version timestamp, and effective conditions. The conflict resolution sequence is then sorted again using the patient's multimodal raw data to obtain the execution order for each conflict tuple. The conflicting tuples are adjudicated one by one according to the execution order described, including four processing methods: retention, downgrading, freezing, and replacement. The rules at the field level in the associated subset after the adjudication are orchestrated for effectiveness, and the triggering stage, triggering conditions, interlocking relationships, and coverage boundaries are marked. Session-level hard constraint sets and soft constraint sets are output. Interlocking relationships refer to rules that need to be used in an interlocking manner with other rules. Hard constraints are deterministic rules that must be strictly satisfied, and once violated, the field is deemed invalid. Soft constraints are probabilistic constraints based on statistical laws or empirical models, used to evaluate the rationality of fields. Additional data collection actions are generated based on information gain, and the field evidence graph is iteratively updated, including: Retrieve medical record fields from the field evidence graph that do not meet the constraint set or whose confidence is lower than the dynamic threshold to form a target field set, and group the target fields according to the priority of medical record fields and the stage of medical treatment; For each target field, a corresponding set of additional collection actions is generated from the preset action library, and the expected requirements are labeled for each action. The action candidate set is evaluated and sorted according to the information gain strategy to generate an additional collection action execution sequence, and actions with dependencies are split and actions that can be parallelized are merged; wherein, when evaluating the information gain of each candidate action, the relative utility is calculated based on the evidence gap that the candidate action can cover, the required modal resources, and the corresponding confidence improvement. Additional collection actions are issued according to the execution sequence to obtain new evidence. The new evidence is then aligned with the time base, verified for channel consistency, and deduplicated for fragments to generate evidence identifiers and evidence fingerprints. The newly added evidence is attached to the field evidence graph, and the supporting relationships, rebuttal relationships, sequential relationships and source relationships are updated. The target field set and action candidate set are calculated based on the updated field evidence graph until the constraint set is satisfied or the confidence level is higher than the dynamic threshold.
2. The automatic medical record generation method based on multimodal data and active perception as described in claim 1, characterized in that, The process of issuing additional data collection actions according to the execution sequence to acquire new evidence, performing time base alignment, channel consistency verification, and fragment deduplication on the new evidence, and generating evidence identifiers and evidence fingerprints includes: According to the execution sequence, additional acquisition actions are issued to each acquisition channel. The additional acquisition action information is encapsulated into control instructions containing source type, channel label and sampling window, and a tracking number is registered for each control instruction. Receive additional collected multimodal raw data, segment it according to the sampling window, generate corresponding evidence fragments, and attach an initial timestamp and channel label to each evidence fragment; The evidence fragments are time-aligned and channel-consistent. Evidence fragments that fail the verification are transferred to the isolation list, and evidence fragments that pass the verification are transferred to the acceptance list. The evidence fragments in the accepted list are deduplicated to generate a unique evidence identifier, and an evidence fingerprint is generated based on the initial timestamp, channel label and evidence identifier.
3. The automatic medical record generation method based on multimodal data and active perception as described in claim 2, characterized in that, The process of performing constraint decoding on the medical record fields in the iteratively updated field evidence graph to obtain the values of each target field and outputting medical records with attached evidence fingerprint chains includes: Based on the iteratively updated field evidence graph, the hard and soft constraint versions of the current session are locked. Based on the latest topology of the field evidence graph, the medical record fields whose target values have not yet been determined are filtered, and combined with the time-series anchor information, these fields are organized into a set of medical record fields to be decoded. At the same time, a constraint reference list is generated for each medical record field in the set of medical record fields to be decoded. The constraint reference list includes the hard constraint entries, soft constraint entries, time-series anchors, and cross-field dependencies associated with the medical record field. The candidate values of the medical record fields to be decoded are preprocessed to achieve consistency, forming a set of decodable candidates; Based on the constraint reference list and the decodable candidate set, a decoding priority sequence is constructed, and a conflict handling order is set. Starting from the beginning of the decoding priority sequence, for each field of the medical record to be decoded, its constraint reference list is called, and the degree of conformity between the candidate value and the hard constraint is compared item by item. If the candidate value satisfies all hard constraint conditions, it is marked as the target candidate value; if the candidate value can only satisfy the soft constraint conditions, or there are some conflicts, the optimal solution is selected by confidence ranking; for fields that cannot find candidate values that satisfy the hard constraint conditions, or for which the evidence is insufficient to complete the parsing, a freeze mark is set, and the re-entry conditions are registered at the same time. Establish a medical record field mapping table according to the field identifier of the target candidate value, register each medical record field with the selected target candidate value, candidate source and constraint reference number, and assign a unique mapping sequence number to each entry; Retrieve evidence fragments associated with the target candidate value, and concatenate them according to the mapping sequence number of the medical record field mapping table to form an evidence chain structure arranged in the order of medical record fields; Add evidence fingerprints to the evidence fragments in the evidence chain structure. The evidence fingerprints include the initial timestamp and channel parameters of the evidence fragments. Merge the evidence fingerprints with the evidence fragments to generate an evidence fingerprint chain. The medical record field mapping table is combined with the evidence fingerprint chain to form a medical record with an attached evidence fingerprint chain.
4. A medical record automatic generation system based on multimodal data and active perception, used to implement the medical record automatic generation method based on multimodal data and active perception as described in any one of claims 1-3, characterized in that, The system includes: a graph construction module, a constraint configuration module, an iterative update module, and a medical record generation module; The graph construction module is used to collect multimodal raw data from patients and doctors and construct field evidence graphs. The constraint configuration module is used to generate a constraint set based on the clinical ontology and time sequence rules, and to configure priority and dynamic threshold for each medical record field. The iterative update module is used to generate additional acquisition actions based on information gain and iteratively update the field evidence graph; The medical record generation module is used to perform constraint decoding on the medical record fields in the iteratively updated field evidence graph to obtain the values of each target field and output the medical record with an attached evidence fingerprint chain.
Citation Information
Patent Citations
Method and system for intelligently generating chronic disease follow-up table based on multi-modal data fusion
CN120783927A
Patient information collection and medical record construction system and method based on multiple rounds of dialogues
CN120998388A