Cardiac ultrasound report automatic generation and auditing method and device

By performing structured analysis and standardized template matching on cardiac ultrasound examination data, combined with natural language generation models and diagnostic rule comparisons, efficient, professional, and consistent generation of cardiac ultrasound reports has been achieved, solving the problems of long report processing time and poor consistency in existing technologies.

CN122067697AInactive Publication Date: 2026-05-19THE FIRST AFFILIATED HOSPITAL OF FUJIAN MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE FIRST AFFILIATED HOSPITAL OF FUJIAN MEDICAL UNIV
Filing Date
2026-04-21
Publication Date
2026-05-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current methods for writing cardiac ultrasound reports rely on manual processes, which are time-consuming and lack consistency. They are difficult to generate high-quality reports in cases of multiple lesions or high diagnostic uncertainty. Furthermore, existing methods lack systematic rule verification mechanisms, making it difficult to guarantee the professional quality and consistency of reports.

Method used

By performing structured analysis of cardiac measurement parameters and physician examination records, an ultrasound examination dataset is constructed. Combined with standardized report templates and case complexity assessments, deviation annotations are generated. Then, by comparing natural language generation models and diagnostic rules, human-machine collaborative review is achieved, resulting in a logically coherent cardiac ultrasound diagnostic report.

Benefits of technology

It improves the efficiency and quality of report generation, ensuring the professionalism and consistency of reports. Through parameter correlation parsing and diagnostic rule verification, it enhances the logical coherence and content reliability of reports, and reduces the time and inconsistencies of manual writing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067697A_ABST
    Figure CN122067697A_ABST
Patent Text Reader

Abstract

The invention discloses a cardiac ultrasound report automatic generation and auditing method and device, and the method comprises the steps: carrying out the structural joint analysis of cardiac measurement parameters and doctor examination records, processing a cross-equipment parameter naming difference and free text semantic alignment problem, and constructing an ultrasound examination data set; in combination with report specification template matching and case complexity evaluation, differential analysis is carried out on the parameter deviation amplitude to generate a deviation mark; fusing the description feature set and examination-measurement semantic deviation information, dynamically adjusting detailed weights of chapters and sections by directional clinical prompt identification, and guiding a natural language generation model to generate a report draft; a revision sequence is generated through diagnosis rule comparison and rare sign recognition, a credible expression identifier is extracted through man-machine collaborative auditing and is integrated with auditing feedback, a cardiac ultrasound diagnosis report with a complete structure and coherent logic is output, and the standardization and auditing efficiency of cardiac ultrasound report generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information processing technology, and in particular to a method and apparatus for automatically generating and reviewing cardiac ultrasound reports. Background Technology

[0002] Echocardiography is an important imaging tool for assessing cardiac structure and function. After the examination, physicians need to synthesize the various quantitative measurement results output by the equipment with their subjective observations during the examination into a standardized diagnostic report. Currently, report writing mainly relies on manual completion. When the examination involves multiple structures or when there are discrepancies or contradictions in the measurement results, report writing becomes more time-consuming. Differences in the expression styles and diagnostic emphases among different physicians also lead to a lack of consistency in reports for similar cases, affecting the reference efficiency of subsequent clinical decisions.

[0003] Existing report generation methods are mostly based on fixed template filling strategies, which are applicable to routine examinations with a single lesion type. However, when faced with multiple lesions or cases with high diagnostic uncertainty, the template adaptation ability is limited, and semantic discrepancies easily arise between the generated content and the actual examination findings. In addition, existing methods generally lack a systematic rule verification mechanism for the generated content and fail to effectively integrate human review opinions with model generation results, making it difficult to establish a collaborative closed loop between generation and review while ensuring the professional quality of the report. Summary of the Invention

[0004] This invention discloses a method and apparatus for automatically generating and reviewing cardiac ultrasound reports. The method constructs an ultrasound examination dataset by performing structured joint analysis of measurement parameters and examination records. It generates deviation annotations by combining report template matching and case complexity assessment. A natural language generation model is driven by fusing descriptive feature sets and clinical prompts to generate a draft report. A revision sequence is generated through diagnostic rule comparison and rare sign identification. Credible expression identifiers are extracted through human-machine collaborative review, and review feedback is integrated to finally generate a structurally complete and logically coherent cardiac ultrasound diagnostic report.

[0005] The first aspect of this invention proposes a method for automatically generating and reviewing cardiac ultrasound reports, comprising the following steps:

[0006] Obtain cardiac measurement parameters and physician examination records, and construct an ultrasound examination dataset based on the cardiac measurement parameters and physician examination records through structured parsing;

[0007] The ultrasound examination dataset is matched with the report standard template to generate a report framework and low matching degree intervals. Based on the low matching degree intervals, the case complexity identifier is extracted. According to the report framework and the case complexity identifier, the deviation of the parameter reference interval is analyzed to generate deviation labels.

[0008] Based on the deviation annotation, parameter feature association parsing is performed to obtain a descriptive feature set. The ultrasound examination dataset is then subjected to diagnostic-measurement semantic deviation feature extraction to generate clinical prompt labels. Based on the descriptive feature set and the clinical prompt labels, a report draft is generated using a natural language generation model.

[0009] Extract the key indicator set from the draft report, perform diagnostic rule comparison based on the key indicator set to generate verification results, extract rare sign identifiers for parameter combinations not covered by the rule base in the key indicator set, and generate a revision sequence based on the priority ranking of the verification results and the rare sign identifiers.

[0010] Based on the revised sequence, human-machine collaborative review is triggered to generate review feedback. Features of unrevised paragraphs are extracted from the review feedback to generate a credible expression identifier. The content is then integrated with the credible expression identifier and the review feedback to generate a cardiac ultrasound diagnostic report.

[0011] A second aspect of this invention provides an automatic generation and review device for cardiac ultrasound reports, comprising:

[0012] The data acquisition module is used to acquire cardiac measurement parameters and physician examination records, and to construct an ultrasound examination dataset based on the cardiac measurement parameters and physician examination records through structured parsing.

[0013] The framework generation module is used to perform report specification template matching on the ultrasound examination dataset to generate a report framework and low matching degree intervals, extract case complexity identifiers based on the low matching degree intervals, and perform parameter reference interval deviation magnitude analysis to generate deviation labels according to the report framework and the case complexity identifiers.

[0014] The draft generation module is used to obtain a descriptive feature set by performing parameter feature association analysis based on the deviation annotation, extract diagnostic-measurement semantic deviation features from the ultrasound examination dataset to generate clinical prompt labels, and generate a report draft based on the descriptive feature set and the clinical prompt labels through a natural language generation model.

[0015] The rule verification module is used to extract the key indicator set of the report draft, perform diagnostic rule comparison based on the key indicator set to generate verification results, extract rare sign identifiers for parameter combinations not covered by the rule base in the key indicator set, and generate a revision sequence by prioritizing the verification results and the rare sign identifiers.

[0016] The report output module is used to trigger human-machine collaborative review based on the revision sequence to generate review feedback, extract unrevised paragraph features from the review feedback to generate a credible expression identifier, and integrate the content with the credible expression identifier and the review feedback to generate a cardiac ultrasound diagnostic report.

[0017] The beneficial effects of this invention are reflected in the following points: First, it performs structured joint analysis of cardiac measurement parameters and physician examination records, establishing a cross-device parameter classification and free text semantic alignment mechanism, thus addressing issues of source differences and non-standard expression during the data construction stage. Combined with report template matching and case complexity assessment, it adopts differentiated parameter deviation analysis strategies for examinations of varying complexity, improving the quality of data conversion from examination data to report generation input. Second, through parameter feature association analysis, it identifies parameter groups with intrinsically related deviation magnitudes as joint description units. Combined with examination-measurement semantic deviation feature extraction, it makes the information difference between measurement data and physician subjective descriptions explicit. It dynamically adjusts the detail weight of each chapter using two directional labels: aggravation tendency and mitigation tendency, ensuring that the generated draft report maintains consistency with actual examination findings in terms of parameter logic and clinical risk emphasis. Finally, the draft report was systematically quality-checked by comparing diagnostic rules and identifying rare signs. The degree of rule violation and the level of signs were combined to generate a revision sequence, guiding review resources to focus on report segments with higher quality risks. Credible expression identifiers were extracted by combining revision behavior analysis to distinguish between actively retained content and overlapping paragraphs in the system deviation area. Differentiated handling strategies were adopted for paragraphs with different credibility levels during the content integration stage, ensuring the overall quality of cardiac ultrasound diagnostic reports in both rule compliance and content reliability. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the automatic generation and review method for cardiac ultrasound reports according to the present invention.

[0019] Figure 2 This is a structural block diagram of an automatic cardiac ultrasound report generation and review device according to the present invention. Detailed Implementation

[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0021] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0022] The technical solutions of the embodiments of this application will be described below.

[0023] like Figure 1 As shown, this embodiment of the invention provides a method for automatically generating and reviewing cardiac ultrasound reports, including the following steps S11-S15:

[0024] Step S11: Obtain cardiac measurement parameters and physician examination records, and construct an ultrasound examination dataset based on structured analysis of cardiac measurement parameters and physician examination records.

[0025] Specifically, cardiac measurement parameters and physician examination records are acquired. Cardiac measurement parameters are output by the ultrasound equipment during the examination according to the examination protocol. Different examination devices may have different naming conventions for the same measurement dimensions. During the acquisition phase, the device model and protocol version information must be recorded simultaneously to ensure the comparability of cardiac measurement parameters across different devices. During the measurement process, some measurement items may not be collected due to insufficient image quality or limited patient cooperation. These items are included in the acquired results with a missing measurement marker instead of a numerical value in the cardiac measurement parameters to avoid silently discarding missing information and distorting the completeness assessment. Physician examination records are entered by the physician in real time during the examination. The entry method varies depending on the level of informatization in medical institutions, including both text input and speech transcription. Speech transcription of physician examination records may introduce issues such as homophone substitution or truncation of professional terminology during recognition. During the acquisition phase, source type annotations are added to speech-based records to strengthen verification of such sources. The language used in physician examination records is relatively free, with significant differences in individual expression habits. Some physicians habitually use custom grading labels such as "MR mild-moderate" to describe the degree of mitral regurgitation. There is a direct semantic gap between this type of expression and the corresponding standardized values ​​in cardiac measurement parameters. The size of the semantic gap is positively correlated with the degree of difference in physician expression style. Both cardiac measurement parameters and physician examination records carry examination timestamps. When the difference between the two timestamps exceeds a set threshold, source verification is triggered to ensure that the two data streams belong to the same complete examination event and to avoid mixing data from different examination batches.

[0026] In some embodiments, the step of constructing an ultrasound examination dataset based on the structured parsing of the cardiac measurement parameters and the physician's examination record includes: classifying and organizing the cardiac measurement parameters according to measurement type to generate a parameter classification set; performing semantic mapping between the parameter classification set and the physician's examination record to generate an alignment annotation set; extracting unstructured features from the parsing-failed fields in the alignment annotation set to generate a parsing residual itemset; and performing reverse tracing and completion of the residual fields based on the parsing residual itemset to construct the ultrasound examination dataset.

[0027] Cardiac measurement parameters were categorized and organized according to measurement type to generate parameter classification sets. The categorization followed the standard division of measurement dimensions in echocardiography guidelines, classifying cardiac measurement parameters into structural parameters, functional parameters, and hemodynamic parameters groups. The structural parameters group included measurements such as heart chamber diameter, ventricular wall thickness, and valve morphology; the functional parameters group included measurements such as ejection fraction, fractional shortening, and diastolic function assessment; and the hemodynamic parameters group included measurements such as blood flow velocity at each valve orifice, differential pressure, and pulmonary artery pressure estimation. The physical dimensions of parameters within each group were relatively consistent with normal reference intervals. Due to different naming conventions of different manufacturers' equipment, the same measurement dimension of cardiac measurement parameters may have multiple labeling forms. For example, left ventricular ejection fraction is labeled as EF on some devices and LVEF on others; left ventricular end-diastolic diameter is labeled as LVDd on some devices and LVEDD on others. When generating the parameter classification sets, these entries were uniformly mapped to standard parameter names before grouping to prevent duplicate entries of the same measurement dimension due to name differences from interfering with subsequent grouping statistics. Entries marked as missing in cardiac measurement parameters are also included in the corresponding groups of the parameter classification set. The missing status is retained separately within each group. When the proportion of missing entries in a functional parameter group exceeds a set threshold, that group is marked as low in integrity across the entire parameter classification set. The low integrity marker also records the number of missing entries and a list of missing parameter names at the time of triggering, for verification of the missing range during subsequent processing. The number of entries in each group of the parameter classification set is verified for integrity after generation; groups with missing key parameters are marked as incomplete.

[0028] Semantic mapping is used to generate an alignment annotation set based on parameter classification sets and physician examination records. The core challenge of semantic mapping lies in the inherent difference in expression hierarchy between the two types of inputs. Parameter classification sets use single quantitative indicators as the smallest granularity, while physician examination records often summarize the joint abnormal state of multiple parameters in a single sentence. Descriptions such as "abnormal ventricular wall motion with left ventricular dilation" span both structural and functional parameter groups. Semantic decomposition must be performed first, and then each item must be located to the corresponding entry in the parameter classification set. The quality of decomposition directly affects the coverage of effective alignment entries in the alignment annotation set. The mapping process compares each group in the parameter classification set with the corresponding semantic unit in the physician examination record. Successfully matched entries generate effective alignment annotations. The alignment annotations record the correspondence between parameters and examination descriptions and the mapping confidence. Entries with confidence scores below a set threshold are marked as low-confidence alignments in the alignment annotation set. Abnormalities in some parameters in physician examination records are implicitly manifested. Phrases like "decreased cardiac function" do not directly name specific parameters, but their semantics refer to multiple entries within the functional parameter group. Mapping requires semantic inference to establish an indirect alignment relationship between the description and the implicit parameters corresponding to the parameter classification set. This type of indirect alignment is distinguished by low confidence in the alignment annotation set. The ratio of valid alignment entries to low-confidence alignment entries in the alignment annotation set reflects the overall parseability of the examination data. When the proportion of low-confidence alignment entries exceeds a set ratio, a low-confidence annotation is added to the entire batch of alignment annotation sets. The low-confidence annotations completely preserve the quality of the entire batch of data in the alignment annotation set.

[0029] Unstructured feature extraction is performed on fields that failed to parse in the alignment annotation set to generate parsing remainder itemsets. The causes of parsing failure determine the approach strategy for feature extraction. Low-confidence alignment entries with mapping confidence below a threshold and free descriptive fragments not covered by the mapping process have different reasons for failure, and therefore require different extraction methods. For low-confidence alignment entries, feature extraction identifies key semantic words and local descriptive patterns from the original text. The extraction results are incorporated into the parsing remainder itemset as free text features and are accompanied by original parameter position annotations in the alignment annotation set, which record the specific location of the entry in the alignment annotation set. For uncovered descriptive fragments, feature extraction identifies independent semantic units from the fragment content. Subjective biases or notes on limitations of examination conditions added by physicians at the end of their examination records, although not mappable to any standard parameters in the parameter classification set, carry important diagnostic contextual information. The parsing remainder itemset assigns a separate descriptive feature category to these fragments for preservation. The alignment annotation set and the parsing remainder itemset maintain a bidirectional positional association. During the reverse tracing phase, this positional association allows for accurate location of the original context of each entry in the alignment annotation set within the parsing remainder itemset, ensuring that the completion operation is performed within the correct parameter context and avoids incorrect attribution across parameter groups. Each entry in the parsing remainder itemset is labeled with a parsing failure reason type, distinguishing between insufficient confidence and missing coverage. Each entry is written into the parsing remainder itemset with the corresponding reason type label.

[0030] An ultrasound examination dataset is constructed by performing reverse source tracing and completion of residual fields based on the parsed residual itemset. Reverse tracing starts with the original parameter position annotations of each item in the parsed residual itemset, traces back to the corresponding parameters in the alignment annotation set, and re-attempts to establish parsing relationships by combining the overall semantic context of the parameter's group. This direction is the opposite of forward semantic mapping, and it can utilize the already aligned parameter groups to provide additional semantic anchors for failed fields. Successfully completed items are written into the ultrasound examination dataset with corrected alignment status. The confidence of corrected aligned items is a weighted composite value of the reverse source matching degree and the original low confidence. Items in the parsed residual itemset that belong to uncovered descriptive fragments are identified as nearest neighbor parameters during reverse tracing through semantic association analysis with existing fields in the ultrasound examination dataset. These parameters are then attached to the nearest neighbor parameters as extended descriptive fields. For compound descriptions such as "mild tricuspid regurgitation with mild pulmonary hypertension," if they are not fully covered during forward mapping, reverse tracing can locate the descriptive fragment to the corresponding extended descriptive position through the aligned valve parameter group, preserving its semantic contribution to the report content. Entries in the parsing set that still cannot establish a valid mapping relationship after reverse tracing are incorporated into the ultrasound examination dataset as independent descriptive fields. These independent fields typically originate from physicians' subjective notes regarding examination limitations or special circumstances, and their clinical significance is undeniable. Retaining them as independent fields in the ultrasound examination dataset ensures that this content is read by the natural language model without being lost during report generation. When the ultrasound examination dataset is completed, the alignment status of all parameters, corrected entries, and independent descriptive fields together form a structurally complete data layer, with each field carrying parsing quality annotations.

[0031] Step S12: Perform report template matching on the ultrasound examination dataset to generate a report framework and low-match intervals. Extract case complexity identifiers based on the low-match intervals. Perform deviation analysis on parameter reference intervals according to the report framework and case complexity identifiers to generate deviation labels.

[0032] Specifically, report frameworks and low-match intervals are generated by matching report templates to the ultrasound examination dataset. The template library pre-sets report structures according to cardiac disease categories, including custom field mapping rules and chapter organization methods for templates such as valvular disease, cardiomyopathy, and congenital heart disease. The matching score S_match is a weighted synthesis of field coverage F_cover and parameter feature similarity P_sim, i.e., S_match=w1×F_cover+w2×P_sim, where F_cover reflects the proportion of valid fields in the ultrasound examination dataset that the template can accommodate, P_sim reflects the closeness of the current parameter value distribution to the typical value range of the template, and w1 and w2 are weight coefficients that satisfy w1+w2=1, with typical values ​​of w1=0.6 and w2=0.4, calibrated based on the contribution of each indicator to the matching accuracy in historical examination data. In a patient's ultrasound examination dataset, the left ventricular ejection fraction was low, the mitral valve area was reduced, and the pulmonary artery pressure was elevated. These parameters simultaneously triggered high-score intervals in both the valvular disease and cardiomyopathy templates. The S_match difference between the two templates was less than 0.1, indicating that the parameter features of the ultrasound examination dataset could not be fully covered by a single template. When multiple high-score candidate templates exist in the ultrasound examination dataset, the report framework is generated by superimposing a main template to carry the main chapter structure and supplementary templates to fill in the parameter chapters not covered by the main template. When two templates conflict in their definition of the same chapter structure, the template with the higher F_cover takes priority. The S_match corresponding to each chapter position in the report framework is recorded. Chapter intervals with scores below a set threshold are marked as low-match intervals. The concentrated occurrence of low-match intervals indicates that the ultrasound examination dataset deviates from the typical coverage boundary of the report specification template within the corresponding parameter range.

[0033] In some embodiments, the step of extracting case complexity identifiers based on the low-matching interval includes: performing multi-template scoring conflict identification based on the low-matching interval to generate diagnostic uncertainty labels; determining the parameter heterogeneity analysis range based on the diagnostic uncertainty labels to generate an analysis segment set; performing parameter heterogeneity analysis on the analysis segment set to generate heterogeneity quantification values; and performing complexity level mapping based on the heterogeneity quantification values ​​to generate case complexity identifiers.

[0034] Multi-template scoring conflict identification is performed on low-match intervals to generate diagnostic uncertainty annotations. The clinical significance of multi-template scoring conflict lies in the fact that when the score difference between multiple candidate templates in the same interval is below the resolution threshold, the parameter characteristics within that interval simultaneously conform to the typical manifestations of multiple disease categories, resulting in substantial ambiguity in diagnostic attribution rather than scoring error. The parameter combination of a slight reduction in mitral valve orifice area accompanied by regurgitation signals generates typical scoring conflicts between mitral stenosis and mitral prolapse templates. The two types of templates have different coverage strategies for this parameter interval, but the S_match difference is close to zero, meaning the diagnostic attribution of low-match intervals cannot be distinguished solely by template scores. Conflict identification is performed sequentially for each section within the low-match interval. The number of template pairs with differences below the resolution threshold and the mean of the differences are used to quantify the uncertainty at that location, and after normalization, this is used as the uncertainty measure for diagnostic uncertainty annotation. When the number of conflicting chapters within a low-match interval exceeds 50% of the total number of chapters in that interval, the corresponding interval is marked as a high-uncertainty interval in the diagnostic uncertainty annotation. In patients with advanced cardiomyopathy and secondary valvular dysfunction, multiple system parameters are simultaneously at the diagnostic boundary, and their low-match intervals often fall entirely into the high-uncertainty interval. Chapter ranges with persistently high uncertainty measures in the diagnostic uncertainty annotation are marked as high diagnostic clarity insufficient intervals in the overall annotation.

[0035] The analysis segment set is generated based on the range of parameter heterogeneity analysis determined by the diagnostic uncertainty annotations. Segment boundaries are defined according to the gradient change of the uncertainty metric sequence in the diagnostic uncertainty annotations. Points where uncertainty metrics rise from low to high are marked as the segment's starting boundary, and points where they fall back from high to low are marked as the segment's ending boundary. The parameter range between these two boundaries constitutes an independent member of the analysis segment set. This division method ensures that the distribution of uncertainty metrics within each segment is relatively concentrated, and that the diagnostic uncertainty characteristics differ between different segments. High uncertainty intervals marked by the diagnostic uncertainty annotations are preferentially included in the analysis segment set. Diastolic function parameter groups in subjects with preserved ejection fraction heart failure often cause persistent conflicts between multiple diastolic function assessment templates. The uncertainty metric sequence at the corresponding chapter position exhibits a high-level plateau shape, and the segments formed therein are marked as priority analysis levels in the analysis segment set. Each segment in the analysis segment set is accompanied by the mean uncertainty metric value from the diagnostic uncertainty annotations; a higher mean value indicates a denser concentration of diagnostic conflicts within the segment. The number of segments in the analysis segment set and the width of the parameter range covered by a single segment together reflect the complexity of the examination data. When there are many segments and the range of a single segment is wide, it indicates that there is a large range of parameter mixing in the ultrasound examination dataset, and the whole dataset is marked as a large range of parameter mixing state in the analysis segment set.

[0036] For example, the step of performing parameter heterogeneity analysis on the analysis segment set to generate heterogeneity quantization values ​​includes: performing parameter distribution statistics within the analysis segment set to generate a parameter distribution map; identifying parameter transition positions between adjacent segments based on the parameter distribution map to generate a transition feature set; performing transition amplitude classification based on the transition feature set to generate local anomaly annotations; and jointly quantizing the local anomaly annotations and the parameter distribution map to generate heterogeneity quantization values.

[0037] Parameter distribution maps are generated by statistically analyzing the parameter distribution within each analysis segment set. The parameter distribution statistics normalize each parameter value according to its deviation from the reference interval and then compare them under a unified dimension. Normalization eliminates the interference of different parameter dimensions on the distribution pattern, allowing millimeter-level values ​​of cardiac chamber diameter and percentage values ​​of ejection fraction to be presented together within the same distribution dimension. Each segment in the analysis segment set generates an independent parameter distribution map, with the parameter deviation degree as the horizontal axis and parameter quantity density as the vertical axis. Zero deviation corresponds to the median position of the reference interval; a positive extension indicates a parameter value above the upper limit of the reference interval, and a negative extension indicates a value below the lower limit. A unimodal distribution pattern indicates that the overall deviation direction of the parameters within the segment is consistent. In subjects with mitral stenosis and aortic regurgitation, the two sets of parameters deviate in opposite directions, with a lower valve orifice area and higher regurgitation volume. The parameter distribution map exhibits a typical bimodal shape, and the width of the trough between the two peaks directly reflects the degree of opposition between the two sets of parameter deviations. The wider the trough, the more significant the diagnostic difference between the two sets of parameters. The peak width of the parameter distribution plot reflects the dispersion of parameter deviation within a segment. A larger peak width indicates a greater difference in the magnitude of deviation among parameters within the segment. Parameters whose deviation exceeds a set extreme value are marked as extreme deviation points on the parameter distribution plot. Ejection fractions in patients with advanced dilated cardiomyopathy often fall within this extreme deviation region. The pattern of these extreme deviation points between adjacent segments is an important basis for identifying transition features. After arranging the parameter distribution plots of each analysis segment set in segment order, areas of significant morphological difference between adjacent segments are recorded as boundary difference markers on the parameter distribution plot.

[0038] Transition feature sets are generated based on the identification of parameter transition locations between adjacent segments using parameter distribution maps. Transition location identification focuses on the boundaries where both peak offset and peak shape similarity change significantly between adjacent segment parameter distribution maps. Relying solely on peak offset cannot rule out the possibility of normal gradual changes being misjudged as transitions; introducing peak shape similarity as a supplementary criterion improves the specificity of transition identification. Peak shape similarity is calculated by determining the morphological distance between adjacent segment parameter distribution maps. This morphological distance comprehensively considers three dimensions: changes in peak number, peak width, and peak position displacement. Transition identification is triggered when any dimension exceeds a threshold and the peak offset is simultaneously significant. Taking a patient with apical hypertrophic cardiomyopathy as an example, the parameter distribution maps of the basal and apical segment thickness parameters show significant peak offsets and a simultaneous decrease in peak shape similarity at the corresponding segment boundaries. The basal segment distribution map exhibits a narrow peak with low deviation, while the apical segment distribution map exhibits a wide peak with high deviation. The morphological distance between the two exceeds the threshold, meeting the transition triggering condition, and the corresponding boundary is recorded as a high-intensity transition location in the transition feature set. When an extreme deviation point in the parameter distribution map appears only on one side of the boundary segment and not in adjacent segments, this boundary is recorded as a unilateral extreme transition. The intensity of a unilateral extreme transition is usually higher than that of a bilateral symmetrical transition caused by the overall displacement of the parameter distribution. The transition feature set consists of all identified transition positions and their quantized intensities. Positions with intensities exceeding a set threshold are marked as significant transition positions. A high spatial density of significant transition positions indicates a sharp abrupt change in parameter heterogeneity concentrated at a few boundaries, while a low density indicates a more uniform distribution of heterogeneity. The two modes are distinguished in the transition feature set by a distribution feature field.

[0039] Local anomaly annotations are generated by classifying the transition amplitude based on the transition feature set. The classification discretizes the intensity score of each transition position in the feature set into three levels: mild, moderate, and severe. Mild anomalies correspond to gradual changes in parameter distribution between adjacent segments, with peak offsets below the global mean and high peak similarity. Moderate anomalies correspond to significant differences in distribution morphology but maintain some continuity, with peak offsets exceeding the global mean but without triggering unilateral occurrences of extreme deviations. Severe anomalies correspond to sudden discontinuities in parameter distribution between adjacent segments, characterized by a sharp drop in peak similarity and unilateral concentration of extreme deviations. Different anomaly weights are assigned to each of the three levels in the local anomaly annotation, and these weights directly affect the final quantification result of the corresponding segment during the joint evaluation stage of heterogeneity quantification. When significant transition features are concentrated at the same segment boundary, local anomaly annotation marks this boundary as a clustered severe transition. Subjects with severely impaired systolic function but only mildly impaired diastolic function often exhibit this clustering pattern at the boundary between systolic and diastolic parameter groups. Ejection fraction and cardiac index simultaneously fall into extreme deviation regions, while diastolic function parameters deviate less. The intensity difference between the two parameter groups on both sides of the boundary forms a typical clustered severe transition. Local anomaly annotation assigns a higher weight to clustered severe transitions than to dispersed transitions of the same level, reflecting the amplified effect of the concentrated effect of parameter distribution abrupt changes on the overall heterogeneity assessment. Transition locations with similar intensity but not exceeding the severe threshold are marked with a blurred boundary in the local anomaly annotation. This blurred boundary annotation triggers the heterogeneity quantification stage, using interval estimation instead of point estimation for this location. Local anomaly annotation summarizes the level information and weights of all transition locations, maintaining a correlation with the parameter distribution map through segment boundary identifiers.

[0040] Heterogeneity quantization is generated based on joint quantization of local anomaly annotations and parameter distribution maps. Joint quantization extracts the mean μ_w and standard deviation σ_w of the boundary transition level weight sequence of each segment from the local anomaly annotations as boundary heterogeneity feature components. Then, it extracts the normalized peak width W_norm and normalized peak number P_norm of each segment from the parameter distribution map as internal heterogeneity feature components. The four components are combined according to preset weights to form the heterogeneity quantization value Q_het, i.e., Q_het = α_h × (μ_w + σ_w) / 2 + β_h × (W_norm + P_norm) / 2, where α_h is 0.6 and β_h is 0.4. The reason α_h is higher than β_h is that the interference of drastic transitions at the segment boundaries on the diagnostic direction is usually greater than that of uniform dispersion within the segment. The complementary coverage of local anomaly annotations and parameter distribution maps reveals two different causes of heterogeneity: concentrated abrupt changes at the boundary and uniform heterogeneity within the segment. Segments exhibiting severe transitions on both sides of the boundary show a Q_het value close to 1. Segments with only a slight bimodal distribution map show a lower Q_het value, but still higher than unimodal segments. Valve parameter segments in patients with rheumatic mitral valve disease exhibit moderate to high-level characteristics in both input types, with a high mean of transition level weight sequences and a bimodal, wide-valley shape in the parameter distribution map. Q_het typically falls between 0.55 and 0.75, mapping to medium or high complexity levels. Segments with additional blurred boundary annotations in the local anomaly annotations are output as confidence intervals during joint quantization. The upper and lower bounds of the intervals correspond to the quantization results when the blurred location is processed as severe and moderate, respectively. The interval width reflects the uncertainty of the final Q_het for that segment.

[0041] Case complexity labels are generated based on a complexity grading mapping performed on heterogeneous quantification values. The grading mapping discretizes the continuous values ​​of heterogeneous quantification into three levels: low complexity, medium complexity, and high complexity. Low complexity corresponds to heterogeneous quantification values ​​below 0.3, medium complexity to values ​​between 0.3 and 0.6, and high complexity to values ​​above 0.6. Different levels trigger different intensities of content enhancement strategies during report generation: low complexity corresponds to concise statements, medium complexity to moderate elaboration, and high complexity to detailed descriptions and the introduction of pathological mechanisms. In subjects with rheumatic heart disease and atrial fibrillation, both mitral stenosis and cardiac enlargement parameter groups exhibited high heterogeneous quantification values ​​in their respective segments, both of which were mapped to the high complexity level. The high complexity segments in the case complexity labels are concentrated in the range of valve and atrial related parameters, indicating that these two types of parameters require more detailed content within the report framework. Segments with heterogeneity quantification values ​​near the grade boundary are marked with boundary uncertainty in the case complexity identifier. Boundary uncertainty marking triggers the deviation annotation stage, which uses a more conservative deviation assessment method for the parameters of that segment. The spatial distribution of high-complexity segments in the case complexity identifier reveals which parameter ranges of the current examination data are most difficult to incorporate into the standardized report structure. Diffusely distributed high-complexity segments suggest that the disease involves multiple cardiac structures, while focal distribution suggests that a specific parameter group has rare or atypical manifestations.

[0042] Deviation annotations are generated by analyzing the deviation magnitude of parameter reference intervals according to the report framework and case complexity indicators. The parameter reference interval is defined by the normal range of each parameter in the report specification template library. The standardized distance between the measured parameter value and the boundary of the reference interval is used as a measure of the deviation magnitude. The direction of deviation is recorded as either above the upper boundary or below the lower boundary. The direction information is retained in the deviation annotation to distinguish between the direction of disease aggravation and the direction of disease mitigation during the report generation stage. The report framework organizes the deviation magnitude analysis by chapter. The chapter-level deviation annotation summarizes the maximum deviation magnitude, parameter deviation ratio, and consistency of deviation direction for all parameters within that chapter. The chapter granularity of the report framework directly affects the information organization method of the deviation annotation. If the granularity is too fine, it will lead to information fragmentation between chapters; if the granularity is too coarse, it will mask the significant deviation of local parameters. Case complexity indicators have a differentiated impact on the generation strategy of deviation annotations. For parameters in high-complexity segments, a tolerance interval correction is introduced during deviation magnitude analysis, appropriately expanding the reference interval. In subjects with heart failure and mitral regurgitation, multiple parameters deviate from their respective reference intervals simultaneously due to the cumulative effect. Decreased ejection fraction and increased mitral regurgitation each generate independently higher deviation records in the deviation annotation. Without tolerance correction, the deviation annotation magnitude significantly overestimates the independent contribution of a single lesion. The deviation annotation obtained by re-measuring with the tolerance-corrected reference interval better reflects the degree of independence of each parameter's deviation. For parameters in low-complexity segments, no tolerance correction is introduced. Deviation annotations in these segments have higher direct reliability. After the deviation annotations for all parameters in the reporting framework are generated, they are arranged in chapter order to form a complete deviation annotation layer.

[0043] Step S13: Based on the deviation annotation, perform parameter feature association analysis to obtain the descriptive feature set, extract the diagnostic-measurement semantic deviation features of the ultrasound examination dataset to generate clinical prompt labels, and generate a draft report based on the descriptive feature set and clinical prompt labels using a natural language generation model.

[0044] Specifically, a descriptive feature set is obtained through parameter feature association analysis based on deviation annotations. The association between parameters is not only reflected in the coordinated deviation at the numerical level, but also in the pathological logic mapped by the combination of deviation magnitudes. The requirements for report descriptions differ drastically between single-parameter deviations and multi-parameter joint deviations. Association analysis identifies parameter groups with consistent pathological logic from the deviation patterns in the deviation annotations. Association analysis uses the deviation magnitude and direction of each parameter in the deviation annotations as input, establishing a correlation assessment between parameter pairs. This assessment considers both the degree of coordinated change in deviation magnitudes and the consistency of deviation directions. Parameter pairs whose deviation directions corroborate each other have a stronger association than parameter pairs with opposing directions. When both low left ventricular ejection fraction and high mitral regurgitation show significant deviations in the deviation annotations, association analysis identifies them as a function-valvular joint deviation pattern and extracts them as a set of associated features. In the report description, this set of features must demonstrate a causal link between decreased systolic function and valvular disease, rather than stating the two measurement results in isolation. The parameter combinations for tricuspid regurgitation with elevated pulmonary artery pressure show strong consistency in their deviation directions and are included in the descriptive feature set with high correlation strength annotations. During report generation, this is used to determine that the two parameters must demonstrate a physiological correlation within the same descriptive unit. Correlation analysis locates each parameter group according to the chapter position of the deviation annotation. Parameter groups spanning different chapters are marked as cross-chapter correlated parameter groups. For example, in patients with diffuse cardiomyopathy, the three parameters—cardiac chamber enlargement, decreased systolic function, and secondary valvular disease—belong to different chapters in the deviation annotation; correlation analysis identifies these three as cross-chapter correlated parameter groups. Each parameter group in the descriptive feature set is accompanied by its correlation strength and the chapter position of the deviation annotation it covers. Parameter groups with high correlation strength require higher logical coherence in the report wording.

[0045] In some embodiments, the step of extracting diagnostic-measurement semantic deviation features from the ultrasound examination dataset to generate clinical prompt labels includes: extracting the semantic vectors of measurement parameters and diagnostic records from the ultrasound examination dataset; performing similarity analysis on the semantic vectors of measurement parameters and diagnostic records to generate a semantic deviation quantity; determining the deviation direction based on the semantic deviation quantity to generate aggravation tendency labels and mitigation tendency labels; and performing clinical risk association mapping based on the aggravation tendency labels and mitigation tendency labels to generate clinical prompt labels.

[0046] Semantic vectors for measurement parameters and examination records are extracted from the ultrasound examination dataset. Semantic vectorization encodes structured measurement values ​​and free text examination descriptions into the same vector space, providing a basis for direct similarity comparison between two types of data that were originally at different expression levels. Simply relying on the surface differences between numerical values ​​and text cannot capture the substantial deviations between them at the clinical semantic level. Semantic vectorization is the key process to fill this gap. The semantic vectors for measurement parameters are based on the normalized values ​​of the deviation of each parameter in the ultrasound examination dataset. Each parameter's initial coordinates in the semantic space are determined by pre-trained embeddings according to its physiological system and measurement type. The normalized values ​​of the deviation are stretched to the final landing point in the corresponding coordinate direction. Examination data with severely reduced ejection fraction fall into the high deviation region of contractile dysfunction in the semantic vector space of measurement parameters, and the vector distance from the same parameter set within the normal range is significantly increased. In contrast, examination data with only mild reduction show a clear gradient in the deviation of the measurement parameter semantic vector. The semantic vectors of examination records are generated from physician examination record text in the ultrasound examination dataset. Physician descriptions such as "slightly weakened myocardial contractility, valvular activity is acceptable" and "significantly reduced cardiac function, severe mitral regurgitation" are significantly different in the semantic vector space of the examination records. The former's vector representation is closer to the normal reference area, while the latter leans towards the area of ​​severe functional impairment. If the semantic vectors of the measurement parameters of these two descriptions deviate significantly in direction from the corresponding semantic vectors of the examination records, a high-amplitude semantic deviation judgment is triggered during the similarity analysis stage. The ultrasound examination dataset contains simplified labeled entries in the examination record text. During the generation of the semantic vectors of the examination records, semantic completion is performed by combining contextual parameter features. The completed vectors are then appended with low-confidence labels, which are written into the vector space along with the corresponding semantic vector pairs.

[0047] A semantic deviation is generated by performing similarity analysis on the semantic vectors of the measurement parameters and the examination records. Cosine distance measures the directional deviation between the two vectors. The semantic deviation is calculated as D_sem = 1 − (v_m·v_d) / (‖v_m‖×‖v_d‖), where v_m is the semantic vector of the measurement parameters and v_d is the semantic vector of the examination records. Since both the semantic vectors of the measurement parameters and the semantic vectors of the examination records are generated based on the normalized deviation value and each dimension component is non-negative, the cosine similarity between the two vectors ranges from 0 to 1, corresponding to a D_sem value ranging from 0 to 1. The closer D_sem is to 1, the more significant the deviation between the measurement data and the examination records at the clinical semantic level. The semantic vector of measurement parameters reflects the clinical semantic location corresponding to the objective quantitative measurement value, while the semantic vector of the examination record reflects the clinical semantic location corresponding to the physician's subjective judgment. The semantic deviation between the two may originate from the physician's subjective amplification of mild abnormalities, or from the physician observing early signs of functional abnormalities when the measurement value falls within the normal boundary. This dual possibility means that the semantic deviation itself does not carry directional information and needs to be further distinguished during the deviation direction determination stage. Significant differences in semantic deviation exist between different parameter groups within the same subject. A high semantic deviation often appears between the measurement indicators of diastolic function parameters and the physician's description of symptom severity because there is an inherent physiological incomplete correspondence between objective diastolic function measurements and the patient's subjective feelings. In contrast, the semantic deviation of cardiac chamber structure parameters is usually low, and the consistency between diameter measurement results and physician descriptions is relatively stable. Parameter groups with semantic deviation exceeding a set threshold are marked as high deviation groups and are prioritized during the deviation direction determination stage. Parameter groups with deviation below the threshold are marked as low deviation groups and written as low deviation in the semantic deviation data.

[0048] Based on semantic deviation, a tendency to aggravate or mitigate the condition is generated. The core issue addressed by this deviation direction determination is whether the physician's description indicates a more severe or milder condition than the measured data suggests. This is determined by the difference in projection of the semantic vector of the examination record relative to the semantic vector of the measured parameters onto the clinical severity axis. A positive projection difference indicates a bias towards a more severe clinical condition, while a negative projection difference indicates a bias towards a milder clinical condition. For example, if measured data shows mild mitral regurgitation but the physician describes it as "moderate to severe regurgitation with elevated left atrial pressure," the projection of the semantic vector of the examination record onto the severity axis is higher than that of the measured parameters, and the semantic deviation of this parameter group generates a tendency to aggravate the condition. Conversely, if measured data shows mild wall motion abnormalities but the physician describes it as "no obvious abnormalities in wall motion," the projection difference is negative, corresponding to a tendency to mitigate the condition, suggesting that the physician's assessment of the clinical significance of the measured findings is conservative. The strong aggravation parameter groups in the aggravation tendency annotation correspond to large semantic deviations and significantly positive projection differences, while the strong mitigation parameter groups in the mitigation tendency annotation correspond to large semantic deviations and significantly negative projection differences. The existence of both types of strong annotations indicates that the information gap between measurement and diagnosis needs to be clearly reflected in the report description. Both aggravation and mitigation tendency annotations together cover all high-deviation parameter groups, and the direction determination result is only applied to the high-deviation groups.

[0049] Clinical risk association mapping is performed based on aggravation and mitigation propensity annotations to generate clinical warning labels. This mapping is built upon a disease progression path knowledge base. Parameters marked in aggravation propensity annotations are mapped to corresponding disease progression nodes, determining whether the semantic gap between measured values ​​and diagnostic descriptions falls within a certain risk threshold. Parameters marked in mitigation propensity annotations are mapped to corresponding stable or improving disease nodes. These two mapping results together form the risk stratification basis for clinical warning labels. Parameters strongly aggravated in aggravation propensity annotations generate high-priority clinical risk signals during mapping. For example, if a pulmonary artery pressure measurement is borderline high while the physician's description emphasizes "significantly increased right ventricular load," the corresponding aggravation propensity annotation generates a clinical warning label indicating a risk of pulmonary hypertension progression after risk mapping, suggesting that the report should provide a detailed description of the pulmonary circulation pressure status in the corresponding section. Parameters strongly mitigated in mitigation propensity annotations generate low-priority signals after mapping. These signals are marked as a cautionary area in clinical warning labels, and the report should avoid using emphatic descriptive language at the corresponding parameter locations to prevent exaggerating the clinical significance of the measurement findings. Each item in the clinical prompt label also carries a direction type derived from either an aggravation tendency label or a mitigation tendency label. The direction type information is written into each item of the clinical prompt label as a direction type field.

[0050] In some embodiments, generating a draft report using a natural language generation model based on the descriptive feature set and the clinical prompt identifier includes: planning chapter content according to the descriptive feature set to generate chapter content configuration; dynamically adjusting the chapter detail weight based on the clinical prompt identifier and the chapter content configuration to generate enhanced content configuration; generating initial report text using a natural language generation model based on the enhanced content configuration; and annotating the initial report text with low-confidence segments to generate a draft report.

[0051] Chapter content planning is performed based on the descriptive feature set to generate chapter content configurations. Chapter content planning transforms the association structure of parameter groups in the descriptive feature set into a content allocation scheme for each chapter of the report. Chapters corresponding to parameter groups with high association strength receive more content expansion space, while chapters corresponding to parameter groups with low association strength are mainly concisely stated. The planning result is solidified into the basic content weight sequence of each chapter in the chapter content configuration. The chapter position covered by each parameter group in the descriptive feature set determines the chapter activation state of the chapter content configuration. Chapters not covered by any parameter group are marked as low-activation, and low-activation chapters are assigned the lowest content weight in the chapter content configuration. When cross-functional system-related parameter groups appear in the descriptive feature set, the chapter content configuration establishes content association links between corresponding chapters. The association between the systolic dysfunction parameter group and the mitral regurgitation parameter group is recorded as a function-valve association link. When generating the report, the wording of the corresponding chapters must reflect the clinical causal relationship between the two rather than stating them independently. The descriptive feature set of subjects with diffuse cardiomyopathy often forms multiple association links between the three groups of parameter groups: cardiac chamber enlargement, functional decline, and secondary valvular disease. The cross-chapter content coordination requirements in the corresponding chapter content configuration are the most complex. The chapter content configuration fully describes the activation status, basic content weights, and distribution of related links for each chapter. The more complex the relationship structure of the feature set, the more cross-chapter relationship links need to be maintained in the chapter content configuration.

[0052] Enhanced content configurations are generated by dynamically adjusting the weighting of chapter details based on clinical prompts and chapter content configurations. The necessity of dynamic adjustment lies in the fact that chapter content configurations only reflect the impact of parameter group association structures on content planning, while the clinical risk information carried by clinical prompts has a weighting effect on the reporting importance of certain parameter groups, independent of association strength. Only by combining these two types of inputs can the enhanced content configuration fully reflect the clinical value emphasis of the current examination. The adjustment formula is W_aug = W_base × (1 + λ × R_risk), where W_aug is the dynamically adjusted weight of the enhanced chapter content, W_base is the base content weight of the corresponding chapter in the chapter content configuration, R_risk is the normalized risk level assigned to the chapter by the clinical prompts, and λ is an adjustment coefficient with a range of ±0.5. Positive values ​​are used for λ corresponding to aggravated tendencies in clinical prompts to amplify the corresponding chapter content weight, while negative values ​​are used for λ corresponding to reduced tendencies to appropriately compress the content expansion of the corresponding chapter. A patient's clinical profile indicated a high-priority aggravated risk signal in the aortic valve section. The basic weight of this section's content configuration was at a medium level. Dynamically adjusted enhanced content configuration pushed this section's weight to a high level, triggering the report generation model to fully describe the degree of aortic valve lesions, hemodynamic impact, and clinical significance. The enhanced content configuration fully preserves the correlation information of the section's content configuration and superimposes the adjusted weights at each link node. The skewness of the weight distribution in the enhanced content configuration reflects the current report's content focus; when the weight is highly concentrated in a certain section, that section is marked as the dominant finding in the enhanced content configuration.

[0053] The initial report text is generated using a natural language generation model based on enhanced content configuration. The model uses the weight distribution of each chapter in the enhanced content configuration as a constraint on the extent of content expansion. Chapters with higher weights correspond to richer sentences, while those with lower weights correspond to concise declarative output. This constraint ensures that the level of detail in the initial report text aligns with the importance ranking of clinical findings, rather than being determined autonomously by the model. The cross-chapter association constraint model in the enhanced content configuration reflects the clinical causal relationship between parameters at the wording level when generating related chapter content. For example, after the systolic function chapter describes "reduced ejection fraction with enlarged cardiac chambers," the mitral valve chapter, which has a related link, expands its description in the initial report text within the context of functional impairment, rather than simply stating valve morphology in isolation. This contextual coherence is the core contribution of the enhanced content configuration's association link information to the quality of the generated initial report text. When the parameter group of weakened wall motion combined with enlarged cardiac chambers is given a high weight, the corresponding chapter in the initial report text includes a graded description of the degree of functional decline and suggestive wording regarding the etiology, rather than merely stating measurement values. This is the substantial impact of the enhanced content configuration's weight design on the depth of report expression. After the initial report text is generated, the structural integrity of the entire text is verified. The verification checks whether each chapter has generated the corresponding content according to the activation state of the enhanced content configuration. Low-activation chapters that have not generated content are filled with standardized placeholder text. Each content fragment in the initial report text carries a generation confidence label, and the confidence label is written into the initial report text along with the corresponding fragment.

[0054] The initial report text is annotated with low-confidence segments to generate a draft report. Low-confidence segment identification is conducted from two dimensions: first, segments with generation confidence levels below a threshold; and second, segments where there is a significant mismatch between the initial report text content and the weight distribution of the enhanced content configuration. The former stems from the model's own uncertainty assessment of the generated results, while the latter arises from the detection of deviations between external structural constraints and the actual generated content. These two types of low-confidence segments are distinguished in the draft report using different annotation types. When a segment in the initial report text contains a descriptive level far exceeding the expected weight of the corresponding chapter in the enhanced content configuration, this segment is marked as a weight mismatch in the draft report, indicating that the descriptive strength of this segment exceeds the range of expression supported by the parameter association structure and clinical risk assessment. During the review phase, the appropriateness of the wording in such segments is given special attention. In the report draft, low-confidence segments are highlighted to distinguish them from normal-confidence content. The annotation information includes the confidence value and the specific reason for the low confidence level. The reason types are categorized into two types: model generation uncertainty and structural constraint mismatch. These two types correspond to different handling methods during the review stage. For model generation uncertainty, physicians are given priority to assess the accuracy of the content; for structural constraint mismatch, the results of parameter correlation analysis are reviewed for potential biases. The density and distribution of low-confidence segments in the report draft reflect the overall quality level of the generated report. A high concentration of low-confidence segments in a particular chapter indicates a significantly low quality in the initial report text for that chapter, serving as an important reference for prioritizing revisions. The report draft is output as an overlay of a low-confidence annotation layer and a normal-confidence text layer. The low-confidence annotation layer does not affect the continuous reading of the normal text in the report draft.

[0055] Step S14: Extract the key indicator set of the report draft, perform diagnostic rule comparison based on the key indicator set to generate verification results, extract rare sign identifiers for parameter combinations not covered by the rule base in the key indicator set, and generate a revision sequence based on the priority ranking of the verification results and rare sign identifiers.

[0056] Specifically, the key indicator set is extracted from the draft report. The core parameters that trigger the expansion of each chapter's content are located in reverse from the generated text of the draft report. This reverse mapping from text to parameters is the opposite of the forward mapping in the report generation stage, aiming to establish a traceable correspondence between the final content of the draft report and the original examination data. Parameters located at the positions of low-confidence segments in the draft report are prioritized for inclusion in the key indicator set. Parameters from low-confidence sources carry uncertainty annotations in the key indicator set, and the verification conclusions of such parameters during the diagnostic rule comparison stage require higher caution. Parameters cross-referenced in multiple chapters in the draft report, such as ejection fraction which is reflected in both the systolic function chapter and the overall cardiac function assessment chapter, are recorded with multi-reference annotations in the key indicator set. The verification conclusions of multi-referenced parameters cover multiple chapters, and any deviation in verification has the greatest impact on the overall accuracy of the draft report. The size of the key indicator set is limited by the complexity of the report draft. The reports of subjects with diffuse cardiomyopathy involve dense cross-references of parameters in each chapter, and the number of key indicator set entries is significantly more than that of subjects with simple valvular lesions. When the size of the key indicator set exceeds the set threshold, a batch verification mechanism is triggered. Each batch is divided according to the order of the chapter to which the parameters belong. The parameter relationships within each batch are fully preserved, and cross-referenced parameters between batches are recorded synchronously in each batch in the form of shared entries.

[0057] The diagnostic rule comparison is performed based on the key indicator set to generate verification results. The pre-set rules in the diagnostic rule base cover three levels: single-parameter normal range constraints, physiological correlation constraints between parameter pairs, and diagnostic conclusion constraints triggered by multiple parameters. The comparison process is executed sequentially according to these levels. Each parameter in the key indicator set is matched against the single-parameter rule in turn. When the left ventricular end-diastolic diameter exceeds the upper limit set by the rule, the rule for left ventricular enlargement is triggered and hit. The hit result, along with the measured deviation, is written into the verification result. The larger the deviation, the higher the severity level. The physiological correlation rules at the parameter pair level check whether there are contradictory numerical combinations in the key indicator set. When the ejection fraction is significantly reduced but the cardiac index is still within the normal range, the verification result marks this combination as a violation of the correlation constraint. The severity of the correlation constraint violation is one level higher than that of the parameter pair with the larger deviation. Multi-parameter joint rules trigger diagnostic conclusions based on the overall pattern of the key indicator set. An inverted ratio of early to late diastolic blood flow velocity (E / A) at the mitral valve orifice, combined with an increased ratio of early diastolic blood flow velocity to annular tissue Doppler velocity (E / e'), satisfies the multi-parameter joint rule for diastolic dysfunction. The verification results simultaneously record the rule name, triggering parameter, and rule confidence level for each hit. The severity of the multi-parameter joint hit is a weighted average of the rule confidence level and the mean deviation of the triggering parameter. The verification results summarize all hit and miss states. Chapters with high hit density indicate sufficient rule support for the corresponding content in the draft report. Missed parameter items are recorded separately in the verification results as uncovered entries, carrying three attributes: parameter value, current deviation magnitude, and nearest neighbor rule trigger distance. The severity levels of the three types of hit items are calculated according to their respective methods and then normalized in the full verification results. The normalized result serves as the normalized severity value for rule violations, used in the priority ranking stage.

[0058] In some embodiments, the step of extracting rare sign identifiers from parameter combinations not covered by the rule base of the key indicator set includes: performing a diagnostic rule base coverage comparison on the key indicator set to generate an uncovered parameter set; performing parameter correlation constraint analysis based on the uncovered parameter set to generate a set of mutually exclusive parameter pairs; performing abnormal co-occurrence pattern identification based on the set of mutually exclusive parameter pairs to generate co-occurrence anomaly labels; and performing sign level assessment based on the co-occurrence anomaly labels to generate rare sign identifiers.

[0059] A diagnostic rule base coverage comparison is performed on the key indicator set to generate an uncovered parameter set. The coverage comparison starts with parameter entries marked as rule misses in the validation results. These entries have completed parameter value recording in the key indicator set but did not trigger any rules during the rule comparison stage. The reasons for not triggering are divided into two categories: either the parameter, when existing alone, does not reach the trigger threshold of any rule, or the combination of parameters exceeds the preset pattern range of the rule base. The former type of entries are marked as threshold not reached in the uncovered parameter set and are accompanied by the nearest neighbor rule trigger distance. The smaller the trigger distance, the closer the parameter is to the known rule activation boundary. When generating the mutually exclusive parameter pair set, the trigger distance is used as an auxiliary correction factor for the mutual exclusion strength score. The latter type is marked as pattern out of bounds. The two types of markings are distinguished by independent fields in the uncovered parameter set. The coverage boundary of the diagnostic rule base is continuously updated with the diagnostic consensus of common diseases, but there are naturally coverage blind spots for complex mixed lesions with extremely low incidence. The probability of parameter combinations involving multiple cardiac structures simultaneously falling into such blind spots is significantly higher than that of single-structure lesion combinations in the key indicator set. Entries whose individual parameters in the key indicator set are within the reference range but exhibit an abnormal pattern when combined with adjacent parameters are also included in the uncovered parameter set. When the tricuspid annulus systolic displacement value is at the lower limit of normal while the concurrent right ventricular outflow tract velocity time integral is low, and no single-parameter rules are triggered, but the combined pattern indicates a marginal state of right ventricular function, this parameter is included in the uncovered parameter set to retain the opportunity for inclusion in correlation constraint analysis. When the number of entries in the uncovered parameter set exceeds 30% of the total number of key indicator sets, the entire uncovered parameter set is marked with a high atypical proportion. The high atypical proportion annotation records the ratio of the number of entries triggered and the chapter distribution information of each parameter.

[0060] Based on the uncovered parameter set, a set of mutually exclusive parameter pairs is generated through parameter correlation constraint analysis. This correlation constraint analysis examines whether there are theoretical coexistence limitations among the parameters in the uncovered parameter set from the perspective of cardiac physiology. A combination of two parameters that should not simultaneously exhibit significant abnormalities under normal physiological conditions is defined as a mutually exclusive parameter pair and included in the mutually exclusive parameter pair set. The determination of mutual exclusivity is based on a physiological constraint knowledge base rather than statistical regularities, ensuring the medical rigor of the mutually exclusive parameter pair set. There is a strong coupling relationship between regurgitation parameters and their corresponding intracardiac pressure parameters in the uncovered parameter set. Severe mitral regurgitation should physiologically drive an increase in left atrial pressure. If both occur simultaneously in the uncovered parameter set but their measured directions are contradictory, they constitute a mutually exclusive parameter pair and are written into the mutually exclusive parameter pair set. This combination suggests one of three possible explanations: measurement error, a special compensatory state, or a rare pathological mechanism. Each parameter pair in the mutually exclusive parameter pair set is accompanied by a mutual exclusion strength score. The mutual exclusion strength is determined by the strictness of the physiological constraints between the two parameters. Parameter pairs that are anatomically directly related have a higher mutual exclusion strength than parameter pairs that are functionally indirectly related. The mutual exclusion strength score is divided into three levels according to the strictness of physiological constraints: strong mutual exclusion, medium mutual exclusion, and weak mutual exclusion. Strong mutual exclusion corresponds to parameter pairs that are anatomically directly related, medium mutual exclusion corresponds to parameter pairs that are functionally indirectly related, and weak mutual exclusion corresponds to parameter pairs that are statistically occasionally co-occurring. The three levels of scores are stored as independent fields in each entry of the mutually exclusive parameter pair set. Isolated parameter entries in the uncovered parameter set that cannot establish a mutual exclusion relationship with any other parameter are retained in the mutually exclusive parameter pair set as single-parameter anomalies. Single-parameter anomaly entries are accompanied by an isolation label, which records the original non-triggered cause type and parameter deviation magnitude of the entry in the uncovered parameter set.

[0061] For example, the step of generating co-occurrence anomaly labels by performing anomaly co-occurrence pattern recognition based on the mutual exclusion parameter pair set includes: performing feature vectorization encoding on the mutual exclusion parameter pair set to generate a feature encoding set; performing rare sign similarity retrieval on the feature encoding set to generate a candidate sign set; performing confidence dispersion evaluation on the candidate sign set to generate a diagnostic boundary fuzzy label; and performing optimal matching selection based on the diagnostic boundary fuzzy label and the candidate sign set to generate co-occurrence anomaly labels.

[0062] A feature encoding set is generated by feature vectorization encoding the set of mutually exclusive parameter pairs. Feature vectorization transforms the measurement attributes of each parameter pair in the set into fixed-dimensional vectors. Each dimension of the vector corresponds to the parameter deviation magnitude, deviation direction, physiological region type to which the parameter belongs, and mutual exclusion strength score, respectively. These four types of features comprehensively express the abnormal co-occurrence characteristics of the parameter pair in the same vector space. The number of dimensions in each vector in the feature encoding set is uniform, providing a basis for direct distance comparison between vectors of different parameter pairs in the same space. The level of mutual exclusion strength score in the set of mutually exclusive parameter pairs affects the weight distribution of the corresponding dimensions in the feature encoding set. Vectors of parameter pairs with high mutual exclusion strength have larger values ​​in the strength dimension, forming a clear numerical stratification with vectors of parameter pairs with low mutual exclusion strength in this dimension. The regional distribution of the two types of parameter pairs in the vector space can thus be distinguished. The parameter pairs for aortic stenosis combined with aortic regurgitation exhibit highly specific coding combinations in terms of deviation magnitude, direction, and region. Vectors of this parameter pair in the feature coding set are located in a region highly similar to rare signs associated with bidirectional valvular disease in the retrieval space. The smaller the distance between vectors, the more consistent the abnormal features of the current parameter pair are with the characteristic pattern of that sign. The feature coding set vectors corresponding to single-parameter abnormal entries in the mutually exclusive parameter pair set are assigned zero values ​​in the mutual exclusion strength dimension. Vectors with zero values ​​are concentrated in the low-value region of the intensity dimension in the vector space, forming a natural spatial separation from the vectors of parameter pairs with high mutual exclusion strength. All vectors in the feature coding set constitute the abnormal feature matrix of the current examination. When feature vectors are concentrated in a certain region, it indicates that the abnormal pattern of this examination has strong homogeneity; when they are dispersed, it indicates that the abnormal features span multiple sign types.

[0063] A candidate feature set is generated by performing rare feature similarity retrieval on the feature encoding set. The similarity retrieval uses each vector in the feature encoding set as query input, searching for the entry with the highest cosine similarity in the feature vector index of the rare feature knowledge base. The results are sorted in descending order of similarity. Entries exceeding a set threshold are included in the candidate feature set. The threshold is set to balance recall and precision; too high a threshold leads to missed detections of truly rare features, while too low a threshold results in a large number of irrelevant entries mixed into the candidate feature set. The error direction is biased towards the lower threshold side when the threshold is calibrated, with the cost of missed detections being higher than the cost of false positives. Multiple vectors from the same subject in the feature encoding set are retrieved independently. When different parameter pairs hit the same knowledge base entry, they are merged into a single record in the candidate feature set, and their support is accumulated. Higher support indicates that more parameter pairs in the current examination data point to the same feature. Cardiac amyloidosis is a clinically underestimated infiltrative cardiomyopathy. Its characteristic ultrasound manifestations involve a combination of parameters including ventricular wall thickening, granular myocardial echo enhancement, and diastolic dysfunction. When the feature encoding vectors of these three parameters achieve high similarity with cardiac amyloidosis entries in the knowledge base, the support of that entry in the candidate sign set accumulates rapidly. The support field ultimately records the number of hits of the three vectors and the average similarity of each hit. Each entry in the candidate sign set carries both the highest similarity value and the support. Entries with high highest similarity but low support indicate that only a few parameter pairs point to the sign, while entries with moderate highest similarity but high support indicate that multiple parameter pairs point to the sign, but the quality of a single match is generally low. The difference between the two types of entries is reflected in the relative size of the support field and the similarity field, both of which are recorded in each entry of the candidate sign set.

[0064] Based on the candidate sign set, a confidence dispersion assessment is performed to generate fuzzy labeling of the diagnostic boundary. The confidence dispersion assessment determines the degree of certainty in the diagnostic indication by quantifying the score difference between the highest and second-highest matching items in the candidate sign set. A large difference indicates a clear abnormal feature pointing to a single sign, while a small difference suggests that multiple signs have similar scores and the diagnosis is uncertain. The comprehensive score of each item is C_score = γ × S_max + (1−γ) × N_norm, where S_max is the highest similarity, and N_norm is the normalized value of the number of candidate sign hits. γ is set to 0.65 and (1−γ) to 0.35 to highlight the dominant role of similarity. The item with the highest comprehensive score in the candidate sign set is denoted as C_score1, and the item with the second-highest score is denoted as C_score2. The difference between the two, ΔC = C_score1−C_score2, constitutes the core quantitative indicator for fuzzy labeling of the diagnostic boundary. The smaller ΔC is, the more intense the competition among candidate signs and the more ambiguous the diagnostic boundary. The ambiguity of the diagnostic boundary is classified into three levels based on ΔC: ΔC below 0.05 is marked as highly ambiguous. Dilated cardiomyopathy and late-stage ischemic cardiomyopathy are highly similar in terms of ultrasound parameters, and the comprehensive score of the two types of signs often shows a ΔC below 0.05. For highly ambiguous entries, the C_score values ​​of the first and second candidate entries, as well as the corresponding S_max and N_norm sub-items, are recorded simultaneously to verify the specific contribution of each sub-item when context constraints are introduced; ΔC between 0.05 and 0.2 is marked as moderately ambiguous. For moderately ambiguous entries, the C_score values ​​of the first and second candidate entries are recorded simultaneously; ΔC above 0.2 is marked as clear boundary. For clear boundary entries, only the identifier and C_score value of the highest-scoring item are retained, without additional suboptimal candidate information.

[0065] Co-occurrence anomaly labels are generated by combining diagnostic boundary fuzzy annotation with a candidate sign set for optimal matching. The optimal matching selection strategy varies depending on the fuzziness level of the diagnostic boundary fuzzy annotation. In cases with clear boundaries, the highest-scoring item in the candidate sign set is used directly as the matching result. In cases of high fuzziness, parameter context constraints are introduced from the two candidate items with similar scores for further filtering. The constraints extract the co-existence probability with other confirmed parameter features of the current subject from the accompanying parameter descriptions of each candidate item, and candidate items with higher co-existence probabilities are selected first. Constrictive cardiomyopathy and pericardial constriction have highly overlapping ultrasound manifestations at the hemodynamic parameter level, and the score difference between the two in the candidate sign set is very small. However, after introducing context constraints of pericardial thickness and myocardial echo characteristics, the ranking of the two types of signs often changes significantly. The labeling of diagnostic boundary fuzzy annotation in these cases ensures that the context constraints are forcibly triggered rather than skipped. In the diagnostic boundary ambiguity annotation, candidates marked as highly or moderately ambiguous are simultaneously recorded as suboptimal matching entries in the co-occurrence anomaly annotation after the optimal match is completed. The closer the scores of the suboptimal and optimal entries are, the lower the recognition confidence of the co-occurrence anomaly annotation. The recognition confidence is linearly mapped from the ΔC value to the range of 0 to 1. When ΔC is 0, the confidence is 0; when ΔC reaches 0.2 or above, the confidence is 1. The upper limit of the mapping result is truncated to 1. The co-occurrence anomaly annotation uses the identifier of the optimal matching symptom entry and the original score of the corresponding candidate symptom set as the core output. The recognition confidence is attached to each record of the co-occurrence anomaly annotation. Each record also retains the feature encoding set vector index that triggered the match to support the source verification of the match during the symptom level assessment stage.

[0066] Rare sign identifiers are generated based on co-occurrence anomaly annotations. The sign level assessment compares the co-occurrence patterns of parameters described by the anomaly annotations with a rare sign knowledge base for cardiac ultrasound. This knowledge base includes special parameter combinations that occur less frequently than a set threshold in literature reports or large-sample examination data. When a co-occurrence anomaly annotation matches an entry in the knowledge base, the corresponding sign level is triggered. The level is determined by the historical incidence rate of the knowledge base entry and the mutual exclusion strength score of the co-occurrence anomaly annotation. Incomplete ventricular compaction is a rare congenital cardiomyopathy characterized by simultaneous abnormalities in trabecularization parameters and compaction layer thickness parameters at the apex. When the co-occurrence anomaly annotation of these two parameters matches a low-incidence entry in the knowledge base, the corresponding rare sign identifier is rated as high-level. The high-level rare sign identifier records the historical incidence rate range and corresponding mutual exclusion strength score level used to trigger the level in the level field. When co-occurrence anomaly labels partially match multiple knowledge base entries, rare feature identifiers are recorded as a candidate list. The entry with the highest matching degree is output as the primary rare feature identifier, and secondary candidate entries are added to the candidate list in descending order of matching degree. Each entry in the list is accompanied by a matching degree value and a historical occurrence rate classification. These two attributes jointly determine the reference weight of the candidate entry in the overall identification result. Each entry in the rare feature identifier also carries an identification confidence field. The confidence field value comes from the ΔC mapping result of the record corresponding to the co-occurrence anomaly label. When high-level feature identifiers coexist with low identification confidence, an unverified label is added to the rare feature identifier. The unverified label records the confidence value and the specific reason type for the low confidence.

[0067] Based on the verification results and rare sign identifiers, a revision sequence is generated by prioritizing the data. The priority score P_revise = α × R_severity + β × G_rarity, where R_severity is the normalized value of the severity of rule violations in the verification results, G_rarity is the normalized value of the sign level of the rare sign identifier, and α and β are weighting coefficients satisfying α + β = 1. Typical values ​​are α = 0.6 and β = 0.4. For examination data with high rule hit density, α is higher than β; for examination data with a high proportion of parameters not covered by the rule, β is higher than α. The contribution weights of the two types of inputs to the revision priority are dynamically adjusted according to the overall characteristics of this examination. In the verification results, the P_revise of items violating associated constraints is generally higher than that of items violating single parameter ranges. This is because data consistency issues corresponding to associated constraint violations affect the presentation logic of multiple chapters in the report draft. If the constraint conflict between left ventricular systolic function parameters and output blood flow parameters is not prioritized, the content of the functional assessment chapter and the hemodynamics chapter in the report will produce contradictory descriptions, which are difficult to detect segment by segment during the review stage. Entries with high-level rare feature identifiers are mapped to high-priority revision tasks. These high-level rare feature identifiers originate from parameter anomalies with extremely low historical occurrence rates. The corresponding draft report content often exceeds the normal training distribution of natural language generation models, making it the most difficult to guarantee in terms of expression quality. Prioritizing these tasks for human-machine collaborative review helps concentrate review resources on the most unreliable report segments. The joint arrangement of validation results and rare feature identifiers is sorted in descending order of P_revise in the revision sequence. Revision tasks with similar scores are secondarily sorted by the order of the draft report chapters, ensuring that the review process maintains chapter-level coherence even when priorities are similar. Each task in the revision sequence includes the contribution source of its priority score. The source field distinguishes whether the score mainly comes from validation results or rare feature identifiers. The source field and the P_revise value of each task are both written into the revision sequence entry.

[0068] Step S15: Based on the revision sequence, trigger human-machine collaborative review to generate review feedback, extract features of unrevised paragraphs from the review feedback to generate credible expression identifiers, and integrate the content with the credible expression identifiers and review feedback to generate a cardiac ultrasound diagnostic report.

[0069] Specifically, the human-machine collaborative review process generates review feedback based on the revision sequence. Tasks in the revision sequence are prioritized and scored sequentially before being pushed to the review interface. This targeted push focuses reviewers' attention on report segments with the most severe rule violations and the highest concentration of rare symptoms, significantly improving overall review efficiency, especially in highly complex reports. Tasks with the highest revision sequence scores are displayed first. Reviewers can perform three types of operations on the corresponding report segments: modification, retention, and annotation. Modification directly replaces the corresponding text in the report draft; retention records the reviewer's explicit approval of the current content; and annotation adds review comments to the segment without changing the text. The results of all three operations are written to the review feedback in real time. When tasks with high rare symptom identification levels are pushed in the revision sequence, typical parameter descriptions of that symptom from the knowledge base are additionally displayed for comparison. For example, the ultrasound manifestations of ejection fraction-preserving heart failure and restrictive cardiomyopathy are highly similar. After comparing the reference descriptions, the accuracy of reviewers' judgments regarding modifications to the relevant statements in the report draft is significantly improved. Upon completion of a revision sequence task, a linked update of related tasks is triggered. When modifications to the shrinking function chapter affect the consistency of the hemodynamics chapter, the review feedback automatically marks the affected chapter as a related pending review area, prompting reviewers to extend the review of related segments after completing the current task. The review feedback records for modification categories save the text before and after modification along with the execution timestamp; records for retention categories save the review approval status and the execution timestamp; and records for annotation categories save the review comments text and the execution timestamp. All three types of records include a corresponding revision sequence task identifier.

[0070] In some embodiments, the step of extracting unrevised paragraph features from the review feedback to generate a credible representation identifier includes: parsing the revision operation records of the review feedback to generate a revision location set; performing high-frequency revision clustering identification based on the revision location set to generate a system deviation area label; extracting the difference set of unrevised content from the review feedback based on the revision location set to generate an unrevised paragraph set; and comparing the credibility of the unrevised paragraph set and the system deviation area label to generate a credible representation identifier.

[0071] The revision operation records of the review feedback are analyzed to generate a revision location set. The value of analyzing the revision operation records lies not only in locating the modified text positions, but also in reconstructing the reviewer's revision intentions from the operation sequence. The operation trajectory of repeatedly modifying the same position and completing a one-time revision correspond to different revision confidence levels. The former indicates that the reviewer is uncertain about the correct expression of that position, while the latter indicates that the revision direction is clear. Each modification operation in the review feedback carries the text before modification, the text after modification, and the execution timestamp. The analysis process extracts the paragraph position identifier corresponding to the modification operation. The number of modification operations and the distribution of modification magnitude within each chapter form the basic statistical layer of the revision location set. The modification magnitude is quantified by the degree of difference between the text before and after modification. The magnitude of deletion modification is measured by the proportion of the deleted text to the original segment length, and the magnitude of addition modification is measured by the semantic difference between the added text and the original text. When the reviewer rewrote "mildly reduced ejection fraction" as "significantly impaired left ventricular systolic function," the modification magnitude was significantly high, and this position was marked as a high-magnitude revision in the revision location set. In the review feedback, the paragraph position corresponding to the retained operation is written to the revision position set with zero modification. The retained record and the modification record together cover the entire position range of the revision position set. Each position in the revision position set is accompanied by an operation timestamp sequence. Positions with a longer dwell time but which are ultimately not modified are marked as review hesitant positions. The hesitant mark records the dwell time of the position and the chapter identifier to which it belongs. When the dwell time exceeds twice the average of other positions in the same chapter, the hesitant mark is marked with a high hesitant level.

[0072] Based on the revision location set, high-frequency revision clustering is performed to generate systematic bias zone annotations. The spatial clustering pattern of high-frequency revisions reveals not individual description quality issues, but rather a systematic generation bias in the natural language generation model's description of certain parameters. The characteristic of systematic bias is that the sequence of modification magnitudes at each location within the cluster is consistently high and tends to align in direction. Clustering is triggered when the sequence of modification magnitudes across chapters in the revision location set shows a continuous high-value interval. The clustering threshold is set when the average modification magnitude at three or more consecutive locations exceeds 1.5 times the global average modification magnitude. Continuous intervals meeting this condition are marked as high-frequency revision clusters. The start and end positions of the clusters, along with the average modification magnitude within the interval, constitute the basic description for systematic bias zone annotation. Reports of mitral valve complex lesions often exhibit clustered high-frequency revisions in the mitral valve function description section. Reviewers repeatedly revise descriptions involving both regurgitation and stenosis. Systematic bias zone annotation marks this entire interval as a systematic bias zone, indicating a systematic weakness in the natural language generation model's language generation strategy for complex valvular lesions. Areas with concentrated revisions and hesitant revisions are also included in cluster identification. The high density of hesitant revisions reflects that the generation quality of that segment is near the review threshold. Systematic deviation area annotations assign a moderate deviation weight to hesitant clusters, lower than that of modified clusters but higher than that of normal segments. Each deviation area in the systematic deviation area annotation carries a range identifier and average modification magnitude for the high-frequency revision cluster. The range identifier field also records the start and end chapter positions covered by the deviation area and the number of hesitant positions within the interval.

[0073] Based on the revision location set, the unrevised content of the review feedback is extracted to generate an unrevised paragraph set. Unrevised content consists of report segments that reviewers either actively choose to retain or passively skip. These two retention methods are distinguished in the unrevised paragraph set by different annotation types. Active retention represents the reviewer's explicit endorsement of the content's accuracy, while passive skipping may stem from review time constraints or the coverage boundaries of the revision sequence's push scope. The difference extraction is based on the full paragraph positions of the report draft, subtracting the positions marked with modification operations from the revision location set. The remaining paragraph content forms the main body of the unrevised paragraph set. Paragraphs marked with explicit retention operations in the review feedback are marked with an active retention annotation in the unrevised paragraph set. The active retention annotation records the timestamp of the retention operation and the corresponding revision sequence task identifier. Report segments outside the revision sequence's push coverage are marked with an "unpushed review" annotation in the unrevised paragraph set. The unpushed review annotation records the historical retention rate statistics and sample size of the chapter to which the segment belongs. The historical retention rate is 1 minus the modification rate of that chapter in historical reviews. The review feedback timestamp sequence shows that when the reviewer spends very little time in a certain chapter, a quick skip mark is added to the corresponding paragraph. The quick skip mark records the ratio of the actual time spent in the chapter to the average review time spent in the chapter. When the ratio is less than 0.3, a high-speed skip level mark is added with C_src set to 0.5. When the ratio is between 0.3 and 0.6, a medium-speed skip level mark is added with C_src set to 0.6.

[0074] Credible representation identifiers are generated based on credibility comparison using an unrevised paragraph set and system deviation area annotations. The core contradiction in credibility comparison lies in the spatial overlap between some content in the unrevised paragraph set and the deviation areas annotated by the system deviation area. These paragraphs have not been modified in this review and are located in known systematically weak generation areas. Simply equating unmodified with credibility would overestimate their true reliability. The credibility score for each paragraph is T_trust = φ × C_src − ψ × B_overlap, where C_src is the initial credibility weight corresponding to the paragraph's source type, with 0.9 for actively retained categories and 0.7 for passively untouched categories. For categories not pushed for review, the historical retention rate is used, which is 1 minus the historical modification rate. C_src is 0.5 for high-speed skipping level and 0.6 for medium-speed skipping level. B_overlap is the normalized value of the overlap between the paragraph's location and the deviation area annotated by the system deviation area, with φ set to 0.7 and ψ set to 0.3. The lower limit of T_trust is 0, and the formula result is uniformly set to 0 when it is negative. The systematic deviation zone is concentrated in the section describing diastolic function. Within this zone, the T_trust values ​​of paragraphs in the unrevised paragraph set are generally low. The credible statement identifier marks these paragraphs as high-risk statements requiring verification. Paragraphs with T_trust values ​​exceeding a set threshold generate high-credibility credible statement identifiers, while those below the threshold generate verification-pending identifiers. Paragraphs with high-credibility credible statement identifiers are included in the original text during integration, while paragraphs with verification-pending identifiers trigger conservative wording processing. Each credible statement identifier entry is written with a T_trust value and a source type identifier. The T_trust value is calculated using a formula and has a lower limit constraint of 0. The source type identifier retains information about the review method for this paragraph's inclusion in the unrevised paragraph set.

[0075] The cardiac ultrasound diagnostic report is generated by integrating content based on credible description markers and review feedback. Content integration uses the chapter positions of the draft report as a coordinate benchmark. Modifications in the review feedback are overwritten according to their chapter positions in the original text. High-credibility paragraphs identified by the credible description markers are filled into the remaining positions not covered by the modified text, with each position having a unique source, eliminating ambiguity caused by multiple versions competing for the same position. High-risk descriptions in the credible description markers undergo conservative wording processing before filling; absolute statements in the original text are appropriately down-toned; and in the description of right ventricular function, "severely impaired function" is changed to "significantly limited function." These processed segments are marked with adjustment tags in the cardiac ultrasound diagnostic report for future review. Positions with added comments in the review feedback but not yet modified are integrated into the corresponding chapters of the cardiac ultrasound diagnostic report as footnotes for clinical reference. After integration, a full-text consistency check is performed. If descriptions of the same parameter across chapters have inconsistent wording, automatic alignment is triggered based on the revised version in the review feedback, ensuring logical coherence and avoiding contradictions across chapters. After verification, a report quality summary is attached and stored in the form of metadata. The summary is not included in the main text of the cardiac ultrasound diagnostic report. The summary fields cover the review modification coverage, the proportion of highly credible paragraphs marked with credible statements, and the consistency verification results of the full text. The summary is continuously written to the summary database using the inspection batch identifier as an index.

[0076] To implement the above-described method embodiments, a method for automatically generating and reviewing cardiac ultrasound reports is provided to achieve the corresponding functionalities and technical effects. See also... Figure 2 , Figure 2 This diagram illustrates a structural block diagram of an automatic cardiac ultrasound report generation and review device according to an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The automatic cardiac ultrasound report generation and review device provided in this embodiment includes:

[0077] Data acquisition module 201 is used to acquire cardiac measurement parameters and physician examination records, and to construct an ultrasound examination dataset based on the cardiac measurement parameters and physician examination records through structured parsing.

[0078] The frame generation module 202 is used to perform report specification template matching on the ultrasound examination dataset to generate a report frame and low matching degree intervals, extract case complexity identifiers based on the low matching degree intervals, and perform parameter reference interval deviation magnitude analysis to generate deviation labels according to the report frame and the case complexity identifiers.

[0079] The draft generation module 203 is used to obtain a descriptive feature set by performing parameter feature association analysis based on the deviation annotation, extract the diagnostic-measurement semantic deviation features of the ultrasound examination dataset to generate clinical prompt labels, and generate a report draft based on the descriptive feature set and the clinical prompt labels through a natural language generation model.

[0080] Rule verification module 204 is used to extract the key indicator set of the report draft, perform diagnostic rule comparison based on the key indicator set to generate verification results, extract rare sign identifiers for parameter combinations not covered by the rule base in the key indicator set, and generate a revision sequence by prioritizing based on the verification results and the rare sign identifiers.

[0081] The report output module 205 is used to trigger human-machine collaborative review based on the revision sequence to generate review feedback, extract unrevised paragraph features from the review feedback to generate a credible expression identifier, and integrate the content with the credible expression identifier and the review feedback to generate a cardiac ultrasound diagnostic report.

[0082] The aforementioned automatic cardiac ultrasound report generation and review device can implement one of the methods described in the above embodiments for automatic generation and review of cardiac ultrasound reports. The options described in the above method embodiments are also applicable to this embodiment and will not be detailed here. The remaining content of this application's embodiments can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.

[0083] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.

Claims

1. A method for automatically generating and reviewing cardiac ultrasound reports, characterized in that, include: Obtain cardiac measurement parameters and physician examination records, and construct an ultrasound examination dataset based on the cardiac measurement parameters and physician examination records through structured parsing; The ultrasound examination dataset is matched with the report standard template to generate a report framework and low matching degree intervals. Based on the low matching degree intervals, the case complexity identifier is extracted. According to the report framework and the case complexity identifier, the deviation of the parameter reference interval is analyzed to generate deviation labels. Based on the deviation annotation, parameter feature association parsing is performed to obtain a descriptive feature set. The ultrasound examination dataset is then subjected to diagnostic-measurement semantic deviation feature extraction to generate clinical prompt labels. Based on the descriptive feature set and the clinical prompt labels, a report draft is generated using a natural language generation model. Extract the key indicator set from the draft report, perform diagnostic rule comparison based on the key indicator set to generate verification results, extract rare sign identifiers for parameter combinations not covered by the rule base in the key indicator set, and generate a revision sequence based on the priority ranking of the verification results and the rare sign identifiers. Based on the revised sequence, human-machine collaborative review is triggered to generate review feedback. Features of unrevised paragraphs are extracted from the review feedback to generate a credible expression identifier. The content is then integrated with the credible expression identifier and the review feedback to generate a cardiac ultrasound diagnostic report.

2. The method according to claim 1, characterized in that, The process of constructing an ultrasound examination dataset based on the cardiac measurement parameters and the physician's examination record through structured analysis includes: The cardiac measurement parameters are categorized and organized according to measurement type to generate a parameter classification set; Based on the parameter classification set and the physician's examination record, a semantic mapping is performed to generate an aligned annotation set; Unstructured feature extraction is performed on the fields that failed to be parsed in the alignment annotation set to generate a parsing remainder itemset; Based on the parsed remaining itemset, the residual fields are traced back and completed to construct the ultrasound examination dataset.

3. The method according to claim 1, characterized in that, The step of extracting case complexity identifiers based on the low-matching intervals includes: Based on the low-matching interval, multi-template scoring conflict identification is performed to generate diagnostic uncertainty labels; Based on the diagnostic uncertainty annotation, the range of parameter heterogeneity analysis is determined, and an analysis segment set is generated; Perform parameter heterogeneity analysis on the analysis segment set to generate heterogeneity quantification values; Based on the heterogeneity quantification value, perform complexity level mapping to generate case complexity identifiers.

4. The method according to claim 1, characterized in that, The step of extracting diagnostic-measurement semantic deviation features from the ultrasound examination dataset to generate clinical prompt labels includes: Extract the semantic vectors of the measurement parameters and the semantic vectors of the examination records from the ultrasound examination dataset; A similarity analysis is performed between the semantic vector of the measurement parameters and the semantic vector of the medical record to generate a semantic deviation. Based on the semantic deviation amount, the deviation direction is determined to generate aggravation tendency labels and mitigation tendency labels; Based on the aggravation tendency label and the mitigation tendency label, a clinical risk association mapping is performed to generate clinical prompt labels.

5. The method according to claim 1, characterized in that, The process of generating a draft report using a natural language generation model based on the descriptive feature set and the clinical prompt identifiers includes: The chapter content configuration is generated by planning the chapter content according to the described feature set; Based on the clinical prompts and the chapter content configuration, the chapter detail weight is dynamically adjusted to generate enhanced content configuration; Based on the enhanced content configuration, an initial report text is obtained by generating text using a natural language generation model. The initial report text is annotated with low-confidence segments to generate a draft report.

6. The method according to claim 1, characterized in that, The extraction of rare feature identifiers for parameter combinations not covered by the rule base of the key indicator set includes: The key indicator set is compared with the coverage of the diagnostic rule base to generate an uncovered parameter set; Based on the uncovered parameter set, perform parameter correlation constraint analysis to generate a set of mutually exclusive parameter pairs; Based on the mutual exclusion parameters, the set is used to identify co-occurrence patterns and generate co-occurrence anomaly labels; Based on the co-occurrence anomaly labeling, a rare phenomenon identifier is generated by assessing the phenomenon level.

7. The method according to claim 1, characterized in that, The step of extracting unrevised paragraph features from the review feedback to generate a credible representation identifier includes: The revision operation records of the audit feedback are parsed to generate a set of revision locations; Based on the revised location set, perform high-frequency revised clustering to identify and generate system deviation region annotations; Based on the set of revision locations, the unrevised content difference set of the review feedback is extracted to generate a set of unrevised paragraphs; Credible representation identifiers are generated by comparing the credibility of the unrevised paragraph set and the system deviation area annotation.

8. The method according to claim 3, characterized in that, The step of performing parameter heterogeneity analysis on the analysis segment set to generate heterogeneity quantification values ​​includes: Parameter distribution maps are generated by statistically analyzing the parameter distribution within the analysis segment set. Based on the parameter distribution map, identify the parameter transition locations between adjacent segments and generate a transition feature set; Based on the aforementioned transition feature set, transition amplitude is graded to generate local anomaly annotations; Heterogeneity quantization values ​​are generated based on the local anomaly annotations and the parameter distribution map.

9. The method according to claim 6, characterized in that, The step of generating co-occurrence anomaly labels by identifying co-occurrence patterns in the set based on the mutual exclusion parameters includes: The set of mutually exclusive parameter pairs is vectorized and encoded to generate a feature encoding set; Rare feature similarity retrieval is performed on the feature encoding set to generate a candidate feature set; Based on the candidate feature set, a confidence dispersion assessment is performed to generate fuzzy labeling of the diagnostic boundary; Based on the fuzzy labeling of the diagnostic boundary and the candidate sign set, optimal matching is performed to generate co-occurrence anomaly labels.

10. A device for automatically generating and reviewing cardiac ultrasound reports, characterized in that, include: The data acquisition module is used to acquire cardiac measurement parameters and physician examination records, and to construct an ultrasound examination dataset based on the cardiac measurement parameters and physician examination records through structured parsing. The framework generation module is used to perform report specification template matching on the ultrasound examination dataset to generate a report framework and low matching degree intervals, extract case complexity identifiers based on the low matching degree intervals, and perform parameter reference interval deviation magnitude analysis to generate deviation labels according to the report framework and the case complexity identifiers. The draft generation module is used to obtain a descriptive feature set by performing parameter feature association analysis based on the deviation annotation, extract diagnostic-measurement semantic deviation features from the ultrasound examination dataset to generate clinical prompt labels, and generate a report draft based on the descriptive feature set and the clinical prompt labels through a natural language generation model. The rule verification module is used to extract the key indicator set of the report draft, perform diagnostic rule comparison based on the key indicator set to generate verification results, extract rare sign identifiers for parameter combinations not covered by the rule base in the key indicator set, and generate a revision sequence by prioritizing the verification results and the rare sign identifiers. The report output module is used to trigger human-machine collaborative review based on the revision sequence to generate review feedback, extract unrevised paragraph features from the review feedback to generate a credible expression identifier, and integrate the content with the credible expression identifier and the review feedback to generate a cardiac ultrasound diagnostic report.