Construction method of composite trauma medical record intelligent quality control system

By fusing and analyzing the text data of the trauma medical record and the medical image data, semantic correlation features are generated, and directed acyclic graph of clinical events is constructed to identify missing nodes and edges, the problems of potential logical defects and data loss in the trauma medical record are solved, efficient quality control and risk assessment are achieved, and the accuracy and efficiency of medical quality control are improved.

CN120088806AActive Publication Date: 2025-06-03北京紫云智能科技有限公司

Patent Information

Application Number
CN202510558912.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and solve the potential logical defects and data loss problems between text data and medical images in trauma medical records, especially in cross-modal semantic modeling, rule mapping and risk assessment automation.

Method used

By obtaining the text data of the diagnosis and treatment record and medical image data, segmenting it into text fragments and image blocks, generating adversarial samples, calculating word vector representations and visual feature representations, and projecting them to the common feature space to generate semantic correlation features. Then, the semantic correlation features are mapped with the quality control rule base, the directed acyclic graph of clinical events is constructed, the timing dependencies are extracted, the causal effect and counterfactual conditional probability are calculated, the missing nodes and edges are identified, the quality control defect data is generated, and the computing resources are allocated according to the risk scores are generated, and the quality control alarm information is generated.

Benefits of technology

The collaborative processing of cross-modal medical data is realized, the comprehensiveness and accuracy of quality control is improved, the timing dependence and potential missing links between clinical events can be automatically identified, the risk of medical errors is reduced, and the quality control efficiency is improved through automated quality control alarms and resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088806A_ABST
    Figure CN120088806A_ABST
Patent Text Reader

Abstract

The invention provides a construction method of a composite trauma medical record intelligent quality control system, and relates to the technical field of multi-modal data intelligent processing, and the method comprises the steps: obtaining trauma medical record data containing a text record and a medical image, generating an image-text confrontation sample based on a shielding recombination mechanism, and carrying out the reconstruction of the image-text confrontation sample; and mapping the text and image features to a unified feature space by adopting a cross-modal embedding method to generate semantic association features. And the system maps the semantic features with a preset quality control rule, constructs a clinical event atlas, identifies missing information through causal effect analysis and anti-factual reasoning, and generates quality control defect data. And then, constructing a risk propagation model according to the number of defects and criticality, forming a thermal distribution matrix and an evolution path, calculating a medical record risk score, and finally realizing quality control alarm output and data writing. The method has the characteristics of high universality, high timeliness and high intelligence, and is suitable for quality control system construction tasks in various complex data environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of intelligent processing of multimodal data, and particularly to a construction method of an intelligent quality control system for composite trauma medical records. Background Art

[0002] With the wide deployment of electronic medical record systems, a large amount of structured diagnosis and treatment records and unstructured image data have been accumulated in the daily operation of medical institutions. In the scenario of trauma cases, there is a strong temporal dependence and semantic coupling relationship between text data and medical images, and traditional quality control processes are difficult to effectively identify potential logical defects and data missing problems therein.

[0003] Existing methods usually perform static verification on structured data based on rule templates, unable to dynamically model semantic associations, image regions of interest, and causal chains between events in the fused data, and also lacking the ability to automatically judge the risk level and generate feedback reports. This situation limits the generality and intelligence level of the quality control system in complex scenarios.

[0004] Therefore, there is an urgent need for a multimodal intelligent quality control system that integrates natural language processing, computer vision, and causal reasoning to realize a complete process of cross-modal semantic modeling, rule mapping, and risk assessment automation. Summary of the Invention

[0005] An embodiment of the present invention provides a construction method of an intelligent quality control system for composite trauma medical records, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiment of the present invention, A construction method of an intelligent quality control system for composite trauma medical records is provided, including: Obtaining trauma medical record data including diagnosis and treatment record text data and medical image data, respectively segmenting them into text fragments and image blocks, generating adversarial samples by randomly masking and recombining the text fragments and image blocks, calculating the word vector representation of the text fragments and the visual feature representation of the image blocks in the adversarial samples, and projecting them into a common feature space to generate semantic association features; Mapping the semantic association features to the rule items in a preset quality control rule library to obtain rule association results, constructing a directed acyclic graph of clinical events based on the rule association results, extracting the temporal dependence relationship between events, and identifying the missing nodes and edges in the directed acyclic graph of clinical events by calculating the local causal effect and counterfactual conditional probability of clinical event nodes to generate quality control defect data; Calculate the medical record risk score based on the number and importance of missing nodes in the quality control defect data, allocate computing resources based on the medical record risk score, generate a quality control warning message when the medical record risk score exceeds the preset risk threshold, write the quality control defect data and the quality control warning message into the quality control cache, generate a quality control analysis report, send the quality control analysis report to the clinical terminal, and store the quality control defect data, the quality control warning message, and the quality control analysis report in the quality control database.

[0007] In an alternative embodiment, Obtain trauma medical record data including diagnostic record text data and medical image data, respectively segment them into text fragments and image patches, and generate adversarial samples by randomly masking and recombining the text fragments and image patches, including: Use a sliding window to segment the diagnostic record text data to generate multiple text fragments, and the text fragments contain diagnostic description information and treatment record information; Construct an attention heat map for the medical image data, calculate the attention scores of the lesion regions according to the attention heat map; determine the target regions based on the attention scores, set a cropping window centered on the target regions, and generate multiple image patches; Count the occurrence frequencies of medical terms in the text fragments, select the medical term with the highest occurrence frequency for masking to generate text masking data; at the same time, mask the non-target regions in the image patches on the premise of keeping the target regions intact to generate image masking data; Pair the text masking data and the image masking data in time series to construct original sample pairs; recombine the text masking data and the image masking data of adjacent time series in the original sample pairs to generate recombined sample pairs, and merge the original sample pairs and the recombined sample pairs to form the final adversarial samples.

[0008] In an alternative embodiment, Calculate the word vector representation of the text fragments and the visual feature representation of the image patches in the adversarial samples, and project them into a common feature space to generate semantic association features, including: Use a pre-trained language model to extract features from the text fragments to obtain word vectors; construct a query matrix and a key-value matrix based on the word vectors; multiply the query matrix by the key-value matrix to obtain attention weights; weight the attention weights with the word vectors to obtain text semantic features; Perform convolutional feature extraction on the image patches to obtain global visual features, perform multi-scale pooling operations on the global visual features to obtain multi-scale visual features, and perform feature fusion on the global visual features and the multi-scale visual features to obtain image semantic features; Calculate the mean vector and covariance matrix of the text semantic features to obtain the text distribution parameters, calculate the mean vector and covariance matrix of the image semantic features to obtain the image distribution parameters, construct a multi-layer perceptron as a feature projection network based on the text distribution parameters and the image distribution parameters, and perform feature projection on the text semantic features and the image semantic features to obtain text projection features and image projection features; Construct a set of sample pairs, including positive sample pairs composed of text projection features and image projection features corresponding in time series, and negative sample pairs composed of text projection features and image projection features not corresponding in time series. Calculate the positive sample similarity of the positive sample pairs and the negative sample similarity of the negative sample pairs respectively, and at the same time divide the text projection features and the image projection features into regions and calculate the local similarity; Combine the positive sample similarity, the negative sample similarity and the local similarity to construct a multi-granularity contrast loss function, and iteratively optimize the multi-granularity contrast loss function based on the gradient descent method to obtain semantic association features.

[0009] In an alternative embodiment, Construct a directed acyclic graph of clinical events based on the rule association results, extract the temporal dependence relationship between events, and identify the missing nodes and edges in the directed acyclic graph of clinical events by calculating the local causal effect and counterfactual conditional probability of the clinical event nodes. The generated quality control defect data includes: Construct the conditional events in the rule association results as parent nodes and the conclusion events as child nodes, obtain the time information of the parent nodes and the child nodes, determine the temporal dependence relationship between the parent nodes and the child nodes based on the time information, calculate the conditional probability value between the parent nodes and the child nodes, and combine the parent nodes, the child nodes, the temporal dependence relationship and the conditional probability value to construct a directed acyclic graph of clinical events; Extract event node pairs from the directed acyclic graph of clinical events, obtain the local causal effect value by performing intervention calculations on the event node pairs, extract the covariate information of the event node pairs, and calculate the counterfactual conditional probability value of the event node pairs based on the covariate information; Determine the events with local causal effect values exceeding the preset causal threshold but without corresponding nodes in the directed acyclic graph of clinical events as missing nodes; determine the event relationships with counterfactual conditional probability values exceeding the preset probability threshold but without corresponding edges in the directed acyclic graph of clinical events as missing edges, and generate quality control defect data based on the missing nodes and missing edges.

[0010] In an alternative embodiment, Extract event node pairs from the directed acyclic graph of the clinical events, calculate the local causal effect value by performing intervention calculation on the event node pairs, extract the covariate information of the event node pairs, and calculate the counterfactual conditional probability value of the event node pairs based on the covariate information, including: Extract the duration information of the events from the clinical data, statistically calculate the duration distribution of each event, and determine the core impact duration of each event based on the duration distribution; Calculate the basic impact intensity of the event based on the core impact duration, substitute the time difference between the event occurrence time and the current time into the exponential decay function to obtain the time decay value, determine whether the time decay value exceeds the core impact duration to obtain the impact decay value, and multiply the basic impact intensity, the time decay value, and the impact decay value to obtain the event impact domain value; Extract event node pairs from the directed acyclic graph of the clinical events, perform an intervention operation on the source event in the event node pair, calculate the conditional probabilities of the target event before and after the intervention of the source event respectively, multiply the conditional probabilities by the event impact domain value and perform cumulative integration in the time dimension to obtain the local causal effect value of the event node pair; Extract the covariate information of the event node pair, calculate the impact domain value of the covariate information based on the core impact duration and the event impact domain value, and multiply the impact domain value of the covariate information by the event impact domain value of the event node pair to obtain the impact evaluation value of the covariate information; Based on the impact evaluation value of the covariate information, weight the conditional probability of the event node pair, multiply the weighted conditional probability by the event impact domain value and integrate to obtain the counterfactual conditional probability value of the event node pair.

[0011] In an alternative embodiment, Calculate the medical record risk score according to the number and importance of the missing nodes in the quality control defect data, allocate computing resources based on the medical record risk score, and generate a quality control warning message when the medical record risk score exceeds the preset risk threshold, including: Extract the medical decision points and the sequence of diagnosis and treatment behaviors from the historical medical record data, calculate the influence probability between the medical decision points and the sequence of diagnosis and treatment behaviors, and construct a clinical event causal chain based on the influence probability; Extract the missing node information from the quality control defect data, obtain the target node corresponding to the missing node in the standard clinical pathway, calculate the data integrity difference value between the missing node and the target node, calculate the node deviation coefficient based on the difference value, analyze the cross-action intensity between the missing node and the nodes on the clinical event causal chain according to the node deviation coefficient, and recursively propagate the cross-action intensity to obtain the risk propagation value; Construct a multi-dimensional heat distribution matrix based on the risk propagation value, calculate the time-series sampling difference of the risk propagation value to obtain the risk evolution rate, calculate the heat diffusion coefficient of the multi-dimensional heat distribution matrix according to the risk evolution rate, and use the heat aggregation value of the multi-dimensional heat distribution matrix as the medical record risk score; Predict the heat diffusion area according to the heat diffusion coefficient, allocate computing resources based on the heat aggregation value, and when the medical record risk score exceeds the preset risk threshold, generate a quality control warning message based on the heat transfer path of the multi-dimensional heat distribution matrix.

[0012] In an alternative embodiment, Constructing a multi-dimensional heat distribution matrix based on the risk propagation value, calculating the time-series sampling difference of the risk propagation value to obtain the risk evolution rate, calculating the heat diffusion coefficient of the multi-dimensional heat distribution matrix according to the risk evolution rate, and using the heat aggregation value of the multi-dimensional heat distribution matrix as the medical record risk score includes: Construct an initial heat distribution matrix by correlating the risk propagation values in time series and spatial distribution, calculate the propagation attenuation of the risk propagation value between medical treatment links to obtain the heat transfer coefficient, and fuse the heat transfer coefficient with the initial heat distribution matrix to construct a multi-dimensional heat distribution matrix; Sample the risk propagation value at regular intervals, calculate the difference between the risk propagation values at adjacent sampling times to obtain the risk evolution rate, and calculate the heat diffusion coefficient of the multi-dimensional heat distribution matrix according to the time-series correlation between the risk evolution rate and the medical treatment links; Analyze the risk sensitivity of different types of medical behaviors based on a medical decision tree, convert the risk sensitivity into a risk weight coefficient, and redistribute the heat in the multi-dimensional heat distribution matrix according to the risk weight coefficient and the heat diffusion coefficient; Perform a summation operation on the redistributed multi-dimensional heat distribution matrix according to each dimension to obtain the heat aggregation value, and use the heat aggregation value as the medical record risk score.

[0013] In the second aspect of the embodiments of the present invention, Provide an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0014] In the third aspect of the embodiments of the present invention, Provide a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0015] In this embodiment, by fusing and analyzing the diagnosis and treatment record text data and medical image data, a composite trauma medical record intelligent quality control system is constructed, realizing the collaborative processing of multi-modal medical data, improving the comprehensiveness and accuracy of quality control, and being able to capture potential problems that are difficult to discover by traditional single-modal quality control systems. By introducing the clinical event directed acyclic graph and causal reasoning mechanism, the temporal dependence relationship and potential missing links between clinical events can be automatically identified. It can not only discover the surface data missing problems, but also deeply dig out the breakpoints in the clinical logic chain, improving the depth and precision of quality control and effectively reducing the risk of medical errors. Based on the medical record risk score, intelligent allocation of computing resources and automatic triggering of quality control alarms are realized. At the same time, a structured quality control analysis report is generated and pushed to the clinical terminal in real time, improving the quality control efficiency, reducing the manual review cost, providing a traceable and quantifiable quality control management system for medical institutions, and helping to continuously improve the medical quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flowchart of a method for constructing a composite trauma medical record intelligent quality control system according to an embodiment of the present invention; Figure 2 It is a bar chart of medical term frequency statistics according to an embodiment of the present invention; Figure 3 It is a heat map of quality control defect data of clinical pathways according to an embodiment of the present invention; Figure 4 It is a heat map of risk propagation value distribution according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0019] Figure 1 It is a schematic flowchart of a method for constructing a composite trauma medical record intelligent quality control system according to an embodiment of the present invention, as Figure 1 shown, the method includes: Obtain trauma medical record data including diagnosis and treatment record text data and medical image data, respectively segment them into text fragments and image patches, generate adversarial samples by randomly masking and recombining the text fragments and image patches, calculate the word vector representation of the text fragments and the visual feature representation of the image patches in the adversarial samples, and project them into a common feature space to generate semantically associated features; Map the semantically associated features to the rule items in the preset quality control rule library to obtain rule association results, construct a directed acyclic graph of clinical events based on the rule association results, extract the temporal dependence relationship between events, and identify the missing nodes and edges in the directed acyclic graph of clinical events by calculating the local causal effect and counterfactual conditional probability of clinical event nodes to generate quality control defect data; Calculate the medical record risk score according to the number and importance of the missing nodes in the quality control defect data, allocate computing resources based on the medical record risk score, generate a quality control warning message when the medical record risk score exceeds the preset risk threshold, write the quality control defect data and the quality control warning message into the quality control cache, generate a quality control analysis report, send the quality control analysis report to the clinical terminal, and store the quality control defect data, the quality control warning message and the quality control analysis report in the quality control database.

[0020] In an alternative implementation, obtaining trauma medical record data including diagnosis and treatment record text data and medical image data, and respectively segmenting them into text fragments and image patches, generating adversarial samples by randomly masking and recombining the text fragments and image patches includes: Use a sliding window to segment the diagnosis and treatment record text data to generate multiple text fragments, and the text fragments contain diagnostic description information and treatment record information; Construct an attention heat map for the medical image data, calculate the attention score of the lesion area according to the attention heat map; determine the target area based on the attention score, set a cropping window centered on the target area, and generate multiple image patches; Count the occurrence frequency of medical terms in the text fragments, select the medical term with the highest occurrence frequency for masking to generate text masking data; at the same time, mask the non-target area in the image patches on the premise of keeping the target area intact to generate image masking data; Pair the text masking data and the image masking data in time series to construct an original sample pair; recombine the text masking data and the image masking data at adjacent time series in the original sample pair to generate a recombined sample pair, and merge the original sample pair and the recombined sample pair to form the final adversarial sample.

[0021] Exemplarily, trauma medical record data is obtained, which includes two main parts: diagnostic record text data and medical image data. The diagnostic record text data includes written records such as the patient's symptom description, physical examination results, diagnostic opinions, and treatment plans. The medical image data includes medical imaging materials such as X-rays, CT scans, and magnetic resonance imaging of the trauma site.

[0022] For the processing of the diagnostic record text data, a sliding window approach is used for segmentation. The sliding window is a text segmentation method that gradually moves a window of a fixed size over the text, generating a text segment each time it moves. The window size can be adjusted according to actual needs, for example, set to 50 characters. The step size of each window movement can also be flexibly set, usually half of the window size, which can ensure a certain overlap between adjacent text segments and maintain semantic coherence. After processing with the sliding window, multiple text segments containing diagnostic description information and treatment record information are obtained.

[0023] For the processing of the medical image data, an attention heatmap needs to be constructed first. The attention heatmap is a technique for highlighting important regions in an image, which calculates the importance of each region in the image through a deep learning model. Specifically, a convolutional neural network can be used to extract image features, and then the attention weights of each pixel position are calculated by gradient weighting to generate the heatmap. The warmer the color in the heatmap, the more important the region is and the more likely it is to be a lesion area.

[0024] Calculate the attention score of the lesion area based on the attention heatmap. The attention score reflects the likelihood that a certain region is judged to be a lesion. The calculation method is to sum and normalize the attention weights of all pixels in the region. Set a threshold. When the attention score of a certain region exceeds the threshold, it is judged as the target region. Set a cropping window centered on the target region. The window size should be determined according to the actual size of the target region, usually set to expand a certain number of pixels outward from the target region boundary. By sliding the cropping window on the image, multiple image patches are generated.

[0025] Process the text segments and count the number of occurrences of each medical term in all text segments. Medical terms include professional vocabulary such as disease names, symptom descriptions, examination items, and treatment methods. Sort these terms in descending order of frequency and select several terms with the highest frequency for masking. The masking method is to replace the selected terms with special markers, such as "[MASK]". The generated text masking data retains the overall structure of the original text, but the key information is masked.

[0026] At the same time, perform masking processing on the image blocks. On the premise of ensuring the integrity of the target area, mask the non-target area. The masking method can be to set the pixel values of the non-target area to zero, or perform Gaussian blur processing, or fill in a fixed background color. The generated image masking data highlights the lesion area, while the information in other areas is suppressed.

[0027] Pair the text masking data and the image masking data in chronological order to construct the original sample pairs. The chronological relationship refers to the chronological order of the acquisition times of the diagnosis and treatment records and the medical images. Then, recombine the text masking data and the image masking data with adjacent chronologies in the original sample pairs to generate recombined sample pairs. Finally, merge the original sample pairs and the recombined sample pairs to form the final adversarial samples.

[0028] Figure 2 For the bar chart of the medical term frequency statistics in the embodiment of the present invention, as Figure 2 shown, this figure shows the statistical results of the occurrence frequencies of medical terms in the trauma medical record text. The horizontal axis represents different medical term categories, and the vertical axis represents the frequency percentage (%). The technical solution of the present invention adopts a context-sensitive term recognition method, which can more accurately identify and count the term frequencies. From the data in the figure, the recognition frequency of the "fracture" term in the technical solution of the present invention is 28.7%, higher than 23.4% of the traditional word frequency statistics; the recognition rate of the "hematoma" term is 19.3%, and the traditional method is 16.8%; the "contusion" term is 15.6%, and the traditional method is 14.2%; the "laceration" term is 12.8%, and the traditional method is 11.3%; the "nerve injury" term is 10.2%, and the traditional method is 8.7%; the "vascular injury" term is 8.9%, and the traditional method is 7.4%; the "inflammation" term is 4.5%, and the traditional method is 9.3%. By considering the context relationship of terms, the technical solution of the present invention improves the recognition accuracy of key diagnostic terms by an average of 4.3 percentage points, while reducing the over-recognition of general descriptive terms (such as "inflammation" etc.), enabling the masking strategy to more accurately target core medical terms, thereby more effectively retaining the key information of the medical record when generating text masking data, and the semantic integrity of the masked text is improved by 23.6%.

[0029] In this embodiment, by processing and reorganizing text and image data, key information in trauma medical records can be effectively identified and extracted. The sliding window segmentation method ensures the coherence of text semantics, making the segmented text fragments still have complete clinical diagnostic value. The construction of the attention heatmap helps to accurately locate the lesion areas in medical images, improving the accuracy of lesion recognition. The medical term frequency statistics and masking strategy highlight the important information in clinical diagnosis, and at the same time, by masking non-critical areas, the impact of interference information is reduced. The method of temporal pairing and recombination of text and images to generate adversarial samples enhances the model's ability to understand different types of medical record data, improving the accuracy and reliability of clinical diagnosis. This processing method can also help medical staff quickly locate key diagnostic information and improve the diagnosis and treatment efficiency. By generating high-quality adversarial samples, this solution can also be used for the training and verification of medical AI models, improving the robustness and generalization ability of the models.

[0030] In an alternative embodiment, the word vector representation of the text fragment and the visual feature representation of the image patch in the adversarial sample are calculated and projected into a common feature space to generate semantic association features, including: Using a pre-trained language model to extract features from the text fragment to obtain word vectors; constructing a query matrix and a key-value matrix based on the word vectors; multiplying the query matrix by the key-value matrix to obtain attention weights; weighting the attention weights with the word vectors to obtain text semantic features; Performing convolutional feature extraction on the image patch to obtain global visual features, performing multi-scale pooling operations on the global visual features to obtain multi-scale visual features, and fusing the global visual features and the multi-scale visual features to obtain image semantic features; Calculating the mean vector and covariance matrix of the text semantic features to obtain text distribution parameters, calculating the mean vector and covariance matrix of the image semantic features to obtain image distribution parameters, constructing a multi-layer perceptron as a feature projection network based on the text distribution parameters and the image distribution parameters, and performing feature projection on the text semantic features and the image semantic features to obtain text projection features and image projection features; Constructing a set of sample pairs, including positive sample pairs composed of temporally corresponding text projection features and image projection features, and negative sample pairs composed of temporally non-corresponding text projection features and image projection features, respectively calculating the positive sample similarity of the positive sample pairs and the negative sample similarity of the negative sample pairs, and at the same time dividing the text projection features and the image projection features into regions and calculating the local similarity; Combining the positive sample similarity, the negative sample similarity, and the local similarity to construct a multi-granularity contrast loss function, and iteratively optimizing the multi-granularity contrast loss function based on the gradient descent method to obtain semantic association features.

[0031] Exemplarily, feature extraction is performed on the text fragment to obtain a word vector representation. Specifically, a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model is used as the feature extractor. This model is pre-trained with large-scale medical text data and can well understand the professional terms and expressions in the medical field. For example, for a text fragment like "The patient has obvious tenderness in the middle segment of the right tibia and local swelling", the model will convert it into a word vector representation with a fixed dimension.

[0032] Based on the obtained word vectors, a query matrix and a key-value matrix are constructed. The query matrix represents the text content that needs to be focused on currently, and the key-value matrix stores the key information in the text. For example, when processing the sentence "The preliminary diagnosis is a fracture of the right tibia", the query matrix will focus on the action of "diagnosis", and the key-value matrix will contain the diagnosis result of "fracture of the right tibia". Multiply the query matrix by the key-value matrix to obtain the attention weights reflecting the correlation degree between different words. The larger the attention weight, the more important the corresponding word is for understanding the text semantics.

[0033] The attention weights are weighted with the word vectors to obtain the text semantic features. This step is equivalent to reassigning weights to the original word vectors, making the important information prominent. For example, when processing treatment records, words related to surgical operations will obtain higher weights. The text semantic features obtained in this way can better reflect the core meaning of the text.

[0034] For the processing of image patches, first, a convolutional neural network is used for feature extraction to obtain global visual features. Network structures such as ResNet, which perform excellently in the field of medical images, can be selected. For example, when processing a fracture X-ray, the network will extract key visual features such as fracture lines and bone density.

[0035] A multi-scale pooling operation is performed on the extracted global visual features to obtain multi-scale visual features. Multi-scale pooling means sampling the feature map using pooling windows of different sizes, which can capture image information at different scales. For example, for a fracture image, large-scale features reflect the overall bone structure, and small-scale features contain fracture details.

[0036] The global visual features and the multi-scale visual features are fused to obtain the image semantic features. Feature fusion can be performed in the way of feature concatenation or weighted summation. The fused features contain both the overall information of the image and retain local details. For example, the fused features can express both the location of the fracture and the specific shape of the fracture at the same time.

[0037] Calculate the mean vector and covariance matrix of the text semantic features to obtain the text distribution parameters. The mean vector reflects the overall trend of the features, and the covariance matrix reflects the correlation between the dimensions of the features. Calculate the distribution parameters of the image semantic features in the same way. These parameters describe the statistical characteristics of the text and image features.

[0038] Construct a multi-layer perceptron based on the distribution parameters of the text and image as the feature projection network. The role of this network is to map the text features and image features to the same feature space. The network structure can include multiple fully connected layers, and use non-linear activation functions to enhance the feature expression ability. Project the text semantic features and image semantic features through this network to obtain text projection features and image projection features with the same dimension.

[0039] Construct a set of sample pairs, including positive sample pairs and negative sample pairs. The positive sample pairs are composed of text and image features corresponding in time sequence, such as the medical record description at the time of admission and the X-ray film features taken at that time. The negative sample pairs are feature combinations that do not correspond in time sequence, such as the admission record and the imaging features at the time of discharge. Calculate the similarity of the positive and negative sample pairs respectively, and the similarity can be measured by the inner product of the feature vectors.

[0040] Perform regional partitioning on the text projection features and image projection features, and calculate the local similarity. Regional partitioning is to divide the feature vector into multiple sub-vectors and calculate the similarity respectively. This can measure the matching degree of the features more carefully. For example, the matching situations of different aspects such as symptom description and treatment plan can be concerned respectively.

[0041] Combine the positive sample similarity, negative sample similarity and local similarity to construct a multi-granularity contrast loss function. Specifically, design a loss function that comprehensively considers the global similarity and local similarity, so that the similarity of the positive sample pairs increases and the similarity of the negative sample pairs decreases. Iteratively optimize the multi-granularity contrast loss function based on the gradient descent method. Specifically, use the Adam optimizer, set the learning rate to 0.0001, the batch size to 64, and the number of training epochs to 50. Update the model parameters through the backpropagation algorithm, and finally obtain a feature representation that can effectively express the semantic association between the text and the image, that is, the semantic association feature.

[0042] In practical applications, if a text describing "the patient's right ankle was swollen and had limited mobility after a fall, and the X-ray showed a fracture of the right calcaneus" and the corresponding X-ray of the calcaneus fracture are input, the system first extracts the word vector representation of the text through a pre-trained language model, and highlights key information such as "fracture of the right calcaneus" through the attention mechanism. At the same time, convolutional feature extraction and multi-scale pooling are performed on the X-ray to obtain a visual representation containing fracture features. After projecting the two types of features into the same space through a feature projection network, their semantic similarity can be accurately calculated to verify the consistency between the text description and the imaging manifestation. In this way, the system can effectively learn the semantic association between medical texts and images, providing support for clinical diagnosis.

[0043] In this embodiment, by using a pre-trained language model and a deep convolutional network to extract text and image features respectively, and combining the attention mechanism to highlight key information, the semantic content of medical texts and imaging data can be accurately captured. The multi-scale pooling operation enables the system to simultaneously focus on the local details and overall structure of the image, improving the comprehensiveness of feature extraction. The feature projection network maps features of different modalities into a unified semantic space, achieving effective alignment of text and image features. The multi-granularity contrast learning strategy takes into account both the overall semantic similarity and the matching degree of local regions, enhancing the accuracy of feature expression. This technical solution can effectively learn the semantic association relationship between medical texts and images, improving the accuracy and efficiency of clinical diagnosis. At the same time, this solution has good scalability and versatility, can adapt to different types of medical data, and provides strong support for intelligent medical diagnosis systems.

[0044] In an alternative implementation, a directed acyclic graph of clinical events is constructed based on the rule association results, the temporal dependence relationship between events is extracted, and by calculating the local causal effect and counterfactual conditional probability of clinical event nodes, the missing nodes and edges in the directed acyclic graph of clinical events are identified, and the quality control defect data generated includes: Construct the conditional events in the rule association results as parent nodes and the conclusion events as child nodes, obtain the time information of the parent nodes and child nodes, determine the temporal dependence relationship between the parent nodes and child nodes based on the time information, calculate the conditional probability value between the parent nodes and child nodes, and combine the parent nodes, child nodes, temporal dependence relationship and the conditional probability value to construct a directed acyclic graph of clinical events; Extract event node pairs from the directed acyclic graph of clinical events, obtain the local causal effect value through intervention calculation on the event node pairs, extract the covariate information of the event node pairs, and calculate the counterfactual conditional probability value of the event node pairs based on the covariate information; Determine the events with local causal effect values exceeding the preset causal threshold but without corresponding nodes in the clinical event directed acyclic graph as missing nodes; determine the event relationships with counterfactual conditional probability values exceeding the preset probability threshold but without corresponding edges in the clinical event directed acyclic graph as missing edges, and generate quality control defect data based on the missing nodes and missing edges.

[0045] This embodiment provides a method for constructing a clinical event directed acyclic graph based on rule association results and extracting temporal dependence relationships. By calculating the local causal effect and counterfactual conditional probability of clinical event nodes, missing nodes and edges in the clinical event directed acyclic graph are identified, and quality control defect data is generated. Exemplarily, first construct the conditional events in the rule association results as parent nodes and the conclusion events as child nodes. For example, for the rule "if the patient's body temperature drops after receiving antibiotic treatment", "antibiotic treatment" can be set as the parent node and "body temperature drops" as the child node. Specifically, extract patient clinical event records from the medical database, and each record contains fields such as event ID, event name, and occurrence time. For the above example, there may be the following records: Event 1 (ID: 1001, name: antibiotic treatment, time: 2023-01-15 09:30) and Event 2 (ID: 1002, name: body temperature drops, time: 2023-01-15 14:20).

[0046] Obtain the time information of the parent node and the child node, and determine the temporal dependence relationship between the parent node and the child node based on the time information. In the above example, by comparing the time of "antibiotic treatment" (09:30) and the time of "body temperature drops" (14:20), it is found that "antibiotic treatment" occurs before "body temperature drops", which conforms to the causal logic, so a directed edge pointing from "antibiotic treatment" to "body temperature drops" is established. If it is found that the occurrence time of the child node event is earlier than that of the parent node event, it is considered illogical and this directed edge is not established.

[0047] Calculate the conditional probability value between the parent node and the child node. The conditional probability P(B|A) represents the probability that event B occurs under the condition that event A occurs. For example, among all patient samples receiving antibiotic treatment, count the proportion of patients with body temperature drops. If 80 out of 100 patients receiving antibiotic treatment subsequently have body temperature drops, then the conditional probability P(body temperature drops|antibiotic treatment) = 0.8.

[0048] Combine the parent node, child node, temporal dependence relationship, and conditional probability value to construct a clinical event directed acyclic graph. In practical applications, there may be multiple nodes and edges. For example, there may also be a path such as "fever" → "antibiotic treatment" → "body temperature drops" → "discontinue antibiotic", forming a complete directed acyclic graph structure.

[0049] Extract event node pairs from the clinical event directed acyclic graph, such as (antibiotic treatment, body temperature decrease), (fever, antibiotic treatment), etc. Calculate the local causal effect value by performing intervention calculations on the event node pairs. This step is achieved by comparing the outcome differences between the intervention group and the non-intervention group. Specifically, select the patient group that performs a certain intervention and the patient group that does not perform this intervention, and compare the differences in the outcome events between the two groups. For example, compare the difference in body temperature decrease between the patient group receiving antibiotic treatment and the patient group not receiving antibiotic treatment. If the body temperature decrease rate in the treatment group is 85% and in the non-treatment group is 25%, then the local causal effect value is 0.6.

[0050] Extract the covariate information of the event node pairs, such as factors that may affect the outcome, such as the patient's age, gender, underlying diseases, etc. For example, for the event pair (antibiotic treatment, body temperature decrease), the covariates may include the patient age group (0 - 18 years old, 19 - 65 years old, > 65 years old), underlying diseases (heart disease, diabetes, etc.). Calculate the counterfactual conditional probability value of the event node pair based on the covariate information, that is, consider how the outcome would change if the patient attributes were different. For example, calculate the probability of the effect of antibiotic treatment on body temperature decrease after controlling covariates such as age and underlying diseases, and the counterfactual conditional probability value may be 0.75.

[0051] Identify the events with local causal effect values exceeding the preset causal threshold but without corresponding nodes in the clinical event directed acyclic graph as missing nodes. Assume the preset causal threshold is 0.4. If it is found that the local causal effect of "antibiotic dose adjustment" on "body temperature decrease" is 0.5, but there is no node of "antibiotic dose adjustment" in the existing graph, then it is identified as a missing node.

[0052] Identify the event relationship with a counterfactual conditional probability value exceeding the preset probability threshold but without a corresponding edge in the clinical event directed acyclic graph as a missing edge. Assume the preset probability threshold is 0.6. If it is found that the counterfactual conditional probability of "antibiotic treatment" on "renal function change" after controlling covariates is 0.7, but there is no edge from "antibiotic treatment" to "renal function change" in the existing graph, then this relationship is identified as a missing edge.

[0053] Generate quality control defect data based on the missing nodes and missing edges. For example, generate the following quality control defect report: Missing node "antibiotic dose adjustment" (local causal effect value: 0.5, recommended to be added between "antibiotic treatment" and "body temperature decrease"); Missing edge "antibiotic treatment → renal function change" (counterfactual conditional probability value: 0.7, recommended to add this monitoring relationship). Medical institutions can improve the clinical pathway and monitoring specifications based on these quality control defect data to improve the medical quality.

[0054] Figure 3This is the heat map of the quality control defect data in the clinical pathway of the embodiment of the present invention. As Figure 3 shown, this figure intuitively shows the degree of quality control defects between each event node in the clinical pathway. The horizontal and vertical coordinate axes respectively represent 11 clinical event nodes from "admission examination" to "life guidance". The gray scale of each cell in the matrix reflects the possible degree of defects between the corresponding event nodes, and the values range from 0.0 to 1.0. It can be clearly seen from the figure that the defect value from "admission examination" to "rehabilitation treatment" is 0.5, and the defect values from "admission examination" to "discharge assessment" and "follow-up plan" are as high as 0.7 and 0.8 respectively, indicating serious data defects. In addition, the two cells marked as "missing" on the right (preliminary diagnosis → life guidance, postoperative recovery → life guidance) indicate that there should be an association between these event nodes but it is completely missing in the current clinical pathway. Particularly noteworthy is the two "missing edges" marked in the figure (drug treatment → surgical intervention, rehabilitation treatment → follow-up plan), which means that according to causal analysis, there should be a direct relationship between these nodes but it is not correctly recorded. The heat map as a whole shows that 36.4% of the node pairs have mild defects (0.1 - 0.3), 14.5% have moderate defects (0.4 - 0.6), and 4.5% have serious defects (0.7 - 1.0), which provides a clear improvement direction for the optimization and quality control of the clinical pathway.

[0055] In this embodiment, by constructing a directed acyclic graph of clinical events, the temporal dependence relationship and conditional probability association between clinical events are effectively established. By performing intervention calculations to obtain local causal effect values and combining covariate information to calculate counterfactual conditional probabilities, potential event omissions and relationship omissions in clinical practice can be accurately identified. This quality control method based on causal reasoning can effectively discover hidden defects in the clinical pathway and avoid the omission of important clinical events or key treatment steps. At the same time, this solution can automatically identify and supplement quality control defects in the clinical diagnosis and treatment process, improving the accuracy and efficiency of medical quality control. By analyzing the causal relationship and conditional probability between event nodes, the system can warn of potential medical risks and provide support for clinical decision-making. This method also has strong interpretability, which helps medical staff understand and improve the clinical diagnosis and treatment process and improve the overall quality of medical services.

[0056] In an alternative embodiment, extracting event node pairs from the directed acyclic graph of clinical events, obtaining local causal effect values by performing intervention calculations on the event node pairs, extracting covariate information of the event node pairs, and calculating the counterfactual conditional probability values of the event node pairs based on the covariate information includes: Extracting the duration information of events from clinical data, statistically calculating the duration distribution of each event, and determining the core impact duration of each event based on the duration distribution; Calculate the basic impact intensity of an event based on the core impact duration. Substitute the time difference between the event occurrence time and the current time into the exponential decay function to obtain the time decay value. Determine whether the time decay value exceeds the core impact duration to obtain the impact decay value. Multiply the basic impact intensity, the time decay value, and the impact decay value to obtain the event impact threshold. Extract event node pairs from the clinical event directed acyclic graph. Apply an intervention operation to the source event in the event node pair, and calculate the conditional probabilities of the target event before and after the intervention of the source event respectively. Multiply the conditional probabilities by the event impact threshold and accumulate the integral in the time dimension to obtain the local causal effect value of the event node pair. Extract the covariate information of the event node pair. Calculate the impact threshold of the covariate information based on the core impact duration and the event impact threshold. Multiply the impact threshold of the covariate information by the event impact threshold of the event node pair to obtain the impact evaluation value of the covariate information. Based on the impact evaluation value of the covariate information, weight the conditional probability of the event node pair. Multiply the weighted conditional probability by the event impact threshold and integrate to obtain the counterfactual conditional probability value of the event node pair.

[0057] Exemplarily, extract the duration information of an event from clinical data. Specifically, for each clinical event, count its duration in different cases to obtain the duration distribution of the event. For example, for the event of "fever", the following distribution may be obtained: 40% for 1 - 3 days, 50% for 4 - 7 days, and 10% for more than 7 days. Based on this distribution, determine the core impact duration of the event, which can be taken as the median or mode of the distribution. In this example, 4 - 7 days can be taken as the core impact duration of fever.

[0058] Calculate the basic impact intensity of an event based on the core impact duration. The following method can be used: Substitute the core impact duration into the preset impact intensity function to obtain a value between 0 and 1 as the basic impact intensity. For example, for a fever event with a core impact duration of 5 days, its basic impact intensity may be 0.7. Then calculate the time decay value of the event. Specifically, take the time difference between the event occurrence time and the current time and substitute it into the exponential decay function. This function can be expressed as: with e as the base and the exponent being the negative product of the time difference and the preset decay coefficient. For example, if a fever event occurred 3 days ago and the decay coefficient is 0.2, then its time decay value is e to the power of -0.6, approximately 0.549.

[0059] Determine whether the time decay value exceeds the core influence duration to obtain the influence decay value. If the time decay value does not exceed the core influence duration, the influence decay value is 1; otherwise, it is a value less than 1, which can be obtained by substituting the exceeded part into another faster decay function. Multiply the base influence intensity, the time decay value, and the influence decay value to obtain the event influence domain value. For example, for the above fever event, its event influence domain value is 0.7×0.549×1 = 0.384.

[0060] Extract event node pairs from the clinical event directed acyclic graph. For example, a pair of event nodes such as "fever" and "elevated white blood cells" may be extracted. Apply an intervention operation to the source event (such as "fever"), and calculate the conditional probabilities of the target event (such as "elevated white blood cells") before and after the intervention respectively.

[0061] The conditional probability before the intervention can be obtained by statistically analyzing the proportion of elevated white blood cells after fever in the clinical data. The conditional probability after the intervention needs to be estimated through causal inference methods such as do-calculus to estimate the probability of elevated white blood cells when all patients are artificially made to have a fever. Multiply the obtained conditional probability by the event influence domain value and perform cumulative integration in the time dimension. Here, the integration can be understood as: within the entire time range affected by the event, multiply the conditional probability at each time point by the event influence domain value at that time point, and then sum them up. Finally, obtain the local causal effect value of the event node pair.

[0062] Extract the covariate information of the event node pair. Covariates refer to other factors that may simultaneously affect the source event and the target event, such as age, gender, etc. For each covariate, calculate the core influence duration and the event influence domain value based on its characteristics. For example, for the covariate "age", its core influence duration may be lifelong, and the event influence domain value may vary with different age groups. The possible values are: children (0 - 14 years old) 0.8, young people (15 - 44 years old) 0.6, middle-aged people (45 - 64 years old) 0.7, elderly people (65 years old and above) 0.9.

[0063] Multiply the influence domain value of the covariate information by the event influence domain value of the event node pair to obtain the influence evaluation value of the covariate information. For example, if the patient is a 70-year-old elderly person, the influence evaluation value of the covariate "age" is 0.9×0.384 = 0.3456.

[0064] Based on the impact evaluation value of covariate information, the conditional probability of the event node pair is weighted. The specific method is as follows: Multiply the impact evaluation value of each covariate by a preset weight, then sum up the weighted results of all covariates to obtain a total weight factor. Use this weight factor to adjust the original conditional probability. Multiply the weighted conditional probability by the event impact domain value and integrate to obtain the counterfactual conditional probability value of the event node pair. This value represents the true causal effect of the source event on the target event after considering the impacts of various covariates.

[0065] Existing clinical event causality analysis usually only considers the chronological order and statistical correlation of event occurrences, ignoring the dynamic decay characteristics of event impacts and complex interference factors in the clinical environment, resulting in inaccurate causality identification. This solution introduces the concept of the core impact duration, combines it with the time decay function to calculate the actual impact intensity of events, which is more in line with the objective law that the impact of medical events gradually weakens in clinical practice. By establishing an event impact domain value evaluation system, the time decay characteristics and impact intensity are organically combined, improving the accuracy of causality identification. When calculating the local causal effect, a cumulative integration method in the time dimension is adopted, fully considering the persistence characteristics of event impacts. At the same time, covariate information is innovatively incorporated into the evaluation system, and through weighted calculation of the impact domain value, the interference factors in the clinical environment are effectively eliminated. This method that comprehensively considers time decay, impact intensity, and environmental factors significantly improves the accuracy and reliability of clinical event causality analysis, providing more accurate decision-making support for medical quality control.

[0066] In an alternative implementation, calculate the medical record risk score according to the number and importance of missing nodes in the quality control defect data, allocate computing resources based on the medical record risk score, and when the medical record risk score exceeds the preset risk threshold, generate quality control warning information including: Extract medical decision points and the sequence of diagnosis and treatment behaviors from historical medical record data, calculate the influence probability between the medical decision points and the sequence of diagnosis and treatment behaviors, and construct a clinical event causal chain based on the influence probability; Extract missing node information from the quality control defect data, obtain the target node corresponding to the missing node in the standard clinical pathway, calculate the difference value of data integrity between the missing node and the target node, calculate the node deviation coefficient based on the difference value, analyze the cross-action intensity between the missing node and the nodes on the clinical event causal chain according to the node deviation coefficient, and recursively propagate the cross-action intensity to obtain the risk propagation value; Construct a multi-dimensional heat distribution matrix based on the risk propagation value, calculate the time-series sampling difference value of the risk propagation value to obtain the risk evolution rate, calculate the heat diffusion coefficient of the multi-dimensional heat distribution matrix according to the risk evolution rate, and use the heat aggregation value of the multi-dimensional heat distribution matrix as the medical record risk score; Predict the heat diffusion region according to the heat diffusion coefficient, allocate computing resources based on the heat aggregation value, and generate a quality control warning message based on the heat transfer path of the multi-dimensional thermal distribution matrix when the medical record risk score exceeds the preset risk threshold.

[0067] Exemplarily, extract medical decision points and treatment behavior sequences from historical medical record data. Specifically, natural language processing techniques can be used to analyze historical medical record texts to identify key medical decision points, such as diagnosis, medication, surgery, etc. At the same time, arrange the treatment behaviors recorded in the medical record in chronological order to form a treatment behavior sequence. For example, for the medical record of a pneumonia patient, decision points and behavior sequences such as "initial diagnosis - chest X-ray examination - antibiotic use - reexamination" may be extracted.

[0068] Calculate the influence probability between medical decision points and treatment behavior sequences. A Bayesian network model can be used, with decision points as parent nodes and treatment behaviors as child nodes, and learn the conditional probabilities between nodes by statistically analyzing a large number of medical record samples. For example, the influence probability of the decision point of "using antibiotics" on the treatment result of "symptom improvement" may be 0.8. Based on these influence probabilities, a clinical event causal chain can be constructed to reflect the causal relationship between decisions and behaviors.

[0069] Extract missing node information from the quality control defect data. Quality control defect data is usually marked by quality control personnel according to specification requirements and contains the problematic parts of the medical record. This method maps these defective parts to the standard clinical pathway and finds the target nodes corresponding to the missing nodes. For example, if "blood routine examination" is missing, find the corresponding "blood routine examination" node in the standard pathway as the target node.

[0070] Then calculate the difference value of data integrity between the missing node and the target node. A scoring standard can be set to quantify the degree of missing. For example, complete missing is recorded as 0 points, partial missing is recorded as 0.5 points, and complete record is recorded as 1 point. By comparing the score of the missing node with the standard score of the target node (usually 1 point), the difference value is obtained. Based on this difference value, the node deviation coefficient can be further calculated. The deviation coefficient can be defined as the difference value divided by the standard score, reflecting the deviation degree between the actual record and the standard requirement.

[0071] Analyze the cross-action intensity between the missing node and the nodes on the clinical event causal chain. The previously constructed causal chain can be used to examine the influence of the missing node on other nodes on the chain. Specifically, the correlation coefficient between the missing node and the causal chain nodes can be calculated. The higher the correlation coefficient, the stronger the cross-action. For example, if the correlation coefficient between the missing node of "blood routine examination" and the causal chain node of "antibiotic use" is 0.7, it indicates a strong cross-action between them.

[0072] By recursively propagating the cross - interaction intensity along the causal chain, the risk propagation value can be obtained. During the propagation process, an attenuation factor can be adopted so that the influence weakens as the propagation distance increases. For example, it can be set that the intensity attenuates by 20% for each step of propagation. In this way, after multiple steps of propagation, the risk propagation value of each node can be obtained.

[0073] Based on the risk propagation value, a multi - dimensional heat distribution matrix is constructed. This matrix can be regarded as a visual representation of the medical record risk. Each dimension of the matrix can correspond to different risk factors, such as examination items, medication conditions, surgical operations, etc. Each element value in the matrix is the risk propagation value at that position.

[0074] To reflect the dynamic change of the risk, the risk evolution rate also needs to be calculated. This can be achieved by sampling the risk propagation value at different time series and calculating the difference between adjacent time points. For example, it can be sampled every 6 hours, and the change in the risk propagation value between two samplings is calculated to obtain the risk evolution rate.

[0075] Based on the risk evolution rate, the heat diffusion coefficient of the multi - dimensional heat distribution matrix can be further calculated. The heat diffusion coefficient reflects how fast the risk spreads between different dimensions. The risk evolution rate can be normalized and used as the heat diffusion coefficient. The faster the rate of a dimension, the larger the diffusion coefficient.

[0076] The heat aggregation value of the multi - dimensional heat distribution matrix can be used as the risk score of the medical record. The heat aggregation value can be obtained by calculating the weighted sum of all element values in the matrix, and the weights can be set according to the importance of different dimensions. For example, a weight of 0.3 may be given to the examination item dimension, 0.5 to the medication condition dimension, and 0.2 to the surgical operation dimension. According to the heat diffusion coefficient, the heat diffusion area can be predicted. Specifically, starting from the high - risk points, according to the magnitude of the diffusion coefficient, the diffusion process of heat in the matrix is simulated to obtain the possible diffusion area. This helps to predict the scope that the risk may affect. Based on the heat aggregation value, that is, the medical record risk score, the computing resources are allocated. Multiple risk levels can be set, such as low, medium, and high levels, which respectively correspond to different resource allocation strategies. For example, low - risk medical records may only require regular review, medium - risk medical records need key attention, and high - risk medical records need to be processed immediately.

[0077] When the medical record risk score exceeds the preset risk threshold, the system will generate quality control warning information. The generation of the warning information is based on the heat transfer path of the multi - dimensional heat distribution matrix. Specifically, starting from the high - risk points, along the direction of heat transfer, the source and possible influence scope of the risk are traced. The warning information should include content such as the risk score, main risk points, possible influence scope, and recommended handling measures.

[0078] Figure 4 This is the heat map of the risk propagation value distribution in the embodiments of the present invention. As Figure 4 shown, this figure shows a multi-dimensional heat distribution matrix constructed based on the node deviation coefficient and the risk propagation path length. Different shades of gray are used in the figure to represent the magnitude of the risk propagation value, and the larger the value, the darker the color. It can be clearly observed from the heat map that when the node deviation coefficient is higher (close to 1.0) and the propagation path length is shorter (close to 0), the risk propagation value reaches the highest, which is 0.92. As the node deviation coefficient decreases and the propagation path length increases, the risk propagation value shows an obvious attenuation trend. When the node deviation coefficient is 0.8 and the propagation path length is 1, the risk propagation value is 0.78; when the node deviation coefficient drops to 0.4 and the propagation path length increases to 3, the risk propagation value drops to 0.24. This attenuation trend conforms to the actual law of medical risk propagation - that is, the risk gradually weakens as the propagation path extends, and the smaller the degree of node deviation, the smaller its contribution to the overall risk. The highest risk area in the figure is concentrated in the upper left corner, indicating that nodes with high deviation coefficients will have the greatest risk impact in short-path propagation. This heat distribution pattern provides important decision-making support for medical quality control, guiding medical institutions to prioritize nodes with high deviation degrees and short chains on key propagation paths, so as to allocate quality control resources more efficiently. By visualizing risks in the form of heat, this technical solution makes complex risk propagation patterns intuitive and easy to understand, effectively improving the accuracy and response speed of medical quality management.

[0079] Through the above steps, this method realizes the accurate evaluation and risk early warning of medical record quality, which helps to improve the efficiency and accuracy of medical quality control. The core advantage of this method lies in combining quality control defects with the clinical decision-making process, and comprehensively evaluating the potential risks of medical records through multi-dimensional risk propagation analysis. At the same time, dynamic risk evolution analysis and resource allocation strategies enable quality control work to be more timely and targeted. This method is not only applicable to the risk assessment of individual medical records, but can also be extended to the quality management system of the entire medical institution, providing strong support for the continuous improvement of medical quality.

[0080] In an alternative embodiment, constructing a multi-dimensional heat distribution matrix based on the risk propagation value, calculating the time-series sampling difference of the risk propagation value to obtain the risk evolution rate, calculating the heat diffusion coefficient of the multi-dimensional heat distribution matrix according to the risk evolution rate, and taking the heat aggregation value of the multi-dimensional heat distribution matrix as the medical record risk score includes: Constructing an initial heat distribution matrix by associating the risk propagation value in time series and spatial distribution, calculating the propagation attenuation of the risk propagation value between diagnosis and treatment links to obtain a heat transfer coefficient, and fusing the heat transfer coefficient with the initial heat distribution matrix to construct a multi-dimensional heat distribution matrix; Timely sample the risk propagation value, calculate the difference in risk propagation values at adjacent sampling times to obtain the risk evolution rate, and calculate the heat diffusion coefficient of the multi-dimensional heat distribution matrix according to the temporal correlation between the risk evolution rate and the diagnosis and treatment process; Analyze the risk sensitivity of different types of medical behaviors based on a medical decision tree, convert the risk sensitivity into a risk weight coefficient, and redistribute the heat in the multi-dimensional heat distribution matrix according to the risk weight coefficient and the heat diffusion coefficient; Perform a summation operation on the redistributed multi-dimensional heat distribution matrix according to each dimension to obtain a heat aggregation value, and use the heat aggregation value as the medical record risk score.

[0081] Exemplarily, the system first constructs an initial heat distribution matrix based on the temporal correlation and spatial distribution of the risk propagation value. For example, for a medical record containing four diagnosis and treatment processes: outpatient, examination, medication, and surgery, a 4-dimensional initial heat distribution matrix can be constructed. Assume that the risk propagation value of the outpatient process is 0.6, the examination process is 0.4, the medication process is 0.7, and the surgery process is 0.8. Then the initial heat distribution matrix can be expressed as [0.6, 0.4, 0.7, 0.8].

[0082] Calculate the propagation attenuation of the risk propagation value between diagnosis and treatment processes to obtain the heat transfer coefficient, and analyze the risk transfer relationship between adjacent diagnosis and treatment processes. For example, the heat transfer coefficient from the outpatient to the examination process may be 0.9, indicating that 90% of the risk will be transferred; the heat transfer coefficient from the examination to the medication process is 0.8; the heat transfer coefficient from the medication to the surgery process is 0.95. These heat transfer coefficients form a transfer matrix [0.9, 0.8, 0.95].

[0083] Fuse the heat transfer coefficient with the initial heat distribution matrix to construct a multi-dimensional heat distribution matrix. The fusion method is to multiply the risk propagation value of each process by its corresponding heat transfer coefficient and consider the correlation between processes. For example, the updated heat value of the outpatient process is 0.6×1 = 0.6, the examination process is 0.4×0.9 + 0.4 = 0.76, the medication process is 0.7×0.8 + 0.7 = 1.26, and the surgery process is 0.8×0.95 + 0.8 = 1.56. The updated multi-dimensional heat distribution matrix is [0.6, 0.76, 1.26, 1.56].

[0084] The system samples the risk propagation value at regular intervals, calculates the risk evolution rate, and determines the heat diffusion coefficient. In actual operation, the system can sample the risk propagation value once per hour. Suppose the risk propagation values sampled at time T1 are [0.6, 0.76, 1.26, 1.56], and the risk propagation values sampled at time T2 are [0.65, 0.8, 1.3, 1.6]. Then the risk evolution rate is [0.05, 0.04, 0.04, 0.04] per hour.

[0085] The system calculates the heat diffusion coefficient of the multi-dimensional thermal distribution matrix based on the temporal correlation between the risk evolution rate and the diagnosis and treatment processes. For example, for the outpatient process with a relatively high risk evolution rate, its heat diffusion coefficient can be set to 0.3; for other processes, it can be set to [0.3, 0.24, 0.24, 0.24] according to the ratio of the risk evolution rate. Then the system analyzes the risk sensitivity of different types of medical behaviors based on the medical decision tree. In actual implementation, the system constructs a medical decision tree by analyzing the occurrence frequencies of risk events of different medical behaviors in historical cases. For example, through decision tree analysis, it can be known that the risk sensitivity of the surgical process is the highest, at 0.4; the risk sensitivity of the medication process is the second highest, at 0.3; the risk sensitivity of the examination process is 0.2; and the risk sensitivity of the outpatient process is the lowest, at 0.1.

[0086] The system converts the risk sensitivity into a risk weight coefficient. The conversion method is to directly use the sensitivity value as the weight coefficient, that is, [0.1, 0.2, 0.3, 0.4]. The system redistributes the heat in the multi-dimensional thermal distribution matrix according to the risk weight coefficient and the heat diffusion coefficient. The specific operation is to first multiply the original thermal value by the heat diffusion coefficient to obtain the diffusion amount; then the diffusion amount is weighted and distributed to each process according to the risk weight coefficient.

[0087] Taking the outpatient process as an example, its diffusion amount is 0.6×0.3 = 0.18; this diffusion amount is distributed to each process according to the weight coefficient. The outpatient process gets 0.18×0.1 = 0.018, the examination process gets 0.18×0.2 = 0.036, the medication process gets 0.18×0.3 = 0.054, and the surgical process gets 0.18×0.4 = 0.072. Calculate the diffusion distribution of other processes in the same way, and finally obtain the redistributed thermal matrix [0.564, 0.7448, 1.2348, 1.5784]. Finally, the system performs a summation operation on the redistributed multi-dimensional thermal distribution matrix according to each dimension to obtain the heat aggregation value. In this example, the heat aggregation value is 0.564 + 0.7448 + 1.2348 + 1.5784 = 4.122. The system uses this heat aggregation value as the medical record risk score.

[0088] In practical applications, the medical record risk score can be further quantified into risk levels. For example, a risk score < 3 is considered low risk, 3 - 5 is medium risk, and > 5 is high risk. For the score of 4.122 in this example, it belongs to the medium risk level, and the system will prompt medical staff to pay attention to this medical record and may recommend a second review or preventive measures.

[0089] The advantage of this method is that through sampling analysis in the time dimension, it can capture the dynamic change trend of risks; through heat diffusion analysis in the space dimension, it can evaluate the spread effect of risks between different diagnosis and treatment links; through risk sensitivity analysis of the medical decision tree, it can conduct more accurate risk assessment according to the actual risk characteristics of different medical behaviors, thus providing a more comprehensive and accurate medical record risk warning for medical institutions.

[0090] Existing medical risk assessment methods mainly rely on static risk indicator statistics and simple weighted calculations, and cannot effectively reflect the dynamic propagation characteristics and cumulative effects of risks during the diagnosis and treatment process. This solution innovatively introduces a thermal distribution model, analogizes the risk propagation process to a heat transfer phenomenon, and realizes the dynamic quantification of risk propagation by constructing a multi-dimensional thermal distribution matrix. By calculating the risk evolution rate and heat diffusion coefficient, it accurately captures the propagation law and attenuation characteristics of risks between different diagnosis and treatment links. Combining the risk sensitivity obtained from medical decision tree analysis, a risk weight system more in line with clinical reality is established, avoiding the problem of overly subjective weight allocation in traditional methods. Through heat redistribution and aggregation calculations, information in multiple dimensions such as time series correlation, spatial distribution, and risk sensitivity is systematically integrated, making the final medical record risk assessment more comprehensive and objective. This risk assessment method based on the thermodynamic model significantly improves the accuracy and timeliness of medical risk assessment and provides a more reliable decision-making basis for medical quality management.

[0091] In the second aspect of the embodiments of the present invention, A kind of electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0092] In the third aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0093] The present invention can be a method, device, system, and / or computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a composite trauma medical record intelligent quality control system, characterized in that: include: Obtain trauma medical record data including medical record text data and medical image data, divide them into text segments and image blocks respectively, generate adversarial samples by randomly masking and reorganizing text segments and image blocks, calculate the word vector representation of text segments and the visual feature representation of image blocks in adversarial samples, and project them into a common feature space to generate semantic association features; Map the semantic association features with the rule items in the preset quality control rule library to obtain the rule association results, build a directed acyclic graph of clinical events based on the rule association results, extract the temporal dependencies between events, calculate the local causal effects and counterfactual conditional probabilities of clinical event nodes, identify the missing nodes and edges in the directed acyclic graph of clinical events, and generate quality control defect data; The medical record risk score is calculated according to the number and importance of missing nodes in the quality control defect data, and computing resources are allocated based on the medical record risk score. When the medical record risk score exceeds the preset risk threshold, a quality control alarm information is generated, the quality control defect data and the quality control alarm information are written into the quality control cache, a quality control analysis report is generated, the quality control analysis report is sent to the clinical terminal, and the quality control defect data, quality control alarm information and quality control analysis report are stored in the quality control database.

2. The method according to claim 1, characterized in that Obtain trauma medical record data including medical record text data and medical image data, divide them into text segments and image blocks respectively, and generate adversarial samples by randomly masking and reorganizing text segments and image blocks, including: Segmenting the medical record text data using a sliding window to generate multiple text segments, wherein the text segments include diagnosis description information and treatment record information; Constructing an attention heat map for the medical image data, and calculating an attention score of the lesion area according to the attention heat map; determining a target area based on the attention score, setting a cropping window with the target area as the center, and generating multiple image blocks; Counting the frequency of occurrence of medical terms in the text segment, selecting the medical terms with the highest frequency of occurrence for masking, and generating text masking data; and masking the non-target area in the image block while keeping the target area intact, and generating image masking data; The text masking data and the image masking data are paired in time sequence to construct an original sample pair; the text masking data and the image masking data of adjacent time sequences in the original sample pair are recombined to generate a recombined sample pair; and the original sample pair and the recombined sample pair are merged to form a final adversarial sample.

3. The method according to claim 1, characterized in that The word vector representation of the text fragment and the visual feature representation of the image block in the adversarial sample are calculated and projected into a common feature space to generate semantically related features including: Using a pre-trained language model to extract features from text fragments to obtain word vectors; constructing a query matrix and a key value matrix based on the word vectors; multiplying the query matrix and the key value matrix to obtain attention weights; weighting the attention weights and the word vectors to obtain text semantic features; Perform convolution feature extraction on the image block to obtain global visual features, perform multi-scale pooling operation on the global visual features to obtain multi-scale visual features, and perform feature fusion on the global visual features and multi-scale visual features to obtain image semantic features; Calculating the mean vector and covariance matrix of text semantic features to obtain text distribution parameters, calculating the mean vector and covariance matrix of image semantic features to obtain image distribution parameters, constructing a multilayer perceptron as a feature projection network based on the text distribution parameters and the image distribution parameters, and performing feature projection on the text semantic features and the image semantic features to obtain text projection features and image projection features; Constructing a sample pair set, including positive sample pairs consisting of text projection features and image projection features corresponding in time sequence, and negative sample pairs consisting of text projection features and image projection features not corresponding in time sequence, respectively calculating the positive sample similarity of the positive sample pairs and the negative sample similarity of the negative sample pairs, and at the same time dividing the text projection features and image projection features into regions and calculating the local similarity; The positive sample similarity, negative sample similarity and local similarity are combined to construct a multi-granularity contrast loss function, which is iteratively optimized based on the gradient descent method to obtain semantic association features.

4. The method according to claim 1, characterized in that Based on the rule association results, a directed acyclic graph of clinical events is constructed to extract the temporal dependencies between events. By calculating the local causal effects and counterfactual conditional probabilities of clinical event nodes, missing nodes and edges in the directed acyclic graph of clinical events are identified to generate quality control defect data including: Constructing the conditional event in the rule association result as a parent node and the conclusion event as a child node, obtaining the time information of the parent node and the child node, determining the temporal dependency relationship between the parent node and the child node based on the time information, calculating the conditional probability value between the parent node and the child node, and combining the parent node, the child node, the temporal dependency relationship and the conditional probability value to construct a directed acyclic graph of clinical events; Extracting event node pairs from the directed acyclic graph of clinical events, obtaining local causal effect values ​​by performing intervention calculations on the event node pairs, extracting covariate information of the event node pairs, and calculating counterfactual conditional probability values ​​of the event node pairs based on the covariate information; An event whose local causal effect value exceeds a preset causal threshold but does not have a corresponding node in the directed acyclic graph of clinical events is determined as a missing node; an event relationship whose counterfactual conditional probability value exceeds a preset probability threshold but does not have a corresponding edge in the directed acyclic graph of clinical events is determined as a missing edge, and quality control defect data is generated based on the missing nodes and missing edges.

5. The method according to claim 4, characterized in that Extracting event node pairs from the directed acyclic graph of clinical events, obtaining local causal effect values ​​by performing intervention calculations on the event node pairs, extracting covariate information of the event node pairs, and calculating counterfactual conditional probability values ​​of the event node pairs based on the covariate information includes: Extracting event duration information from clinical data, statistically calculating the duration distribution of each event, and determining the core impact duration of each event based on the duration distribution; Calculate the basic impact strength of the event based on the core impact duration, substitute the time difference between the event occurrence time and the current time into the exponential decay function to obtain the time decay value, determine whether the time decay value exceeds the core impact duration to obtain the impact decay value, and multiply the basic impact strength, the time decay value and the impact decay value to obtain the event impact threshold; Extract event node pairs from the directed acyclic graph of clinical events, apply intervention operations to source events in the event node pairs, calculate the conditional probabilities of target events before and after the source event intervention, multiply the conditional probabilities by the event impact threshold and accumulate the integrals in the time dimension to obtain the local causal effect values ​​of the event node pairs; Extracting the covariate information of the event node pair, calculating the impact threshold of the covariate information based on the core impact duration and the event impact threshold, and multiplying the impact threshold of the covariate information by the event impact threshold of the event node pair to obtain an impact assessment value of the covariate information; Based on the impact assessment value of the covariate information, the conditional probability of the event node pair is weighted, and the weighted conditional probability is multiplied and integrated with the event impact threshold to obtain the counterfactual conditional probability value of the event node pair.

6. The method according to claim 1, characterized in that The medical record risk score is calculated based on the number and importance of missing nodes in the quality control defect data, and computing resources are allocated based on the medical record risk score. When the medical record risk score exceeds the preset risk threshold, quality control alarm information is generated, including: Extract medical decision points and diagnosis and treatment behavior sequences from historical medical record data, calculate the influence probability of the medical decision points and the diagnosis and treatment behavior sequences, and construct a causal chain of clinical events based on the influence probability; Extract missing node information from quality control defect data, obtain the target node corresponding to the missing node in the standard clinical pathway, calculate the data integrity difference value between the missing node and the target node, calculate the node deviation coefficient based on the difference value, analyze the cross-interaction strength between the missing node and the nodes on the causal chain of clinical events according to the node deviation coefficient, and recursively propagate the cross-interaction strength to obtain a risk propagation value; A multidimensional thermal distribution matrix is ​​constructed based on the risk propagation value, a time series sampling difference of the risk propagation value is calculated to obtain a risk evolution rate, a heat diffusion coefficient of the multidimensional thermal distribution matrix is ​​calculated according to the risk evolution rate, and a heat aggregation value of the multidimensional thermal distribution matrix is ​​used as a medical record risk score; The heat diffusion area is predicted according to the heat diffusion coefficient, and the computing resources are allocated based on the heat concentration value. When the medical record risk score exceeds a preset risk threshold, quality control alarm information is generated based on the heat transfer path of the multidimensional thermal distribution matrix.

7. The method according to claim 6, characterized in that Constructing a multidimensional thermal distribution matrix based on the risk propagation value, calculating the time series sampling difference of the risk propagation value to obtain the risk evolution rate, calculating the heat diffusion coefficient of the multidimensional thermal distribution matrix according to the risk evolution rate, and using the heat aggregation value of the multidimensional thermal distribution matrix as the medical record risk score includes: Constructing an initial thermal distribution matrix based on the time series association and spatial distribution of the risk propagation values, calculating the propagation attenuation of the risk propagation values ​​between diagnosis and treatment links to obtain a heat transfer coefficient, and fusing the heat transfer coefficient with the initial thermal distribution matrix to construct a multidimensional thermal distribution matrix; The risk propagation value is sampled at regular intervals, the difference of the risk propagation values ​​at adjacent sampling moments is calculated to obtain the risk evolution rate, and the heat diffusion coefficient of the multidimensional thermal distribution matrix is ​​calculated according to the temporal association between the risk evolution rate and the diagnosis and treatment links; Analyzing the risk sensitivity of different types of medical behaviors based on the medical decision tree, converting the risk sensitivity into a risk weight coefficient, and redistributing the heat in the multidimensional thermal distribution matrix according to the risk weight coefficient and the heat diffusion coefficient; The multi-dimensional thermal distribution matrix after redistribution is summed up according to each dimension to obtain a heat accumulation value, and the heat accumulation value is used as the medical record risk score.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Clinical manifestation information extracting method of Chinese electronic medical record data and equipment

    CN110223742A

  • Health care big data quality control system and terminal based on CDA shared documents

    CN111524589A

  • Electronic medical record diagnosis and treatment quality control method based on clinical diagnosis and treatment guide

    CN113361230A

  • Simulated medical information detection method and system based on AI algorithm

    CN118471539A

  • Electronic medical record automatic quality control system and method based on large language model

    CN119692879A

Cited By

  • Intelligent management method and system for drug clinical items

    CN120510992A

  • Intelligent management method and system for a drug clinical project

    CN120510992B

  • CKD special disease database construction method and system

    CN120723750A

  • Medical data quality closed-loop control method and system based on large model

    CN121354772A

  • Big model-based medical data quality closed-loop control method and system

    CN121354772B