A multi-source medical data-oriented quality intelligent evaluation method and system

By preprocessing, spatiotemporal semantic alignment, and joint representation learning of multi-source medical data, combined with multi-dimensional quality indicators and end-to-end credibility auditing, the problems of spatiotemporal mismatch and anomaly detection of multi-source medical data are solved, achieving efficient data repair and transparent data traceability, and optimizing the flexibility and performance of data processing.

CN122264627APending Publication Date: 2026-06-23FUTURE VALLEY TECHNOLOGY (TIANJIN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUTURE VALLEY TECHNOLOGY (TIANJIN) CO LTD
Filing Date
2026-03-31
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies lack intelligent processing methods for spatiotemporal alignment, quality assessment, and anomaly detection of multi-source medical data, resulting in large data assessment errors, low efficiency in processing abnormal data, difficulty in tracing data sources, and a lack of transparent reporting.

Method used

By collecting and preprocessing multi-source medical data, we drive the spatiotemporal semantic alignment and joint representation learning of cross-modal data, integrate spatiotemporal correlation quantities and multi-dimensional quality indicators for data quality evaluation, perform abnormal data repair and trusted dataset archiving, implement end-to-end trustworthiness auditing and root cause analysis, and optimize model parameters to adapt to environmental changes.

Benefits of technology

It achieves high-precision alignment of multi-source medical data, efficient repair of abnormal data, improved full-link visualization and transparency of data traceability, and enhanced flexibility and performance of the optimization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122264627A_ABST
    Figure CN122264627A_ABST
Patent Text Reader

Abstract

The application discloses a kind of quality intelligent evaluation methods and systems for multi-source medical data, it is related to big data processing technical field.The quality intelligent evaluation methods and systems for multi-source medical data, including S1, acquisition multi-source medical data is preprocessed, store and construct medical database;S2, drive the spatiotemporal semantic alignment and joint representation learning of cross-modal data, construct medical data optimization model and output spatiotemporal alignment value;S3, fusion spatiotemporal correlation and multidimensional quality index, data quality evaluation and abnormal analysis are carried out;S4, according to data credible result executes the traceability path generation and explainability report output operation;S5, through user feedback data and model evolution state analysis optimization effect.It solves the problem that multi-source medical data exists in spatiotemporal alignment, quality evaluation and abnormal detection, spatiotemporal mismatch, data inconsistency and low quality repair efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data processing technology, specifically to a method and system for intelligent quality assessment of multi-source medical data. Background Technology

[0002] With the continuous development of medical informatization, the types and quantities of medical data have experienced explosive growth. Modern medicine involves various data sources, including imaging data, clinical text data, laboratory test data, and patient monitoring data. These data often originate from different devices and systems, exhibiting significant differences in data format, acquisition standards, and timestamps. Effectively assessing and processing the quality of this multi-source, heterogeneous data has become crucial for improving the application effectiveness of technologies such as medical decision support systems and intelligent diagnosis and treatment systems.

[0003] For example, the invention patent with publication number CN120087554A discloses a big data-based system for predicting and evaluating scientific and technological achievements. Specifically, it includes: a research and development process data collection module, an information technology achievement prediction module, and an information technology achievement comprehensive evaluation module. The research and development process data collection module is used to collect historical research data, feature data, and network transmission performance data of scientific research projects in real time during the research and development process. The information technology achievement prediction module is used to establish a prediction model for the feature data of the research process of scientific and technological achievements, set research process status indicators based on feature data, and screen network transmission performance data within a stable time range for further training and optimization of the LSTM prediction model. The information technology achievement comprehensive evaluation module is used to comprehensively evaluate the stability of the research process of information technology achievements. By predicting real-time network transmission performance data, it sets information research fluctuation indicators to predict the stability of the research process.

[0004] For example, the invention patent with publication number CN109615204B discloses a method, device, equipment, and readable storage medium for quality assessment of medical data. The method includes: reading basic data tables, medical insurance data tables, and medical treatment data tables for each patient, and setting these tables as group elements to form a medical data group; performing data rationality testing, data correspondence testing, and correlation testing on each group element in the medical data group, generating test results; determining the target quality level corresponding to the generated test results based on a preset correspondence between the test results and quality levels, and conducting a quality assessment of the medical data group based on the target quality level, serving as the basis for medical data quality assessment, thus making the assessment more accurate and improving its efficiency and automation.

[0005] However, most traditional data quality assessment methods still rely on manual verification or rule-driven strategies, lacking intelligent processing methods for spatiotemporal alignment, data repair, and optimization of multi-source data. Existing technologies primarily assess the quality of single data sources, lacking cross-modal, multi-dimensional data quality assessment frameworks. Regarding spatiotemporal alignment, current technologies only handle data alignment issues through simple timestamp matching and spatial location annotation, failing to accurately address the complex relationships between imaging data and clinical records. Furthermore, with the surge in data volume, manual verification and repair become extremely labor-intensive, inefficient, and prone to errors.

[0006] Therefore, in order to address the above issues, there is an urgent need for a quality intelligent assessment method and system for multi-source medical data. Summary of the Invention

[0007] Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a method and system for intelligent quality assessment of multi-source medical data, which solves the problems of spatiotemporal mismatch, data inconsistency, and low efficiency of quality repair in spatiotemporal alignment, quality assessment, and anomaly detection of multi-source medical data.

[0008] Technical solution To achieve the above objectives, this invention provides the following technical solution: a method and system for intelligent quality assessment of multi-source medical data, comprising: S1, collecting multi-source medical data, performing preprocessing operations, storing and constructing a medical database; S2, driving spatiotemporal semantic alignment and joint representation learning of cross-modal data, constructing a medical data optimization model and outputting spatiotemporal alignment values; S3, integrating spatiotemporal correlation quantities and multi-dimensional quality indicators to perform data quality evaluation and anomaly analysis, and performing anomaly data repair and trusted dataset archiving operations based on the data quality analysis results; S4, implementing defect root cause tracing analysis based on data end-to-end trustworthiness audit, and performing tracing path generation and interpretability report output operations based on the data trustworthiness results; S5, analyzing the optimization effect through user feedback data and model evolution status, and performing parameter dynamic adjustment, model version management, and adaptive evolution operations based on the optimization analysis results.

[0009] Further, the specific steps for collecting, preprocessing, storing, and constructing a medical database from multiple sources of medical data are as follows: Collecting image acquisition timestamps, clinical event times, equipment spatial coordinates, patient position identifiers, and clinical anatomical locations; acquiring medical image sequence files and clinical record streams; and calculating the total number of data points. By comparing the patient ID and examination type in the image tags with the corresponding information in the clinical records, the matching degree is calculated using string similarity to obtain the image-medical record matching degree. The time compliance degree is obtained by calculating the time difference between the image acquisition time and the clinical event, and judging the rationality of the time logic according to clinical guidelines. The spatial matching degree of each data point is obtained by comparing the spatial coordinates in the image data with the anatomical locations in the medical record. Finally, the required fields for each data point are checked for blank or invalid values. The information completeness is obtained by calculating the proportion of the entire field. The reasonable range of data values, the logical sequence of verification time, and the conformity of coding to the standard glossary are checked for each data point. The effective data quantity is obtained by calculating the proportion of data that passes all verifications. The initial association of all collected data is completed by patient unique identifier and acquisition session number, and written into temporary acquisition records. The image acquisition timestamp and clinical event time are unified with a time reference and converted into a timestamp format with the same sampling starting point. The spatial reference of equipment spatial coordinates, patient position identifiers, and clinical anatomical positions is unified, and standardized coding and structured terminology mapping are completed to eliminate ambiguity in expression. The maximum and minimum value normalization method is used to scale multi-source medical data to a unified range to achieve dimensionlessness. The standardized and normalized medical data are then stored with the patient unique identifier and acquisition session number attached, and a medical database is constructed.

[0010] Furthermore, the specific steps for driving the spatiotemporal semantic alignment and joint representation learning of cross-modal data, constructing a medical data optimization model, and outputting spatiotemporal alignment values ​​are as follows: Using image sequence files and clinical record streams as basic inputs, timestamp matching records, event continuity records, and temporal conflict marker records are generated in the medical database based on image acquisition timestamps and clinical event times; a spatial mapping input set is constructed by combining device spatial coordinates, patient position identifiers, and clinical anatomical location descriptions, and the patient's unique identifier and acquisition session number are written as association traceability fields; a multimodal feature joint embedding model is used to perform spatiotemporal semantic alignment calculations, learning images within a unified spatiotemporal embedding space. The joint learning representation of feature vectors and clinical record feature vectors is used. During training, image sequences and clinical records from the same patient in the same acquisition session form positive samples to form matched data pairs. Non-matched pairs are generated by mismatched pairings between different acquisition sessions of the same patient, between different patients, or artificially constructed. This minimizes the Mahalanobis feature distance of matched data pairs and maximizes the Mahalanobis feature distance of non-matched pairs. Adaptive attention weight adjustment is driven by temporal conflict markers and spatial matching differences to improve the discriminative power of key alignment factors. A medical data optimization model is constructed to output spatiotemporal alignment values, matching pair determination accuracy, and version identifiers. The model complexity is obtained by reading the model's hierarchical depth.

[0011] Furthermore, the specific steps for data quality evaluation and anomaly analysis by integrating spatiotemporal correlation parameters and multi-dimensional quality indicators are as follows: Obtain the image-medical record matching degree, temporal compliance, spatial matching degree, and spatiotemporal alignment value of the i-th data point; sum the image-medical record matching degree, temporal compliance, and spatial matching degree of the i-th data point to obtain the cumulative value of the three-dimensional quality features of the data point; sum the cumulative values ​​of the three-dimensional quality features of all data points to obtain the total sum of quality feature values; divide the total sum of quality feature values ​​by the total number of data points to obtain the average three-dimensional quality feature value; multiply the average three-dimensional quality feature value by the spatiotemporal alignment value of each data point to obtain the data quality value of each data point.

[0012] Furthermore, the specific steps for performing abnormal data repair and trusted dataset archiving based on the data quality analysis results are as follows: By comparing the data quality value and the quality threshold in real time, when the data quality value is less than the quality threshold, the data repair process is initiated: checking and re-collecting missing data in missing images or clinical records; comparing the timestamp differences between image data and clinical data, adjusting data that does not conform to the time sequence to correct the timeliness of the data; using image registration technology to align the image data to the standard anatomical position to ensure that the images are consistent with the anatomical position of the clinical records; using isolated forest to mark potential abnormal data and re-collecting data; after completing the data repair, recalculating the data quality value, if it is still less than the quality threshold, issuing a manual verification instruction and generating an early warning report; when the data quality value is greater than or equal to the quality threshold, confirming that the data has met the quality standards, constructing a traceability dataset, calculating the total number of traceability data points, and archiving it to the medical database.

[0013] Furthermore, the specific steps for implementing defect root cause tracing analysis based on data end-to-end credibility audit are as follows: obtain the information completeness, data validity, and data quality value of the j-th traceable data point; multiply the information completeness and data validity of the j-th traceable data point to obtain the complete and valid product value of the data point; sum the complete and valid product values ​​of all traceable data points to obtain the total complete and valid product value; divide the total complete and valid product value by the total number of traceable data points L to obtain the average complete and valid product value; multiply the average complete and valid product value by the data quality value to obtain the data verification value.

[0014] Furthermore, the specific steps for generating traceability paths and outputting interpretability reports based on the data credibility results are as follows: By comparing data verification values ​​and verification thresholds in real time, when the data verification value is less than the verification threshold, a traceability link graph from the collection source to the current state is constructed and presented based on the data flow log, identifying the metadata and transformation records of each link; a re-verification request containing specific missing fields is sent to the source data based on the traceability link graph to verify and improve the traceability information; the completed collection source information is updated in the medical database, and a structured interpretability report of the traceability analysis conclusion is output; when the data verification value is greater than or equal to the verification threshold, the permission flag of the data point is updated to "verified"; the permission flag is updated in the database to remove access restrictions on the batch data; and a metadata update message is written to the message queue to send an update notification to the application service subscribing to this data type.

[0015] Further, the specific steps for analyzing the optimization effect through user feedback data and model evolution status are as follows: By performing the same indicator alignment evaluation on the previous and current versions of the medical data optimization model at version evaluation snapshots, and subtracting the values ​​from the two evaluations to obtain the model performance change; by collecting data problem reports submitted by clinicians and statistically analyzing the frequency of various problems, the clinical problem feedback frequency is obtained; data verification values, model performance change, clinical problem feedback frequency, and model complexity are obtained; the product of the model performance change adjustment coefficient and the absolute value of the model performance change is calculated to obtain the performance change weighting term; the product of the user feedback intensity adjustment coefficient and the clinical problem feedback frequency is calculated to obtain the user feedback weighting term; the performance change weighting term and the user feedback weighting term are added together and then one is added to obtain the optimization gain term; the data verification value is multiplied by the optimization gain term to obtain the retrospective weighting term; the product of the model complexity adjustment coefficient and the model complexity is calculated and then one is added to obtain the complexity suppression term; the retrospective weighting term is divided by the complexity suppression term to obtain the optimization feedback value.

[0016] Furthermore, the specific steps for performing dynamic parameter adjustment, model version management, and adaptive evolution based on the optimization analysis results are as follows: By comparing the optimization feedback value and the optimization feedback threshold in real time, when the optimization feedback value is less than the optimization feedback threshold, if the performance change is less than zero or the frequency of clinical problem feedback is greater than the feedback threshold, all model parameters are frozen. Combined with tracing the link graph to locate the performance indicators on the validation set that have decreased K times consecutively with a decrease greater than ϵ, a root cause analysis report is generated. In an isolated sandbox computing environment, a recently verified subset of data is selected as training samples, the learning rate is reduced, the model is retrained, the parameter freeze is lifted, and each time... Gradually implement the gray-scale recovery of online traffic by increasing p% of traffic, while monitoring the accuracy, recall, and throughput of the medical data optimization model; if the recalculated optimization feedback value is still less than the optimization feedback threshold, roll back the model version to the previous stable state, and encrypt and archive all context information of failed cases to the medical knowledge base; when the optimization feedback value is greater than or equal to the optimization feedback threshold, confirm that the optimization measures have achieved the expected results, maintain the validated medical data optimization model parameters and data processing rules; start the continuous monitoring program, and arrange incremental learning tasks based on newly accumulated qualified quality data to ensure the continuous evolution of evaluation capabilities.

[0017] Furthermore, a second aspect of this invention provides a quality intelligent assessment system for multi-source medical data, applying a quality intelligent assessment method for multi-source medical data, comprising: a medical data acquisition and preprocessing module for acquiring multi-source medical data, performing preprocessing operations, storing and constructing a medical database; a spatiotemporal alignment module for driving spatiotemporal semantic alignment and joint representation learning of cross-modal data, constructing a medical data optimization model and outputting spatiotemporal alignment values; a quality assessment and anomaly detection module for fusing spatiotemporal correlation quantities and multi-dimensional quality indicators to perform data quality evaluation and anomaly analysis, and performing anomaly data repair and trusted dataset archiving operations based on the data quality analysis results; a data traceability module for implementing defect root cause tracing analysis based on the data end-to-end trustworthiness audit, and performing traceability path generation and interpretability report output operations based on the data trustworthiness results; and a feedback and optimization module for analyzing the optimization effect through user feedback data and model evolution status, and performing parameter dynamic adjustment, model version management, and adaptive evolution operations based on the optimization analysis results.

[0018] Beneficial effects The present invention has the following beneficial effects: (1) This invention improves the spatiotemporal matching accuracy between image data and clinical records by introducing spatiotemporal alignment values ​​and using a multimodal feature joint embedding model for spatiotemporal semantic alignment calculation. This achieves high-precision multi-source medical data alignment and optimization, effectively solving the problem of data evaluation errors caused by spatiotemporal mismatch in the prior art.

[0019] (2) This invention assesses data quality by combining spatiotemporal correlation quantities and multi-dimensional quality indicators, and performs abnormal data repair and trusted dataset archiving operations based on the assessment results. This achieves efficient automatic repair of abnormal data and effectively solves the problems of inefficiency and manual dependence in abnormal data processing in the prior art.

[0020] (3) This invention, by establishing a data end-to-end credibility audit mechanism and implementing root cause analysis, can accurately locate the source of data problems and generate an interpretable report. This achieves end-to-end visualization and improved transparency of data traceability, effectively solving the problems of difficult data source traceability and lack of transparent reports in existing technologies.

[0021] (4) This invention, through a comprehensive evaluation of model performance changes and user feedback intensity, performs version management and adaptive evolution of the medical data optimization model, ensuring continuous improvement of the medical data optimization model in a dynamic environment. This achieves efficient adjustment and performance improvement of the optimization process, effectively solving the problems of lag and inflexibility in the optimization process in existing technologies.

[0022] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0023] Figure 1 This is a flowchart of a quality intelligent assessment method for multi-source medical data according to the present invention. Figure 2 This is a structural diagram of a quality intelligent assessment system for multi-source medical data according to the present invention; Figure 3 This is a flowchart of the medical data verification and traceability control process of the present invention; Figure 4 This is a medical data quality feedback surface plot based on multi-parameter collaborative optimization, as presented in this invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Please see Figures 1-4 This invention provides a technical solution: a method and system for intelligent quality assessment of multi-source medical data, comprising: S1, collecting multi-source medical data, performing preprocessing operations, storing and constructing a medical database; S2, driving spatiotemporal semantic alignment and joint representation learning of cross-modal data, constructing a medical data optimization model and outputting spatiotemporal alignment values; S3, fusing spatiotemporal correlation quantities and multi-dimensional quality indicators to perform data quality evaluation and anomaly analysis, and performing anomaly data repair and trusted dataset archiving operations based on the data quality analysis results; S4, implementing defect root cause tracing analysis based on the data end-to-end trustworthiness audit, and performing tracing path generation and interpretability report output operations based on the data trustworthiness results; S5, analyzing the optimization effect through user feedback data and model evolution status, and performing parameter dynamic adjustment, model version management and adaptive evolution operations based on the optimization analysis results.

[0026] Specifically, the steps for collecting, preprocessing, storing, and constructing a medical database from multiple sources of medical data are as follows: First, collect image acquisition timestamps, clinical event times, equipment spatial coordinates, patient position identifiers, and clinical anatomical locations. Second, acquire medical image sequence files and clinical record streams, and statistically determine the total number of data points. Third, calculate the image-to-medical-record matching degree by comparing the patient ID and examination type in the image tags with the corresponding information in the clinical records, reflecting the consistency between image data and clinical records. Fourth, calculate the time difference between the image acquisition time and the clinical event, and determine the rationality of the time logic according to clinical guidelines to obtain the time compliance degree, indicating whether the time data conforms to clinical standards. Fifth, obtain the spatial matching degree of each data point by comparing the spatial coordinates in the image data with the anatomical locations in the medical records, reflecting the consistency between image data and clinical anatomical location information. Sixth, check the relevant data in the required fields of each data point. The process involves checking for blank or invalid values, calculating the proportion of complete fields to determine information completeness, indicating whether the data covers all necessary record items; checking the reasonable range of data values, verifying the logical order of time, and checking whether the coding conforms to the standard glossary, calculating the proportion of data that passes all checks to determine the data validity, reflecting the data's effectiveness; initially associating all collected data with the patient's unique identifier and acquisition session number, and writing them into temporary acquisition records; unifying the time base for image acquisition timestamps and clinical event times, converting them into timestamps from the same sampling starting point; unifying the spatial base for device spatial coordinates, patient position identifiers, and clinical anatomical locations, completing standardized coding and structured terminology mapping to eliminate ambiguity; scaling multi-source medical data to a unified range using maximum and minimum value normalization methods to achieve dimensionlessness; and storing the standardized and normalized medical data with the patient's unique identifier and acquisition session number to construct a medical database.

[0027] This implementation plan achieves efficient organization and standardization of medical data through the collection and processing of multi-dimensional data, including image acquisition timestamps, clinical event times, equipment spatial coordinates, patient positioning identifiers, and clinical anatomical locations. By comparing image data with clinical records, the matching degree of image-medical records, temporal compliance, and spatial matching degree were calculated to ensure data consistency and temporal logical rationality. Secondly, the required fields and numerical ranges of each data point were checked, and the information completeness and data validity were statistically determined, ensuring data integrity and validity. Standardization and normalization processes eliminated ambiguities in data representation and inconsistencies in units of measurement, ensuring comparability between different data sources. All processed data is uniformly stored in a medical database, providing reliable basic data support for subsequent analysis and modeling.

[0028] Specifically, the steps to drive spatiotemporal semantic alignment and joint representation learning of cross-modal data, construct a medical data optimization model, and output spatiotemporal alignment values ​​are as follows: Using image sequence files and clinical record streams as basic inputs, and combining image acquisition timestamps and clinical event times, timestamp matching records, event continuity records, and temporal conflict marker records are generated in the medical database. A spatial mapping input set is constructed by combining device spatial coordinates, patient position identifiers, and clinical anatomical location descriptions. The patient's unique identifier and acquisition session number are stored as associated traceability fields to ensure data traceability and reliability. A multimodal feature joint embedding model is used for spatiotemporal semantic alignment calculation, learning a joint learning representation of image feature vectors and clinical record feature vectors in a unified spatiotemporal embedding space. During training, image sequences and clinical records from the same patient and within the same acquisition session constitute positive samples, generating matched data pairs. Mismatched pairs between different acquisition sessions of the same patient, between different patients, or artificially constructed mismatched pairs constitute non-matched data pairs. This method ensures that the Mahalanobis feature distance of matched data pairs is minimized, and the Mahalanobis feature distance of non-matched data pairs is maximized, thereby optimizing the contrastive learning effect of the medical data optimization model. To ensure the effectiveness of this process, a loss function is defined to measure the distinction between positive and negative samples, and a temperature parameter is introduced to adjust the sensitivity of contrastive learning, thereby improving the discriminative ability of the medical data optimization model. Driven by temporal conflict markers and spatial matching differences, an adaptive attention mechanism is employed to further enhance the discriminative power of key alignment factors. The medical data optimization model is constructed and outputs spatiotemporal alignment values, matching pair determination accuracy, and version identifiers. The model complexity is obtained by reading the hierarchical depth of the medical data optimization model.

[0029] In this implementation scheme, the correlation and matching degree between image data and clinical records are significantly improved through spatiotemporal semantic alignment of multimodal data. By employing a strategy of minimizing Mahalanobis feature distance and maximizing the Mahalanobis feature distance of mismatched pairs, matching and mismatched data can be accurately distinguished, thereby optimizing the alignment effect. Simultaneously, by introducing an adaptive attention mechanism, the discriminative power of key alignment factors is further enhanced, ensuring the effective handling of temporal conflicts and spatial matching differences. This process not only improves the accuracy of data matching but also enhances the learning ability of the medical data optimization model on complex medical data, ultimately providing the medical data optimization model with accurate spatiotemporal alignment values ​​and efficient matching determination.

[0030] Specifically, the steps for data quality evaluation and anomaly analysis by integrating spatiotemporal correlation parameters and multi-dimensional quality indicators are as follows: Obtain the image-medical record matching degree, temporal compliance, spatial matching degree, and spatiotemporal alignment value for the i-th data point; sum the image-medical record matching degree, temporal compliance, and spatial matching degree of the i-th data point to obtain the cumulative three-dimensional quality feature value of the data point; sum the cumulative three-dimensional quality feature values ​​of all data points to obtain the total cumulative quality feature value; divide the total cumulative quality feature value by the total number of data points to obtain the average three-dimensional quality feature value; multiply the average three-dimensional quality feature value by the spatiotemporal alignment value of each data point to obtain the data quality value of each data point. This process comprehensively evaluates the quality of each data point in spatiotemporal alignment by weighted summation of image-medical record matching degree, temporal compliance, and spatial matching degree, thereby generating a comprehensive data quality value. The average three-dimensional quality feature value can effectively quantify the quality of multi-source medical data. This is reflected in the weighted approach that balances the quality of data across different dimensions, reducing the potential bias caused by a single indicator and enhancing the comprehensiveness and accuracy of data quality assessment. Based on the multi-dimensional attributes of actual medical data, and using a parametric weighted calculation method, it ensures that the quality assessment of each data point considers both temporal and spatial accuracy, as well as the matching between images and clinical records. By accumulating and averaging the quality features of all data points, global consistency is achieved, enabling adaptation to changes and updates in different datasets and ensuring a comprehensive and accurate assessment of data quality.

[0031] The specific calculation method for data quality values ​​is as follows: ; In the formula, This represents the data quality value, reflecting the overall quality level of the data, and is the output value of the module. This represents the total number of data points and is a basic data indicator obtained from the data preprocessing and integration module. This represents the image-medical record matching degree of the i-th data point, which measures the degree of matching between image data and clinical records. This represents the time compliance of the i-th data point, measuring the compliance of clinical data in the time dimension; This represents the spatial matching degree of the i-th data point, which measures the spatial information matching degree between the image and the clinical record. This represents the spatiotemporal alignment value, reflecting the spatiotemporal matching degree between imaging data and clinical data.

[0032] In this implementation plan, a data quality value for each data point is derived by comprehensively calculating its image-medical record matching degree, temporal compliance, spatial matching degree, and spatiotemporal alignment value. This process weights and accumulates multiple dimensions of quality features, reflecting the overall quality of the data in terms of spatiotemporal alignment, thereby achieving quantitative evaluation of multi-source medical data. By summing the accumulated quality feature values ​​of all data points and calculating the average, the consistency of global data quality is further ensured. This method, by comprehensively considering multiple factors such as time, space, and image matching, effectively improves the accuracy and comprehensiveness of data quality assessment, ensures data reliability, and provides a solid data foundation for subsequent analysis, modeling, and decision-making.

[0033] Specifically, the steps for performing abnormal data repair and trusted dataset archiving based on data quality analysis results are as follows: By comparing data quality values ​​and quality thresholds in real time, when the data quality value is less than the quality threshold, the data repair process is initiated: checking and re-acquiring missing image data or missing parts in clinical records; comparing the timestamp differences between image data and clinical data, adjusting data that does not conform to the time sequence, correcting the timeliness of the data, and ensuring the consistency of data in the time dimension; using image registration technology to align the image data to the standard anatomical position, ensuring that the image data is consistent with the anatomical position of the clinical record, and avoiding quality issues caused by spatial mismatch. The process involves: using an isolated forest to mark potentially anomalous data and re-collecting the anomalous portions to ensure accuracy and completeness; after data repair, recalculating the data quality value; if it remains below the quality threshold, sending a manual verification instruction; this instruction includes detailed repair logs, a specific description of the problematic data, fields requiring manual verification, and related data records; the instruction is in structured data format for easy tracking and verification by manual operators; and generating an early warning report, which includes: the specific data points where repair failed, the execution status of each repair operation, an analysis of why the planned repair measures did not achieve the desired effect, and suggestions for the next steps. The early warning report is output in a structured format to facilitate further technical review and manual intervention. The report also includes detailed traceability information for the data points, error types, repair measures, and improvement suggestions to help data managers assess the effectiveness of current repair measures and decide whether to proceed with further operations. When the data quality value is greater than or equal to the quality threshold, the data is confirmed to meet the quality standards. At this point, the system constructs a traceability dataset and counts the total number of traceability data points that meet the quality standards. A traceability data point refers to a valid data point that is ultimately confirmed to meet the quality standards by comparing the data quality value with the quality threshold during the quality assessment process. These data points are filtered through multi-dimensional quality assessment criteria such as spatiotemporal alignment and information matching degree, and are archived into the medical database to ensure the integrity and availability of the data.

[0034] In this implementation plan, by comparing data quality values ​​and quality thresholds in real time, this step can effectively detect and repair data quality issues. If the data quality value is lower than the quality threshold, the system will automatically initiate a data repair process, including re-collecting missing data, correcting temporal order and spatial consistency, performing spatial repair using image registration technology, and using isolated forests to label and correct potentially abnormal data. After repair, the data quality value will be recalculated, and if the data quality still does not meet the standard, a manual verification instruction will be triggered and an early warning report will be generated for further processing. When the data quality value is greater than or equal to the quality threshold, the system confirms that the data meets the quality standards, constructs a traceability dataset, and archives it into the medical database, ensuring the accuracy, integrity, and consistency of the data and providing reliable data support for subsequent processing.

[0035] Specifically, the steps for implementing root cause analysis of defects based on end-to-end data credibility auditing are as follows: Obtain the information completeness, data validity, and data quality value of the j-th traceable data point; multiply the information completeness and data validity of the j-th traceable data point to obtain the complete and valid product value; sum the complete and valid product values ​​of all traceable data points to obtain the total complete and valid product value; divide the total complete and valid product value by the total number of traceable data points L to obtain the average complete and valid product value; multiply the average complete and valid product value by the data quality value to obtain the data verification value. This process comprehensively evaluates the quality of each data point by combining the product of information completeness and data validity to ensure consistency in data completeness and validity. During the calculation, the summation and average value of the complete and valid product values ​​of all traceable data points further ensures the consistency and accuracy of the data globally. Finally, this average value is combined with the data quality value, and the data verification value is used to more accurately evaluate the overall data quality, ensuring that the traceability dataset meets the predetermined quality standards.

[0036] The specific calculation method for the data verification value is as follows: ; In the formula, This represents the data verification value, reflecting the level of data integrity and consistency, and is the module output value; L represents the total number of traceable data points; This represents the information completeness of the j-th data point, measuring the degree of information completeness of a single data point. Q represents the effective data quantity of the j-th data point, measuring the practical application value and logical rationality of a single data point; Q represents the data quality value, reflecting the overall quality level of the data.

[0037] In this implementation plan, the integrity, validity, and quality of each traceability data point are comprehensively evaluated to calculate the product of integrity and validity. The average product of integrity and validity for the entire traceability dataset is then obtained through global summation and averaging. Multiplying this average by the data quality value yields the data verification value. This process effectively provides a unified assessment of data integrity, validity, and quality, offering a more precise quantitative basis for quality control of the entire dataset, ensuring that traceability data meets quality standards, and providing reliable support for subsequent data analysis and decision-making.

[0038] Specifically, the steps for generating traceability paths and outputting interpretability reports based on data reliability results are as follows: Figure 3 This is a flowchart illustrating the medical data verification and traceability control process in this embodiment. By comparing data verification values ​​and thresholds in real time, when a data verification value is less than the threshold, a traceability link graph from the data collection source to the current state is constructed and presented based on the data flow log, identifying the metadata and transformation records of each link. The data flow log records all data flow information from the data collection source to the final storage process, including data source, data type, operation time, and processing steps. The traceability link graph consists of multiple nodes and edges. Each node represents a stage or link in data processing, and each edge represents a data flow path. Each node stores a set of metadata fields for the corresponding link, including key fields such as data source, processing method, timestamp, and version number. Based on the traceability link graph, a re-verification request containing specific missing fields is sent to the source data system. This request includes a set of data fields such as missing image data fields, missing clinical record fields, and timestamp inconsistencies. The request format is a structured data format, facilitating automatic processing and verification of missing information by the system. By sending requests, the system verifies and improves the traceability information during the data collection process, ensuring the integrity and accuracy of the data chain. It updates the completed source information in the medical database and generates a structured, interpretable report. The report includes data verification analysis conclusions, a detailed list of missing fields, data integrity scores, and remedial measures. When the data verification value is greater than or equal to the verification threshold, the system updates the data point's permission flag to "verified," indicating that the data point has been verified and meets quality standards. The system also updates the permission flag in the database, removing access restrictions on the batch data and allowing it to proceed to subsequent processing and use. Simultaneously, it writes metadata update messages to the message queue and sends update notifications to application services subscribed to this data type, ensuring that relevant systems and applications can promptly obtain the latest verified medical data for subsequent analysis and processing.

[0039] In this implementation plan, by comparing data verification values ​​and verification thresholds in real time, the system can effectively identify and repair data quality issues. When a data verification value falls below the verification threshold, the system constructs a traceability map to identify each stage of data processing, ensuring data integrity and traceability. By sending a re-verification request and supplementing missing information, the system can improve data traceability records, update the source information in the medical database, and generate a structured, interpretable report. If the data verification value reaches or exceeds the verification threshold, the system updates the data point's status to verified, removes data access restrictions, and sends an update notification to relevant application services, thereby ensuring that the data can proceed to subsequent processing and usage stages. Through this process, the system achieves comprehensive control over data quality, guarantees the accuracy and integrity of medical data, and provides reliable data support.

[0040] Specifically, the steps for analyzing the optimization effect through user feedback data and model evolution status are as follows: First, perform aligned evaluations of the previous and current versions of the medical data optimization model on the same metrics at version evaluation snapshots, and obtain the model performance change by subtracting the values ​​from the two evaluations. Second, obtain the clinical problem feedback frequency by collecting data problem reports submitted by clinicians and statistically analyzing the frequency of various problems. Third, obtain the data verification value, model performance change, clinical problem feedback frequency, and model complexity. Fourth, obtain the model performance change adjustment coefficient by fitting the correlation analysis between the model performance change and optimization feedback values ​​in the system's historical optimization records, with a value range of 0 to 2. Fifth, obtain the user feedback intensity adjustment coefficient by performing logistic regression correlation analysis between the user feedback intensity value and the optimization effect in historical clinical feedback data, with a value range of 0 to 2. Sixth, analyze historical medical data... In the data optimization model, a negative exponential relationship between the number of model layers and the convergence efficiency of the optimization feedback value is fitted to obtain the model complexity adjustment coefficient, with a value ranging from 0 to 0.1. The performance change weighting term is obtained by multiplying the model performance change adjustment coefficient by the absolute value of the model performance change. The user feedback weighting term is obtained by multiplying the user feedback intensity adjustment coefficient by the frequency of clinical problem feedback. The optimization gain term is obtained by adding the performance change weighting term and the user feedback weighting term together. The traceability weighting term is obtained by multiplying the data verification value by the optimization gain term. The complexity suppression term is obtained by multiplying the model complexity adjustment coefficient by the model complexity and adding one. The optimization feedback value is obtained by dividing the traceability weighting term by the complexity suppression term. This method effectively improves the quality of medical data by dynamically adjusting the parameters and feedback mechanism of the medical data optimization model, achieving continuous optimization of data quality. Especially when facing data that does not meet quality standards, the data processing strategy can be adjusted based on the optimization feedback value, thereby effectively improving the completeness, timeliness, and accuracy of medical data, ensuring that medical data meets the needs of clinical decision support, and providing reliable basic data support for medical research and diagnosis.

[0041] The specific calculation method for the optimized feedback value is as follows: ; In the formula, F represents the optimization feedback value, which reflects the strength of the system optimization effect; This represents the data verification value, reflecting the completeness and reliability of the data; It represents the change in model performance, reflecting the degree of improvement or degradation in the performance of the medical data optimization model during the system optimization process; δ represents the frequency of clinical problem feedback, reflecting the frequency of problems reported by clinicians; a higher value indicates more feedback problems. β represents the model performance change adjustment coefficient, used to control the impact of ΔP on system optimization. γ represents the user feedback intensity adjustment coefficient, used to control the impact of clinical problem feedback frequency on system optimization. δ represents the model complexity adjustment coefficient, used to control the inhibitory effect of model complexity on system optimization. C represents the model complexity, reflecting the structural hierarchy characteristics of the medical data optimization model; a higher value indicates a more complex medical data optimization model.

[0042] In this embodiment, the data validation value for evaluation example 1 is 0.85, the model performance change adjustment coefficient is 1.1, the user feedback intensity adjustment coefficient is 0.9, the model complexity adjustment coefficient is 0.018, the model complexity is 45, the model performance change is 0.12, the user feedback intensity value is 0.22, and the optimization feedback value is 0.78; the data validation value for evaluation example 2 is 0.72, the model performance change adjustment coefficient is 0.8, the user feedback intensity adjustment coefficient is 0.7, the model complexity adjustment coefficient is 0.024, the model complexity is 60, the model performance change is 0.10, the user feedback intensity value is 0.20, and the optimization feedback value is 0. 65; Evaluation Example 3: Data validation value is 0.60, model performance change adjustment coefficient is 1.3, user feedback intensity adjustment coefficient is 1.2, model complexity adjustment coefficient is 0.030, model complexity is 75, model performance change is 0.15, user feedback intensity is 0.25, and optimization feedback value is 0.58; Evaluation Example 4: Data validation value is 0.82, model performance change adjustment coefficient is 1.8, user feedback intensity adjustment coefficient is 1.7, model complexity adjustment coefficient is 0.022, model complexity is 50, model performance change is 0.08, user feedback intensity is 0.18, and optimization feedback value is 0.71.

[0043] Table 1 Optimization Configuration Table for Multi-Source Medical Data Quality Assessment Parameters like Figure 4 The image shown is a medical data quality feedback surface diagram based on multi-parameter collaborative optimization provided in an embodiment of this application. Combined with... Figure 4As shown in Table 1, the optimized feedback value comprehensively reflects the overall performance level of the medical data quality assessment system. Different system configurations correspond to different regions on the surface, and the optimized feedback value shows a clear correlation with the system's operating status. Taking assessment example 1, the data validation value is 0.85, the model performance change is 0.12, the user feedback intensity value is 0.22, and the optimized feedback value is 0.78. The corresponding region in the figure is located at the central peak position of the surface, indicating that the system has high data quality and stable operation, meeting the needs of clinical data applications. The optimized feedback value of assessment example 4 is 0.71. Although the data verification value is relatively high at 0.82, the corresponding area in the graph is located on the right descending edge of the surface, indicating that the system is over-adjusted and needs to enter the parameter optimization loop. The optimization feedback value of evaluation example 3 is 0.58, and the corresponding area in the graph is located on the lower left side of the surface, reflecting that when the data verification value is low, the overall system performance is insufficient, and even with an aggressive adjustment strategy, it is still difficult to meet clinical use standards. The optimization feedback value of evaluation example 2 is 0.65, and the corresponding area in the graph is located in the middle transition area of ​​the surface, indicating that the conventional medical system can still maintain basic operating requirements under conservative parameter configurations. The comparison results of the graph and table together illustrate that the optimization feedback value can uniformly quantify the multi-dimensional influencing factors of medical data quality, providing a direct basis for system parameter tuning, data quality judgment, and triggering of abnormal data processing procedures. By adjusting the system parameters to the optimal configuration corresponding to the peak area of ​​the surface, higher optimization feedback values ​​can be obtained directly and stably, thereby systematically improving the overall quality level of medical data and reducing clinical problem feedback.

[0044] This implementation plan effectively quantifies the impact of medical data optimization on improving medical data quality by comparing the performance changes, clinical problem feedback frequency, data verification values, and model complexity of the medical data optimization model in real time. By calculating the optimization gain term and complexity suppression term, the final optimization feedback value is derived, providing the system with a comprehensive evaluation index to reflect the effectiveness of optimization measures. Furthermore, by integrating various factors, it reflects the overall impact of medical data optimization on improving medical data quality and ensures the correlation between feedback and actual data quality improvements.

[0045] Specifically, the steps for performing dynamic parameter adjustment, model version management, and adaptive evolution based on the optimization analysis results are as follows: By comparing the optimization feedback value and the optimization feedback threshold in real time, when the optimization feedback value is less than the optimization feedback threshold, if the performance change is less than zero or the frequency of clinical problem feedback is greater than the feedback threshold, all parameters are frozen. A root cause analysis report is generated by tracing the performance indicators on the validation set and determining if the performance indicators have decreased K times consecutively with a decrease greater than ϵ. Here, K represents the number of consecutive decreases in the performance indicator, defined as K=3, meaning that if three consecutive performance decreases meet the condition, root cause analysis is initiated; ϵ represents the performance decrease threshold, defined as ϵ=0.05, meaning that each performance decrease must exceed 5% to trigger the corresponding operation. The generated root cause analysis report is used to deeply analyze the reasons for performance changes in the medical data optimization model. Next, in an isolated sandbox computing environment, a recently verified subset of data is selected as training samples. The learning rate is reduced to retrain the medical data optimization model. The parameters are unfrozen, and online traffic is gradually restored by increasing traffic by p% each time. Simultaneously, the accuracy, recall, and throughput of the medical data optimization model are monitored. Here, p represents the percentage increase in traffic during the gradual restoration process, defined as p=10%, meaning an increase of 10% traffic each time. If the recalculated optimization feedback value is still less than the optimization feedback threshold, the model version is rolled back to the previous stable state, and all context information of failed cases is encrypted and archived in the medical knowledge base. When the optimization feedback value is greater than or equal to the optimization feedback threshold, it is confirmed that the optimization measures have achieved the expected effect. The validated medical data optimization model parameters and data processing rules are maintained. A continuous monitoring program is initiated, and incremental learning tasks are arranged based on newly accumulated qualified quality data to ensure the continuous evolution of evaluation capabilities.

[0046] In this implementation plan, by comparing the optimized feedback value with the optimized feedback threshold in real time, the model parameters of the medical data optimization model can be frozen in a timely manner when performance degradation or clinical problems occur. Root cause analysis can then be performed, and necessary retraining measures can be taken to restore the stability of the medical data optimization model. During the recovery process, traffic is gradually increased through a gradual recovery phase, while the accuracy, recall, and throughput of the medical data optimization model are monitored to ensure a gradual recovery to the optimal state while maintaining stability. If the optimization measures do not achieve the expected results, the model will revert to the previous stable version, and the context information of all failed cases will be recorded for subsequent analysis. When the optimized feedback value is greater than or equal to the optimized feedback threshold, the expected results are confirmed, and continuous monitoring and incremental learning tasks are initiated to ensure the continuous optimization and evolution of the evaluation capability of the medical data optimization model.

[0047] Specifically, this embodiment provides a quality intelligent assessment system for multi-source medical data, applied to a quality intelligent assessment method for multi-source medical data, including: a medical data acquisition and preprocessing module for acquiring multi-source medical data, performing preprocessing operations, storing and constructing a medical database; a spatiotemporal alignment module for driving spatiotemporal semantic alignment and joint representation learning of cross-modal data, constructing a medical data optimization model and outputting spatiotemporal alignment values; a quality assessment and anomaly detection module for fusing spatiotemporal correlation quantities and multi-dimensional quality indicators to perform data quality evaluation and anomaly analysis, and performing anomaly data repair and trusted dataset archiving operations based on the data quality analysis results; a data traceability module for implementing defect root cause tracing analysis based on the data end-to-end trustworthiness audit, and performing traceability path generation and interpretability report output operations based on the data trustworthiness results; and a feedback and optimization module for analyzing the optimization effect through user feedback data and the evolution status of the medical data optimization model, and performing dynamic parameter adjustment, medical data optimization model version management, and adaptive evolution operations based on the optimization analysis results.

[0048] In this implementation plan, through the collaborative work of various modules, the system can efficiently collect, preprocess, align, evaluate, and optimize medical data. In the data acquisition and preprocessing module, the system processes multi-source medical data and establishes a medical database, providing an accurate and unified data foundation for subsequent steps. The spatiotemporal alignment module ensures the consistency of imaging data and clinical records in time and space, outputting high-quality spatiotemporal alignment values. The quality assessment and anomaly detection module, through the fusion of multi-dimensional quality indicators, promptly identifies and corrects anomalies in the data, ensuring data integrity and accuracy. The data traceability module, through end-to-end reliability auditing, provides detailed root cause analysis and interpretability reports, ensuring data traceability and reliability. The feedback and optimization module performs real-time optimization based on user feedback and the performance of the medical data optimization model, ensuring continuous improvement of data processing strategies. Overall, through the close cooperation of these modules, the system achieves a comprehensive improvement in the quality of medical data, providing reliable data support for medical decision-making and research.

[0049] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0050] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A multi-source medical data oriented quality intelligent evaluation method, characterized in that, Includes the following steps: S1 collects multi-source medical data, performs preprocessing operations, stores and builds a medical database; S2 drives the spatiotemporal semantic alignment and joint representation learning of cross-modal data, constructs a medical data optimization model, and outputs spatiotemporal alignment values; S3 integrates spatiotemporal correlation quantities and multi-dimensional quality indicators to perform data quality evaluation and anomaly analysis, and performs anomaly data repair and trusted dataset archiving operations based on the data quality analysis results. S4, based on the data end-to-end credibility audit, implement defect root cause tracing analysis, and perform tracing path generation and interpretability report output operations according to the data credibility results; S5 analyzes and optimizes the results based on user feedback data and model evolution status, and performs dynamic parameter adjustment, model version management, and adaptive evolution operations based on the optimization analysis results.

2. The quality intelligent evaluation method for multi-source medical data according to claim 1, characterized in that: The specific steps for collecting multi-source medical data, performing preprocessing, storing, and constructing a medical database are as follows: The system collects image acquisition timestamps, clinical event times, equipment spatial coordinates, patient positioning identifiers, and clinical anatomical locations, and obtains medical image sequence files and clinical record streams to calculate the total number of data points. It calculates the image-medical record matching degree by comparing the patient ID and examination type in the image tags with the corresponding information in the clinical records using string similarity. It calculates the time difference between the image acquisition time and the clinical event, and judges the rationality of the time logic according to clinical guidelines to obtain the time compliance degree. It obtains the spatial matching degree for each data point by comparing the spatial coordinates in the image data with the anatomical locations in the medical record. It checks for blank or invalid values ​​in the required fields of each data point, and calculates the proportion of complete fields to obtain the information completeness degree. It checks the reasonable range of data values, verifies the time sequence logic, and checks whether the coding conforms to the standard glossary for each data point, and calculates the proportion of data that passes all checks to obtain the effective data quantity. Finally, it completes the initial association of all acquired data by the patient's unique identifier and acquisition session number, and writes it into a temporary acquisition record. A unified time reference is established for image acquisition timestamps and clinical event timestamps, converting them into timestamp formats with the same sampling starting point. A unified spatial benchmark is established for equipment spatial coordinates, patient positioning identifiers, and clinical anatomical locations. Standardized coding and structured terminology mapping are completed to eliminate ambiguity in expression. The maximum and minimum value normalization method is used to scale multi-source medical data to a unified range to achieve dimensionlessness. The standardized and normalized medical data are then appended with unique patient identifiers and collected session numbers before being stored and used to construct a medical database.

3. The intelligent quality assessment method for multi-source medical data according to claim 1, characterized in that: The specific steps for driving the spatiotemporal semantic alignment and joint representation learning of cross-modal data, constructing a medical data optimization model, and outputting spatiotemporal alignment values ​​are as follows: Using image sequence files and clinical record streams as basic inputs, timestamp matching records, event continuity records, and temporal conflict marker records are generated in the medical database based on image acquisition timestamps and clinical event times. A spatial mapping input set is constructed by combining device spatial coordinates, patient position identifiers, and clinical anatomical location descriptions, and the patient's unique identifier and acquisition session number are written as associated traceability fields. A multimodal feature joint embedding model is used to perform spatiotemporal semantic alignment calculations. The joint learning expression of image feature vectors and clinical record feature vectors is learned in a unified spatiotemporal embedding space. During training, image sequences and clinical records from the same patient in the same acquisition session form positive samples to form matching data pairs. Non-matching pairs are generated by mismatching pairs between different acquisition sessions of the same patient, between different patients, or artificially constructed, so that the Mahalanobis feature distance of matching data pairs is minimized and the Mahalanobis feature distance of non-matching pairs is maximized. Adaptive attention weight adjustment is driven by temporal conflict markers and spatial matching differences to improve the discriminative power of key alignment factors; The medical data optimization model is constructed to output spatiotemporal alignment values, matching pair determination accuracy, and version identifiers, and the model complexity is obtained by reading the model's hierarchical depth.

4. The intelligent quality assessment method for multi-source medical data according to claim 1, characterized in that: The specific steps for data quality evaluation and anomaly analysis by integrating spatiotemporal correlation quantities and multi-dimensional quality indicators are as follows: Obtain the image-medical record matching degree, temporal compliance degree, spatial matching degree, and spatiotemporal alignment value of the i-th data point; add the image-medical record matching degree, temporal compliance degree, and spatial matching degree of the i-th data point to obtain the cumulative value of the three-dimensional quality features of the data point; The sum of the three-dimensional quality features of all data points is obtained by summing the total sum of the quality features. The average three-dimensional quality feature value is obtained by dividing the sum of the accumulated quality feature values ​​by the total number of data points. The average three-dimensional quality feature value is multiplied by the spatiotemporal alignment value of each data point to obtain the data quality value of each data point.

5. The intelligent quality assessment method for multi-source medical data according to claim 1, characterized in that: The specific steps for performing abnormal data repair and trusted dataset archiving based on data quality analysis results are as follows: By comparing data quality values ​​and quality thresholds in real time, when a data quality value falls below the quality threshold, a data repair process is initiated: This includes checking and re-acquiring missing data from incomplete images or clinical records; comparing timestamp differences between image data and clinical data, adjusting data that does not conform to the chronological order to correct data timeliness; using image registration technology to align image data to standard anatomical positions, ensuring consistency between images and anatomical locations in clinical records; using isolated forests to mark potentially abnormal data and re-acquiring data; and recalculating the data quality value after repair. If the value is still below the quality threshold, a manual verification instruction is issued and an early warning report is generated. When the data quality value is greater than or equal to the quality threshold, it is confirmed that the data meets the quality standards. A traceability dataset is constructed, the total number of traceability data points is calculated, and the data is archived into the medical database.

6. The intelligent quality assessment method for multi-source medical data according to claim 1, characterized in that: The specific steps for implementing defect root cause analysis based on end-to-end data credibility audit are as follows: Obtain the information completeness, data validity, and data quality value of the j-th traceability data point; multiply the information completeness and data validity of the j-th traceability data point to obtain the complete and valid product value of the data point; Sum the complete and valid product values ​​of all traceable data points to obtain the total sum of complete and valid product values; Divide the sum of complete and valid product values ​​by the total number of traceable data points L to obtain the average complete and valid product value; multiply the average complete and valid product value by the data quality value to obtain the data verification value.

7. The intelligent quality assessment method for multi-source medical data according to claim 1, characterized in that: The specific steps for generating a traceability path and outputting an interpretability report based on the data reliability results are as follows: By comparing data verification values ​​and verification thresholds in real time, when the data verification value is less than the verification threshold, a traceability link graph from the data collection source to the current state is constructed and presented based on the data flow log, identifying the metadata and transformation records of each link; based on the traceability link graph, a re-verification request containing specific missing fields is sent to the source data to verify and improve the traceability information; The collected source information is updated and supplemented in the medical database, and a structured and interpretable report of the traceability analysis conclusions is output. When the data verification value is greater than or equal to the verification threshold, the permission flag of the data point is updated to "verified". Update the permission flag in the database to remove access restrictions on batch data; write metadata update messages to the message queue to send update notifications to application services that subscribe to this data type.

8. The intelligent quality assessment method for multi-source medical data according to claim 1, characterized in that: The specific steps for analyzing the optimization effect through user feedback data and model evolution status are as follows: The model performance change was obtained by aligning the previous and current versions of the medical data optimization model with the same metrics on the version evaluation snapshot and subtracting the values ​​from the two evaluations; the clinical problem feedback frequency was obtained by collecting data problem reports submitted by clinicians and statistically analyzing the frequency of occurrence of various problems. Obtain data validation values, model performance changes, clinical problem feedback frequency, and model complexity; calculate the product of the model performance change adjustment coefficient and the absolute value of the model performance change to obtain the performance change weighting term. The user feedback intensity adjustment coefficient is calculated by multiplying it by the frequency of clinical problem feedback to obtain the user feedback weighting term; the performance change weighting term is added to the user feedback weighting term and then one is added to obtain the optimization gain term; the data verification value is multiplied by the optimization gain term to obtain the retrospective weighting term; the model complexity adjustment coefficient is calculated by multiplying it by the model complexity and then one is added to obtain the complexity suppression term; the retrospective weighting term is divided by the complexity suppression term to obtain the optimization feedback value.

9. The intelligent quality assessment method for multi-source medical data according to claim 1, characterized in that: The specific steps for performing dynamic parameter adjustment, model version management, and adaptive evolution operations based on the optimization analysis results are as follows: By comparing and optimizing feedback values ​​and feedback thresholds in real time, if the performance change is less than zero or the frequency of clinical problem feedback is greater than the feedback threshold when the optimized feedback value is less than the optimized feedback threshold, all model parameters are frozen. In addition, the performance indicators on the validation set are located to show a continuous decrease of K times with a decrease of greater than ϵ, and a root cause analysis report is generated. In an isolated sandbox computing environment, a recently verified subset of data is selected as training samples, the learning rate is reduced, the model is retrained, the parameter freeze is lifted, and the online traffic is gradually restored by increasing the traffic by p% each time. At the same time, the accuracy, recall, and throughput of the medical data optimization model are monitored. If the recalculated optimization feedback value is still less than the optimization feedback threshold, the model version is rolled back to the previous stable state, and all context information of the failed cases is encrypted and archived to the medical knowledge base. When the optimization feedback value is greater than or equal to the optimization feedback threshold, it is confirmed that the optimization measures have achieved the expected results. The validated medical data optimization model parameters and data processing rules are maintained. A continuous monitoring program is initiated, and based on the newly accumulated qualified quality data, incremental learning tasks are arranged to ensure the continuous evolution of evaluation capabilities.

10. A quality intelligent assessment system for multi-source medical data, employing the quality intelligent assessment method for multi-source medical data as described in any one of claims 1-9, characterized in that... ,include: The medical data acquisition and preprocessing module is used to collect multi-source medical data, perform preprocessing operations, store and build a medical database. The spatiotemporal alignment module is used to drive the spatiotemporal semantic alignment and joint representation learning of cross-modal data, build a medical data optimization model, and output spatiotemporal alignment values. The quality assessment and anomaly detection module is used to integrate spatiotemporal correlation quantities and multi-dimensional quality indicators to perform data quality evaluation and anomaly analysis, and to perform anomaly data repair and trusted dataset archiving operations based on the data quality analysis results. The data traceability module is used to perform defect root cause analysis based on the credibility audit of the entire data chain, and to perform traceability path generation and interpretability report output operations based on the data credibility results; The feedback and optimization module is used to analyze the optimization effect through user feedback data and model evolution status, and to perform dynamic parameter adjustment, model version management and adaptive evolution operations based on the optimization analysis results.

Citation Information

Patent Citations

  • Methods, devices, equipment and readable storage media for quality assessment of medical data

    CN109615204B

  • Scientific and technological achievement prediction and evaluation system based on big data

    CN120087554A