Intelligent management method and system for multi-source and multi-mode clinical research data

By vectorizing and feature expansion fusion of multi-source multi-modal clinical research data, and combining the in-order and arrangement standards for analysis, the problem of difficulty in integrating multi-source multi-modal data is solved, and effective correlation analysis of multi-modal data and the accuracy of subject screening is achieved.

CN120089401APending Publication Date: 2025-06-03HEFEI UNIV OF TECH
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510254086.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The difficulty in integrating multi-source and multimodal clinical research data has led to the inability to achieve effective correlation analysis, which limits the comprehensiveness of subject evaluation and leads to inaccurate subject screening.

Method used

By obtaining the multimodal clinical data of the patient (including text, image and gene data), it is vectorized and feature expansion fusion, obtaining the fusion feature vector, combining the inclusion and ranking standards for analysis, calculating the degree of matching differences, and screening the patient according to the preset screening threshold, obtaining a list of potential subjects that meet the requirements.

Benefits of technology

Effective correlation analysis of multimodal data is realized, which improves the comprehensiveness and scientificity of data utilization and analysis, makes the matching difference more accurate and improves the accuracy of subject screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089401A_ABST
    Figure CN120089401A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent management method and system for multi-source and multi-mode clinical research data, and relates to the technical field of data processing. The method comprises the following steps: acquiring an in-and-out standard of a subject and multi-modal clinical data of a patient; respectively vectorizing the text data, the image data and the gene data to obtain a text feature vector, a high-dimensional feature vector and a gene feature vector, so as to facilitate subsequent data integration; performing feature extension fusion on the text feature vector, the high-dimensional feature vector and the gene feature vector to obtain a fused feature vector, and effectively associating multi-modal data; according to the fusion feature vector and the in-arrangement feature vector of the in-arrangement standard, the matching difference degree is obtained, so that the multi-modal data play a mutual synergistic role when the matching difference degree is analyzed, and the matching difference degree is more accurate; according to the matching difference degree and a preset screening threshold value, the patients are screened, a list of potential subjects meeting requirements is obtained, and clinical researchers can make accurate decisions conveniently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to an intelligent management method and system for multi-source and multi-modal clinical research data. Background Art

[0002] In clinical research, subjects are the core components of the research, and they are irreplaceably necessary. They are the core driving forces for scientific hypothesis verification, medical knowledge accumulation, new therapy verification, and patient health improvement. Through strict subject screening, the scientific nature, safety, and social value of the research can be ensured.

[0003] However, currently, clinical data comes from a wide range of sources, covering multiple types of platforms such as hospital information systems (HIS), laboratory information systems (LIS), and picture archiving and communication systems (PACS). The data forms are diverse, including structured (such as electronic medical record forms), semi-structured (such as test reports), and unstructured (such as images, free text records) data. The degree of data standardization is insufficient, and the data formats between different systems vary significantly (such as field naming, unit specifications, and coding systems are not unified), resulting in difficulties in data integration and a sharp increase in management complexity.

[0004] Existing solutions are difficult to effectively integrate structured, semi-structured, and unstructured data from different systems, lacking a unified term mapping system and cross-modal fusion mechanism, resulting in isolated storage of multi-modal data such as text, images, and genes, and unable to achieve effective correlation analysis, making the multi-modal data unable to work together, restricting the comprehensiveness and scientific nature of subject evaluation and clinical research, and resulting in limited accuracy of intelligent decision-making. Summary of the Invention

[0005] The problem to be solved by the present invention is the difficulty in integrating multi-source and multi-modal data, unable to achieve effective correlation analysis, restricting the comprehensiveness of subject evaluation, and resulting in inaccurate subject screening.

[0006] To solve the above problems, in a first aspect, the present invention provides an intelligent management method for multi-source and multi-modal clinical research data, including:

[0007] Obtain the inclusion and exclusion criteria of the subject and the multi-modal clinical data of the patient, where the multi-modal clinical data includes text data, image data, and gene data;

[0008] Vectorize the text data, the image data, and the gene data respectively to obtain a text feature vector, a high-dimensional feature vector, and a gene feature vector;

[0009] Perform feature extension and fusion on the text feature vector, the high-dimensional feature vector, and the gene feature vector to obtain a fusion feature vector;

[0010] Based on the inclusion and exclusion feature vectors of the fusion feature vector and the inclusion and exclusion criteria, the matching difference degree is obtained;

[0011] According to the matching difference degree and a preset screening threshold, the patients are screened to obtain a list of potential subjects meeting the requirements.

[0012] Optionally, the step of performing feature extension and fusion on the text feature vector, the high-dimensional feature vector, and the gene feature vector to obtain a fusion feature vector includes:

[0013] Inputting the text feature vector, the high-dimensional feature vector, and the gene feature vector into a two-layer neural network MLP respectively to obtain corresponding feature weights;

[0014] Performing cross-modal extended convolution processing on the text feature vector, the high-dimensional feature vector, and the gene feature vector to obtain feature enhanced vectors, where the feature enhanced vectors include a text feature enhanced vector, a high-dimensional feature enhanced vector, and a gene feature enhanced vector;

[0015] According to the feature weights and the feature enhanced vectors, the features are weighted and spliced to obtain a multi-modal spliced vector;

[0016] Inputting the multi-modal spliced vector into a multi-layer fully connected layer for non-linear transformation to obtain a fusion feature vector.

[0017] Optionally, the step of obtaining the inclusion and exclusion criteria of the subjects and the multi-modal clinical data of the patients includes:

[0018] Extracting the subject requirement text and the data requirement text from the clinical research protocol document;

[0019] Extracting logical rules from the subject requirement text, parsing numerical rules, identifying text rules, and mapping non-standard terms to standardized codes to obtain an inclusion and exclusion criteria table;

[0020] Extracting trial design parameters from the data requirement text, mapping unstructured requirements to standardized data elements, and associating with the data source system to obtain a data requirement table;

[0021] According to the data types in the inclusion and exclusion criteria table, obtaining the multi-modal clinical data of the patients from multiple data source systems.

[0022] Optionally, after obtaining the inclusion and exclusion criteria of the subjects and the multi-modal clinical data of the patients, the intelligent management method for multi-source multi-modal clinical research data further includes multi-modal clinical data preprocessing; the multi-modal clinical data preprocessing includes:

[0023] Using the HanLP toolkit and named entity recognition NER to preprocess the text data to obtain structured text;

[0024] The ITK toolkit is used to perform gray level equalization and spatial normalization on the image data to obtain an image standardization file;

[0025] Based on BioPython, the gene data is standardized in terms of expression level and wavelet denoising is performed to obtain a gene standard format file.

[0026] Optionally, the intelligent management method for multi-source and multi-modal clinical research data further includes subject recruitment and screening; the subject recruitment and screening include:

[0027] According to the inclusion and exclusion criteria table and the data requirement table, a recruitment poster for the clinical research project is generated;

[0028] The recruitment poster for the clinical research project is sent to the patient clients in the list of potential subjects in a targeted manner, and the recruitment poster for the clinical research project is published on the network for subject recruitment;

[0029] According to the registration information collected from the feedback, the registered patients are preliminarily screened;

[0030] The list of registered patients after preliminary screening is compared with the list of potential subjects to obtain a highly matching list;

[0031] According to the matching difference degree, the potential subjects in the highly matching list are ranked in priority;

[0032] If the matching difference degree is less than the preset passing threshold, the potential subject corresponding to the matching difference degree is exempt from review, and the potential subject corresponding to the matching difference degree is included in the list of subjects passing the review;

[0033] The list of subjects passing the review, the list of registered patients after preliminary screening and the registration information are pushed to the researcher client.

[0034] Optionally, the intelligent management method for multi-source and multi-modal clinical research data further includes generating a clinical research data collection plan; the generating of the clinical research data collection plan includes:

[0035] According to the data requirement table, a data collection framework is constructed, where the data collection framework includes a core data domain, data collection rules and data quality control rules;

[0036] According to the data types and data collection rules within the core data domain, data collection plan documents for multiple data source systems are generated;

[0037] According to the data collection plan documents, multi-source and multi-modal data sources are determined, and a data collection link is established.

[0038] Optionally, the intelligent management method for multi-source and multi-modal clinical research data further includes multi-modal data fusion analysis during the clinical research of the subject; the multi-modal data fusion analysis during the clinical research of the subject includes:

[0039] Obtain video data, audio data, and text expression data during the clinical research of the subject;

[0040] Extract features from the video data and audio data respectively to obtain video spatio-temporal features and audio spatio-temporal features;

[0041] According to the video spatio-temporal features and audio spatio-temporal features, use the attention mechanism to obtain video enhanced features and audio enhanced features;

[0042] Extract features from the text expression data to obtain a global text vector;

[0043] Perform vector splicing processing on the video spatio-temporal features, audio spatio-temporal features, video enhanced features, audio enhanced features, and global text vector, and pass through the classification layer to obtain the predicted emotion category of the subject;

[0044] Push the predicted emotion category to the researcher client so that the researcher can master the reaction of the subject during the clinical research.

[0045] In a second aspect, the present invention also provides an intelligent management system for multi-source and multi-modal clinical research data, including:

[0046] A historical data acquisition module for acquiring the inclusion and exclusion criteria of the subject and the multi-modal clinical data of the patient, where the multi-modal clinical data includes text data, image data, and gene data;

[0047] A data vectorization module for vectorizing the text data, the image data, and the gene data respectively to obtain a text feature vector, a high-dimensional feature vector, and a gene feature vector;

[0048] A data fusion module for performing feature extension and fusion on the text feature vector, the high-dimensional feature vector, and the gene feature vector to obtain a fusion feature vector;

[0049] A matching analysis module for obtaining the matching difference degree according to the fusion feature vector and the inclusion and exclusion feature vector of the inclusion and exclusion criteria;

[0050] A patient screening module for screening the patients according to the matching difference degree and a preset screening threshold to obtain a list of potential subjects meeting the requirements.

[0051] The present invention provides an intelligent management method and system for multi-source and multi-modal clinical research data. Compared with the prior art, it has the following beneficial effects:

[0052] By performing vectorization processing on the multi-modal clinical data of patients, quantifying structured, semi-structured, and unstructured data, text feature vectors, high-dimensional feature vectors, and gene feature vectors are obtained, facilitating subsequent integration of the data. The quantified data is expanded and fused to obtain fused feature vectors, effectively correlating the multi-modal data. Analyzing and calculating using the fused feature vectors and the inclusion and exclusion feature vectors of the inclusion and exclusion criteria to obtain the matching difference degree, thereby enabling the multi-modal data to play a synergistic role with each other when analyzing the matching difference degree, improving the data utilization rate, making the data analysis more comprehensive and scientific, making the obtained matching difference degree more accurate. Finally, according to the matching difference degree and the preset screening threshold, patients are screened to obtain a list of potential subjects that meet the requirements, improving the accuracy of the list, and pushing the list of potential subjects to the client of clinical researchers for clinical researchers to make accurate decisions. Brief Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is a schematic flowchart of an intelligent management method for multi-source multi-modal clinical research data provided by an embodiment of the present invention;

[0055] Figure 2 It is a schematic full-process flowchart of an intelligent management method for multi-source multi-modal clinical research data provided by an embodiment of the present invention;

[0056] Figure 3 It is a schematic data flow diagram of an intelligent management method for multi-source multi-modal clinical research data provided by an embodiment of the present invention;

[0057] Figure 4 It is a schematic flowchart of an intelligent management method during the clinical research of subjects provided by an embodiment of the present invention;

[0058] Figure 5 It is a schematic structural diagram of an intelligent management system for multi-source multi-modal clinical research data provided by an embodiment of the present invention. Detailed Embodiments

[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described clearly and completely. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0060] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0061] As Figure 1 shown, an intelligent management method for multi-source and multi-modal clinical research data provided by an embodiment of this application includes:

[0062] S100: Obtain the inclusion and exclusion criteria of the subject and the multi-modal clinical data of the patient, where the multi-modal clinical data includes text data, image data, and gene data.

[0063] S200: Vectorize the text data, the image data, and the gene data respectively to obtain a text feature vector, a high-dimensional feature vector, and a gene feature vector.

[0064] S300: Perform feature extension and fusion on the text feature vector, the high-dimensional feature vector, and the gene feature vector to obtain a fused feature vector.

[0065] S400: Obtain a matching difference degree according to the fused feature vector and the inclusion and exclusion feature vector of the inclusion and exclusion criteria.

[0066] S500: Screen the patient according to the matching difference degree and a preset screening threshold to obtain a list of potential subjects meeting the requirements.

[0067] In this embodiment, by performing vectorization processing on the multi-modal clinical data of the patient, quantifying structured, semi-structured, and unstructured data, etc., a text feature vector, a high-dimensional feature vector, and a gene feature vector are obtained, which is convenient for subsequent data integration. The quantified data is extended and fused to obtain a fused feature vector, effectively associating the multi-modal data. Using the fused feature vector and the inclusion and exclusion feature vector of the inclusion and exclusion criteria for analysis and calculation, a matching difference degree is obtained. Furthermore, the multi-modal data plays a synergistic role with each other when analyzing the matching difference degree, improving the data utilization rate, making the data analysis more comprehensive and scientific, making the obtained matching difference degree more accurate. Finally, according to the matching difference degree and the preset screening threshold, the patient is screened to obtain a list of potential subjects meeting the requirements, improving the accuracy of the list, and pushing the list of potential subjects to the client of the clinical research personnel for the clinical research personnel to make accurate decisions.

[0068] As Figure 2 and Figure 3 shown below, the entire process of the intelligent management method for multi-source and multi-modal clinical research data will be described in detail.

[0069] Step 1: Extraction of Clinical Research Protocol Requirements

[0070] Obtain the inclusion and exclusion criteria of the subjects. The extraction of clinical research protocol requirements is the starting point of the entire intelligent management process of clinical research data, aiming to transform the key information in the clinical trial protocol into structured and computable logical rules. The specific process is as follows:

[0071] (1) Acquisition of Clinical Research Protocol

[0072] Extract the subject requirement text and data requirement text from the clinical research protocol document. For electronic clinical research protocol documents (such as PDF / Word), use OCR technology to extract the text content and parse the key fields based on NLP algorithms, including subject requirement-related fields (such as inclusion criteria, exclusion criteria, subject characteristic descriptions) and data requirement-related fields (such as trial design type, data collection requirements, data quality standards). The parsed content is stored in a MySQL database in a structured manner, and the main table fields cover protocol ID, protocol name, version number, original subject requirement text, and original data requirement text.

[0073] (2) Extraction of Subject Requirements

[0074] Extract the logical rules from the subject requirement text, parse the numerical rules, identify the text rules, and map the non-standard terms to normalized codes to obtain the inclusion and exclusion criteria table. Extract the logical rules from the original text based on the rule engine (Drools) and regular expressions. The numerical rules automatically parse the age range (such as 18 ≤ age ≤ 65) or biochemical index threshold (such as HbA1c ≥ 6.5%). The text rules identify disease names (such as "diabetes") or treatment history (such as "not receiving insulin treatment"). By integrating the ICD-10 coding library and the SNOMED CT terminology set, map the non-standard terms to normalized codes (for example, "hyperglycemia" is mapped to ICD-10:E11.9, and "glycated hemoglobin" is mapped to LOINC:4548-4), and at the same time establish a mapping relationship table (fields include rule ID, original term, standard code, mapping confidence). Finally, output the structured inclusion and exclusion criteria table (fields include rule ID, rule type, logical expression, associated medical terms), support JSON / CSV format export, and be called by the subsequent potential subject intelligent mining module.

[0075] The mapping confidence is a quantitative indicator that measures the accuracy of the match between non-standardized medical terms and terms in the standardized coding library. Calculation methods: (1) Exact match: directly match the main name in the term library (e.g., "E11.9" in ICD-10 corresponds to "type 2 diabetes"), and the confidence is 100%; (2) Synonym expansion: use the synonym list of the term library (e.g., map "HbA1c" to "glycated hemoglobin"), which is achieved through preset synonym rules, and the confidence is usually 80%-95%; (3) Semantic similarity algorithm: for terms that cannot be exactly matched, use NLP models (such as BERT, UMLS MetaMap) to calculate the semantic similarity with the standard terms, and output the probability value as the confidence.

[0076] The logical expression is a machine-readable conditional rule extracted from the inclusion and exclusion criteria text used to describe the subject screening logic, supporting nested structures and operators. Example of its formal expression: (age ≥ 18 AND age ≤ 65) AND (diagnosis = "diabetes" OR HbA1c ≥ 6.5%).

[0077] (3) Data requirement extraction

[0078] Extract the trial design parameters from the data requirement text, map the unstructured requirements to standardized data elements, and associate them with the data source system to obtain the data requirement table.

[0079] Classify and extract the key parameters of the trial design, including the trial design type (such as randomized controlled trial, observational study), visit plan (such as baseline visit, follow-up cycle "once every 4 weeks"), and data collection points (such as vital signs, laboratory tests). Map the unstructured requirements to standardized data elements according to the CDISC SDTM standard (for example, map "baseline fasting blood glucose" to the numerical result SDTM.LB.LBSTRESN under the standardized unit), and associate with data source systems such as HIS, LIS, and PACS. The generated structured data requirement table (fields including data element ID, element name, source system, collection frequency, SDTM mapping) is output in JSON format for subsequent automatic generation of the data collection plan.

[0080] Step 2: Data source selection and data acquisition

[0081] Obtain the multi-modal clinical data of the patient. According to the data types in the inclusion and exclusion criteria table, obtain the multi-modal clinical data of the patient from multiple data source systems. The multi-modal clinical data includes text data, image data, and gene data.

[0082] (1) Precise data source matching based on subject requirements

[0083] Based on the inclusion and exclusion criteria generated in "1. Clinical Research Protocol Requirement Extraction", dynamically select the types and scopes of data sources to be collected. For diabetes drug trials, according to the extracted rules (such as "HbA1c≥6.5%" and "not using insulin treatment"), it is necessary to obtain the diabetes duration, medication records, and demographic data (age, gender) of patients from the hospital information system (HIS), and simultaneously retrieve blood glucose, glycated hemoglobin, and liver and kidney function indicators (text data) from the laboratory information management system (LIS); in the cardiac pacemaker trial, based on exclusion criteria such as "bradycardia", it is necessary to screen the electronic medical records (text data) of patients with cardiac disease characteristics from the HIS, and extract cardiac ultrasound image data (image data) from the picture archiving and communication system (PACS) to evaluate key functional parameters such as ventricular ejection fraction; for gene therapy trials, based on the inclusion requirement of "specific gene mutations", it is necessary to obtain mutation site information (gene data) from the gene database, and at the same time associate the clinical symptom records (such as muscle atrophy, movement disorder) in the HIS to ensure the consistency of genotype and phenotype.

[0084] (2) Technical Adaptation and Efficient Acquisition of Heterogeneous Data Sources

[0085] According to the differences in data source types and protocols, design protocol adaptation middleware that supports custom field mapping templates, allowing researchers to dynamically adjust interface configurations according to the characteristics of local data sources, and deploy targeted technical solutions to achieve data connection and extraction. The HIS system uses the HL7 standard interface to dock with the RESTful API to obtain structured patient information (such as name, medical history) and semi-structured electronic medical record texts; the LIS system directly extracts numerical results of test items (such as blood glucose values, creatinine levels) through ODBC / JDBC database connection technology; the PACS system transmits standardized medical images based on the DICOM communication protocol to ensure image resolution and metadata integrity; the gene database encrypts and transmits gene sequencing data through a dedicated API interface authenticated by OAuth 2.0. Structured data (such as test values, patient age) is batch exported through SQL queries, and unstructured data (such as DICOM images, free-text medical records) respectively use image processing tools (such as ITK-SNAP) and natural language processing (NLP) models to complete format conversion and key information extraction.

[0086] Step 3: Preprocessing of Multi-modal Clinical Data of Potential Subjects

[0087] (1) Requirement Analysis of Multi-modal Data Fusion

[0088] The requirement analysis of multi-modal data fusion is driven by the inclusion and exclusion criteria of the clinical trial protocol, and the triggering condition is when the inclusion and exclusion criteria involve multi-dimensional medical characteristics.

[0089] Based on the inclusion and exclusion criteria defined in "1. Extraction of Clinical Research Protocol Requirements" (such as specific disease characteristics, biochemical thresholds, gene mutation requirements), clarify the necessity of multi-source data integration. For example, in a diabetes drug trial, it is necessary to synchronously analyze the patient medication records (text data) in HIS, the blood glucose indicators (numeric data) in LIS, and the complication images (image data) in PACS; in gene therapy research, it is necessary to combine gene sequencing data (sequence data) with the clinical symptom descriptions (text data) in HIS. By fusing multi-modal data such as text, images, and genes, the limitations of single data dimension in subject evaluation can be compensated (such as the inability to quantify lesion morphology relying solely on EHR text), ensuring the comprehensiveness and accuracy of subsequent intelligent mining.

[0090] (2) Preprocessing of multi-source heterogeneous data

[0091] Adopt adapted preprocessing methods for different data types. Use the HanLP toolkit and named entity recognition NER to preprocess text data to obtain structured text; specifically, for text data (such as electronic health records EHR), perform lexical analysis (word segmentation, part-of-speech tagging) and dependency syntax parsing through the HanLP tool, and use named entity recognition (NER) to accurately extract disease names (such as "type 2 diabetes"), symptoms (such as "polydipsia and polyuria"), and treatment methods (such as "oral metformin"), and output a structured entity table. Use the ITK toolkit to perform gray level equalization and spatial normalization on image data to obtain a standardized image file; specifically, for medical image data, use the ITK toolkit to perform gray level equalization (enhance image contrast) and spatial normalization (unify pixel resolution), eliminate the influence of device differences on image analysis, and generate a standardized DICOM (International Standard for Medical Digital Imaging and Communications) file. Based on BioPython, perform expression level normalization and wavelet denoising on gene data to obtain a gene standard format file; specifically, for gene data, perform expression level normalization (Z-score correction) and wavelet denoising (filter out high-frequency sequencing noise) based on BioPython, retain low-frequency mutation signals related to the target disease, and output a quality-controlled FASTQ / VCF file.

[0092] (3) Extraction of multi-modal data feature vectors

[0093] Extract features from the standardized data to provide input for intelligent mining of potential subjects. Vectorize text data, image data, and gene data respectively to obtain text feature vectors, high-dimensional feature vectors, and gene feature vectors. Vectorization of text data: For text data such as electronic medical records and follow-up records, use the pre-trained medical language model BioBERT for word segmentation, and combine the medical knowledge graph to convert the text into a feature vector Z EHR; For image feature extraction, the DenseNet model is used to extract multi-level features from the preprocessed images. By means of dense connections, the morphological details of the lesions (such as the diameter of lung nodules and the ejection fraction of the heart) are retained, and a high-dimensional feature vector Z is generated. image ; For gene feature extraction, k-mer combined with the CountVectorizer tool is used to mine the non-linear associations between gene mutations and phenotypes, and the gene feature vector Z is output. genetic . Vectorize the multi-modal data to facilitate subsequent data integration, make full use of the multi-modal data, and avoid the problem of data silos.

[0094] Step 4. Intelligent mining of potential subjects

[0095] Based on the extracted features and the inclusion and exclusion criteria of the trial project, a deep multi-modal fusion model based on the attention mechanism is constructed for intelligent mining of potential subjects. The specific steps are as follows:

[0096] (1) Multi-modal attention weight assignment

[0097] The model automatically learns the importance of different features through the attention mechanism. The text feature vector, the high-dimensional feature vector, and the gene feature vector are respectively input into a two-layer neural network MLP to obtain the corresponding feature weights. Taking the EHR text as an example, assuming the EHR feature vector Z EHR contains information such as patient age, medical history, and medication records, and these features are analyzed through a two-layer neural network (MLP). The first layer uses a non-linear activation function (ReLU) to mine complex relationships (such as the association between medication history and blood glucose levels).

[0098] h 1 = ReLU(W 1 Z EHR + b 1 )

[0099] where W 1 and b 1 are the parameters of the first-layer neural network MLP. W 1 represents the weight matrix of the first-layer neural network MLP, and b 1 represents the bias vector of the first-layer neural network MLP. W 1 serves as a d mid × d EHR weight matrix, responsible for linearly transforming the features of Z EHR , and h 1 is d midThe 1D offset vector is used to adjust the transformed eigenvalues, while the ReLU activation function introduces non-linearity, enhancing the model's ability to mine complex feature relationships in EHR data, identifying information that has an important impact on subject screening, and preparing for subsequent calculation of the attention weight vector. The second layer outputs the weight of each feature.

[0100]

[0101] Among them, W 2 and b 2 are the parameters of the second-layer neural network MLP, which are learnable parameters of the model (randomly initialized at the start of training and automatically updated and optimized through training data). W 2 represents the weight matrix of the second-layer neural network MLP, and b 2 represents the offset vector of the second-layer neural network MLP. To make the attention weight vector accurately reflect the feature importance, the Softmax function is used to normalize α EHR to obtain Each element represents the relative importance of the corresponding feature in the EHR modality for subject screening. For example, when judging subjects in a diabetes clinical trial, blood glucose-related features may be assigned higher weights. For the imaging feature vector Z image and the gene feature vector Z genetic the same steps are performed to obtain and The calculation process is similar to that of the EHR modality.

[0102] The summarized feature weights are:

[0103] Among them, h 1 = ReLU(W 1 Z j + b 1 ), W 1 , W 2 , b 1 and b 2 are the parameters of the neural network MLP. W 1 represents the weight matrix of the first-layer neural network MLP, and b 1 represents the offset vector of the first-layer neural network MLP. W 2 represents the weight matrix of the second-layer neural network MLP, and b 2 represents the offset vector of the second-layer neural network MLP. ReLU represents the activation function, and Z j represents the feature vector of the j-th modality, represents the feature weight of the feature vector of the j-th modality.

[0104] (2) Cross-modal Feature Interaction Modeling

[0105] Single modality data cannot fully reflect the patient's condition, so cross-modal associations need to be established. Perform cross-modal extended convolution processing on the text feature vector, high-dimensional feature vector, and gene feature vector to obtain a feature enhancement vector, where the feature enhancement vector includes a text feature enhancement vector, a high-dimensional feature enhancement vector, and a gene feature enhancement vector. For example, in lung cancer clinical trials, the "smoking history" in the EHR may be related to the "lung nodule size" in the image, while the "EGFR mutation" in the genetic data may affect the strength of the association between the two. First, dimensional alignment is performed to adjust features of different structures (such as text sequences, image matrices, and gene vectors) to the same dimension. For example, the EHR text features are expanded into a matrix that matches the image features for joint analysis. Cross-modal convolution: Use a convolution kernel to simultaneously scan text features (text data) and image features (image data) to capture spatial and semantic associations. Assume that the convolution kernel is K and the dimension is k h ×k w ×C in ×C out , the convolution output feature map is:

[0106]

[0107] in Represents the EHR feature map and image feature map after splicing and expansion in the channel dimension, and then performing convolution calculation. For example, in lung cancer screening, the convolution operation may find a strong correlation between "smoking for more than 10 years" and "right upper lobe nodule diameter > 3cm". Feature map F based on cross-modal convolution output E-I Generate query matrix Q = W through linear transformation Q ·F E-I Sum key matrix K = W K ·F E-I , where W Q and W K is a learnable weight matrix (automatically updated without manual setting). Then, by calculating the similarity between features at different positions, we get the attention matrix A E-I

[0108]

[0109] where d k is the dimension of the key vector, which serves as a scaling factor to prevent the dot product value from being too large, resulting in unstable gradients. The weighted sum of features based on the attention matrix further strengthens the interaction between the two modalities and mines more comprehensive potential subject feature information. For example, first weight the image feature B to obtain A E-I B represents the enhanced part of the text feature on the image feature. Add the weighted output to the original text feature D, retain the original information and enhance the interaction, and get the text feature enhancement vector D+AE-I ·B, finally obtain the cross-modal enhanced text feature enhancement vector, which retains the features of the original text feature vector while enhancing its expressive ability.

[0110] Similarly, if necessary, similar cross-modal interaction operations can be performed on EHR and gene data, and imaging and gene data. By calculating the cross-modal attention matrix, the key associations are further strengthened. For example, the "ALK fusion mutation" in gene data may enhance the association weight between "mediastinal lymph node enlargement" in imaging and "cough symptom" in EHR.

[0111] (3) Feature fusion and matching decision

[0112] The ultimate goal is to screen out the subjects who fully meet the inclusion and exclusion criteria. According to the feature weights and feature enhancement vectors, the features are weighted and concatenated to obtain a multi-modal concatenated vector. First, perform feature fusion, and concatenate the weighted features of each modality (such as high-weight blood glucose values, tumor imaging features, gene mutations) into a unified vector (i.e., multi-modal concatenated vector):

[0113]

[0114] Among them, V concat represents the multi-modal concatenated vector, represents the weight of the text feature vector, represents the text feature enhancement vector, represents the weight of the high-dimensional feature vector, represents the high-dimensional feature enhancement vector, represents the weight of the gene feature vector, represents the gene feature enhancement vector. Input the multi-modal concatenated vector into a multi-layer fully connected layer for non-linear transformation to obtain the fused feature vector. Specifically, the multi-modal concatenated vector is used as the input of the fusion network composed of multi-layer fully connected layers. Through the non-linear transformation of the multi-layer fully connected layer, the potential features after multi-modal data fusion are deeply mined, enabling the model to express the features of the subjects more comprehensively and accurately. Finally, the final fused feature vector Z total is obtained through the output layer. Secondly, for the inclusion and exclusion standard matching, compare the patient's fused feature vector Z total with the inclusion and exclusion feature vector Z criteria of the inclusion and exclusion criteria. According to the fused feature vector and the inclusion and exclusion feature vector of the inclusion and exclusion criteria, obtain the matching difference degree:

[0115]

[0116] Among them, C represents the matching difference degree, Z total (i) represents the i-th fused feature vector, Z criteria(i) represents the i-th entry feature vector, and the i-th fused feature vector corresponds to the i-th entry feature vector.

[0117] According to the matching difference and the preset screening threshold, the patients are screened to obtain a list of potential subjects that meet the requirements. For example, if the matching difference is less than the preset screening threshold (such as 0.5), it is considered a match and a list of potential subjects is automatically generated for review by researchers.

[0118] For example, the long-term blood glucose monitoring data of diabetic patients cannot be automatically associated with the imaging features of their complications (such as retinopathy), resulting in fragmented analysis of the disease progression; gene mutation information (such as BRCA1 mutation) and clinical symptom descriptions (such as breast cancer mass morphology) lack dynamic interaction, making it difficult to explore the potential association between genotype and phenotype, limiting the comprehensiveness of subject evaluation and affecting the scientific nature of research conclusions. Using the above method, multimodal data is fused to achieve cross-modal data analysis effects, enhance cross-modal data interaction modeling capabilities, and achieve semantic-level feature fusion. Subject screening will be more accurate and efficient, providing a higher-quality subject group for clinical research.

[0119] Step 5: Digital Subject Recruitment

[0120] (1) Digital release of clinical research projects

[0121] Generate a clinical research project recruitment poster based on the inclusion criteria table and data requirements table. Automatically generate a clinical research project poster based on the clinical research plan obtained in the first step, including the research purpose, research institution, basic inclusion conditions, privacy protection statement, etc., with an embedded online registration QR code and jump link.

[0122] The clinical research project recruitment poster is sent to the patient clients in the potential subject list, and the clinical research project recruitment poster is published on the Internet for subject recruitment. After the potential subject list generated above is confirmed by the clinical researchers, the recruitment poster is pushed to the patients in the potential subject list, and the subject recruitment advertisement is published on the Internet through the social media platform.

[0123] (2) Online registration and informed consent of subjects

[0124] According to the registration information collected through feedback, the registered patients are preliminarily screened. After the subjects enter the online registration interface, they fill out an electronic form containing basic information, medical history and medication records. The system verifies the data format in real time and automatically filters those who obviously do not meet the requirements to achieve preliminary pre-screening. The paper version of the informed consent form is converted into a scrollable web page format, supports font enlargement / reduction, adds a floating explanation box to professional terms, and adds a mandatory reading mechanism (the subjects need to slide the document to the bottom and read for more than 5 minutes before they can proceed to the next step) to prevent the subjects from failing to read the whole document.

[0125] (3) Online signature and face recognition of the subject

[0126] Provide a handwritten signature area to generate a PDF signature file with a timestamp, and pop up a secondary prompt: "You have signed the Informed Consent Form for XX Research. Do you confirm the submission?" After confirmation, the subject is required to complete a random action according to the screen instructions for face recognition authentication. After successful recognition, a unique subject ID is generated.

[0127] Step 6: Subject screening with human-machine collaboration

[0128] (1) Intelligent and rapid screening

[0129] Compare the preliminary screened registration list with the list of potential subjects to obtain a highly matching list. Conduct a rapid preliminary screening of the subjects who have completed digital recruitment, and automatically compare the recruitment list with the recommended list generated in Step 4 (Intelligent mining of potential subjects). If a subject exists in the recommended list, mark it as "highly matching" to identify candidates with a high degree of matching.

[0130] Rank the potential subjects in the highly matching list according to the matching difference degree. If the matching difference degree is less than the preset passing threshold, the potential subject corresponding to the matching difference degree is exempt from review, and the potential subject corresponding to the matching difference degree is included in the list of subjects passing the review.

[0131] Specifically, conduct priority ranking based on the matching difference degree. For subjects with a matching difference degree exceeding the preset passing threshold (for example, the matching difference degree < 0.3), directly enter the "automatically pre-approved" queue and push them to the clinical researchers for final confirmation.

[0132] (2) Manual review and confirmation

[0133] Push the list of subjects passing the review, the preliminary screened registration list, and the registration information to the researcher client. For the subjects in the "automatically pre-approved" queue, clinical researchers can view the complete files of the subjects in one stop on the review interface, including: basic information, medical records, test results, system judgment basis, etc., and conduct batch confirmation by the clinical researchers. At the same time, classify the subjects from other channels (such as online recruitment, community recruitment) or subjects determined by the system as complex cases (with a relatively high matching difference degree) into the "to be reviewed" queue. Clinical researchers compare the materials submitted in Step 5 (Digital recruitment of subjects) combined with on-site physical examination materials with the clinical research protocol in detail, and conduct manual review and verification of the subjects. After the review is completed, the system automatically combines the subjects who have passed the automatic pre-review and the manual review to generate a confirmation list for final confirmation of the inclusion list.

[0134] Such as Figure 4As shown, after completing subject screening, data on the subjects are collected during the clinical study and the collected data are analyzed.

[0135] Step 7: Generate clinical research data collection plan

[0136] (1) Data collection framework generation

[0137] Step 71: According to the data requirements table, construct a data collection framework, where the data collection framework includes core data domains, data collection rules and data quality control rules. Based on the data requirements of the clinical research plan extracted in step 1, the key data collection elements are automatically parsed according to the clinical trial plan, and a collection framework containing three types of elements is constructed: 1) Core data domains: vital signs, laboratory indicators, adverse events, etc., and the data sources are marked (HIS, LIS, PACS, ECG recorders, non-contact intelligent devices, etc.); 2) Collection rules: define data collection time points (such as baseline period + visit window ± 3 days), equipment technical parameters (such as DanaRS-3 type dynamic ECG recorder sampling frequency ≥ 128Hz); 3) Quality control rules: missing value tolerance, data abnormality threshold (such as systolic blood pressure> 180mmHg requires data verification). Generate a standardized data collection framework that meets GCP specifications to support clinical researchers to solidify the version after online adjustment.

[0138] (2) Supplementary data collection forms

[0139] Step 72: Generate data collection plan documents for multiple data source systems based on the data types and data collection rules in the core data domain. On the one hand, generate institutional-level data docking plans for system / device docking: 1) Hospital system docking: map EMR test item codes (such as LOINC→local LIS codes) through HL7 / FHIR protocol; 2) IoT device networking: formulate data transmission frequency and encryption standards for Bluetooth 5.0 / BLE devices (such as Medtronic CGMS); 3) Third-party data channel: configure RESTful API field mapping rules for platforms such as LabCorp. On the other hand, for special inspection items (such as dynamic electrocardiogram, continuous blood glucose monitoring), automatically associate preset device parameter templates (sampling frequency: DanaRS-3 type equipment defaults to 128Hz, Medtronic CGMS every 5 minutes / time), and generate special collection guidelines including device operation SOP and abnormal data processing rules.

[0140] A structured clinical research data collection plan document is formed, which supports online collaborative revision and optimization. The final version (Word and JSON formats) can be pushed to all personnel involved in clinical research data collection (PI, CRC, nurses, data managers, etc.) for reference with one click.

[0141] Step 8: Automatic acquisition of multi-source and multi-modal data

[0142] (1) Integrated acquisition of data sources

[0143] Step 81: According to the data acquisition plan document, determine the multi-source and multi-modal data sources and establish a data acquisition link. Based on the data acquisition plan generated by the personalized clinical research plan, determine the multi-source and multi-modal data sources for this clinical research and establish an acquisition link. For the acquisition of the subject's historical data, connect to systems such as HIS, LIS, and PACS within the hospital through system interfaces and communication protocols to obtain relevant text, image, and waveform data uniquely associated with the subject ID. For example, for the HIS system, use the HL7 standard interface to dock with the RESTful API to obtain structured patient information (such as name, medical history) and semi-structured electronic medical record text; for the LIS system, directly extract the numerical results of test items (such as blood glucose value, creatinine level) through ODBC / JDBC database connection technology; for the PACS system, transmit standardized medical images based on the DICOM communication protocol to ensure image resolution and metadata integrity.

[0144] Obtain the video data, audio data, and text expression data of the subject during the clinical research. For the data acquisition during the subject's clinical research, video and audio acquisition devices need to be added to the original data sources to continuously record the dynamic information of the subject's behavior, expression, language, actions, etc. during the experiment, making up for the intermittency and subjectivity of traditional manual observation. For example, sudden events (such as epileptic seizures, emotional fluctuations) may be missed by manual records, but the video can be fully traced back. In addition, the subject's body language (such as tremors, abnormal gait), micro-expressions (such as physiological manifestations of pain or anxiety), and voice characteristics (such as speech rate, intonation changes) may reflect potential physiological or psychological states, and these information are difficult to quantify through questionnaires or single sensors. Combining physiological indicators (such as heart rate, electroencephalogram) with audio-visual behavior data can establish a more comprehensive bio-behavioral association model. For example, in the study of mental diseases, the synchronous changes in speech pauses and heart rate variability may indicate specific symptoms. Collect video, audio, and other data of the subject during the clinical research through non-contact intelligent devices integrated with visible light cameras, thermal infrared cameras, microphones, and other sensors.

[0145] (2) Secure data transmission

[0146] To achieve multi-source and multi-modal data collection, it is necessary to ensure the security and transmission efficiency of data across systems. Structured data such as basic information of subjects and test values (e.g., age, blood glucose value) is encrypted using AES-256, and the key is encrypted twice using the RSA public key. For the time-series data with high-frequency updates in the LIS system (such as continuous blood glucose monitoring values), the lightweight national cryptographic algorithm SM4-CBC mode is enabled to balance encryption speed and security. For large-volume unstructured data such as DICOM images and video streams, block hybrid encryption is adopted: Visible light video: It is divided into blocks of 128×128 pixels, and a pseudo-random sequence is dynamically generated in combination with the chaotic mapping algorithm to scramble pixel values and append a check code; Thermal infrared video: For temperature matrix data (such as 32-bit floating-point type), sparse matrix compression based on ZigZag scanning is enabled (compression rate ≥ 60%), and then encrypted using the AES-GCM mode; Audio data: For voice-sensitive information (such as subject conversations), frequency-domain watermark embedding technology is adopted to insert irreversible noise interference in the 3000-4000 Hz frequency band to achieve de-identification. The video stream is fragmented according to H.265 encoding (10 MB / slice), combined with FLAC lossless compression of audio data, and the parallel transmission delay ≤ 50 ms. Associate HIS text, LIS numerical values, and video behavior data to ensure the integrity and temporal consistency of multi-modal data.

[0147] (3) Data integration and storage

[0148] Based on the standard terminology libraries such as ICD-10 established in Step 1, construct a multi-level term mapping system: Diagnostic terms: Map the localized diagnostic names of each medical institution to the ICD-10 standard codes (e.g., type 2 diabetes mellitus → E11.9); Drug data: Standardize the trade names and generic names using the WHO Drug Dictionary; Test indicators: Unify the reference ranges and units of different detection methods based on LOINC codes (e.g., HbA1c is mapped to 4548-4); Procedure terms: Convert the surgical procedure names into ICD-9-CM standard codes. Establish a unit conversion rule library to automatically convert the unit systems of different data sources: Standardize laboratory units (e.g., unify blood glucose values to mmol / L or mg / dL); Convert imaging parameters (e.g., unify CT values to Hounsfield Unit); Align time units (convert all time data to the ISO 8601 format in the UTC+8 time zone); Standardize gene data (e.g., unify RS numbers to the NCBI RefSeq format). Associate the subject identifiers of different systems through a fuzzy matching algorithm (e.g., HIS patient ID and gene database sample number); Establish a subject-data relationship graph to record the source system, collection time, version number, and other traceability information of each data point; Generate the core key values associated with the globally unique subject UUID data. Design a hierarchical storage system: Metadata warehouse: Store the standardized structured data (MySQL cluster); Object storage system: Save the original files and preprocessed unstructured data (e.g., store DICOM images in a MinIO cluster); Time series database: Specifically store high-frequency time series data such as vital sign monitoring (e.g., InfluxDB); Graph database: Maintain the complex association relationships between data (e.g., store the subject-disease-treatment relationship network in Neo4j). Construct a complete metadata management system: Technical metadata: Record technical attributes such as data format, storage location, access interface, etc.; Business metadata: Label business attributes such as the clinical meaning of data, collection scenarios, association schemes, etc.; Management metadata: Maintain management attributes such as data responsible person, privacy level, retention period, etc.; Process metadata: Record each operation log during the data conversion process (including cleaning rules, conversion parameters, etc.).

[0149] (4) Data quality verification

[0150] To ensure the quality and compliance of multi-source and multi-modal data, a full-process automated verification mechanism and a dynamic feedback system need to be established. First, conduct a comprehensive review of the data through multi-level verification rules: Integrity verification is based on predefined templates to verify the existence of required fields (such as subject ID, timestamp), and to check that unstructured data (such as DICOM images, video streams) is not damaged during transmission; Logical consistency verification relies on standard terminology libraries (such as ICD-10, LOINC) and rule engines to verify the accuracy of diagnostic codes and test item units. At the same time, detect the temporal offset of multi-modal data through the dynamic time warping (DTW) algorithm, such as the millisecond-level synchronization of electroencephalogram signals and behavioral videos; Compliance verification further conducts de-identification verification on sensitive data (such as facial videos, voice recordings) to ensure that privacy protection measures comply with GDPR requirements.

[0151] On this basis, the system realizes dynamic correction and optimization through an automated feedback mechanism. The real-time error correction module automatically calls predefined rules (such as converting "mg / dL" to "mmol / L") for repairable problems (such as unit conversion errors, term mapping errors), and records the correction log; The abnormal data classification and warning mechanism triggers different responses according to the severity: Abnormalities in non-keyword fields (such as insufficient image resolution) generate to-do tasks and push them to researchers, while serious errors (such as subject ID conflicts or temporal breaks exceeding 5 seconds) suspend the acquisition process until manual intervention; For irreversible errors (such as video stream frame loss rate > 5%), the system starts redundant device re-acquisition or interpolates and reconstructs missing temporal data based on the LSTM model to minimize data loss.

[0152] Step 9: Multi-modal data feature fusion analysis during the subject's clinical research

[0153] Step 91: Obtain video data, audio data, and text expression data during the subject's clinical research. Multi-modal feature fusion includes audio-visual feature extraction and fusion, and text feature extraction and fusion. Since the correlation between audio and video data is closely related to practical applications, a model for feature extraction and feature fusion is constructed. Through joint learning by extracting features from videos and audio, introducing spatio-temporal self-attention mechanisms and multi-modal cross-attention mechanisms, the accuracy and generalization ability of the model are balanced. The BRET model is used for feature extraction and learning of text data. Finally, multi-modal fusion of audio, video, and text features is performed.

[0154] (1) Audio-visual modal feature fusion

[0155] Step 92: Extract features from video data and audio data respectively to obtain video spatio-temporal features and audio spatio-temporal features.

[0156] For video analysis, when video clips are regarded as a form of 3D data, 3D convolutional neural networks (3DCNNs) can be used to extract spatial features. Capturing spatial features from video clips using 3D-CNNs, representing these features as

[0157] For temporal feature extraction, we first train a 2D-CNN using individual video frames, with RMSE as the loss function. The output of the final fully connected layer is considered the encoded representation of a single frame. Note that the training of the 2D-CNN is conducted independently, and each frame is labeled according to the corresponding label of the video. Therefore, each video clip can be converted into a sequence of vectors. This vector sequence is input into a gated recurrent unit (GRU) to extract temporal dynamics, and the result is also known as Subsequently, we obtain Q v .

[0158]

[0159] where "⊕" represents matrix multiplication operation, capturing and the correlation between the features of each frame, used to integrate spatio-temporal information and generate spatio-temporal attention weights Q v , representing the temporal features of video data, representing the spatial features of video data.

[0160] Consisting of and Q v Combining to obtain the result of the spatio-temporal attention mechanism, the video spatio-temporal feature

[0161] Finally, M v Passes through the fully connected layer and the output region pooling layer of the last layer. It should be emphasized that during the model training process, the 2DCNN is first trained independently, and then the 3D-CNN, GRU, and fully connected layer are jointly trained, all using RMSE as the loss function. This phased training method ensures the optimization of each component in the model.

[0162] For audio extraction, its operation is similar to video feature extraction, and then the audio spatio-temporal feature M A is obtained.

[0163] Step 93: According to the video spatio-temporal feature and the audio spatio-temporal feature, using the attention mechanism, obtain the video enhanced feature and the audio enhanced feature.

[0164] In clinical practice, multimedia often provides video and audio data. Therefore, by fusing audio-visual features, complementary information between different modalities is captured. First, we summarize the performance characteristics of video and audio over a period of time as and We adopt an attention mechanism between the video spatio-temporal summary features and the audio spatio-temporal summary features, and can obtain the video enhanced features under the modality interaction attention mechanism and the audio enhanced features respectively as follows:

[0165]

[0166] Among them, represents the video spatio-temporal summary features, represents the audio spatio-temporal summary features, represents the i-th audio spatio-temporal feature, represents the i-th video spatio-temporal feature, s i represents the similarity weight between the i-th video spatio-temporal feature and the i-th audio spatio-temporal feature, and i represents the number of segments of video / audio spectrum segmentation.

[0167] (2) Text modality feature fusion

[0168] Step 94: Extract features from the text expression data to obtain the global text vector.

[0169] To make full use of the information embedded in the text data, we adopt the BERT-BiGRU model. The architecture of the BERT-BiGRU model consists of three main components. The initial stage includes preprocessing the original text, which includes removing stop words and punctuation marks to reduce interference, and then converting the cleaned text into tokens. The second stage is dedicated to training the word vectors of the classification documents, using the BERT model to obtain the comprehensive text vectors encapsulating the richness of language semantics. In the final stage, the text vectors with the most significant weights are identified and then input into the BiGRU classifier. Then the output probabilities of the model are interpreted to determine the final classification result. The above method ensures a nuanced understanding of the text, resulting in robust and accurate classification.

[0170] Pre-training is a key stage of the BERT model, which is trained through a large corpus. The text vectors are obtained through BERT.

[0171]

[0172] The bidirectional GRU neural network model absorbs the advantages of the bidirectional RNN and LSTM models and makes further corrections. The algorithm of the final output formula of the bidirectional GRU neural network is:

[0173] M t = g{U[f(W i + V + b)] + c}

[0174] where W i represents the total text vector, represents the text vector of the nth segment of text, V and U are weight matrices, and b and c are bias matrices.

[0175] Step 95: Concatenate the video spatio-temporal features, audio spatio-temporal features, video enhancement features, audio enhancement features, and global text vector, and pass through the classification layer to obtain the predicted emotion category of the subject.

[0176] In the feature fusion stage, we emphasize the effectiveness of the extraction process by directly connecting three different feature sets:

[0177]

[0178] where X represents the concatenated fusion vector, represents the video enhancement feature, represents the audio enhancement feature, M t represents the global text vector.

[0179] This concatenation is followed by a dropout layer, which helps to reduce the differences between feature types. Subsequently, a linear classification layer is used to obtain the predicted emotion category y.

[0180] (3) Subject status monitoring

[0181] Step 96: Push the predicted emotion category to the researcher client so that the researcher can grasp the subject's reactions during the clinical study.

[0182] Paying attention to the status of subjects (such as emotions, cognition, etc.) during the clinical research process is a core link to ensure the scientific nature and ethical compliance of the trial. The multi-modal data feature fusion of video, audio, and text can provide an objective basis for clinical researchers to detect the status of subjects. According to the predicted emotion categories, the responses of subjects during the clinical research period can be quickly and clearly grasped, which can directly reflect the feelings of subjects towards drugs or treatment plans, facilitating researchers to make timely adjustments according to the predicted emotion categories and enabling researchers to flexibly adjust the plans for different subjects in a targeted manner. Non-contact intelligent devices collect multi-modal data of subjects such as video, audio, and text: through video, micro-facial expressions (such as brief frowning, twitching of the corners of the mouth), and body movements (such as restlessness, trembling) are captured; through audio, the intonation, speech rate, pause frequency of speech (such as the speech rate accelerating when anxious, the intonation flattening when depressed), and acoustic features (such as abnormal vocal cord vibration frequency) are analyzed; through text, the content of patient conversations and self-reports (such as the frequency of using negative words) is captured to mine the emotional tendency. In the traditional clinical research process, researchers' assessment of the status of subjects usually relies on patients' self-reports, scale filling, or doctors' subjective observations, which are easily affected by individual expression differences, social desirability bias (such as concealing negative emotions), or the experience of evaluators.

[0183] For researchers who already include the analysis of subjects' status in the design of the clinical research plan, the above multi-modal data feature fusion method provides them with a basic model framework for non-contact monitoring of subjects' status, which can be improved and applied according to the specific content of the plan. For researchers whose clinical research plan design does not include the analysis of subjects' status, the results of multi-modal feature fusion generated by the above multi-modal feature fusion method (such as cross-modal correlation feature vectors, emotion recognition types) can be used as the analysis basis for drug safety assessment, biomarker association, ethics and protection of subjects' rights and interests, compliance with regulatory and compliance requirements, and improvement of data quality and experimental efficiency.

[0184] (4) Fusion data back storage

[0185] The purpose of fusion data back storage is to structurally store the results of multi-modal feature fusion (such as cross-modal correlation feature vectors, emotion recognition types) in the clinical research database and deeply integrate them with the original data and patient files to provide multi-dimensional support for subsequent analysis. Fine-grained access permissions are set according to the roles of researchers (such as PI, data administrator). For example, only authorized personnel are allowed to view the emotion classification results, and downloading of the original audio and video data is prohibited.

[0186] Hierarchical storage architecture: Structured feature table: Store the fused feature vectors (such as emotion classification labels, spatio-temporal attention weight matrix) in a time-series database (such as InfluxDB), indexed by subject ID and timestamp, supporting efficient retrieval by research stage or event type. Unstructured feature files: The original features of video-audio fusion (such as spatial features output by 3D-CNN, time series extracted by GRU) are stored in an object storage system (such as MinIO) in HDF5 format, retaining complete metadata (sampling rate, model version, preprocessing parameters). Association graph construction: Use a graph database (Neo4j) to establish a subject-feature-data source relationship network, recording the logical associations between features (such as the causal chain between "emotional fluctuations" and "speech pause frequency"), supporting complex inference queries.

[0187] Data consistency guarantee: Term and coding mapping: Map the emotion recognition results (such as "anxiety" and "depression") to a standardized terminology library (such as SNOMED CT codes "48694002" and "35489007"), and dynamically associate them with the inclusion and exclusion criteria in the clinical research protocol. Unit and format alignment: Perform secondary standardization on the fused features. For example, unify the Hz unit of the audio spectrum to a logarithmic scale, and normalize the video feature resolution to 1080p to ensure the comparability of cross-modal data.

[0188] As Figure 5 shown, an intelligent management system for multi-source and multi-modal clinical research data provided by an embodiment of the present application includes:

[0189] A historical data acquisition module 100, configured to acquire the inclusion and exclusion criteria of a subject and the multi-modal clinical data of a patient, where the multi-modal clinical data includes text data, image data, and gene data.

[0190] A data vectorization module 200, configured to vectorize the text data, the image data, and the gene data respectively to obtain a text feature vector, a high-dimensional feature vector, and a gene feature vector.

[0191] A data fusion module 300, configured to perform feature extension and fusion on the text feature vector, the high-dimensional feature vector, and the gene feature vector to obtain a fused feature vector.

[0192] A matching analysis module 400, configured to obtain a matching difference degree according to the fused feature vector and the inclusion and exclusion feature vector of the inclusion and exclusion criteria.

[0193] A patient screening module 500, configured to screen the patient according to the matching difference degree and a preset screening threshold to obtain a list of potential subjects meeting the requirements.

[0194] The beneficial effects of the intelligent management system for multi-source and multi-modal clinical research data are similar to the beneficial effects of the intelligent management method for multi-source and multi-modal clinical research data described above, and will not be elaborated here.

[0195] The intelligent management system for multi-source and multi-modal clinical research data further includes:

[0196] The subject screening module 600 is used for subject recruitment and screening. The data acquisition plan generation module 700 is used for generating clinical research data acquisition plans. The real-time data automatic acquisition module 800 is used for automatic acquisition of multi-source multi-modal data. The acquisition data fusion analysis module 900 is used for multi-modal data fusion analysis during the subject clinical study.

[0197] The historical data acquisition module 100 includes:

[0198] The document extraction unit 110 is used to extract the subject requirement text and the data requirement text from the clinical research protocol document. The inclusion and exclusion criteria extraction unit 120 is used to extract logical rules from the subject requirement text, parse numerical rules, identify text rules, and map non-standard terms to standardized codes to obtain an inclusion and exclusion criteria table. The data requirement extraction unit 130 is used to extract experimental design parameters from the data requirement text, map unstructured requirements to standardized data elements, and associate them with the data source system to obtain a data requirement table. The historical data extraction unit 140 is used to obtain the patient's multimodal clinical data from multiple data source systems according to the data type in the inclusion and exclusion criteria table.

[0199] The subject screening module 600 includes:

[0200] The recruitment poster generation unit 610 is used to generate a clinical research project recruitment poster according to the inclusion and exclusion standard table and the data requirement table. The poster push unit 620 is used to send the clinical research project recruitment poster to the patient client in the potential subject list, and publish the clinical research project recruitment poster to the network for subject recruitment. The preliminary screening unit 630 is used to perform preliminary screening of the registered patients according to the registration information collected through feedback. The list comparison unit 640 is used to compare the registration list after preliminary screening with the potential subject list to obtain a highly matched list. The priority sorting unit 650 is used to prioritize the potential subjects in the highly matched list according to the matching difference. The review unit 650 is used to exempt the potential subject corresponding to the matching difference from review if the matching difference is less than the preset passing threshold, and include the potential subject corresponding to the matching difference in the audited list. The list push unit 660 is used to push the audited list and the registration list after preliminary screening and the registration information to the researcher client.

[0201] The data acquisition scheme generation module 700 includes:

[0202] An acquisition framework construction unit 710, configured to construct a data acquisition framework according to the data requirement table, where the data acquisition framework includes a core data domain, a data acquisition rule, and a data quality control rule. An acquisition document generation unit 720, configured to generate data acquisition scheme documents for multiple data source systems according to the data types and data acquisition rules within the core data domain.

[0203] The acquired data fusion and analysis module 900 includes:

[0204] A real-time data retrieval unit 910, configured to obtain video data, audio data, and text expression data during the clinical research of the subject. A spatio-temporal feature extraction unit 920, configured to perform feature extraction on the video data and audio data respectively to obtain video spatio-temporal features and audio spatio-temporal features. An enhanced feature analysis unit 930, configured to use an attention mechanism according to the video spatio-temporal features and audio spatio-temporal features to obtain video enhanced features and audio enhanced features. A text vector extraction unit 940, configured to perform feature extraction on the text expression data to obtain a global text vector. An emotion prediction unit 950, configured to perform vector splicing on the video spatio-temporal features, audio spatio-temporal features, video enhanced features, audio enhanced features, and global text vector, and obtain the predicted emotion category of the subject through a classification layer. A predicted emotion push unit 960, configured to push the predicted emotion category to the researcher client so that the researcher can master the reactions of the subject during the clinical research.

[0205] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An intelligent management method for multi-source and multi-modal clinical research data, characterized in that: include: Obtaining the inclusion and exclusion criteria of the subject and the multimodal clinical data of the patient, wherein the multimodal clinical data includes text data, imaging data and gene data; Vectorizing the text data, the image data and the gene data respectively to obtain a text feature vector, a high-dimensional feature vector and a gene feature vector; Performing feature expansion and fusion on the text feature vector, the high-dimensional feature vector and the gene feature vector to obtain a fused feature vector; According to the fusion feature vector and the entry and exit feature vector of the entry and exit criteria, the matching difference degree is obtained; The patients are screened according to the matching difference and a preset screening threshold to obtain a list of potential subjects that meet the requirements.

2. The intelligent management method of multi-source and multi-modal clinical research data according to claim 1, characterized in that: The feature expansion and fusion of the text feature vector, the high-dimensional feature vector and the gene feature vector to obtain the fused feature vector comprises: Inputting the text feature vector, high-dimensional feature vector and gene feature vector into a two-layer neural network MLP respectively to obtain corresponding feature weights; Performing cross-modal extended convolution processing on the text feature vector, the high-dimensional feature vector and the gene feature vector to obtain a feature enhancement vector, wherein the feature enhancement vector includes a text feature enhancement vector, a high-dimensional feature enhancement vector and a gene feature enhancement vector; According to the feature weights and the feature enhancement vectors, weighted splicing is performed on the features to obtain a multimodal splicing vector; The multimodal concatenated vector is input into a multi-layer fully connected layer for nonlinear transformation to obtain a fused feature vector.

3. The intelligent management method of multi-source and multi-modal clinical research data according to claim 2, characterized in that: The feature weights are: Where h1 = ReLU(W1Z j +b1), W1, W2, b1 and b2 are the parameters of the neural network MLP, W1 represents the weight matrix of the first layer of neural network MLP, b1 represents the bias vector of the first layer of neural network MLP, W2 represents the weight matrix of the second layer of neural network MLP, b2 represents the bias vector of the second layer of neural network MLP, ReLU represents the activation function, Z j represents the eigenvector of the jth mode, The feature weight representing the feature vector of the jth mode; The multimodal splicing vector is: Among them, V concat represents the multimodal concatenation vector, represents the weight of the text feature vector, represents the text feature enhancement vector, represents the weight of the high-dimensional feature vector, represents a high-dimensional feature enhancement vector, represents the weight of the gene feature vector, represents the gene feature enhancement vector; The matching difference is: Among them, C represents the matching difference, Z total (i) represents the i-th fused feature vector, Z criteria (i) represents the i-th entry feature vector, and the i-th fused feature vector corresponds to the i-th entry feature vector.

4. The intelligent management method of multi-source and multi-modal clinical research data according to claim 1, characterized in that: The method of obtaining the inclusion and exclusion criteria of the subject and the multimodal clinical data of the patient includes: Extract subject requirement text and data requirement text from clinical research protocol documents; Extract logical rules from the subject's demand text, parse numerical rules, identify text rules, and map non-standard terms to standardized codes to obtain the inclusion and exclusion criteria table; Extract the experimental design parameters from the data requirement text, map the unstructured requirements into standardized data elements, and associate them with the data source system to obtain the data requirement table; According to the data types in the inclusion and exclusion criteria table, the multimodal clinical data of the patient is obtained from multiple data source systems.

5. The intelligent management method of multi-source and multi-modal clinical research data according to claim 1, characterized in that: After obtaining the subject's inclusion and exclusion criteria and the patient's multimodal clinical data, the method further includes preprocessing the multimodal clinical data; the multimodal clinical data preprocessing includes: The HanLP toolkit and named entity recognition (NER) are used to preprocess text data to obtain structured text. The ITK toolkit is used to perform grayscale equalization and spatial normalization on the image data to obtain the image standardization file; Based on BioPython, the gene data was normalized and denoised by wavelet, and the gene standard format file was obtained.

6. The intelligent management method of multi-source and multi-modal clinical research data according to claim 4, characterized in that: It also includes subject recruitment and screening; the subject recruitment and screening include: Generate a clinical research project recruitment poster based on the inclusion and exclusion criteria table and the data requirements table; Directly sending the clinical research project recruitment poster to patient clients in the potential subject list, and publishing the clinical research project recruitment poster to the Internet for subject recruitment; Conduct preliminary screening of registered patients based on the registration information collected through feedback; Compare the initial screening registration list with the list of potential subjects to obtain a highly matched list; Prioritizing the potential subjects in the highly matched list according to the matching difference; If the matching difference is less than the preset passing threshold, the potential subject corresponding to the matching difference is exempted from review, and the potential subject corresponding to the matching difference is included in the review passing list; The list of applicants who have passed the review, the list of applicants who have undergone preliminary screening, and the registration information will be pushed to the researcher client.

7. The intelligent management method of multi-source and multi-modal clinical research data according to claim 1, characterized in that: It also includes the generation of a clinical research data collection plan; the generation of the clinical research data collection plan includes: According to the data requirement table, a data collection framework is constructed, wherein the data collection framework includes a core data domain, data collection rules and data quality control rules; Generate data collection plan documents for multiple data source systems based on the data types and data collection rules in the core data domain; According to the data collection solution document, determine the multi-source and multi-modal data source and establish a data collection link.

8. The intelligent management method of multi-source and multi-modal clinical research data according to claim 1, characterized in that: It also includes multimodal data fusion analysis during the subjects’ clinical studies; The multimodal data fusion analysis during the clinical study of the subject includes: Obtain video data, audio data, and text expression data of subjects during clinical research; Extracting features from the video data and the audio data respectively to obtain video spatiotemporal features and audio spatiotemporal features; According to the video spatiotemporal features and the audio spatiotemporal features, an attention mechanism is used to obtain video enhancement features and audio enhancement features; Performing feature extraction on the text expression data to obtain a global text vector; The video spatiotemporal features, audio spatiotemporal features, video enhancement features, audio enhancement features and global text vector are subjected to vector concatenation processing, and the predicted emotion category of the subject is obtained through a classification layer; The predicted emotion categories are pushed to the researcher client so that the researcher can understand the subject's reaction during the clinical study.

9. The intelligent management method of multi-source and multi-modal clinical research data according to claim 8, characterized in that: The video enhancement features are: The audio enhancement features are: in, represents the spatiotemporal summary features of the video, represents the spatiotemporal summary features of audio, represents the i-th audio spatiotemporal feature, represents the spatiotemporal features of the i-th video, s i represents the similarity weight between the i-th video spatiotemporal feature and the i-th audio spatiotemporal feature; Among them, X represents the splicing fusion vector, represents the video enhancement feature, represents the audio enhancement feature, M t Represents the global text vector.

10. An intelligent management system for multi-source and multi-modal clinical research data, characterized in that: include: A historical data acquisition module, used to acquire the inclusion and exclusion criteria of the subject and the multimodal clinical data of the patient, wherein the multimodal clinical data includes text data, image data and gene data; A data vectorization module, used for vectorizing the text data, the image data and the gene data respectively to obtain a text feature vector, a high-dimensional feature vector and a gene feature vector; A data fusion module is used to perform feature expansion and fusion on the text feature vector, the high-dimensional feature vector and the gene feature vector to obtain a fused feature vector; A matching analysis module, used to obtain a matching difference degree according to the fusion feature vector and the entry and exit feature vector of the entry and exit criteria; The patient screening module is used to screen the patients according to the matching difference and a preset screening threshold to obtain a list of potential subjects that meet the requirements.

Citation Information

Cited By

  • Intelligent follow-up visit nursing information management system and method

    CN120496878A

  • Hyperuricemia and gout data processing method, equipment and medium

    CN121034515A

  • A method, device and medium for processing data of hyperuricemia and gout

    CN121034515B

  • Clinical test data query method and system

    CN121092611A

  • Clinical test-oriented subject whole-process intelligent recruitment platform and method

    CN121460034A