Data processing method and device based on artificial intelligence, electronic equipment and medium
By extracting image and clinical feature vectors through deep learning models and generating data processing reports, the problems of low accuracy and low efficiency in traditional data processing are solved, and standardized and efficient data processing is achieved.
Patent Information
- Application Number
- CN202510754577.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional data processing relies on manual analysis, lacks unified standards, has low accuracy, low efficiency, and is time-consuming and labor-intensive.
A deep learning model is used to extract image and clinical feature vectors, data processing results are generated through weighted fusion and data fusion modules, and reports are generated using a large language model.
It improves the accuracy and efficiency of data processing, achieves standardized processing, reduces manual intervention, and improves the consistency and comparability of results.
Smart Images

Figure CN120656657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to artificial intelligence-based data processing methods, devices, electronic devices, and media. Background Art
[0002] In the traditional data processing process, data analysis and processing mainly rely on specific processing personnel, who usually rely on their long-term accumulated experience to analyze the data and ultimately generate corresponding reports.
[0003] However, traditional experience relies on historical data and scenarios. Due to differences in experience and understanding among processing personnel, analytical results lack unified standards and comparability, resulting in low accuracy. Furthermore, traditional data processing relies on manual, step-by-step processing, which is time-consuming, labor-intensive, and inefficient. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide an artificial intelligence-based data processing method, device, electronic device and medium, which can adopt a deep learning model to improve accuracy and efficiency, achieve standardized processing, reduce manual intervention, and improve the consistency and comparability of results.
[0005] In a first aspect, an embodiment of the present application provides a data processing method based on artificial intelligence, the method comprising:
[0006] Extracting an image feature vector and a clinical feature vector of the first user; the image feature vector is obtained by weighted fusion of global features and local features of the target image data; the weights corresponding to the global features and the weights corresponding to the local features are determined based on corresponding statistical indicators;
[0007] Inputting the image feature vector and the clinical feature vector into a data fusion module to perform weighted fusion on the image feature vector and the clinical feature vector based on the correlation between the image feature vector and the clinical feature vector to obtain a fusion feature;
[0008] Inputting the fusion features into a data processing module to obtain a data processing result;
[0009] The data processing results are input into a large language model to obtain a data processing report.
[0010] In a possible implementation, the image feature vector is obtained by the following steps:
[0011] Inputting the target image data into a global extraction branch in an image feature extraction module to obtain global features of the target image data;
[0012] Inputting the target image data into a local extraction branch in an image feature extraction module to obtain local features of the target image data;
[0013] Determining a weight corresponding to the global feature and a weight corresponding to the local feature based on the statistical index of the global feature and the statistical index of the local feature;
[0014] The global features and the local features are fused based on the weights to obtain an image feature vector.
[0015] In one possible implementation, the clinical feature vector is obtained by the following steps:
[0016] Inputting the text data in the clinical data into a text extraction unit in a clinical feature extraction module to obtain a text feature vector;
[0017] Inputting the structured data in the clinical data into the structured extraction unit in the clinical feature extraction module to obtain a structured feature vector;
[0018] Inputting the knowledge graph data in the clinical data into the knowledge graph extraction unit in the clinical feature extraction module to obtain a knowledge graph feature vector;
[0019] Among them, the clinical feature vector includes the text feature vector, the structured feature vector and the knowledge graph feature vector.
[0020] In one possible implementation, the method further includes:
[0021] Performing three-dimensional rendering on the initial image data of the first user to obtain three-dimensional image data; the target image data is obtained by dividing the initial image data into blocks;
[0022] Obtaining a target heat map based on the local features and attention weights of the target image data obtained by the image feature extraction module;
[0023] The three-dimensional image data and the target heat map are sent to a second user.
[0024] In a possible implementation, before extracting the image feature vector and clinical feature vector of the first user, the method further includes:
[0025] receiving a first model parameter value or a second model parameter value sent by a second server;
[0026] Determine the first model parameter value as the corresponding value of the model parameter in the initial data processing model to obtain a target data processing model; the target data processing model includes the image feature extraction module, the clinical feature extraction module, the data fusion module and the data processing module;
[0027] or, initializing the initial data processing model based on the second model parameter value;
[0028] Iterate and update the corresponding values of the model parameters in the initialized initial data processing model;
[0029] The updated iterative model parameter corresponding values are sent to the second server, so that the second server determines new second model parameters or first model parameters based on all the model parameter corresponding values sent by the first server.
[0030] In a possible implementation, the iterative updating of the corresponding values of the model parameters in the initialized initial data processing model includes:
[0031] Acquire a data sample pair and actual data processing results of the data sample pair; the data sample pair includes an imaging data sample and a clinical data sample of a second user;
[0032] The data sample pairs are used as samples, and the actual data processing results are used as labels, and the corresponding values of the model parameters in the initial data processing model are updated and iterated.
[0033] In one possible implementation, the method further includes:
[0034] receiving a data processing result obtained after the second user adjusts the data processing report;
[0035] The target data processing model is updated based on the adjusted data processing results, the image feature vector and the clinical feature vector.
[0036] In a second aspect, an embodiment of the present application further provides an artificial intelligence-based data processing device, the device comprising:
[0037] an extraction module, configured to extract an image feature vector and a clinical feature vector of the first user; the image feature vector is obtained by weighted fusion of global features and local features of the target image data; the weights corresponding to the global features and the weights corresponding to the local features are determined based on corresponding statistical indicators;
[0038] an input module, configured to input the image feature vector and the clinical feature vector into a data fusion module, so as to perform weighted fusion on the image feature vector and the clinical feature vector based on the correlation between the image feature vector and the clinical feature vector to obtain a fusion feature;
[0039] The input module is further used to input the fusion features into the data processing module to obtain data processing results;
[0040] The input module is further used to input the data processing results into the large language model to obtain a data processing report.
[0041] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the artificial intelligence-based data processing method as described in any one of the first aspects.
[0042] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the artificial intelligence-based data processing method as described in any one of the first aspects are executed.
[0043] The embodiment of the present application provides a data processing method, device, electronic device and medium based on artificial intelligence, the method comprising: extracting the image feature vector and clinical feature vector of the first user; the image feature vector is obtained by weighted fusion of the global features and local features of the target image data; the weight corresponding to the global feature and the weight corresponding to the local feature are determined based on the corresponding statistical indicators; the image feature vector and the clinical feature vector are input into the data fusion module, and the image feature vector and the clinical feature vector are weightedly fused based on the correlation between the image feature vector and the clinical feature vector to obtain a fusion feature; the fusion feature is input into the data processing module to obtain a data processing result; the data processing result is input into a large language model to obtain a data processing report. Through this application, a deep learning model can be adopted to improve accuracy and efficiency, achieve standardized processing, reduce manual intervention, and improve the consistency and comparability of results. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0045] Figure 1 A flowchart of an artificial intelligence-based data processing method provided in an embodiment of the present application is shown;
[0046] Figure 2 A flowchart of another data processing method based on artificial intelligence provided by an embodiment of the present application is shown;
[0047] Figure 3 A schematic diagram of the structure of an artificial intelligence-based data processing method provided in an embodiment of the present application is shown;
[0048] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.
[0050] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0051] To enable those skilled in the art to use the contents of this application, the following embodiments are provided in conjunction with a specific application scenario, the "data processing field." Those skilled in the art will appreciate that the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of this application. Although this application is primarily described in the "data processing field," it should be understood that this is merely an exemplary embodiment.
[0052] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0053] The following is a detailed description of an artificial intelligence-based data processing method provided in an embodiment of the present application.
[0054] Reference Figure 1 FIG. 1 is a flow chart of an artificial intelligence-based data processing method provided in an embodiment of the present application. The exemplary steps of the embodiment of the present application are described below:
[0055] S101: Extracting an image feature vector and a clinical feature vector of a first user.
[0056] In an embodiment of the present application, initial image data and clinical data of a first user uploaded by a second user are received; the initial image data are preprocessed; the preprocessed initial image data are divided into blocks to obtain at least one target image data to reduce the amount of subsequent calculations; clinical feature vectors and image feature vectors are extracted from the clinical data and each target image data, respectively; wherein the image feature vector is obtained by weighted fusion of global features and local features of the target image data; the weights corresponding to the global features and the weights corresponding to the local features are determined based on corresponding statistical indicators.
[0057] The initial image data can be a CT image of a first user (e.g., a patient) uploaded by a second user (e.g., a doctor) in DICOM format, containing multiple 2D slice images. The DICOM (Digital Imaging and Communications in Medicine) format is an international standard for digital medical imaging and communications, used for storing, exchanging, and transmitting medical images and related information.
[0058] Clinical data may include text data, structured data, and knowledge graph data. The text data may include text content such as the name, age, medical history, and laboratory tests of the first user. The structured data includes multiple examination items of the first user stored in a preset storage format and the detection value of each examination item; the preset storage format may be the storage order between multiple examination items, etc. The knowledge graph data is used to characterize the relationship between the medical history of the first user and each nodule type; the relationship between the medical history of the first user and each nodule type is annotated by experts. The nodule type can be classified according to the nodule morphology, such as burr type, ground glass nodule, solid nodule, etc.; it can also be classified according to the nature of the impact on the patient's health, such as benign, malignant, etc.
[0059] The initial image data is preprocessed, including: standardizing each 2D slice image to adjust each 2D slice image to a fixed voxel size (e.g., 1mm×1mm×1mm) to ensure compatibility with the initial image data captured by various devices; denoising each 2D slice image after standardization; and increasing the local contrast of each 2D slice image after denoising through adaptive histogram equalization to optimize detail visibility, thereby improving the accuracy of subsequent feature extraction.
[0060] Here, Contrast Limited Adaptive Histogram Equalization (CLAHE) is an algorithm used for image enhancement, aiming to improve the contrast of an image, especially in contrast enhancement of a local area. The parameter cliplimit of the adaptive histogram equalization can be set to =2.0.
[0061] Each two-dimensional slice image is standardized, including: performing voxel normalization on each two-dimensional slice image so that the voxel values and spatial resolutions of all two-dimensional slice images after voxel normalization are consistent; adjusting the spatial position and direction of each two-dimensional slice image after voxel normalization through radial transformation according to the spatial position and direction of the preset standard image so that the two-dimensional slice image is aligned with the preset standard image; in order to avoid the target point in the adjusted two-dimensional slice image not being on the voxel grid of the two-dimensional slice image before adjustment, it is also necessary to calculate the voxel value of the target point through trilinear interpolation to obtain the two-dimensional slice image after the standardization. Among them, the target point refers to a point in the adjusted two-dimensional slice image, which corresponds to a position in the two-dimensional slice image after voxel normalization. The spatial position of the target point in the new image is determined by the affine transformation matrix, which maps the coordinate system of the two-dimensional slice image after voxel normalization to the coordinate system of the adjusted two-dimensional slice image.
[0062] Each two-dimensional slice image is subjected to voxel normalization, including normalization by Hounsfield Unit (HU) value (a unit used to measure the density of human tissue in CT images), and mapping the grayscale range of the two-dimensional slice image to a preset grayscale range (such as [-1000, 3000]) to enhance image contrast and highlight differences in tissue structure.
[0063] Each two-dimensional slice image after normalization is subjected to denoising, including: performing multi-scale wavelet decomposition on the normalized two-dimensional slice image to obtain approximate coefficients and detail coefficients of different scales; applying bilateral filtering to the approximate coefficients and detail coefficients of each scale to remove high-frequency noise; further applying non-local mean filtering to the approximate coefficients and detail coefficients after bilateral filtering to enhance the self-similarity of the image; performing threshold processing on the detail coefficients after non-local mean filtering (NLM) to remove small coefficient values; and reconstructing the image based on the approximate coefficients after bilateral filtering and the detail coefficients after threshold processing to obtain a denoised two-dimensional slice image.
[0064] Here, the embodiment of the present application adopts bilateral filtering and non-local mean filtering (NLM), and combines wavelet transform with multi-scale noise reduction to improve the visibility of low-contrast nodules.
[0065] The pre-processed initial image data is divided into blocks, including: using a sliding window to divide all the initial image data into blocks based on overlapping sub-regions to obtain at least one target image data. The target image data is a three-dimensional image block (3DPatch) of a fixed size (e.g., 64×64×64 voxels).
[0066] The overlapping sub-region method means that there will be a certain overlap between the target image data after being divided.
[0067] Obtain the image feature vector of each target image data through the following steps:
[0068] Step 1: Input the target image data into the global extraction branch in the image feature extraction module to obtain the global features of the target image data.
[0069] In this application, the global extraction branch is a module of the 3DVisionTransformer (ViT) that incorporates a multi-head self-attention mechanism. Specifically, 3DViT is used to process the target image data, extracting global features through a multi-head self-attention mechanism with 12 heads and an embedding dimension of 768. This captures long-range dependency information in the image to improve subsequent data processing capabilities.
[0070] For example, in the nodule detection scenario, 3DViT is used to process the target image data and extract global features through the multi-head self-attention mechanism, which can capture the long-range dependency information in the image and improve the ability to classify large nodules.
[0071] Step 2: Input the target image data into the local extraction branch in the image feature extraction module to obtain the local features of the target image data.
[0072] In this embodiment of the present application, the local extraction branch includes a CNN unit (convolution kernel 3×3×3, number of layers 18, stride 1) based on the ResNet3D architecture (i.e., a convolutional neural network unit 3D-CNN) and a first fusion unit. The 3D-CNN extracts features from the target image data to obtain initial local features at multiple scales (including scales of 1 / 4, 1 / 8, 1 / 16, etc.); the first fusion unit is used to fuse the initial local features at all scales to obtain the target local features.
[0073] For example, in the nodule detection scenario, local features refer to the edges and textures of the nodules. Using 3D-CNN to extract local, fine-grained features from target image data can improve the accuracy of small nodule detection. Furthermore, the present embodiment, through the feature pyramid structure that extracts local features at multiple scales, can enhance the classification capabilities of nodules of different sizes.
[0074] Here, the first fusion unit introduces the Enhanced Nodule Attention Module (ENAM), which combines sequence attention and spatial attention. Sequence attention is a technique that combines a long short-term memory (LSTM) network (with a hidden layer dimension of 128) and an attention mechanism to dynamically assign attention weights based on statistical indicators. Spatial attention is a technique that combines 3D convolution operations and an attention mechanism to dynamically assign attention weights to each spatial position.
[0075] For example, in the nodule detection scenario, by introducing the sequence-spatial attention module, an adaptive weight distribution mechanism is provided (based on Softmax calculation of attention scores), which dynamically highlights the characteristics of small nodule areas (diameter <5mm), can enhance the recognition effect of small nodules, and increase the attention to nodule features in the target image data, thereby improving the model's detection accuracy for tiny abnormal areas.
[0076] Step 3: Based on the statistical indicators of the global features and the statistical indicators of the local features, determine the weights corresponding to the global features and the weights corresponding to the local features.
[0077] In the embodiment of the present application, the statistical indicator may be the mean or variance of the feature. The weight corresponding to the global feature and the weight corresponding to the local feature have a numerical range of 0-1.
[0078] In addition, in order to optimize the processing efficiency of the image feature extraction module, reinforcement learning is adopted. The data processing accuracy is used as the reward function, and the Q-learning algorithm (discount factor γ = 0.9, learning rate 0.01) is used to adjust the strategy, so as to take into account both computational efficiency and processing accuracy.
[0079] Step 4: Fusion of global features and local features based on weights to obtain image feature vectors.
[0080] The clinical feature vector is obtained by the following steps:
[0081] Step 1: Input the text data in the clinical data into the text extraction unit in the clinical feature extraction module to obtain a text feature vector.
[0082] In an embodiment of the present application, the text extraction unit converts text data into a high-dimensional feature vector (with a dimension of 768) based on BERT (Bidirectional Encoder Representations from Transformers), and maps it into a 512-dimensional text feature vector through a feature compression module so as to align it with other modal features in subsequent fusion.
[0083] Among them, BERT is pre-trained using medical corpus to ensure its semantic accuracy.
[0084] Step 2: Input the structured data in the clinical data into the structured extraction unit in the clinical feature extraction module to obtain a structured feature vector.
[0085] In an embodiment of the present application, the structured extraction unit encodes the structured data based on a Transformer encoder containing a 6-layer, 8-head attention mechanism, and preliminarily extracts a 256-dimensional structured feature vector; the 256-dimensional structured feature vector is converted into a 512-dimensional structured feature representation through a feature mapping module to obtain the final structured feature vector to achieve dimensional alignment with other modal features in the fusion stage.
[0086] Step 3: Input the knowledge graph data in the clinical data into the knowledge graph extraction unit in the clinical feature extraction module to obtain a knowledge graph feature vector; wherein the clinical feature vector includes a text feature vector, a structured feature vector and a knowledge graph feature vector.
[0087] In the embodiment of the present application, the knowledge graph extraction unit extracts features from the knowledge graph data through a two-layer convolutional neural network (hidden layer dimension 128) to generate an initial knowledge graph feature vector. Subsequently, this initial knowledge graph feature vector is uniformly mapped to a final 512-dimensional knowledge graph feature vector through a fully connected layer to achieve dimensional alignment with other modal features and improve the effect of multi-source information fusion.
[0088] In the adjacency matrix representation of knowledge graph data, the Pearson correlation coefficient is used to measure the strength or similarity of the relationship between nodes. Knowledge graph data is modeled using graph neural networks (GCN).
[0089] S102 : Based on the correlation between the image feature vector and the clinical feature vector, perform weighted fusion on the image feature vector and the clinical feature vector to obtain a fusion feature.
[0090] In the embodiment of the present application, the image feature vector (dimension is 64×64×64×128) is mapped to a 512-dimensional image feature vector through global pooling and feature transformation; the image feature vector (dimension is 512) and the clinical feature vector (dimension is 512) after feature transformation mapping are aligned by the feature alignment unit in the cross-modal attention module; the attention matrix between the aligned image feature vector and the aligned clinical feature vector is calculated by QKV (Query, Key, Value) using a multi-head attention mechanism unit with 8 heads and 24 dimensions in the cross-modal attention module; the attention matrix includes the correlation between the aligned image feature vector and the aligned clinical feature vector. Based on the attention matrix, the aligned image feature vector and the aligned clinical feature vector are fused to obtain the initial fusion feature; the initial fusion feature is input into the gated mechanism unit (GatedFusion) in the cross-modal attention module, and the sigmoid-activated gate unit (dimension 512) is used to filter the high-correlation information in the initial fusion feature, and the final fusion feature (dimension 512) is output.
[0091] QKV is a core concept in the Transformer architecture, representing: Query (query vector): used to represent the features of interest; Key (key vector): used to represent the "identity" of other features; Value (value vector): represents the actual content of other features.
[0092] Here, in order to enhance the correlation between modalities (imaging feature vectors and clinical feature vectors), the cross-modal attention module adopts a feature alignment method based on contrastive learning and optimizes the feature space consistency with InfoNCE loss (temperature parameter τ = 0.07).
[0093] In addition, when classifying nodules into subtypes (such as spicule type, ground glass nodules, and solid nodules), it is also necessary to extract the imaging genomics features of the imaging data (including gray-level co-occurrence matrix, shape factor, etc., with a dimension of approximately 50), fuse the nodule subtype classification, imaging feature vectors, and clinical feature vectors to obtain the fusion features.
[0094] S103: Input the fusion features into the data processing module to obtain the data processing results.
[0095] In an embodiment of the present application, the fused features are input into the nodule region detection layer of the data processing module to obtain the nodule region features; all nodule region features are input into the nodule classification layer of the data processing module to obtain the target nodule category of each nodule region feature; the nodule region features and each target nodule category are input into the nodule category probability layer of the data processing module to obtain the probability corresponding to each target nodule category; the probability of each nodule region feature corresponding to the target nodule category and the nodule size corresponding to each nodule region feature are input into the risk calculation layer of the data processing module to calculate the health risk score corresponding to each nodule region feature.
[0096] Each nodule region corresponds to a nodule region feature. Nodule region features refer to the characteristics of the nodule region in the image data. They are used to characterize the nodule's contour, location, and image characteristics (such as the average density, standard deviation, texture features (such as the gray-level co-occurrence matrix), and diameter). These features are very important in medical image analysis, especially in nodule detection and classification tasks. By combining multi-scale feature extraction, attention mechanisms, and deep learning models, the accuracy of nodule detection and classification can be significantly improved.
[0097] Furthermore, the fused features are input into the nodule region detection layer of the data processing module to obtain the nodule region features, including: inputting the fused features into the YOLOv5 network with the addition of CBAM (Convolutional Block Attention Module) to obtain the nodule candidate box; combining the FCOS (Fully Convolutional One-Stage) algorithm and full convolutional regression (step size 1, feature dimension 256) to refine the positioning of the nodule candidate box to obtain the nodule target box to improve the detection accuracy; using Mask R-CNN (the backbone network is ResNet-50-FPN) to perform semantic segmentation on the content of the area corresponding to the nodule target box, generate pixel-level masks, and obtain nodule region features, thereby realizing accurate contour extraction of the nodule area.
[0098] The input resolution of the YOLOv5 network is 512*512*Z. Mask R-CNN uses ResNet-50 as the backbone network and combines it with a Feature Pyramid Network (FPN) to extract features. This configuration effectively handles multi-scale nodule detection and segmentation tasks. The output resolution of the data processing module is 512*512*Z.
[0099] Furthermore, the nodule classification layer is constructed based on the XGBoost model, or it can be constructed using an ensemble learning model (e.g., a combination of random forest and LightGBM, with a tree depth of 200). The tree depth of the XGBoost model is 6, the learning rate η = 0.1, the number of iterations is 500, the L1 regularization parameter is 0.1, and the L2 regularization parameter is 1.0. The objective function L of the XGBoost model uses logarithmic loss, which is expressed as follows:
[0100]
[0101] Among them, y i is the actual nodule category of the i-th sample, P i is the probability of the target nodule category predicted by XGBoost for the i-th sample, and N is the number of samples. Each sample includes imaging data and clinical data.
[0102] Furthermore, the health risk score corresponding to each nodule area feature is calculated, including:
[0103] Step 1: Input the probability of each nodule region feature corresponding to the target nodule category and the nodule size corresponding to each nodule region feature into the following formula to obtain the health risk factor R corresponding to each nodule region feature.
[0104]
[0105] Among them, w1 and w2 are parameters determined by sample fitting (the final selection of w1 = 0.65 and w2 = 0.35 increases the classification AUC of the data processing module by 3.5%), Q is the probability that the nodule region feature corresponds to the target nodule category, S is the nodule region feature corresponding to the nodule size, S max is the preset maximum nodule size.
[0106] Here, classification AUC refers to the area under the receiver operating characteristic curve (ROC curve) for a binary classification problem. AUC is an important metric for evaluating the performance of classification models, especially when working with imbalanced datasets. Higher AUC values indicate better classification performance. AUC has advantages such as robustness, comprehensiveness, and comparability, and is widely used in many fields.
[0107] Step 2: Use Softmax to normalize the health risk factor R corresponding to each nodule area feature to obtain the health risk score corresponding to each nodule area feature.
[0108] In addition, the data processing module can also divide health risk levels based on the health risk scores, such as low, medium, high, etc.
[0109] S104: Input the data processing results into the large language model to obtain a data processing report.
[0110] In the embodiment of the present application, in order to meet clinical standards, the large language model has built-in international standardized templates such as Lung-RADS and Fleischner Guidelines. The large language model matches the report template for the data processing results through a template classification model (based on a CNN classifier, with 10 output categories); and generates a data processing report based on the matched report template and data processing results.
[0111] The data processing results include one or more of the following information: nodule area image, nodule location, nodule size, nodule area characteristics, nodule category, health risk score, and health risk level.
[0112] Here, Lung-RADS is an existing classification system for standardizing low-dose CT (LDCT) lung cancer screening reporting and management recommendations. It is designed to improve the performance of lung cancer screening and provide clinicians with clear guidance on the management of lung nodules. Fleischner GuidelinesFleischner Guidelines
[0113] In addition, this method can also personalize the report template according to nodule size, nodule edge characteristics, etc. to generate more accurate data processing reports.
[0114] Here, the large language model in the embodiment of the present application realizes adaptive diagnostic text generation, combines imaging data and the second user's decision history (historical sample size is about 1000 cases), and automatically generates professional, logically clear and easy-to-read structured reports through the BERT+GPT joint architecture (BERT encoding dimension 768, GPT decoding layer number 12); in order to improve accuracy and interpretability, a medical knowledge graph is introduced (the number of nodes is about 5000, and the edges are built based on medical literature), optimizes term selection (such as the description of the associated risks of "ground glass nodules"), and supports controlled text generation (Controlled Text Generation), allowing the second user to set the report detail level (brief / detailed) and tone (formal / mild) through the interface, with parameter ranges of [1-3] and [0-1] respectively.
[0115] Furthermore, the method also includes: receiving a data processing result obtained after the second user adjusts the data processing report; and updating the target data processing model based on the adjusted data processing result, the image feature vector, and the clinical feature vector.
[0116] In an embodiment of the present application, the target data processing model includes an image feature extraction module, a clinical feature extraction module, a data fusion module and a data processing module. The second user can adjust the data processing results (such as nodule boundary adjustment or classification change). The adjustment method can be to modify the content or supplement the annotation. The adjusted content is generated into an adjustment triplet, the first entity in the adjustment triplet is the original text, the second entity in the adjustment triplet is the adjusted text, and the edge in the adjustment triplet is the reason for the adjustment; when the number of adjustment triples reaches the first preset number, the target data processing model is updated based on the latest second preset number of adjustment triples.
[0117] Here, online learning of the target data processing model is achieved through incremental updates driven by second-user feedback. Reinforcement learning (with a reward function based on modification frequency and a discount factor of γ = 0.95) can be used to adaptively adjust the text generation weights (initial range [0.2, 0.8]) to improve the accuracy of subsequent reports. Adaptive optimization strategies (such as the AdamW optimizer, with an initial learning rate of 1e-4 and a weight decay of 0.01) can also be used to adjust the weight parameters of the target data processing model in real time to adapt to dynamic changes in clinical practice while improving convergence speed and stability.
[0118] Furthermore, the second user can not only adjust the data processing results in the data processing report, but also delete or add content, content expression, etc. in the data processing report.
[0119] In addition, the embodiment of the present application also provides a data processing report with multi-language support. Through a pre-trained translation model (supporting Chinese and English, BLEU score > 0.9), the data processing report is automatically translated to meet the needs of different regions. At the same time, a patient-friendly report is generated, using plain language (such as "small nodules, low risk") to explain nodule types, health risk score assessments and recommendations, thereby improving the first user's understanding and compliance.
[0120] Furthermore, the method further comprises:
[0121] Step 1: Perform three-dimensional rendering on the initial image data of the first user to obtain three-dimensional image data; the target image data is obtained by dividing the initial image data into blocks.
[0122] In an embodiment of the present application, visualization of the initial image data is achieved through voxel-level 3D reconstruction, and the MarchingCubes algorithm is used to convert 2D slices into high-resolution 3D volume models (resolution 512×512×Z, grid step size 1mm). Combined with volume rendering (Volume Rendering, based on ray marching) and surface rendering (Surface Rendering, based on triangle patches) technology, the second user can observe the morphology and growth of lung nodules from any angle; to support real-time interaction, WebGL or Unity3D is used to build a visualization interface, providing zoom (range 0.5-5 times), rotation (360° free rotation) and cutting (arbitrary plane clipping) functions, while enhancing the visibility of nodules in complex lung structures through ray casting (RayCasting, step size 0.5mm) and transparency control (range 0-1), which is convenient for the second user to analyze.
[0123] Step 2: Obtain a target heat map based on the local features and attention weights of the target image data obtained by the image feature extraction module.
[0124] In an embodiment of the present application, an initial heat map is generated based on the local features of the target image data extracted by the local extraction branch in the image feature extraction module; the attention weights corresponding to the target image data generated by the global extraction branch in the image feature extraction module are fused into the denoised initial heat map to obtain a target heat map.
[0125] Specifically, the initial heat map is generated based on the local features of the target image data extracted by the local extraction branch in the image feature extraction module, including:
[0126] Substitute the local features of the target image data into the following formula to obtain the Grad-CAM initial heat map:
[0127]
[0128] in, is the weight of the importance of the target image data to the nodule category c, Z is the total number of elements in the feature map, k represents the kth channel of a specific convolutional layer in the local extraction branch in the image feature extraction module, is the activation value of the kth channel in the i-th row and j-th column of the local feature, y c is the category score of nodule category c, and c is the nodule category in the data processing result corresponding to the target image data.
[0129]
[0130] in, is the initial heat map, ReLU is the activation function, and ReLU is used to ensure that only features that have positive contributions to category c are retained. k It is the local feature generated by the local extraction branch in the image feature extraction module.
[0131] Here, Grad-CAM calculates the gradient contribution of category c to the local features output by the local branch Generate a saliency heatmap.
[0132] Specifically, the initial heat map is subjected to SmoothGrad denoising using the following formula:
[0133]
[0134] Where N is the number of sampling times, N(0,σ 2 ) has a mean of 0 and a variance of σ 2 noise, is the initial heat map after denoising, and x is the target image data.
[0135] Here, by adding noise to the initial heat map and performing multiple sampling, the random noise of the gradient is reduced and the readability of the heat map is improved.
[0136] Specifically, the attention weight corresponding to the target image data generated by the global extraction branch in the image feature extraction module is fused into the denoised initial heat map to obtain the target heat map, including:
[0137] The following formula is used to combine the initial heat map after denoising and the attention weight A of the target image data V Fusion to obtain the target heat map
[0138]
[0139]
[0140] Among them, Q is the query vector of the target image data generated by the global extraction branch in the image feature extraction module, K is the key vector of the target image data generated by the global extraction branch in the image feature extraction module, T is the transpose, d k is the dimension of the key vector, ∈ = 1e-6, ∈ is used to avoid numerical instability and improve computational stability, λ1 and λ2 are adjustable hyperparameters (e.g., λ1 = 0.7, λ2 = 0.3).
[0141] Optionally, the target heat map is mapped to a two-dimensional space using the following formula to obtain the final target heat map:
[0142]
[0143] Among them, P ij Target heat map The similarity between the i-th point and the j-th point in F i Target heat map The eigenvalue of the i-th point in F j Target heat map The eigenvalue of the jth point in P, σ is 0.5 (original value 1.0). ij It is also the value of the point in the i-th row and j-th column in the final target heat map.
[0144] Here, the above formula is used to reduce cluster feature folding and improve visualization resolution.
[0145] Step 3: Send the three-dimensional image data and the target heat map to the second user.
[0146] Furthermore, the method further comprises:
[0147] Further, refer to Figure 2 FIG. 1 is a flow chart of another artificial intelligence-based data processing method provided in an embodiment of the present application. Before extracting the image feature vector and clinical feature vector of the first user, the method further includes:
[0148] S201: Receive a first model parameter value or a second model parameter value sent by a second server.
[0149] In the embodiment of the present application, the first model parameter refers to the model parameter sent by the second server and directly applied. The second model parameter value refers to the model parameter sent by the second server and needs to be updated based on local data.
[0150] S202: Determine the second model parameter value as the corresponding value of the model parameter in the initial data processing model to obtain the target data processing model.
[0151] In the embodiment of the present application, the target data processing model includes an image feature extraction module, the clinical feature extraction module, a data fusion module and a data processing module. The model structure of the initial data processing model is consistent with that of the target data processing model.
[0152] S203, or, initialize the initial data processing model based on the first model parameter value.
[0153] S204: updating and iterating the corresponding values of the model parameters in the initialized initial data processing model.
[0154] In an embodiment of the present application, a data sample pair and an actual data processing result of the data sample pair are obtained; the data sample pair includes an imaging data sample and a clinical data sample of a second user; the data sample pair is used as a sample, and the actual data processing result is used as a label to update and iterate the corresponding values of the model parameters in the initial data processing model.
[0155] Specifically, the data sample pairs are obtained through the following steps:
[0156] Step 1: Obtain an initial data sample pair of the second user; the initial data sample pair includes an initial imaging data sample and a clinical data sample.
[0157] In the embodiment of the present application, the proportion of samples of various nodule categories c in the initial data sample pairs of round t is calculated by the following formula:
[0158]
[0159] in, is the sample proportion of nodule category c in round t. β = 0.1 ensures that the weight is not overly amplified when the category proportion is extremely low. The sample proportion of nodule category c in round t refers to the ratio of the number of original sample data of nodule category c in round t to the number of all original sample data.
[0160] Here, the initial sample size of nodule category c accounts for λ0 Where C is the number of nodule categories.
[0161] Step 2: pre-process the initial image data sample; and divide the pre-processed initial image data sample into blocks to obtain at least one first intermediate image data sample.
[0162] In the embodiment of the present application, the process of pre-processing and dividing the initial image data samples is the same as that of the initial image data in step S101, and will not be repeated here.
[0163] Step 3: Calculate the sampling probability of each intermediate image data sample.
[0164] i. Let the volume of the initial image data sample be V, and the volume of the nodule area in the nth intermediate image data sample be V n , then the ratio R of the nodule area to the total volume V in the nth intermediate image data sample is n for:
[0165]
[0166] ii. Since lung nodules usually occupy a small space, we define the initial sampling probability of the nth intermediate image data sample as Pn for:
[0167]
[0168] in, is the sampling enhancement adjustment parameter for the tth round, β t Controls the relative weight of the entire non-nodule region for round t.
[0169] In the embodiment of the present application, during the training process, gradually increase Reduce β to keep P in the early stage of training n Low to balance sampling and ensure more nodules participate in training in the middle and late stages of training.
[0170] iii. The initial sampling probability P of the nth intermediate image data sample n The product of the nth intermediate image data sample and the preset multiple is determined as the baseline sampling rate P0 of the nth intermediate image data sample; and based on the baseline sampling rate P0 of the nth intermediate image data sample, the sampling probability of the nth intermediate image data sample in the tth round is calculated
[0171]
[0172] Where t is the number of training rounds and γ controls the growth rate of the sampling weights.
[0173] Step 4: Based on the sampling probabilities of all intermediate image data samples, sampling is performed in the intermediate image data to obtain a second intermediate image data sample.
[0174] Step 5: Extract the nodule region in each second intermediate image data sample to obtain a third intermediate image data sample.
[0175] In an embodiment of the present application, a U-Net network is used to segment the second intermediate image data sample, and a candidate region (ROI) containing the nodule is extracted to obtain a candidate region image in each second intermediate image data sample; the candidate region image in each second intermediate image data sample is input into a deep learning model of a cascaded fully convolutional neural network (CascadeFCN) to further accurately locate the nodule contour in the candidate region image; in order to remove artifacts and improve segmentation accuracy, morphological operations (such as opening operations combined with region growing methods) are used to post-process the nodule contour in the candidate region image; an edge optimization method based on an active contour model (Active Contour Model) is introduced, and the edge of the post-processed nodule contour is optimized by iteratively adjusting the contour curve (energy function weights α=0.5, β=1.0) to enhance the clarity and accuracy of the nodule boundary, thereby obtaining a third intermediate image data sample.
[0176] Step 6: Perform data enhancement on the third intermediate image data sample to obtain a target image data sample.
[0177] In an embodiment of the present application, to address the problem of sparse distribution of lung nodules in CT images, the present method uses WassersteinGAN-GP (WGAN-GP) to process the third intermediate image data sample to obtain a fourth intermediate image data sample; uses self-attention GAN (SAGAN) to process the fourth intermediate image data sample to obtain a fifth intermediate image data sample; uses rotation (angle range ±30°), mirroring (horizontal / vertical), random cropping (cropping ratio 0.8-1.0), affine transformation (scaling range 0.9-1.1) and illumination change (brightness adjustment ±20%) to process the fifth intermediate image data sample to obtain a sixth intermediate image data sample; and further processes the sixth intermediate image data sample through Mixup (α=0.2) and CutMix technology to obtain a target image data sample.
[0178] Here, WGAN-GP is primarily used to improve the stability of generated data and avoid mode collapse. SAGAN, combined with an attention mechanism, further enhances the quality of structural details in nodule samples. Through training, the GAN model learns the morphological characteristics of real lung nodules and generates realistic virtual nodules, including samples of different morphologies (such as burr, lobule, and calcification) and sizes, to improve model robustness. This application also increases the diversity of training samples, thereby improving the model's generalization ability.
[0179] In addition, during the training of the data processing model, pseudo-label-based semi-supervised learning was used, combining labeled clinical data with unclearly labeled cases. Training samples were generated using pseudo-labeling technology (confidence threshold 0.9), and the contribution of unsupervised data was optimized using consistency regularization (loss weight 0.5) and contrastive learning (InfoNCE loss, temperature parameter τ = 0.07). To improve the training efficiency of the first server, knowledge distillation was introduced. Through the teacher-student model framework (teacher model parameters 100 million, student model parameters 10 million), the small model inherits the diagnostic capabilities of the large model. The distillation temperature T = 2 significantly improves diagnostic efficiency.
[0180] S205: Send the updated and iterated model parameter corresponding values to the second server, so that the second server determines a new first model parameter or a new second model parameter based on all the model parameter corresponding values sent by the first server.
[0181] Here, this application adopts privacy-preserving decentralized learning, which can jointly update model parameters through multiple first servers. It can ensure that model training is completed without sharing original data between first servers. Moreover, when the model parameters are transmitted between the first server and the second server (such as the first model parameter value or the second model parameter value sent by the second server to the first server, and the updated model parameters sent by the first server to the second server), synchronous encryption and differential privacy technology are combined to ensure data security.
[0182] Specifically, the Paillier encryption scheme is used to perform homomorphic encryption (HE) on the data feature x of the first user locally:
[0183] Enc(x)=g x modn 2 ;
[0184] Among them, Enc(x) is the encrypted data feature, g and n are public key parameters.
[0185] Here, in this way, secure aggregation calculation is performed on the second server without decryption.
[0186] Specifically, Differential Privacy (DP), specifically, adding Laplace noise after each round of training:
[0187]
[0188] Among them, Δf is the gradient sensitivity, ∈ is used to control the degree of privacy protection (such as 1.0-2.0), are the model parameters after adding noise, w are the model parameters before adding noise, and Lap is the Laplace algorithm.
[0189] Specifically, the second server determines new first model parameters or second model parameters based on the corresponding values of the model parameters sent by all first servers, including: substituting the model parameters sent by all first servers into the following formula to obtain new first model parameters or second model parameters.
[0190]
[0191] Among them, w t The new first model parameter or the new second model parameter determined for the second server, N is the number of first servers, w k The corresponding value of the model parameter sent by the kth first server, n k is the number of samples of the kth first server, and n is the sum of the number of samples of all first servers.
[0192] Here, we combine personalized federated learning (Per-FedAvg) and introduce adaptive weighting:
[0193]
[0194] in, The first model parameter or the second model parameter sent by the second server to the i-th first server in the t+1 round, α is the weight coefficient for controlling the degree of personalization (the value range is adjusted between 0.1 and 0.3, and is flexibly set according to the data distribution of different institutions), w t a new first model parameter or a new second model parameter determined for the second server, is the model parameter sent by the i-th first server to the first server in the t-th round.
[0195] Here, to optimize the performance of the cross-institutional model, multi-task learning is adopted to build a generalized target data processing model with institution-specific losses (such as classification and segmentation tasks, with a weight ratio of 1:1). At the same time, lightweight model fine-tuning (with approximately 5 million parameters and a learning rate of 1e-5) is run on the first server of each institution to fully utilize local data characteristics and improve computing and data utilization efficiency.
[0196] Based on the same inventive concept, the embodiments of the present application also provide an artificial intelligence-based data processing device corresponding to the artificial intelligence-based data processing method. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the above-mentioned artificial intelligence-based data processing method in the embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0197] Reference Figure 3 FIG. 1 is a schematic diagram of an artificial intelligence-based data processing method provided in an embodiment of the present application, wherein the device includes:
[0198] Extraction module 301, configured to extract an image feature vector and a clinical feature vector of the first user; the image feature vector is obtained by weighted fusion of global features and local features of target image data; the weights corresponding to the global features and the weights corresponding to the local features are determined based on corresponding statistical indicators;
[0199] An input module 302 is configured to input the image feature vector and the clinical feature vector into a data fusion module, so as to perform weighted fusion on the image feature vector and the clinical feature vector based on the correlation between the image feature vector and the clinical feature vector to obtain a fusion feature;
[0200] The input module 302 is further used to input the fusion features into the data processing module to obtain data processing results;
[0201] The input module 302 is further configured to input the data processing result into the large language model to obtain a data processing report.
[0202] like Figure 4 As shown, an electronic device 400 provided in an embodiment of the present application includes: a processor 401, a memory 402 and a bus, wherein the memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device is running, the processor 401 communicates with the memory 402 through the bus, and the processor 401 executes the machine-readable instructions to perform the steps of the above-mentioned artificial intelligence-based data processing method.
[0203] Specifically, the above-mentioned memory 402 and processor 401 can be general-purpose memory and processor, which are not specifically limited here. When the processor 401 runs the computer program stored in the memory 402, it can execute the above-mentioned artificial intelligence-based data processing method.
[0204] Corresponding to the above-mentioned artificial intelligence-based data processing method, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above-mentioned artificial intelligence-based data processing method are executed.
[0205] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0206] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0207] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0208] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the information processing method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.
[0209] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A data processing method based on artificial intelligence, characterized in that: The method is applied to a first server, and includes: Extracting an image feature vector and a clinical feature vector of the first user; the image feature vector is obtained by weighted fusion of global features and local features of the target image data; the weights corresponding to the global features and the weights corresponding to the local features are determined based on corresponding statistical indicators; Inputting the image feature vector and the clinical feature vector into a data fusion module to perform weighted fusion on the image feature vector and the clinical feature vector based on the correlation between the image feature vector and the clinical feature vector to obtain a fusion feature; Inputting the fusion features into a data processing module to obtain a data processing result; The data processing results are input into a large language model to obtain a data processing report.
2. The data processing method based on artificial intelligence according to claim 1, characterized in that: Obtain the image feature vector through the following steps: Inputting the target image data into a global extraction branch in an image feature extraction module to obtain global features of the target image data; Inputting the target image data into a local extraction branch in an image feature extraction module to obtain local features of the target image data; Determining a weight corresponding to the global feature and a weight corresponding to the local feature based on the statistical index of the global feature and the statistical index of the local feature; The global features and the local features are fused based on the weights to obtain an image feature vector.
3. The data processing method based on artificial intelligence according to claim 2, characterized in that: The clinical feature vector is obtained by the following steps: Inputting text data in the clinical data into the text extraction unit in the clinical feature extraction module to obtain a text feature vector; Inputting the structured data in the clinical data into the structured extraction unit in the clinical feature extraction module to obtain a structured feature vector; Inputting the knowledge graph data in the clinical data into the knowledge graph extraction unit in the clinical feature extraction module to obtain a knowledge graph feature vector; Among them, the clinical feature vector includes the text feature vector, the structured feature vector and the knowledge graph feature vector.
4. The data processing method based on artificial intelligence according to claim 2, characterized in that: The method further comprises: Performing three-dimensional rendering on the initial image data of the first user to obtain three-dimensional image data; the target image data is obtained by dividing the initial image data into blocks; Obtaining a target heat map based on the local features and attention weights of the target image data obtained by the image feature extraction module; The three-dimensional image data and the target heat map are sent to a second user.
5. The data processing method based on artificial intelligence according to claim 3, characterized in that: Before extracting the image feature vector and clinical feature vector of the first user, the method further includes: receiving a first model parameter value or a second model parameter value sent by a second server; Determine the first model parameter value as the corresponding value of the model parameter in the initial data processing model to obtain a target data processing model; the target data processing model includes the image feature extraction module, the clinical feature extraction module, the data fusion module and the data processing module; or, initializing the initial data processing model based on the second model parameter value; Iterate and update the corresponding values of the model parameters in the initialized initial data processing model; The updated iterative model parameter corresponding values are sent to the second server, so that the second server determines new second model parameters or first model parameters based on all the model parameter corresponding values sent by the first server.
6. The artificial intelligence-based data processing method according to claim 5, characterized in that: The updating and iterating of the corresponding values of the model parameters in the initialized initial data processing model includes: Acquire a data sample pair and actual data processing results of the data sample pair; the data sample pair includes an imaging data sample and a clinical data sample of a second user; The data sample pairs are used as samples, and the actual data processing results are used as labels, and the corresponding values of the model parameters in the initial data processing model are updated and iterated.
7. The artificial intelligence-based data processing method according to claim 5, characterized in that: The method further comprises: receiving a data processing result obtained after the second user adjusts the data processing report; The target data processing model is updated based on the adjusted data processing results, the image feature vector and the clinical feature vector.
8. A data processing device based on artificial intelligence, characterized in that: The device comprises: an extraction module, configured to extract an image feature vector and a clinical feature vector of the first user; the image feature vector is obtained by weighted fusion of global features and local features of the target image data; the weights corresponding to the global features and the weights corresponding to the local features are determined based on corresponding statistical indicators; an input module, configured to input the image feature vector and the clinical feature vector into a data fusion module, so as to perform weighted fusion on the image feature vector and the clinical feature vector based on the correlation between the image feature vector and the clinical feature vector to obtain a fusion feature; The input module is further used to input the fusion features into the data processing module to obtain data processing results; The input module is further used to input the data processing results into the large language model to obtain a data processing report.
9. An electronic device, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the storage medium communicate via the bus, and the processor executes the machine-readable instructions to perform the steps of the artificial intelligence-based data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the artificial intelligence-based data processing method according to any one of claims 1 to 7.