CT image target feature classification method and device, electronic equipment and storage medium

By using a multimodal classification network in emergency scenarios, combining CT images and clinical report text data, high-precision classification of target features of CT images is achieved, solving the problem of low classification efficiency of CT images in emergency scenarios, and significantly improving diagnostic efficiency.

CN120197019APending Publication Date: 2025-06-24JILIN UNIV FIRST HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510257850.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The classification efficiency of CT images in emergency scenarios is low, and the diagnosis is delayed due to factors such as excessive burden on doctors, unstable image quality and time-consuming traditional interpretation processes.

Method used

A multimodal classification network is adopted, and through pre-training and binary classification fine-tuning, combining visual branch networks and text branch networks, CT images and clinical report text data are input to achieve high-precision binary classification of target features.

Benefits of technology

It significantly improves the classification efficiency of CT images, and can achieve high-precision classification in the case of high noise images or blurred edge structures, providing clinicians with more reliable assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197019A_ABST
    Figure CN120197019A_ABST
Patent Text Reader

Abstract

The invention discloses a CT image target feature classification method and device, electronic equipment and a storage medium, and belongs to the technical field of image data classification, and the CT image target feature classification method comprises the steps: obtaining a CT image and clinical report text data of a patient in an emergency wound scene; inputting the CT image and the clinical report text data into the trained multi-mode classification network to obtain a binary classification result of the target feature; wherein the trained multi-modal classification network is obtained through pre-training and binary classification fine tuning, and the trained multi-modal classification network comprises a visual branch network and a text branch network. According to the method, a visual branch network, a text branch network and clinical report text data are organically combined through a classification network fusing image-text multi-modal features, and through introduction of priori knowledge of doctors, the network can realize high-precision classification under the condition of images with large noise or fuzzy edge structures. According to the method, the classification efficiency of the CT images is remarkably improved, and more reliable assistance can be provided for clinicians.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image data classification, and particularly relates to a classification method, device, electronic device and storage medium for CT image target features. Background Art

[0002] Quick and accurate diagnosis and treatment are crucial for evaluating the patient's vital status and formulating subsequent medical plans.

[0003] In the emergency environment, CT (Computed Tomography) has become the preferred examination method for most emergency diseases. CT images can directly reflect the morphological and functional pathological changes of diseases in organs, providing important diagnostic basis for doctors. However, the classification efficiency of CT images in the emergency scenario is often limited by the following factors:

[0004] Overburdened doctors: Emergency department doctors usually face high-intensity work, and excessive fatigue may lead to misdiagnosis or missed diagnosis;

[0005] Image quality and clarity of target features: The quality of CT images is affected by multiple factors such as the patient's condition and scanning conditions, which may affect the recognition of target features;

[0006] Interpretation efficiency: The emergency scenario requires quick decision-making, but the traditional process of relying on manual experience for image interpretation is time-consuming and may delay subsequent treatment. Summary of the Invention

[0007] The purpose of this application is to provide a classification method, device, electronic device and storage medium for CT image target features to improve the classification efficiency of CT images of patients in the emergency trauma scenario.

[0008] According to the first aspect of the embodiments of this application, a classification method for CT image target features is provided. The method may include:

[0009] Obtain the CT image and clinical report text data of a patient in the emergency trauma scenario;

[0010] Input the CT image and clinical report text data into a trained multi-modal classification network to obtain a binary classification result of the target features;

[0011] Among them, the trained multi-modal classification network is obtained through pre-training and binary classification fine-tuning. The trained multi-modal classification network includes: a visual branch network and a text branch network.

[0012] In some alternative embodiments of this application, the pre-training process of the trained multi-modal classification network includes:

[0013] Pre-train the visual encoder of the visual branch network and the text encoder of the text branch network using paired image-text training data to obtain the model weights of the multi-modal classification network;

[0014] Among them, the paired image-text training data includes 2D image-text paired data and 3D image-text paired data.

[0015] In some alternative embodiments of the present application, the pre-training process of the trained multi-modal classification network further includes:

[0016] Perform min-max normalization on the image data in the paired image-text training data, and uniformly adjust the three-dimensional image to the standard size of 128×128×128.

[0017] In some alternative embodiments of the present application, the binary classification fine-tuning process of the trained multi-modal classification network includes:

[0018] Freeze the visual encoder and the text encoder;

[0019] Supervisedly train the 3D perceptron of the visual branch network using paired image-text training data to obtain the trained multi-modal classification network;

[0020] The supervised training is performed using the AdamW optimizer and enables the bf16 mixed precision training strategy with the help of the DeepSpeed framework.

[0021] In some alternative embodiments of the present application, input the CT image and clinical report text data into the trained multi-modal classification network to obtain the binary classification result of the target feature, including:

[0022] Input the CT image into the visual branch network to obtain CT feature markers;

[0023] Input the clinical report text data into the text branch network to obtain the core text information.

[0024] In some alternative embodiments of the present application, the visual branch network includes: a spatio-temporal Transformer module and a decoder;

[0025] Input the CT image into the visual branch network to obtain CT feature markers, including:

[0026] Divide the CT image into multiple image blocks according to the coronal plane, sagittal plane, and axial plane;

[0027] Input the image blocks into the spatio-temporal Transformer module for feature extraction to generate low-dimensional CT feature markers;

[0028] Reconstruct the low-dimensional CT feature markers through the decoder to obtain CT feature markers;

[0029] When the CT image is a 2D image, the CT image is subjected to repeated slicing processing to expand a depth dimension so as to be adapted to a three-dimensional input format.

[0030] In some alternative embodiments of the present application, the text branch network includes: a BERT text encoder;

[0031] Inputting the clinical report text data into the text branch network to obtain core text information, including:

[0032] Using the BERT text encoder to convert the clinical report text data into a unified 512-token format, and each token is represented by a 768-dimensional vector;

[0033] Normalizing all the tokens and mapping them to a compact 512-dimensional projection layer through a linear transformation to obtain the core text information.

[0034] According to the second aspect of the embodiments of the present application, there is provided a classification device for CT image target features, and the device may include:

[0035] An acquisition module, configured to acquire the CT image and the clinical report text data of a patient in an emergency trauma scenario;

[0036] A classification module, configured to input the CT image and the clinical report text data into a trained multi-modal classification network to obtain a binary classification result of the target features;

[0037] Wherein, the trained multi-modal classification network is obtained through pre-training and binary classification fine-tuning, and the trained multi-modal classification network includes: a visual branch network and a text branch network.

[0038] According to the third aspect of the embodiments of the present application, there is provided an electronic device, and the electronic device may include:

[0039] A processor;

[0040] A memory for storing processor-executable instructions;

[0041] Wherein, the processor is configured to execute instructions to implement the classification method of CT image target features as shown in any one of the embodiments of the first aspect.

[0042] According to the fourth aspect of the embodiments of the present application, there is provided a storage medium, when the instructions in the storage medium are executed by a processor of an information processing device or a server, so that the information processing device or the server implements the classification method of CT image target features as shown in any one of the embodiments of the first aspect.

[0043] The above technical solutions of the present application have the following beneficial technical effects:

[0044] The method of the embodiment of the present application combines the visual branch network, the text branch network and the clinical report text data organically through a classification network that fuses multi-modal features of images and texts, enabling the network to achieve high-precision classification in the case of images with high noise or blurred edge structures. This method not only significantly improves the classification efficiency of CT images, but also provides more reliable assistance for clinicians. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a schematic flowchart of a method for classifying target features of CT images in an exemplary embodiment of the present application;

[0046] Figure 2 is a multi-modal classification network diagram in an exemplary embodiment of the present application;

[0047] Figure 3 is the test result of binary classification of the multi-modal classification network on the CT dataset in an exemplary embodiment of the present application;

[0048] Figure 4 is a schematic structural diagram of a classification device for target features of CT images in an exemplary embodiment of the present application;

[0049] Figure 5 is a schematic structural diagram of an electronic device in an exemplary embodiment of the present application;

[0050] Figure 6 is a schematic hardware structure diagram of an electronic device in an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present application. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0052] The schematic diagram of the layer structure according to the embodiment of the present application is shown in the accompanying drawings. These figures are not drawn to scale, where for the purpose of clarity, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are only exemplary. In practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes and relative positions according to actual needs.

[0053] Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0054] In the description of this application, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.

[0055] In addition, the technical features involved in different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0056] The following will, with reference to the accompanying drawings, through specific embodiments and their application scenarios, provide a detailed description of the method, apparatus, electronic device, and storage medium for classifying CT image target features provided by the embodiments of this application.

[0057] As Figure 1 shown, in the first aspect of the embodiments of this application, a method for classifying CT image target features is provided. The method may include:

[0058] S110: Obtain the CT image and clinical report text data of a patient in an emergency trauma scenario;

[0059] S120: Input the CT image and clinical report text data into a trained multi-modal classification network to obtain a binary classification result of the target feature;

[0060] Among them, the trained multi-modal classification network is obtained through pre-training and binary classification fine-tuning. The trained multi-modal classification network includes: a visual branch network and a text branch network.

[0061] The method of this embodiment, through a classification network that fuses graphic and text multi-modal features, organically combines the visual branch network, the text branch network, and the clinical report text data. By introducing the prior knowledge of doctors, this network can achieve high-precision classification in the case of images with high noise or blurred edge structures. This method not only significantly improves the classification efficiency of CT images but also provides more reliable assistance for clinicians.

[0062] In this embodiment, the image-text pre-training network consists of two branches: vision and text. The vision branch takes the 3D images of CT as input and extracts visual features through a 3D Vision Transformer (3D ViT). As a three-dimensional vision Transformer model, 3D ViT can extract low-dimensional feature tokens from high-dimensional CT volume data and reconstruct the CT volume data through these tokens. Specifically, in this embodiment, a vision encoder method is designed to divide the CT volume data into uniform image patches of size 30×30×15 in the coronal, sagittal, and axial planes. This operation transforms the high-dimensional CT data into a one-dimensional sequence input form suitable for the Transformer model, thereby capturing local information and modeling global relationships through the global attention mechanism. In addition, the chunking strategy significantly reduces the computational complexity and makes the high-resolution CT image processing task feasible.

[0063] In the processing flow, these image patches are first input into the spatial and temporal Transformer module for feature extraction to generate low-dimensional CT feature tokens, and then the 3D CT volume data is reconstructed through the decoder. On the other hand, the text branch pre-trains by extracting clinical report data and combining contrastive learning. In the encoder part of the pre-trained model, the CT image patches are processed to generate low-dimensional feature tokens. These feature tokens are flattened and mapped to a 512-dimensional projection space through a linear transformation, and finally, the encoded CT feature tokens are generated.

[0064] The second branch processes the input clinical report text data. In this embodiment, a pre-trained BERT text encoder is adopted, which is a Transformer-based language model pre-trained on medical imaging reports. This encoder can process up to 512 tokens, thus completely covering the "impression" and "findings" parts in the radiology report. For reports with fewer than 512 tokens, this embodiment uses <pad>The tokens are padded to standardize the length. Exemplarily, the text encoder is responsible for processing the "Impression" and "Findings" sections in radiology reports and converting them into a unified 512-token format. Each token is represented as a 768-dimensional vector. These tokens are then normalized through a summation operation and mapped to a compact 512-dimensional projection layer through a linear transformation, thereby condensing the core text information in the report. Figure 2 Intuitively demonstrates this encoding process.

[0065] In addition, based on the aligned image-text pairs {vi, ti}, a cross-modal contrast loss function is also designed in this embodiment to align the text features and visual features.

[0066] Specifically, first, the visual features and text features are mapped into a shared embedding space, making the distance between matching positive samples closer and their similarity greater; while the distance between negative samples is farther and their similarity is smaller, thereby achieving the alignment of visual and language features. The calculation method of the cross-modal contrast loss is as follows:

[0067]

[0068] Among them, L contrastive refers to the cross-modal contrast loss, which realizes the alignment of visual features and text features; log(.) represents the logarithmic function; in addition, p(v i , t i ) is obtained through the following formula:

[0069]

[0070] Among them, N refers to the number of pre-trained samples; P(v i , t i ) refers to the overall similarity probability value calculated between the i-th visual feature and the text feature; S(v i , t i ) is to calculate the cosine similarity between v i and t i ; v i and t i refer to the i-th visual and text feature matrices; in addition, here τ is a predefined temperature parameter, and S(v i , t i ) is the matching score or similarity between v i and t i , which can be obtained through the following formula:

[0071]

[0072] Among them, ||v i || and ||v i || represent the vectors v i and v i of the norm length (L2 norm). The cross-modal contrastive loss method of this branch optimizes the mapping of visual and speech features in the shared embedding space through the contrastive learning mechanism, making the matching samples closer, thereby improving the alignment ability, generalization ability, and robustness of the multi-modal features of CT images.

[0073] In some embodiments, a fine-tuning network for the binary classification task of positive and negative of emergency disease characteristics is provided. Using the model weights of the visual and text encoders pre-trained in the above embodiments, training is performed by freezing the network parameters to extract the visual and text features of the binary classification task. For the input 3D CT data, visual features are first extracted through the pre-trained 3D ViT encoder, and the high-dimensional CT volume data is converted into low-dimensional feature tokens. These feature tokens are flattened and mapped to a 512-dimensional projection layer through a linear transformation, thereby generating the encoded CT feature representation.

[0074] For the input text prompt (TextPrompt), two text prompt formats are designed in this embodiment: one is a positive prompt with a diagnosis result, and the other is a negative prompt without a diagnosis result. These text prompts are processed by the pre-trained text encoder and converted into a unified 512-token format. Each token is represented by a 768-dimensional vector. Subsequently, these tokens are normalized through a summation operation and mapped to a compact 512-dimensional projection layer through a linear transformation, thereby generating the encoded text feature representation. Finally, the obtained mapped feature matrix is input into a cross-modal attention module for high-order feature fusion with the visual mapped feature matrix of the first branch. The cross-modal attention mechanism is described as follows. Suppose the feature map is represented as F∈R C×H×W , F v and F t come from the visual and language mapped feature maps respectively. The cross-modal attention between the two is calculated to capture the correlation between them. The overall formula is:

[0075]

[0076] where, W Q represents the weight matrix of the linear transformation for making Query on the feature map F t , W K represents the weight matrix of the linear transformation for making Key on the feature map F t , W V represents the weight matrix of the linear transformation for making Value on the feature map F t , and all the weight matrices are learned through network training. d k is the dimension of the Key, and sm(.) represents the softmax function.

[0077] In the fine-tuning stage, the model is trained through supervised learning. Between the encoded CT features and text features and their corresponding ground truth labels, optimization is carried out through the Binary Cross-Entropy (BCE) loss. The calculation formula of the BCE loss is as follows:

[0078]

[0079] where N is the number of training samples, y i is the binary label 0 or 1, and p(y i ) is the probability value that the model's predicted output belongs to the y i label; L BCE refers to the binary cross-entropy loss function during the model training process, which is used to measure the difference between the model's prediction result and the ground truth label in the current state; during the training process, the L BCE loss function guides the model to adjust parameters to optimize the prediction result through the mechanisms of backpropagation and gradient descent; log(.) represents the logarithmic function; N refers to the number of training samples; y i is the binary classification label 0 or 1; p(y i ) represents the probability value that the model's predicted output belongs to the y i label. This formula evaluates the model's performance by calculating the logarithmic difference between the predicted probability and the ground truth label. By minimizing this loss function, the model can gradually improve its prediction accuracy.

[0080] The overall training process of the network can be as follows: first, set the input 3D CT data as I ∈ R C×D×H×W , CD, H, and W respectively represent the number of channels, depth, height, and width of the input image. The corresponding classification label is y. First, for the paired image-text {vi, ti}, pre-training is performed through a multi-level visual encoder and a large language model to achieve the alignment of visual-speech features. The present invention uses Min-Max Normalization to preprocess the three-dimensional CT image to normalize the input data. In addition, the three-dimensional image is uniformly adjusted to the standard size of 128×128×128. In the multi-level visual encoder, in this embodiment, a model based on the three-dimensional Vision Transformer (3D ViT) can be used. This model consists of 12 layers of Transformers, and the patch size is 4×16×164. The output feature embedding dimension is 2048×768, representing 2048 feature tokens, and each token contains 768-dimensional features. After being processed by the three-dimensional spatial pooling perceptron, the finally generated visual feature token dimension is 256×768. In the pre-training stage of the large language model, the method of this embodiment uses BERT containing 12 layers of Transformers as the language encoder, and the maximum text length it supports is 128 characters. In addition, the method of this embodiment can be trained in parallel on 4 GPUs, the batch size is set to 12, the learning rate is 10-4, and the warm-up and cosine decay learning rate scheduling strategy is used.

[0081] Exemplarily, the fine-tuning stage can achieve the positive and negative binary classification tasks of emergency disease features. Specifically, first freeze the visual encoder and the text encoder, and only fine-tune the three-dimensional perceptron (3D Perceiver). The training data is the image-text pair to achieve supervised training for binary classification. The training settings are a batch size of 12 and a learning rate of 10-5, and the warm-up and cosine decay strategies are also used. All models of the present invention are trained using the AdamW optimizer and enable the bf16 mixed-precision training strategy (Mixed-Precision Training Strategy) with the help of the DeepSpeed framework to improve the training efficiency.

[0082] In one embodiment, fine-tuning and verification are performed based on a multi-modal dataset containing 19,708 cases of emergency data, among which there are 4,324 healthy cases and 15,384 disease cases. In the fine-tuning experiment of the binary classification task, as Figure 3 As shown, the method of this embodiment demonstrates excellent classification performance and achieves the following metrics: AUC is 99.16%, Recall is 97.20%, Precision is 97.71%, F1-score is 97.46%, and Accuracy is 96.04%. The experimental results show that the model can accurately judge "diseased" or "disease-free" in the emergency scenario, thus significantly improving the doctor's diagnosis efficiency. It provides a new solution for target feature classification and subsequent accurate diagnosis in the emergency scenario.

[0083] In some embodiments, the pre-training process of the trained multi-modal classification network includes:

[0084] Using paired image-text training data to pre-train the visual encoder of the visual branch network and the text encoder of the text branch network to obtain the model weights of the multi-modal classification network;

[0085] Among them, the paired image-text training data includes 2D image-text paired data and 3D image-text paired data.

[0086] In some embodiments, the pre-training process of the trained multi-modal classification network further includes:

[0087] Performing min-max normalization on the image data in the paired image-text training data and uniformly adjusting the three-dimensional images to the standard size of 128×128×128.

[0088] In some embodiments, the binary classification fine-tuning process of the trained multi-modal classification network includes:

[0089] Freezing the visual encoder and the text encoder;

[0090] Using paired image-text training data to perform supervised training on the 3D perceptron of the visual branch network to obtain the trained multi-modal classification network;

[0091] The supervised training is performed using the AdamW optimizer and enables the bf16 mixed precision training strategy with the help of the DeepSpeed framework.

[0092] In some embodiments, inputting the CT image and clinical report text data into the trained multi-modal classification network to obtain the binary classification result of the target feature, including:

[0093] Inputting the CT image into the visual branch network to obtain the CT feature marker;

[0094] Inputting the clinical report text data into the text branch network to obtain the core text information.

[0095] In some embodiments, the visual branch network includes: a spatio-temporal Transformer module and a decoder;

[0096] Input the CT image into the visual branch network to obtain CT feature tokens, including:

[0097] Divide the CT image into multiple image patches according to the coronal plane, sagittal plane, and axial plane;

[0098] Input the image patches into the spatio-temporal Transformer module for feature extraction to generate low-dimensional CT feature tokens;

[0099] Reconstruct the low-dimensional CT feature tokens through the decoder to obtain CT feature tokens;

[0100] When the CT image is a 2D image, perform repeated slicing on the CT image to expand a depth dimension, so as to adapt to a three-dimensional input format.

[0101] In some embodiments, the text branch network includes: a BERT text encoder;

[0102] Input the clinical report text data into the text branch network to obtain core text information, including:

[0103] Use the BERT text encoder to convert the clinical report text data into a unified 512-token format, and each token is represented by a 768-dimensional vector;

[0104] Normalize all tokens and map them to a compact 512-dimensional projection layer through a linear transformation to obtain core text information.

[0105] This application aims to design a classification network based on pre-training of a multi-modal large model of text and images for the classification problem of images in the emergency trauma scenario. By combining the clinical reports given by doctors for the CT image data of each patient, the target features of emergency diseases can be accurately classified using the multi-modal large model of text and images and the 3D Vision Transformer network, and the precision, AUC, sensitivity, and specificity of the classification results can be automatically quantitatively calculated. In the test stage, the sklearn.metrics machine learning library can be used to automatically quantitatively calculate the values of precision, AUC, sensitivity, and specificity by inputting the predicted values and the true labels.

[0106] The above embodiments innovatively propose a multi-level visual encoder. This structure integrates spatial Transformer and temporal Transformer to enhance the spatial feature extraction ability and can more effectively learn lesion features at different scales. In addition, through the global attention mechanism, this method effectively suppresses local noise interference and enhances the cross-frame information fusion ability, so that in the CT image environment with high noise, it can still maintain a high classification accuracy. By projecting visual features and text features into a shared embedding space, efficient alignment of cross-modal features is achieved. This method optimizes the distribution of visual and text features by reducing the distance between matching positive samples to enhance their similarity, while increasing the distance between negative samples to reduce their similarity, so as to improve the multi-modal feature learning ability of the model. Using a large language model as the encoder of clinical diagnosis reports, combining the encoded text feature vectors with CT visual feature vectors, and introducing a cross-modal attention mechanism to achieve deep fusion of visual-text features, effectively enhancing the model's understanding ability of lesion areas, thereby improving the accuracy of the CT image binary classification task.

[0107] The method of the above embodiments is based on the actual clinical scenario, specifically optimizes and solves the problem of emergency disease classification based on CT images, and performs quantitative calculation of key indicators, which will assist clinicians to quickly and accurately judge diseases, saving a large amount of manpower and material resources for subsequent rapid treatment. By designing a multi-modal classification network that integrates Transformer and large language models, the model can achieve accurate classification on CT images with high noise based on the prior knowledge prompts of doctors. Furthermore, it saves diagnostic time for patients and improves the diagnostic efficiency of doctors.

[0108] It should be noted that for the classification method of CT image target features provided in the embodiments of the present application, the execution subject may be a classification device for CT image target features, or a control module in the classification device for CT image target features that executes the method for classifying CT image target features. In the embodiments of the present application, the classification method of CT image target features executed by the classification device for CT image target features is taken as an example to illustrate the classification device for CT image target features provided in the embodiments of the present application.

[0109] As Figure 4 shown, in the second aspect of the embodiments of the present application, a classification device for CT image target features is provided, and the device may include:

[0110] An acquisition module 410, configured to acquire CT image and clinical report text data of a patient in an emergency trauma scenario;

[0111] A classification module 420, configured to input the CT image and clinical report text data into a trained multi-modal classification network to obtain a binary classification result of target features;

[0112] Among them, the trained multi-modal classification network is obtained through pre-training and binary classification fine-tuning. The trained multi-modal classification network includes a visual branch network and a text branch network.

[0113] The classification device for CT image target features in the embodiments of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0114] The classification device for CT image target features in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0115] The classification device for CT image target features provided in the embodiments of the present application can implement Figure 1 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0116] Optionally, as Figure 5 shown, the embodiments of the present application further provide an electronic device 500, including a first processor 501, a first memory 502, a program or instruction stored on the first memory 502 and executable on the first processor 501. When the program or instruction is executed by the first processor 501, it implements each process of the above-mentioned classification method embodiments of CT image target features and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0117] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0118] Figure 6 A schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0119] The hardware structure 600 of the electronic device includes, but is not limited to, components such as a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a second memory 609, and a second processor 610.

[0120] Those skilled in the art can understand that the hardware structure 600 of the electronic device may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the second processor 610 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 6 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0121] It should be understood that in the embodiments of the present application, the input unit 604 may include a Graphics Processing Unit (GPU) 6041 and a microphone 6042. The graphics processor 6041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 606 may include a display panel 6061, and the display panel 6061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 607 includes a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. The other input devices 6072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here. The second memory 609 can be used to store software programs and various data, including but not limited to application programs and operating systems. The second processor 610 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the second processor 610.

[0122] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the classification method for CT image target features and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0123] Among them, the processor is the processor in the electronic device described in the foregoing embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0124] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the foregoing embodiment of the CT image target feature classification method, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0125] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, a system chip, a chip system, or a system-on-chip, etc.

[0126] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal (which may be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0128] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.< / pad>

Claims

1. A method for classifying target features of CT images, characterized in that: include: Obtain CT images and clinical report text data of patients in emergency trauma scenarios; Inputting the CT image and the clinical report text data into a trained multimodal classification network to obtain a binary classification result of the target feature; The trained multimodal classification network is obtained through pre-training and binary classification fine-tuning, and the trained multimodal classification network includes: a visual branch network and a text branch network.

2. The method for classifying target features of CT images according to claim 1, characterized in that: The pre-training process of the trained multimodal classification network includes: Pre-training the visual encoder of the visual branch network and the text encoder of the text branch network using paired image-text training data to obtain model weights of the multimodal classification network; The paired image-text training data includes 2D image-text pairing data and 3D image-text pairing data.

3. The method for classifying target features of CT images according to claim 2, characterized in that: The pre-training process of the trained multimodal classification network further includes: The image data in the paired image-text training data are subjected to minimum-maximum normalization processing, and the three-dimensional images are uniformly adjusted to a standard size of 128×128×128.

4. The method for classifying target features of CT images according to claim 2, characterized in that: The binary classification fine-tuning process of the trained multimodal classification network includes: Freezing the visual encoder and the text encoder; Performing supervised training on the three-dimensional perceptron of the visual branch network using paired image-text training data to obtain the trained multimodal classification network; The supervised training is performed using the AdamW optimizer, and the bf16 mixed precision training strategy is enabled with the DeepSpeed ​​framework.

5. The method for classifying target features of CT images according to claim 1, characterized in that: The CT image and the clinical report text data are input into a trained multimodal classification network to obtain a binary classification result of the target feature, including: Inputting the CT image into the visual branch network to obtain a CT feature label; The clinical report text data is input into the text branch network to obtain core text information.

6. The method for classifying target features of CT images according to claim 5, characterized in that: The visual branch network includes: a spatial and temporal Transformer module and a decoder; The step of inputting the CT image into the visual branch network to obtain a CT feature label comprises: Dividing the CT image into a plurality of image blocks according to the coronal plane, the sagittal plane and the axial plane; Inputting the image block into the spatial and temporal Transformer module for feature extraction to generate low-dimensional CT feature markers; Reconstructing the low-dimensional CT feature marker through the decoder to obtain a CT feature marker; When the CT image is a 2D image, the CT image is repeatedly sliced ​​to expand a depth dimension so as to adapt to a three-dimensional input format.

7. The method for classifying target features of CT images according to claim 1, characterized in that: The text branch network includes: a BERT text encoder; The step of inputting the clinical report text data into the text branch network to obtain core text information includes: Using the BERT text encoder to convert the clinical report text data into a unified 512-token format, each token is represented by a 768-dimensional vector; All tags are normalized and mapped to a compact 512-dimensional projection layer through linear transformation to obtain the core text information.

8. A device for classifying target features of CT images, characterized in that: include: The acquisition module is used to obtain CT images and clinical report text data of patients in emergency trauma scenarios; A classification module, used for inputting the CT image and the clinical report text data into a trained multimodal classification network to obtain a binary classification result of the target feature; The trained multimodal classification network is obtained through pre-training and binary classification fine-tuning, and the trained multimodal classification network includes: a visual branch network and a text branch network.

9. An electronic device, characterized in that: include: A processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method for classifying target features of CT images as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the method for classifying CT image target features as described in any one of claims 1 to 7 are implemented.