Full-automatic unmanned CT scanning control method, storage medium and electronic equipment
Through the fully automatic unmanned CT scanning control method, the multimodal large language model and computer vision technology are used to realize the automation of CT inspection, solve the problem of manual operation dependence, and improve the inspection efficiency and accuracy.
Patent Information
- Application Number
- CN202510117889.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing CT examinations require manual operation assistance, which have problems such as complexity, high professionalism, high training costs and cross-infection risk.
The fully automatic unmanned CT scanning control method is adopted to obtain the detector's checklist information through a multi-modal large language model, output guide signals, identify the personnel status in the scanning space, obtain multi-modal image data, identify the detector's posture and human joint node position, and control the CT machine movement to achieve automated scanning.
The automation of CT examinations has been achieved, which reduces the dependence on radiation technicians, improves the inspection efficiency and accuracy, and reduces the risk of cross-infection.
Smart Images

Figure CN120036803A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fully automatic intelligent unmanned CT scanning detection, and particularly relates to a fully automatic unmanned CT scanning control method, a storage medium, and an electronic device. Background Art
[0002] As a non-invasive imaging technology widely used in medical diagnosis, Computed Tomography (CT) can provide high-resolution three-dimensional images and has extremely high value in clinical applications. However, in actual applications, due to differences in the examination site and patient body type, operators need to adjust the patient's position according to personalized scanning conditions, and this operation highly depends on the experience and skills of radiographers. However, this operation not only has problems such as high complexity and professionalism, but also requires long-term training and high-cost investment. In addition, manual operation will increase the risk of cross-infection. Therefore, there is an urgent need for an intelligent and precise unmanned CT scanning technology to solve the above problems. Summary of the Invention
[0003] The technical problem to be solved by this application lies in the problems existing in the need for manual operation assistance in CT examinations in the prior art, and further provides a fully automatic unmanned CT scanning control method, a storage medium, and an electronic device.
[0004] In a first aspect, the technical solution of this application provides a fully automatic unmanned CT scanning control method, including:
[0005] Obtain the examination form information of the detector, and perform the following interaction and control steps based on the multimodal large language model:
[0006] Output a guiding signal, where the guiding signal is used to guide the detector into the scanning space and remove items that affect the detection result;
[0007] Identify the personnel status in the scanning space. If the personnel status does not meet the detection requirements, output a reminder signal, where the reminder signal is used to remind the personnel to act according to the detection requirements, and the personnel include the detector and the accompanying personnel;
[0008] Obtain the multimodal image data of the detector, where the multimodal data includes RGB images and depth maps;
[0009] Based on the RGB image, identify the pose and human joint point positions of the detector. If the pose does not meet the scanning requirements, send a reminder signal again; determine the joint point spatial positions and the distance data required for isocenter positioning according to the depth map;
[0010] Quantify the lifting height and the bed-in distance of the scanning bed according to the distance data, and control the operation of the CT machine according to the lifting height and the bed-in distance;
[0011] Control the CT machine to scan and obtain the CT image of the examinee.
[0012] In some solutions, the fully automatic unmanned CT scan control method
[0013] The obtaining of the examination form information of the examinee includes:
[0014] Collect the examination form picture of the examinee;
[0015] Use OCR character recognition to obtain the examination form information recorded in the examination form picture, and the examination form information includes examinee information and examination site.
[0016] In some solutions, the fully automatic unmanned CT scan control method
[0017] The multi-modal large language model is obtained in the following way:
[0018] Select a multi-modal pre-trained model, the parameters of the pre-trained model are θ, and pre-tune the parameters θ in the pre-trained model according to the CT scan scenario, including:
[0019] Obtain a multi-modal data set of CT scan data where x i represents the multi-modal input, including the CT scan image and the related scan description of the CT scan image; y i represents the corresponding label, including the lung target result and the scan control decision; N represents the number of CT scan data;
[0020] The loss function of the pre-trained model is expressed as: L finetune =L language (θ)+L vision (θ)+L control (θ); where, represents the loss related to text understanding and generation, including diagnostic report generation or CT examination history analysis; L vision (θ) represents the loss related to image understanding, including CT scan scenario picture understanding; represents the loss related to control decision, including scan parameter adjustment or decision-making during automatic adjustment of the scan process;
[0021] Add a regularization term to the loss function to obtain the loss function with reduced model complexity where, λ is the regularization coefficient, is the L2 norm of the model parameter θ;
[0022] By minimizing the loss function with reduced model complexity Obtain the fine-tuned parameter θ * :
[0023]
[0024] Substitute the fine-tuned parameter θ * into the pre-trained model to execute the interaction and control steps, evaluate the result obtained from the interaction and control steps according to a preset evaluation metric, and if the evaluation passes, use the model obtained by substituting the fine-tuned parameter θ * into the pre-trained model as the multimodal large language model.
[0025] In some solutions, the fully automatic unmanned CT scan control method
[0026] Based on the RGB image, identify the pose and human joint point positions of the detector. If the pose does not meet the scanning requirements, send a reminder signal again; determine the joint point spatial positions and the distance data required for isocenter positioning according to the depth map, including:
[0027] Obtain the human joint point positions according to the RGB image in combination with the key point detection algorithm; where: The RGB image is represented as: The joint point positions are represented as: where M is the number of joint points, (s i , u i ) represents the coordinates in the plane coordinate system, i represents the joint serial number, H represents the height of the RGB image, W represents the width of the RGB image, C represents the number of channels of the RGB image, represents the feature space;
[0028] Calculate the weighted feature of the RGB image to obtain the weighted feature representation of the RGB image as:
[0029] Embed the joint point positions into a high-dimensional feature vector: where: d is the embedding dimension, and Embed represents the embedding function;
[0030] Calculate the attention weights using the high-dimensional feature vector:
[0031]
[0032] where, W Q , W K , W V are learnable weight matrices;
[0033] Calculate the attention scores according to the attention weights:
[0034] Weight the high-dimensional feature vector using the attention weights to generate guiding features:
[0035]
[0036] Remap the guiding features back to the original feature dimension: where W O is the weight matrix of the output mapping;
[0037] Fuse the guiding features with the original RGB image to obtain an RGB image with enhanced joint features: F′ = ReLU(F + F att′ ); where: ReLU is the linear activation function;
[0038] Determine the pixel range required for isocenter positioning based on the RGB image with enhanced joint features, and calculate distance data in combination with the depth map data to identify the positioning film scanning range;
[0039] Based on the captured positioning film image, segment the lung field boundary, identify the upper and lower boundaries of the lung field, and determine the CT scanning range according to the upper and lower boundaries of the lung field in combination with the safety distance. The CT scanning range is limited from the lung apex to below the costophrenic angle.
[0040] In some solutions, the fully automatic unmanned CT scanning control method
[0041] The method of segmenting the lung field boundary, identifying the upper and lower boundaries of the lung field based on the captured positioning film image, and determining the CT scanning range according to the upper and lower boundaries of the lung field in combination with the safety distance, where the CT scanning range is limited from the lung apex to below the costophrenic angle, includes:
[0042] Extract the high-dimensional feature representation of the positioning film image: H′ and W′ are the height and width of the feature map of the positioning film image respectively, and C′ is the number of feature channels;
[0043] Obtain the prior anatomical region mask: P ∈ {0, 1} H′×W′ ;
[0044] Generate the spatial attention map corresponding to the high-dimensional feature representation through convolution operation and activation function: where Conv 7×7 represents a 7x7 convolution operation, and σ is the Sigmoid activation function;
[0045] Apply the spatial attention map to the attention image feature representation: where represents element-wise multiplication;
[0046] For each channel, generate a channel attention vector through global average pooling and a fully connected layer:
[0047]
[0048] Among them, GlobalAvgPool represents global average pooling, FC is a fully connected layer, W 1 and W 2 are learnable weight matrices, δ is the ReLU activation function, and σ is the Sigmoid activation function;
[0049] Applying the channel attention vector to the feature image enhanced by spatial attention is expressed as:
[0050]
[0051] Combining the prior anatomical region mask with the feature image enhanced by spatial attention, we get: where ψ is a hyperparameter that controls the influence of the prior mask;
[0052] The updated feature image enhanced by spatial attention is expressed as:
[0053] The updated feature images enhanced by spatial attention corresponding to all channels are integrated into the channel attention enhanced feature image F Z ′;
[0054] Fusing the channel attention enhanced feature image with the original feature image enhanced by spatial attention, the image feature highlighting the key area is expressed as: The key area includes: ribs, lung fields.
[0055] In some solutions, the fully automatic unmanned CT scan control method
[0056] The algorithm for identifying the lung field boundary also includes a loss function:
[0057]
[0058] Among them, is the cross-entropy loss, is the Dice loss, is the regularization loss of the attention mechanism, and α, β, γ are weight coefficients;
[0059] Perform L2 regularization on each weight matrix: where λ dropout and λ L2 are regularization parameters, Dropout is the random inactivation function, and W represents each weight matrix.
[0060] In some solutions, the fully automatic unmanned CT scan control method
[0061] Controlling the CT machine to scan and obtain the CT image of the detected person includes:
[0062] Calling the CT machine control interface to automatically complete the imaging of the positioning film;
[0063] Segmenting the lung fields on the positioning film, ensuring the positioning frame information for setting the tomographic CT scanning range and parameters, and calling the tomographic scanning interface to obtain the CT image;
[0064] The method further includes:
[0065] Automatically calling a pre-stored lung nodule diagnosis model to identify the lung nodule information in the CT image, and generating a preliminary report based on the lung nodule information.
[0066] In a second aspect, the technical solution of the present application provides a computer-readable storage medium, in which program information is stored, and after the computer reads the program information, it executes the steps of the fully automatic unmanned CT scanning control method described in any item of the first aspect.
[0067] In a third aspect, the technical solution of the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the fully automatic unmanned CT scanning control method described in any item of the first aspect are implemented.
[0068] In a fourth aspect, the technical solution of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the fully automatic unmanned CT scanning control method described in any item of the first aspect
[0069] The above technical solution provided by the present application, compared with the prior art, has the following technical effects:
[0070] The full-automatic unmanned CT scan control method, storage medium, and electronic device provided by this application deeply integrate multi-modal large language models, computer vision, and intelligent hardware control technologies, providing a general and extensible CT examination intelligent agent solution. After obtaining the information of the detector, the examination site, and the detection requirements, the solution of this application performs the following interactions and controls based on the multi-modal large language model. By automatically outputting guiding signals, it guides the detector into the scanning space and removes items that affect the detection results. By identifying the personnel status in the scanning space, it ensures that only the detector is on the scanning bed and in the detection posture. According to the multi-modal image data, it further determines the detector's posture and the positions of human joint points, and finally ensures that both the posture and the positions of human joint points meet the detection requirements, and determines the distance data required for isocenter positioning. Quantify the lifting height and the bed-in distance of the scanning bed according to the distance data, and control the actions of the CT machine according to the lifting height and the bed-in distance; control the CT machine to scan and obtain the positioning film image of the detector, call the lung field segmentation model to determine the spiral CT scanning range, and obtain the final scanned CT image. The above solution of this application can achieve dynamic perception and adaptive control of the CT examination scenario, and finally enable the CT examination to achieve the goal of high-efficiency and high-precision imaging without human intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is a flowchart of the full-automatic unmanned CT scan control method according to an embodiment of this application;
[0072] Figure 2 It is a schematic structural diagram of the full-automatic unmanned CT scan device according to an embodiment of this application;
[0073] Figure 3 It is a core architecture diagram of the CTopAgent (CT Examination Intelligent Agent - CT Operator Agent) system according to an embodiment of this application;
[0074] Figure 4 It is a schematic diagram of the detection of the joint points of the subject's anatomical structure based on the improved RTMPose model according to an embodiment of this application;
[0075] Figure 5 It is a large lung field segmentation model based on the improved SAM according to an embodiment of this application;
[0076] Figure 6 It is a schematic diagram of the hardware connection relationship of the electronic device that executes the full-automatic unmanned CT scan control method according to an embodiment of this application;
[0077] Figure 7 It is a system scan flowchart executed by the electronic device of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] The specific implementation manners of the present application will be further described below with reference to the accompanying drawings.
[0079] It is easy to understand that, according to the technical solution of the present application, under the condition of not changing the essence of the present application, there are various structural manners and implementation manners that can be mutually replaced by those of ordinary skill in the art. Therefore, the following specific implementation manners and the accompanying drawings are only exemplary descriptions of the technical solution of the present application, and should not be regarded as all of the present application or as a limitation or restriction on the technical solution of the application.
[0080] The orientation terms such as up, down, left, right, front, back, front side, back side, top, bottom, etc. mentioned or possibly mentioned in this specification are defined relative to the structures shown in the respective drawings, and they are relative concepts. Therefore, they may change accordingly according to their different positions and different usage states. Therefore, these or other orientation terms should not be construed as restrictive terms either.
[0081] This embodiment provides a full-automatic unmanned CT scan control method as Figure 1 shown, which can be applied to the CTopAgent system / the control system of the CT scan device. As Figure 2 shown, the CT scan device includes an OCR system 1, an RGB-D camera 2, a detector 3, a scan bed 4, a display screen 5, a language interaction device 6, and a CTopAgent system 7. As Figure 3 shown, the core architecture of the CTopAgent system 7 includes four major parts: (a) a knowledge base, (b) an input, (c) a tool set, and (d) an agent reply. Among them, the knowledge base is used to implement file parsing, and tools are configured therein to be able to implement the parsing of text, tables, and images of the input data. The input data includes files recording the scan process, files representing the patient's reaction, accident summary files, and diagnostic standard files, etc. The input section is used to give prompt words, etc. The tool set section integrates various tools related to full-automatic unmanned CT scan control, such as scene perception tools, patient scan parameter calculation tools, device control tools, and diagnostic system tools, etc. In the agent reply section, it can implement the intelligent reply function for the entire process of the detector entering the scan room, lying on the scan bed, setting scan parameters, scanning positioning films, tomographic scanning, CT imaging, generating a diagnostic report, and guiding the detector to leave, etc.
[0082] Specifically, when implemented, as Figure 1 shown, the method includes the following steps:
[0083] S100: Obtain the examination list information of the detector.
[0084] In this step, the checklist information may include the information of the detector, the examination site, and the detection requirements. The checklist information of the detector can be obtained through the detector information input module and then sent to the CTopAgent system.
[0085] The CTopAgent system performs the following interaction and control steps based on the multi-modal large language model (Multi-modal Large Language Models, MLLMs):
[0086] S200: Output a guiding signal, which is used to guide the detector into the scanning space and remove the items that affect the detection results.
[0087] In this step, the detector can complete the removal of coats and metal items by playing the guiding signal through the AI voice guiding instruction and the display screen to ensure the imaging quality.
[0088] S300: Identify the status of the people in the scanning space. If the status of the people does not meet the detection requirements, output a reminder signal, which is used to remind the people to act according to the detection requirements. The people include the detector and the accompanying person.
[0089] In this step, the multi-modal large language model senses and identifies the status of the detector and the accompanying person in the scanning space to ensure that only the detector lies on the scanning bed and the accompanying person is not in the scanning room area.
[0090] S400: Obtain the multi-modal image data of the detector. The multi-modal data includes RGB images and depth maps.
[0091] When this step is implemented, the detector can be photographed by an RGB-D camera supporting the CT scanning device. The RGB-D camera is a 3D TOF camera module with RGB function and can obtain the multi-modal data of the detector, including RGB images and depth images.
[0092] S500: Identify the posture and the position of the human body joint points of the detector based on the RGB image. If the posture does not meet the scanning requirements, send a reminder signal again; determine the spatial position of the joint points and the distance data required for isocenter positioning according to the depth map.
[0093] In this step, the multi-modal large language model senses and identifies the posture of the detector based on the RGB image to ensure that the detector meets the scanning posture requirements. For example, Figure 4 as shown, for chest CT, the patient needs to raise both hands above the head and lie flat on the scanning bed. Specifically, the position of the human body joint points is identified by combining the improved RTMpose algorithm, and the spatial position information of the joint points and the distance data required for isocenter positioning are calculated by combining the depth map.
[0094] S600: Quantify the lifting height and the bed - in distance of the scanning bed according to the distance data, and control the operation of the CT machine based on the lifting height and the bed - in distance.
[0095] In this step, the lifting height and the bed - in distance of the scanning bed are quantified to accurately guide the detector to be placed in the required position.
[0096] S700: Control the CT machine to scan and obtain the CT image of the detector.
[0097] The solution in the above embodiments of this application deeply integrates multi - modal large - language models, computer vision, and intelligent hardware control technologies, providing a general and scalable intelligent agent solution for CT examinations. After obtaining the detector information, examination site, and detection requirements, the solution of this application performs the following interactions and controls based on the multi - modal large - language model: By automatically outputting guiding signals, guide the detector into the scanning space and remove items that affect the detection results. By identifying the personnel status in the scanning space, ensure that only the detector is on the scanning bed and in the detection posture. According to the multi - modal image data, further determine the detector posture and the positions of human joint points, and determine the distance data required for isocenter positioning. Quantify the lifting height and the bed - in distance of the scanning bed according to the distance data, and control the operation of the CT machine based on the lifting height and the bed - in distance; control the CT machine to scan and obtain the CT image of the detector. The above solution of this application can achieve dynamic perception and adaptive control of the CT examination scenario, and finally enable the CT examination to achieve the goal of high - efficiency and high - precision imaging without human intervention.
[0098] Preferably, in the above solution, in step S100, the obtaining of the examination form information of the detector includes: collecting a picture of the detector's examination form; using OCR character recognition to obtain the examination form information recorded in the examination form picture, and the examination form information includes detector information and examination site. Specifically, when implementing, after obtaining the detector examination form information through the OCR character recognition model, before / after the detector gets on the scanning bed, the voice interaction module is used to confirm the name of the detector, and after the identity check is correct, the detector can enter the next scanning process.
[0099] Correspondingly, during the entire process of interaction and control, the MLLMs can dynamically perceive the situation in the CT scanning room and the states of the detector and the accompanying personnel, and can dynamically select and call the corresponding tool modules according to the different states of the detector and the accompanying personnel to ensure that each step of the CT examination process is completed accurately and efficiently.
[0100] In specific implementation, the MLLMs model in the embodiments of the present application is mainly implemented based on LLaMA-3.2, and can be fine-tuned and trained according to the CT examination scenario, so that it has the dynamic perception of the complex scenario of CT scanning and the reasoning decision-making of CT scanning control. Preferably, the multimodal large language model is obtained in the following manner:
[0101] S101: Select a multimodal pre-trained model, the parameters of the pre-trained model are θ, and the parameters θ in the pre-trained model are fine-tuned in advance according to the CT scanning scenario, including: obtaining a multimodal dataset of CT scanning data where x i represents the multimodal input, including the CT scan image and the relevant scan description of the CT scan image; y i represents the corresponding label, including the lung target result and the scan control decision; N represents the number of CT scan data.
[0102] S102: The loss function of the pre-trained model is expressed as: L finetune =L language (θ)+L vision (θ)+L control (θ); where, represents the loss related to text understanding and generation, including diagnostic report generation or CT examination history analysis; L vision (θ) represents the loss related to image understanding, including CT examination scenario image recognition; represents the loss related to control decision-making, including scan parameter adjustment or decision-making during automatic adjustment of the scan process.
[0103] In this step, in order to enable LLaMA-3.2 to make better reasoning decisions in the CT examination scenario, the loss function can be designed as a multi-task loss function, including tasks of language understanding (such as CT scan process reasoning), image understanding (such as patient posture image recognition), and control reasoning (such as CT scan control decision). Among them:
[0104] Usually, the cross-entropy loss is adopted:
[0105]
[0106] where, p i,language is the predicted probability of the model for the i-th sample in the language task, and y i,language is the true label.
[0107] Usually, the standard cross-entropy loss for image understanding is used, :
[0108]
[0109] Among them, p i,vision is the predicted probability of the model for the i-th sample in the image understanding task, and y i,vision is the true label.
[0110] It is usually achieved through reinforcement learning or policy optimization. Assume that the model predicts the control policy a i (such as scanning angle, scanning time, etc.), then the control loss can be defined as the difference from the expected reward of the policy:
[0111]
[0112] Among them, R(a i ) represents the reward value obtained after executing the control policy a i (for example, scanning quality, efficiency, etc.).
[0113] S103: Add a regularization term to the loss function to obtain the loss function after reducing the model complexity Among them, λ is the regularization coefficient, is the L2 norm of the model parameter θ. That is, in order to prevent overfitting and ensure stability during the fine-tuning process, an L2 regularization term is added to the loss function to control the complexity of the model.
[0114] S104: Obtain the fine-tuned parameter θ by minimizing the loss function after reducing the model complexity * :
[0115] S105: Substitute the fine-tuned parameter θ * into the pre-trained model to execute the interaction and control steps, evaluate the results obtained from the interaction and control steps according to the preset evaluation metrics, and if the evaluation passes, substitute the fine-tuned parameter θ * into the pre-trained model to obtain the model as the multi-modal large language model.
[0116] For the solution in the above embodiment, load the pre-trained LLaMA-3.2 model and prepare a multi-modal dataset. Through the above loss function, optimize the parameters of LLaMA-3.2 to make it adapt to the dynamic perception and control inference tasks in the CT examination scenario. Use appropriate evaluation metrics (such as the decision accuracy and segmentation accuracy of the model in the CT scanning scenario) to evaluate the performance of the fine-tuned model.
[0117] The CTopAgent system drives the behavior guidance and intelligent Q&A interaction of the detector through the optimized multi-modal large language model. On the premise of ensuring the implementation of the standardized inspection process, according to the current state of the detection value, the MLLMs makes decisions on the behavior guidance behavior and communication of the detector, realizing multi-dimensional personalized support for the detector.
[0118] Further preferably, in the above solution, in step S500, based on the RGB image, the posture and the position of the human body joints of the detector are recognized. If the posture does not meet the scanning requirements, a reminder signal is sent again; according to the depth map, the spatial position of the joints and the distance data required for isocenter positioning are determined. The goal is to use the joint information obtained from the improved RTMPose joint detection algorithm to guide the CT positioning posture attention mechanism and enhance the attention to key areas, thereby improving the system's recognition and analysis capabilities in specific postures (such as supine position, arms raised above the head and holding the head). Specifically, it includes:
[0119] S501: Obtain the position of the human body joints according to the RGB image in combination with the key point detection algorithm; where: The RGB image is represented as: The joint position is represented as: M is the number of joints, (s i , u i ) represents the coordinates in the plane coordinate system, i represents the joint serial number, H represents the height of the RGB image, W represents the width of the RGB image, C represents the number of channels of the RGB image, represents the feature space.
[0120] Specifically, first, the improved RTMPose joint detection algorithm is used to analyze the RGB image to obtain the joint information of the detector. In the case of the supine position and the detector's arms raised above the head and holding the head, the C1, C5, C7, L3, T11, L5, Eyebrow, shoulder_L, shoulder_R, hip, knee, and foot joints are calculated through the detected human body joint recognition algorithm. At the same time, the depth map provides accurate depth information to assist the system in calculating the moving distance and positioning parameters of the scanning bed inside the CT device. In addition, the solution of this application can also use computer vision and deep learning technologies to estimate the height and weight of the detector from the image data captured by the RGB-D camera.
[0121] S502: Calculate the weighted RGB image to obtain the weighted feature representation of the RGB image as:
[0122]
[0123] S503: Embed the joint position into a high-dimensional feature vector: Where: d is the embedding dimension, Embed represents the embedding function, such as a fully connected layer or positional encoding. The embedding step is to convert the two-dimensional coordinate information of the joint points into high-dimensional features for interaction with the image features.
[0124] S504: Calculate the attention weights using the high-dimensional feature vectors to guide the attention distribution in the feature map:
[0125]
[0126] Where, W Q , W K , W V are learnable weight matrices. By calculating the similarity between the image features and the joint point embeddings, attention weights are generated to emphasize the regions related to the joint points.
[0127] S505: Calculate the attention scores according to the attention weights:
[0128] S506: Weight the high-dimensional feature vectors using the attention weights to generate guided features:
[0129]
[0130] S507: Remap the guided features back to the original feature dimension: Where, W O is the weight matrix of the output mapping.
[0131] S508: Fuse the guided features with the original RGB image to obtain the RGB image features after enhancing the joint features: F′ = ReLU(F + F att′ ); Where: ReLU is the linear activation function. Add the guided features to the original features to enhance the information related to the joint points while retaining the original information.
[0132] S509: Determine the distance data required for isocentric positioning according to the depth map after enhancing the joint features, identify the scanning range of the positioning film, and identify the lung field boundaries of the positioning film. Determine the CT scanning range according to the upper and lower boundaries of the lung field in combination with the safety distances (the safety distance for the upper boundary is 15 mm, and the safety distance for the lower boundary is 20 mm). Specifically, the above steps can be accurately identified for the target anatomical region through the image segmentation technology of the SAM (Segment Anything Model) model as Figure 5 shown. This model can efficiently identify and delimit the upper and lower boundaries of the lung field, thereby guiding the scanning range of spiral CT imaging.
[0133] In this step, the Target Attention Module is designed to focus on key anatomical regions such as ribs and lung fields, so as to improve the accuracy and efficiency of image segmentation and achieve precise attention to specific anatomical structures. Through the attention mechanism, the feature representations of key anatomical regions such as ribs and lung fields in the image are highlighted, thereby improving the recognition ability of the segmentation model for these regions. The Target Attention Module combines spatial attention and channel attention to enhance the features of the target region in multiple dimensions. Specifically, the Target Attention Module performs the following steps:
[0134] S5091: Extract the high-dimensional feature representation of the positioning film image: H' and W' are the height and width of the features of the positioning film image respectively, and C' is the number of feature channels.
[0135] S5092: Obtain the prior anatomical region mask: P ∈ {0, 1} H′×W′ 。
[0136] S5093: Generate the spatial attention map corresponding to the high-dimensional feature representation through convolution operation and activation function: where Conv 7×7 represents a 7x7 convolution operation, and σ is the Sigmoid activation function.
[0137] In this step, the spatial attention mechanism is used to emphasize the positional relationship of key anatomical regions in the image, such as the boundaries between ribs and lung fields. Through convolution operation, the spatial information of the key regions can be captured to generate a spatial attention map, emphasizing the positional features of anatomical regions such as ribs and lung fields.
[0138] S5094: Apply the spatial attention map to the attention image feature representation: where represents element-wise multiplication.
[0139] S5095: For each channel, generate the channel attention vector through global average pooling and fully connected layers:
[0140]
[0141] where GlobalAvgPool represents global average pooling, FC is the fully connected layer, W 1 and W 2 are learnable weight matrices, δ is the ReLU activation function, and σ is the Sigmoid activation function.
[0142] In this step, the channel attention mechanism is used to dynamically adjust the importance of each channel, further enhancing the feature representation of key anatomical regions. In this step, the importance of each channel is dynamically adjusted to strengthen the channel features related to key anatomical regions and suppress the interference of irrelevant channels.
[0143] S5096: Applying the channel attention vector to the feature image enhanced by spatial attention is expressed as:
[0144]
[0145] S5097: Combining the prior anatomical region mask with the feature image enhanced by spatial attention to obtain: where ψ is a hyperparameter that controls the influence of the prior mask. The prior anatomical region mask is used to further guide the attention mechanism, ensuring that the model focuses on predefined key regions and improving the accuracy of segmentation.
[0146] S5098: The updated feature image enhanced by spatial attention is expressed as:
[0147] The updated feature images enhanced by spatial attention corresponding to all channels are integrated into the channel attention enhanced feature image F Z ′.
[0148] S5099: Fusing the channel attention enhanced feature image with the original feature image enhanced by spatial attention to obtain an image feature representing the highlighted key area as: The key areas include: ribs, lung fields. Fusing the attention enhanced features with the original features retains the rich information of the original image while highlighting the feature representation of the key areas.
[0149] Further preferably, the key point detection algorithm includes a loss function:
[0150]
[0151] where is the cross-entropy loss, is the Dice loss, is the regularization loss of the attention mechanism, and α, β, γ are weight coefficients; a combined loss function is used to optimize the segmentation performance.
[0152] To prevent overfitting, a Dropout layer is introduced in the attention module, and L2 regularization is performed on the weight matrices: L2 regularization is performed on each weight matrix: where λ dropout and λ L2 are regularization parameters, Dropout is the random inactivation function, and W represents each weight matrix.
[0153] In some embodiments, controlling the CT machine to scan and obtain the CT image of the detector includes: calling the CT machine control interface to automatically complete the imaging of the localization film; segmenting the lung field on the localization film to ensure the localization frame information for setting the tomographic CT scan range and parameters, and calling the tomographic scan interface to obtain the CT image; the method further includes: automatically calling a pre-stored lung nodule diagnosis model to identify the lung nodule information in the CT image, and generating a preliminary report based on the lung nodule information. The lung nodule diagnosis model is not the main improvement of the solution of this application and can be implemented using existing models, so it will not be introduced in detail in this embodiment. Specifically, the lung field on the localization film is segmented by combining with the improved Segment Anything Model to further refine the localization frame information for setting the tomographic CT scan range and parameters, and calling the tomographic scan interface for CT imaging. After the scanning is completed, the CTopAgent system automatically calls the lung nodule diagnosis model to obtain the lung nodule information in the CT imaging and generate a preliminary diagnosis report, and a radiologist reviews the generated preliminary report.
[0154] In the above solution of this application, the CTopAgent system deeply integrates multi-modal large language models (MLLMs), computer vision, and intelligent hardware control technologies, providing a general and scalable intelligent agent solution for CT examinations. The core technical idea is to build a professional medical knowledge base based on multi-modal large language models and combine RGB-D cameras, human joint point recognition, 3D modeling, and automated tool modules to achieve dynamic perception and adaptive control of the CT examination scenario, and finally enable the CT examination to achieve the goal of high-efficiency and high-precision imaging without human intervention.
[0155] Some embodiments of this application also provide a computer-readable storage medium, in which program information is stored, and after the computer reads the program information, it executes the steps of the fully automatic unmanned CT scan control method described in any one of the above method embodiments.
[0156] Some embodiments of this application also provide a computer program product, including computer programs / instructions, characterized in that when the computer programs / instructions are executed by a processor, they implement the steps of the fully automatic unmanned CT scan control method described in any one of the above method embodiments.
[0157] Some embodiments of this application also provide an electronic device, such as Figure 6As shown, the electronic device includes at least one processor 61 and at least one memory 62. Program information is stored in at least one of the memories 62. After reading the program information, at least one of the processors 61 executes the full-automatic unmanned CT scan control method described in any of the above method embodiments. The device may further include: an input device 63 and an output device 64. The processor 61, the memory 62, the input device 63, and the output device 64 can be communicatively connected. The memory 62, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. By running the non-volatile software programs, instructions, and modules stored in the memory 62, the processor 61 executes various functional applications and data processing, that is, implements the full-automatic unmanned CT scan control method provided in any of the above solutions. The memory 62 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the full-automatic unmanned CT scan control method, etc. In addition, the memory 62 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 62 may optionally include a memory remotely set relative to the processor 61, and these remote memories can be connected to the device for executing the full-automatic unmanned CT scan control method through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The input device 63 can receive input user clicks and generate signal inputs related to user settings and function controls of the full-automatic unmanned CT scan control method. The output device 64 may include a display device such as a display screen. When the one or more modules are stored in the memory 62 and run by the one or more processors 61, the full-automatic unmanned CT scan control method in any of the above method embodiments is executed.
[0158] As Figure 7 As shown, during the operation of the above electronic device, the CT scan of the detector is mainly divided into four stages: pre-scan preparation, detector position placement, scan execution, and diagnosis, covering the key links from the interaction between the detector and the device to the generation of the final auxiliary diagnosis report. The important nodes of the entire process are shown in the figure, and the detailed description is as follows:
[0159] OCR recognition: The inspection requirements are obtained through the tester information entry module, and the data of the subject is extracted to the CT system. Voice confirmation: The identity of the tester is confirmed through video, voice and interaction with the tester. The voice interaction module guides the subject to enter the scanning room and guides the tester to prepare for the scan, and evaluates the tester's preparation through a multimodal large language model, including confirming whether the tester is dressed in accordance with the requirements and removing all metal objects that may affect the scan. After completing the pre-scan inspection, assist the tester to adjust the body position on the scanning bed. The RGB-D camera captures RGB images and depth images; RGB images are used in combination with the improved RTMPose model to accurately locate the key points of the tester's body in 3D modeling and human posture recognition; then, by combining the depth map data, the real space distance between the joint point and the zero point of the scanning bed is calculated, so as to automatically calculate and optimize the scanning parameters and determine the optimal starting point of the localizer scan. After the tester is ready, the scanning process is started through the equipment automation control interface. The system controls the movement of the scanning bed and performs the localizer scan. The target anatomical area is accurately identified through the image segmentation technology based on the SAM model. After adjusting the CT scan range according to the recognition results, the CT scan is performed to determine the range of CT imaging and further optimize the scanning efficiency and image quality. After the scan is completed, the lung nodule recognition software is called to generate a preliminary diagnosis report for the radiologist to review. At the same time, the examinee is guided to leave the CT examination room, and the examination process is successfully completed.
[0160] By integrating the above technologies and processes, this application can significantly improve the efficiency of CT examinations and diagnostic report generation. This application scheme covers the entire operational process from tester preparation to final report generation, realizing unmanned and intelligent CT examinations throughout the entire process. It not only promotes technological progress in the field of CT imaging, but also takes an important step in the realization of personalized medicine, becoming a milestone innovation in the field of precision medicine.
[0161] As needed, the above technical solutions can be combined to achieve the best technical effect.
[0162] The above are only the principles and preferred embodiments of the present application. It should be noted that, for ordinary technicians in this field, on the basis of the principles of the present application, several other modifications can be made, which should also be considered as the protection scope of the present application.
Claims
1. A fully automatic unmanned CT scanning control method, characterized in that: include: Obtain the tester's checklist information and perform the following interaction and control steps based on the multimodal large language model: Outputting a guidance signal, wherein the guidance signal is used to guide the inspector to enter the scanning space and remove objects that affect the inspection results; Identify the status of personnel in the scanning space, and output a reminder signal if the status of the personnel does not meet the detection requirements, wherein the reminder signal is used to remind the personnel to act according to the detection requirements, and the personnel include the detector and the accompanying personnel; Acquire multimodal image data of the subject, wherein the multimodal data includes an RGB image and a depth map; Based on the RGB image, the posture and the position of the joints of the subject are identified. If the posture does not meet the scanning requirements, a reminder signal is issued again. The spatial position of the joints and the distance data required for the isocenter positioning are determined according to the depth map. quantifying the lifting height and bed entry distance of the scanning bed according to the distance data, and controlling the movement of the CT machine according to the lifting height and bed entry distance; The CT machine is controlled to scan and obtain a CT image of the subject.
2. The fully automatic unmanned CT scanning control method according to claim 1, characterized in that: The obtaining of the tester's checklist information includes: Collect pictures of the tester's checklist; The inspection form information recorded in the inspection form image is acquired by using OCR text recognition, wherein the inspection form information includes information of the inspector and the inspection part.
3. The fully automatic unmanned CT scanning control method according to claim 1, characterized in that: The multimodal large language model is obtained in the following manner: A multimodal pre-trained model is selected, wherein the parameter of the pre-trained model is θ, and the parameter θ in the pre-trained model is fine-tuned in advance according to a CT scanning scenario, including: Acquire multimodal datasets of CT scan data where x i Represents multimodal input, including CT scan images and related scan descriptions of CT scan images; y i represents the corresponding labels, including lung target results and scan control decisions; N represents the number of CT scan data; The loss function of the pre-trained model is expressed as: L finetune =L language (θ)+L vision (θ)+L control (θ); where Represents losses related to text comprehension and generation, including diagnostic report generation or CT examination history analysis; L vision (θ) represents the loss related to image understanding, including CT scan scene image understanding; Represents losses associated with control decisions, including scan parameter adjustments or automatic adjustments during the scan process; Add a regularization term to the loss function to obtain a loss function with reduced model complexity Where λ is the regularization coefficient, is the L2 norm of the model parameter θ; By minimizing the loss function after reducing the model complexity Get the fine-tuned parameter θ * : The fine-tuned parameter θ * Substitute it into the pre-trained model to execute the interaction and control steps, evaluate the results of the interaction and control steps according to the preset evaluation index, and if the evaluation passes, adjust the fine-tuned parameter θ * The model obtained after substituting into the pre-trained model is used as the multimodal large language model.
4. The fully automatic unmanned CT scanning control method according to claim 1, characterized in that: The posture and the position of the joints of the human body of the subject are identified based on the RGB image, and if the posture does not meet the scanning requirements, a reminder signal is issued again; Determine the spatial position of the joint point and the distance data required for isocenter positioning according to the depth map, including: The position of human joints is obtained according to the RGB image combined with the key point detection algorithm; wherein: the RGB image is represented as: The joint point position is expressed as: Where M is the number of joint points, (s i ,u i ) represents the coordinates in the plane coordinate system, i represents the joint number, H represents the RGB image height, W represents the RGB image width, C represents the number of RGB image channels, represents the feature space; The weighted feature representation of the RGB image obtained by weighted calculation of the RGB image is: Embed the joint point positions into high-dimensional feature vectors: Where: d is the embedding dimension, Embed represents the embedding function; The attention weight is calculated using the high-dimensional feature vector: Among them, W Q ,W K ,W V is a learnable weight matrix; Calculate the attention score according to the attention weights: The high-dimensional feature vector is weighted using the attention weight to generate a guiding feature: Remap the guided features back to the original feature dimension: Among them, W O is the weight matrix of the output mapping; The guide feature is fused with the original RGB image to obtain an RGB image with enhanced joint features: F′=ReLU(F+F att′ );Where: ReLU is the linear activation function; Determine the pixel range required for isocenter positioning according to the RGB image after enhancing joint features, calculate distance data in combination with the depth map data, and identify the scanning range of the positioning film; According to the captured localizer image, the lung field boundary is segmented, the upper and lower boundaries of the lung field are identified, and the CT scanning range is determined according to the upper and lower boundaries of the lung field combined with a safety distance. The CT scanning range is limited from the apex of the lung to below the costophrenic angle.
5. The fully automatic unmanned CT scanning control method according to claim 4, characterized in that: According to the captured localizer image, the lung field boundary is segmented, the upper and lower boundaries of the lung field are identified, and the CT scanning range is determined according to the upper and lower boundaries of the lung field combined with the safety distance. The CT scanning range is limited from the apex of the lung to below the costophrenic angle, including: Extract the high-dimensional feature representation of the topographic image: H′ and W′ are the height and width of the positioning film image feature map, respectively, and C′ is the number of feature channels; Get the prior anatomical region mask: P∈{0,1} H′×W′ ; The spatial attention map corresponding to the high-dimensional feature representation is generated through convolution operations and activation functions: Among them, Conv 7×7 represents a 7x7 convolution operation, σ is the Sigmoid activation function; Apply the spatial attention map to the attention image feature representation: in, Represents element-wise multiplication; For each channel, a channel attention vector is generated through global average pooling and a fully connected layer: Among them, GlobalAvgPool represents global average pooling, FC is the fully connected layer, W1 and W2 are learnable weight matrices, δ is the ReLU activation function, and σ is the Sigmoid activation function; The channel attention vector is applied to the feature image after spatial attention enhancement as: Combining the prior anatomical region mask with the spatial attention enhanced feature image, we obtain: Among them, ψ is a hyperparameter that controls the influence of the prior mask; The updated feature image after spatial attention enhancement is expressed as: The updated spatial attention-enhanced feature images corresponding to all channels are integrated into the channel attention-enhanced feature image F Z ′; The channel attention enhanced feature image is fused with the original spatial attention enhanced feature image to obtain the image feature representation of the highlighted key area: The key areas include: ribs and lung fields.
6. The fully automatic unmanned CT scanning control method according to claim 5, characterized in that: The algorithm for identifying lung field boundaries also includes a loss function: in, is the cross entropy loss, is the Dice loss, is the regularization loss of the attention mechanism, α, β, γ are weight coefficients; Perform L2 regularization on each weight matrix: Among them, λ dropout and λ L2 is the regularization parameter, Dropout is the random inactivation function, and W represents the weight matrix.
7. The fully automatic unmanned CT scanning control method according to any one of claims 1 to 6, characterized in that: The controlling the CT machine to scan and obtain a CT image of the subject includes: Call the CT machine control interface to automatically complete the imaging of the positioning film; Segment the lung field on the positioning film, ensure the positioning frame information to set the tomographic CT scanning range and parameters, and call the tomographic scanning interface to obtain the CT image; The method further comprises: A pre-stored lung nodule diagnosis model is automatically called to identify the lung nodule information in the CT image, and a preliminary report is generated based on the lung nodule information.
8. A computer-readable storage medium, characterized in that: The storage medium stores program information, and the computer reads the program information and executes the steps of the fully automatic unmanned CT scanning control method according to any one of claims 1 to 7.
9. A computer program product, characterized in that The method comprises a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the steps of the fully automatic unmanned CT scanning control method according to any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the fully automatic unmanned CT scanning control method according to any one of claims 1 to 7.