Infrared thermogram auxiliary diagnosis system and method based on multi-modal large model
Through a multimodal large-scale infrared thermal image-assisted diagnosis system that integrates an infrared collector, a visible light optical camera, and a voice module, automated infrared thermal image diagnosis is achieved, solving the problem of traditional reliance on physician experience and improving diagnostic efficiency and accuracy, especially in distinguishing between rheumatoid arthritis and common joint inflammation.
Patent Information
- Application Number
- CN202510860626.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional disease diagnosis based on infrared thermal images relies on professional physicians, has low diagnostic efficiency and depends on physician experience, making it difficult to achieve automation and intelligence.
An infrared thermal image-assisted diagnosis system based on a multimodal large model is used, integrating an infrared collector, a visible light optical camera, a voice module, and a vital sign detection module. Image preprocessing, three-dimensional reconstruction, mapping matching, and disease diagnosis are performed through the multimodal large model, and automated disease diagnosis is achieved by combining the patient's chief complaint and pain information.
It achieves efficient, accurate and intelligent disease diagnosis, significantly improves diagnostic efficiency, is suitable for large-scale screening scenarios, and reduces misdiagnosis rates, especially in distinguishing rheumatoid arthritis from common joint inflammation.
Smart Images

Figure CN120753608A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence and medical equipment, and in particular to an infrared thermal image-assisted diagnosis system and method based on a multimodal large model. Background Art
[0002] With the continuous advancement of medical technology and people's growing concern for health, the application prospects of infrared thermal imaging technology in the medical field have become increasingly broad. However, traditional disease diagnosis based on infrared thermal images mainly relies on professional physicians, resulting in low diagnostic efficiency and dependence on physician experience. Therefore, the consideration and research of disease diagnosis methods based on infrared thermal images has certain practical value and progressive significance for improving the automation and intelligence of diagnosis, and enhancing diagnostic efficiency and effectiveness. Summary of the Invention
[0003] In view of this, the present application provides an infrared thermal image-assisted diagnosis system and method based on a multimodal large model, aiming to overcome the limitations of traditional infrared thermal image disease diagnosis and achieve efficient, accurate and intelligent disease diagnosis effects for patients.
[0004] In the first aspect, the present application provides an infrared thermal image-assisted diagnosis system based on a multimodal large model:
[0005] The infrared thermal image-assisted diagnosis system includes:
[0006] Infrared collector, used to collect infrared thermal images of the human body;
[0007] Visible light optical camera, used to collect visible light images of the human body from multiple angles;
[0008] A voice module is used to obtain the patient's chief complaint and pain information;
[0009] The vital signs detection module is used to detect the diseased area;
[0010] A large multimodal model, including a preprocessing module, a 3D reconstruction module, a mapping and matching module, and a disease diagnosis module;
[0011] The preprocessing module is used to preprocess the infrared thermal image of the human body;
[0012] The three-dimensional reconstruction module is used to construct a three-dimensional point cloud image based on visible light images of the human body collected from multiple angles;
[0013] The disease diagnosis module is used to determine the diseased area based on the pre-processed human infrared thermal image and the main complaint information;
[0014] The mapping and matching module is used to map the diseased area to a three-dimensional point cloud image to obtain a three-dimensional infrared thermal map;
[0015] The disease diagnosis module is further configured to receive pain information collected by the voice module and fed back by the patient based on the detection action, thereby determining the pain area based on the pain information; and output a diagnosis result based on the pain area.
[0016] Optionally, the pre-processing module includes a view selection module, an image restoration module and an image enhancement module;
[0017] The view selection module is used to select one or more infrared thermal images of the human body from the front and rear data captured by the infrared thermal imaging scanner according to the main complaint information;
[0018] The image restoration module is used to determine the temperature difference between the affected side and the normal side of the patient based on the main complaint information, and restore the human body infrared thermal image selected by the view selection module into an RGB pseudo-color infrared image based on the central temperature and temperature difference;
[0019] The image enhancement module is used to convert the RGB pseudo-color infrared image into an HSV image, enhance the hue and saturation channels, and then convert it into an RGB image.
[0020] Optionally, the view selection module consists of a pre-trained first BERT model and a text classification model; the pre-trained first BERT model converts the main complaint information into a first embedding vector; and the text classification model predicts the human body infrared thermal image to be selected based on the first embedding vector.
[0021] Optionally, the image restoration module includes a grayscale image restorer, a human posture recognition model, a pre-trained second BERT model and a pseudo-color image generation network;
[0022] The grayscale image restorer is used to map the temperature data in the selected human body infrared thermal image to the RGB pixel space 0 to 255 to generate a grayscale image;
[0023] The human posture recognition model is used to identify the key points of the human body in the grayscale image, including the left upper arm aa, left forearm ab, left hand ac, right upper arm ba, right forearm bb, right hand bc, left thigh ca, left calf cb, left foot cc, right thigh da, right calf db, right foot dc, left chest ea, right chest eb, left abdomen ec, right abdomen ed, left buttock ee, right buttock ef, left brain fa, and right brain fb;
[0024] The pre-trained second BERT model is used to encode the chief complaint information and generate the second embedding vector annotation to guide the localization of the affected side area;
[0025] The pseudo-color image generation network is used to generate temperature data into RGB pseudo-color images with the goal of maximizing the color difference between the affected side and the normal side.
[0026] Optionally, the pseudo-color image generation network is provided with two learnable variables, namely, core temperature (CT) and temperature width (TW); during training, the pseudo-color image generation network adopts self-supervised learning to adjust the values of core temperature and temperature width with the goal of maximizing the color difference between the affected side and the normal side.
[0027] Optionally, the image enhancement module includes:
[0028] Set two learnable parameters Hz, Ho and Sz, So for hue and saturation respectively, where H represents hue, S represents saturation, z represents scaling factor, and o represents offset; input a set of RGB pseudo-color images and convert the RGB pseudo-color images into HSV images using the following formula:
[0029] H'=H*H z +H o
[0030] S'=S*S z +S o
[0031] Where H' and S' represent the enhanced hue and saturation respectively;
[0032] Finally, the HSV image is restored to an RGB image.
[0033] Optionally, the 3D reconstruction model includes a feature extraction network, a feature alignment module, a multi-view fusion network, and a PointNet decoder;
[0034] The feature extraction network extracts the 2D semantic features of the visible light image of the human body at each viewing angle and generates the corresponding feature vector;
[0035] The feature alignment module is used to ensure that multi-view features are spatially consistent by minimizing the alignment loss, thereby achieving cross-view consistency of multi-view fusion features.
[0036] The multi-view fusion network fuses the image features of multiple views into a unified 3D human body description through feature encoding;
[0037] The PointNet decoder is used to decode the target human body mesh to obtain a three-dimensional point cloud image that can represent the human body 3D model.
[0038] Optionally, the mapping and matching module includes a 3D-2D projection module, a posture alignment module and a 2D-3D matching module;
[0039] The 3D-2D projection module is used to project the vertices of the 3D point cloud image to the infrared collector view;
[0040] The posture alignment module is used to calculate the posture transformation matrix to align the posture of the 3D point cloud image to the human posture space of the human infrared thermal image;
[0041] The 2D-3D matching module is used to match the 2D points of the diseased area to the 3D projection points to obtain a three-dimensional infrared thermal map.
[0042] In the second aspect, the present application provides an infrared thermal image-assisted diagnosis method based on a multimodal large model:
[0043] The infrared thermal image-assisted diagnosis method includes:
[0044] Collect the patient's chief complaint information, human infrared thermal images, and multi-angle visible light images of the human body;
[0045] Construct a 3D point cloud image based on visible light images of the human body from multiple angles;
[0046] Preprocessing of human body infrared thermal images;
[0047] Determine the diseased area based on the pre-processed human infrared thermal image and the chief complaint information;
[0048] Mapping the diseased area to a three-dimensional point cloud image to obtain a three-dimensional infrared thermal map;
[0049] Perform detection actions on the diseased area;
[0050] receiving pain information fed back by the patient based on the detection action, and determining a pain area based on the pain information;
[0051] Outputs diagnosis results based on the pain area.
[0052] In a third aspect, the present application provides a diagnostic device comprising a processor, a memory, and a communication bus, wherein the communication bus is used to realize a communication connection between the processor and the memory, and the processor is used to execute a computer program stored in the memory to realize the infrared thermal image assisted diagnosis method as described in the second aspect above.
[0053] In a fourth aspect, the present application further provides a computer-readable storage medium storing a computer program; the computer program can be executed by a processor to implement the infrared thermal image-assisted diagnosis method as described in the second aspect above.
[0054] In a fifth aspect, the present application also provides a computer program product, including a computer program, which can be executed by a processor to implement the infrared thermal image-assisted diagnosis method as described in the second aspect above.
[0055] This application includes at least the following beneficial technical effects:
[0056] 1. Traditional infrared thermal image diagnosis requires manual analysis of image features by doctors, which is time-consuming and depends on the professional ability and work experience of doctors; the multi-modal large model of the application can complete image preprocessing, feature extraction and disease probability prediction within a few seconds.
[0057] 2. The application is suitable for large-scale screening scenarios (such as occupational disease screening and chronic disease follow-up), and the efficiency is improved by tens of times, which significantly alleviates the problem of tight medical resources.
[0058] 3. The application combines multi-modal technology to cross-analyze infrared thermal images and patient clinical symptoms, which can more accurately distinguish rheumatoid arthritis from ordinary arthritis, and can effectively reduce the misdiagnosis rate. BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 A schematic diagram of an infrared thermal image auxiliary diagnosis system framework based on a multi-modal large model is provided for an embodiment of the application;
[0060] Figure 2 A schematic diagram of the arrangement of a visible light optical camera and a sign detection module in an infrared thermal image auxiliary diagnosis system is provided for an embodiment of the application;
[0061] Figure 3 A schematic diagram of an infrared thermal image auxiliary diagnosis process based on a multi-modal large model is provided for an embodiment of the application;
[0062] Figure 4 A schematic diagram of the module structure of an infrared thermal image auxiliary diagnosis system based on a multi-modal large model is provided for an embodiment of the application;
[0063] Figure 5 A schematic diagram of a human infrared thermal image preprocessing process is provided for an embodiment of the application;
[0064] Figure 6 A schematic diagram of human region segmentation is provided for an embodiment of the application;
[0065] Figure 7 A schematic diagram of three-dimensional point cloud image construction is provided for an embodiment of the application;
[0066] Figure 8 A schematic diagram of pain area calculation is provided for an embodiment of the application;
[0067] Figure 9 A schematic diagram of diagnosis result output is provided for an embodiment of the application;
[0068] Figure 10 A schematic diagram of an infrared thermal image auxiliary diagnosis method based on a multi-modal large model is provided for an embodiment of the application;
[0069] Figure 11A schematic diagram of the structure of a diagnostic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0071] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of this application and the appended claims, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to also include plural expressions, unless the context clearly indicates otherwise. It should also be understood that the terms "and / or" used in this application refer to and include any or all possible combinations of one or more listed items. The term "exemplary" means "serving as an example, embodiment or illustrative", and any embodiment described here as "exemplary" is not necessarily interpreted as being superior to or better than other embodiments. The terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of these features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" is two or more.
[0072] In view of the limitations of current traditional infrared thermal image disease diagnosis, the infrared thermal image-assisted diagnosis system and method based on a multimodal large model provided in the embodiments of the present application can well solve the above technical problems and achieve efficient, accurate and intelligent disease diagnosis effects for patients.
[0073] The infrared thermal image-assisted diagnosis system and method based on the multimodal large model provided in this application can be used as follows: Figure 1 The system architecture implementation shown mainly includes:
[0074] Infrared thermal imaging scanner, visible light optical camera array, bed frame, vital sign detection module, multimodal large model and voice module; the visible light optical camera array is arranged around the bed frame, and the cameras are all aimed at the bed frame; the vital sign detection module includes a self-sensing manipulator, which is set on the side of the bed frame.
[0075] The multimodal large model performs feature fusion and analysis on the infrared thermal image of the human body taken by the infrared thermal imaging scanner and the patient's main complaints collected by the voice module, and obtains the abnormal temperature area in the infrared thermal image as the diseased area; the visible light optical camera array takes multi-angle visible light images of the patient lying on the bed frame; the multimodal large model generates a three-dimensional model of the patient (i.e., a three-dimensional point cloud image) based on the multi-angle visible light image through the three-dimensional reconstruction model; the multimodal large model maps the diseased area in the infrared thermal image to the three-dimensional point cloud image through the mapping and matching module to obtain a three-dimensional infrared thermal image, and the diseased area in the three-dimensional infrared thermal image is recorded as Ω; the vital sign detection module controls the robotic arm to hammer and press the corresponding diseased area of the patient for the Ω area in the three-dimensional infrared thermal image, and the voice module records whether the patient has pain or abnormality to determine the pain area; the multimodal large model outputs the diagnosis results according to the patient's response; for example, including the type of disease, the diseased area and the treatment plan, where the treatment plan includes the treatment method and treatment point.
[0076] The infrared thermal image-assisted diagnosis system based on the multimodal large model provided in this application can realize the automatic collection, analysis and processing of disease information, and output the diagnosis results, which solves the limitations of traditional infrared thermal image disease diagnosis and achieves efficient, accurate and intelligent disease diagnosis effects for patients.
[0077] In an optional embodiment of this application, refer to Figure 2 The visible light optical camera array contains at least 5 synchronized cameras, which can be arranged in the four directions of the head, tail, left and right of the bed frame and above the bed frame to realize the full-dimensional collection of visible light images of the human body from 5 directions; specifically, the upper camera is located 2m above the bed frame, and the remaining cameras are 1m away from the edge of the bed frame and 0.5m above the bed frame, while maintaining a downward tilt of 30 degrees to ensure that the entire area of the human body can be covered; an FPGA synchronization board is used to ensure that multiple cameras can shoot simultaneously.
[0078] In an optional embodiment of the present application, the visible light optical camera array is calibrated using the Zhang Zhengyou calibration method before use, and the camera focal length, principal point, and distortion coefficient are calculated using a high-precision checkerboard to determine the camera intrinsic and extrinsic parameters.
[0079] In an optional embodiment of the present application, the vital sign detection module includes two sets of robotic arms, one on each side of the bed frame, and the other on each side. The robotic arms include a six-axis robotic arm and a self-sensing manipulator. According to the patient's affected area, a trajectory planning algorithm is used to control the motion trajectory of the robotic arm to ensure that the robotic arm moves to the patient's affected area without collision. The self-sensing manipulator is installed at the end of the six-axis robotic arm and can perform both pressing and hammering actions. The self-sensing manipulator has tactile sensors on its fingers and ulnar side, which can sense the force of pressing or hammering to avoid harming the patient.
[0080] In an optional embodiment of the present application, the infrared thermal imaging scanner is composed of an infrared darkroom and an infrared camera. The infrared camera performs a human body scan on the patient, and the patient is photographed by turning around to obtain the infrared thermal images of the front, back, left and right of the body. After the photographing is completed, the infrared thermal image data of the human body is uploaded to the multi-modal large model.
[0081] The voice module collects the chief complaint information and pain information described by the patient and converts them into text, and then uploads the text information to the multi-modal large model. The voice module can also receive the text information output by the multi-modal large model, convert it into a voice signal and play it to remind the patient to describe more detailed complaints or remind the patient of the next step.
[0082] The bed frame is used to assist in diagnosing the patient's signs, and the voice module can remind the patient of the posture (supine, lateral, prone) and direction (head of the bed, tail of the bed) on the bed frame according to the diagnosis needs.
[0083] As shown in Figure 3 , the present embodiment provides an infrared thermal image assisted diagnosis process based on a multi-modal large model, and the specific process includes:
[0084] S1: The patient enters the infrared thermal imaging scanner to take infrared thermal images of the front, back, left and right of the body, and transmits the infrared thermal images of the human body to the multi-modal large model.
[0085] S2: The patient describes the chief complaint, and the voice module converts the chief complaint into text information and transmits it to the multi-modal large model. The multi-modal large model judges the patient's disease type and affected area according to the infrared thermal images of the human body and the chief complaint.
[0086] S3: The patient lies on the bed frame, and the visible light optical camera array takes a full-body photograph of the patient. The three-dimensional reconstruction model is used to perform three-dimensional reconstruction on the photographed visible light images of the human body to obtain a three-dimensional point cloud image of the patient's body.
[0087] S4: The mapping matching module maps the affected area of the infrared thermal image of the human body to the three-dimensional point cloud image to obtain the affected area in the three-dimensional infrared thermal image.
[0088] S5: The sign detection module performs pressing or hammering detection actions on the affected area, and the voice module records the voice feedback data of the patient to the detection actions and converts them into text and then transmits them to the multi-modal large model in real time.
[0089] S6: The multi-modal large model combines the infrared thermal image, the patient's chief complaint and the sign information to output a diagnosis result, including the patient's treatment plan, such as moxibustion, acupuncture and other treatment methods and treatment points.
[0090] During the disease diagnosis process, the multimodal large model integrates the patient's human infrared thermal map and chief complaint information through the disease diagnosis module, and infers the diseased area in the patient's human infrared thermal map; in addition, the patient's physical sign information (pain area) can also be integrated to further obtain recommended treatment methods and treatment points; the three-dimensional reconstruction module constructs the patient's three-dimensional point cloud image by calculating the visible light image of the human body taken by the visible light optical camera array; the mapping matching module integrates the diseased area calculated by the disease diagnosis module and the patient's three-dimensional point cloud image data constructed by the three-dimensional reconstruction module, and calculates the diseased area in the patient's three-dimensional infrared thermal map; the physical sign detection module plans and calculates the motion trajectory according to the diseased area in the three-dimensional infrared thermal map and its own spatial position through the manipulator control algorithm, and controls the manipulator to move to the patient's diseased area to perform detection actions such as hammering or pressing to detect the patient's physical signs and obtain physical sign detection results (i.e., pain area); the disease diagnosis module outputs the diagnosis result based on the pain area.
[0091] refer to Figure 4 The present invention provides an infrared thermal image-assisted diagnosis system based on a multimodal large model, including:
[0092] Infrared collector, used to collect infrared thermal images of the human body;
[0093] Visible light optical camera, used to collect visible light images of the human body from multiple angles;
[0094] A voice module is used to obtain the patient's chief complaint and pain information;
[0095] The vital signs detection module is used to detect the diseased area;
[0096] The multimodal large model includes a preprocessing module, a three-dimensional reconstruction module, a mapping and matching module, and a disease diagnosis module; the preprocessing module is used to preprocess the human body infrared thermal image; the three-dimensional reconstruction module is used to construct a three-dimensional point cloud image based on the visible light image of the human body collected from multiple angles; the disease diagnosis module is used to determine the diseased area based on the preprocessed human body infrared thermal image and the chief complaint information; the mapping and matching module is used to map the diseased area to the three-dimensional point cloud image to obtain a three-dimensional infrared thermal map; the disease diagnosis module is also used to receive the pain information collected by the voice module based on the patient's detection action feedback, thereby determining the pain area based on the pain information; and output the diagnosis result based on the pain area.
[0097] Optionally, the multimodal model incorporates knowledge of disease physiology during pre-training to enhance the accuracy and expertise of the model in diagnosing diseases. The multimodal model first uses a pre-processing module to filter, visualize, and enhance features of infrared thermal images. It then combines the pre-processed infrared thermal images with the patient's primary complaint to deduce the disease type and affected area.
[0098] Optionally, during the pre-training phase, the multimodal large model uses real patient complaints and infrared thermal images as input, and uses disease types and disease areas diagnosed by professional physicians as standard outputs. Using a fine-tuning approach, the large model learns methods related to disease diagnosis. It should be noted that the output of the disease area is not directly generated as an image, but rather as coordinate points in text form, which are then converted into closed disease regions through post-processing.
[0099] By fine-tuning the multimodal large model, we can make the large model output what we want. Specifically, when fine-tuning the model, we need to prepare infrared thermal images + chief complaints as input, and disease types + disease areas as standard outputs. The output of the model in the early stages of training will be different from the standard output. At this time, the model will be "punished" (that is, the loss function), forcing the model to output closer to the standard output. After a certain amount of fine-tuning training, the output of the model is basically consistent with the standard output (at this time, the model has acquired the ability to output disease types + disease areas). In principle, the large model uses an image encoder and a text encoder to encode infrared thermal images and chief complaints respectively (the text encoder generates dynamic text feature vectors by capturing inter-word dependencies, and the image encoder generates image feature vectors by extracting spatial features of the image, such as object contours and textures). Then, a cross-attention mechanism is used to fuse the above two features. Finally, after receiving the multimodal features, the text decoder generates text word by word in an autoregressive manner.
[0100] refer to Figure 5 The preprocessing module includes a view selection module, an image restoration module and an image enhancement module; the view selection module is used to select one or more human body infrared thermal images from the front and rear data taken by the infrared thermal imaging scanner according to the main complaint information; the image restoration module is used to judge the temperature difference between the patient's affected side and the normal side and the affected area and the adjacent area according to the main complaint information, and restore the human body infrared thermal image selected by the view selection module to an RGB pseudo-color infrared image based on the central temperature and temperature difference; the image enhancement module is used to convert the RGB pseudo-color infrared image into an HSV image, enhance the hue and saturation channels, and then convert it into an RGB image.
[0101] Optionally, the view selection module consists of a pre-trained first BERT model and a text classification model; the pre-trained first BERT model converts the chief complaint information into a first embedding vector; the text classification model predicts the human infrared thermal image to be selected based on the first embedding vector. Specifically, the first BERT model first segments the text information (i.e., the chief complaint), and each subword after segmentation is used as a token. Each token is then converted into a 768-dimensional vector, i.e., the input vector, which contains Token Embedding (i.e., lexical semantic encoding), Segment Embedding (i.e., sentence attribution division), and Position Embedding (i.e., position information). Finally, a multi-layer Transformer encoder is used to calculate the dependency between different input vectors and generate a first embedding vector.
[0102] Optionally, the text classification model consists of a Transformer decoder based on an autoregressive mechanism, which uses self-attention and cross-attention mechanisms to calculate the probability distribution of the next word in the output sentence according to the first embedding vector generated by the first BERT model, selects the word with the highest probability, and finally forms an infrared heat map selection result.
[0103] Optionally, the pre-trained first BERT model and text classification model use the patient's chief complaint as input in the pre-training stage, use the infrared thermal image selection results of professional physicians as standard output, and utilize a fine-tuning-based method to enable the pre-trained first BERT model and text classification model to learn relevant methods for infrared thermal image selection.
[0104] Optionally, the image restoration module includes a grayscale image restorer, a human posture recognition model, a pre-trained second BERT model and a pseudo-color image generation network; the grayscale image restorer is used to map the temperature data in the selected human infrared thermal image to the RGB pixel space 0 to 255 to generate a grayscale image; the human posture recognition model is used to identify the key points of the human body in the grayscale image, and then divide the human body into regions according to the key point information, including left upper arm aa, left forearm ab, left hand ac, right upper arm ba, right forearm bb, right hand bc, left thigh ca, left calf cb, left foot cc, right thigh da, right calf db, right foot dc, left chest ea, right chest eb, left abdomen ec, right abdomen ed, left buttock ee, right buttock ef, left brain fa, and right brain fb, with reference to Figure 6The pre-trained second BERT model is used to encode the chief complaint and generate a second embedding vector to guide the localization of the affected area. The pseudo-color image generation network is used to generate RGB pseudo-color images from the temperature data, aiming to maximize the color difference between the affected and normal sides, as well as between the affected area and its adjacent areas. For example, if a patient reports three days of left abdominal pain, the affected side (affected area) is the left abdominal ec, the normal side is the right abdominal ed, and the adjacent areas are the left chest ea and left buttock ee.
[0105] Optionally, the pre-trained second BERT model uses real patient complaints in the pre-training stage, uses professional physicians' judgments on the patient's affected side and normal side as standard output, and utilizes a fine-tuning method to enable the pre-trained second BERT model and the region selection module of the pseudo-color image generation network to learn relevant methods for the selection of the affected side and the normal side.
[0106] Optionally, the human posture recognition model uses a pre-trained OpenPose model, which is trained using 10,000 grayscale human images and can recognize 18 key points of the human body, including nose 0, neck 1, right shoulder 2, right elbow 3, right wrist 4, left shoulder 5, left elbow 6, left wrist 7, right hip 8, right knee 9, right ankle 10, left hip 11, left knee 12, left ankle 13, right eye 14, left eye 15, right ear 16, and left ear 17.
[0107] Optionally, the pseudo-color image generation network consists of a region selection module and a temperature parameter adjustment module. The region selection module consists of a Transformer decoder based on an autoregressive mechanism, which decodes the embedding vector generated by the pre-trained second BERT model into text information containing the region selection results. The region selection results include the location information of the patient's affected and normal sides, as well as the affected area and adjacent areas. It should be noted that the affected side and the affected area refer to the same location.
[0108] Optionally, the temperature parameter adjustment module of the pseudo-color image generation network is set with two learnable variables, namely, central temperature (CT) and temperature width (TW); during training, the pseudo-color image generation network adopts self-supervised learning to adjust the values of central temperature and temperature width with the goal of maximizing the color difference between the affected side and the normal side and between the affected area and its adjacent areas.
[0109] Optionally, the pre-trained second BERT model uses real patient complaints in the pre-training stage, uses professional physicians' judgments on the patient's affected side and normal side as standard output, and utilizes a fine-tuning method to enable the pre-trained second BERT model and the region selection module of the pseudo-color image generation network to learn relevant methods for the selection of the affected side and the normal side.
[0110] Optionally, the image enhancement module includes:
[0111] Set two learnable parameters Hz, Ho and Sz, So for hue and saturation respectively, where H represents hue, S represents saturation, z represents scaling factor, and o represents offset; input a set of RGB pseudo-color images and convert the RGB pseudo-color images into HSV images using the following formula:
[0112] H'=H*H z +H o
[0113] S'=S*S z +S o
[0114] Where H' and S' represent the enhanced hue and saturation respectively;
[0115] Finally, the HSV image is restored to an RGB image. For example, the hsv_to_rgb function of the kornia library can be used to restore the HSV image to an RGB image.
[0116] Optionally, the 3D reconstruction model includes a feature extraction network, a feature alignment module, a multi-view fusion network and a PointNet decoder; the feature extraction network extracts the 2D semantic features of the image at each viewpoint and generates the corresponding feature vector; the feature alignment module is used to ensure that the multi-view features are consistent in space by minimizing the alignment loss, and to achieve cross-view consistency of the multi-view fusion features; the multi-view fusion network fuses the image features of multiple viewpoints into a unified 3D human body description through feature encoding; the PointNet decoder is used to decode the target human body mesh to obtain a 3D point cloud image that can represent the 3D model of the human body. Figure 7 shown.
[0117] Optionally, the feature extraction network consists of a ResNet network, which is encoded by the following formula:
[0118]
[0119] Among them, φ feat is the feature extraction network, h=H / s, w=W / s is the spatial resolution after downsampling, c is the number of channels, N is the number of views, I i is the i-th image.
[0120] Optionally, the feature alignment module aligns each feature map F i Alignment is performed by transforming the camera pose into a shared 3D space coordinate. Specifically, for a point in an image, select another image containing the point, record them as I1 and I2 respectively, detect the point by feature matching on the corresponding pixels (u1, v1) and (u2, v2), and estimate its spatial coordinates P by triangulation:
[0121]
[0122] P=(x,y,z)
[0123] where π i (·) is the camera projection function, z is the depth. Then the feature map F i The point corresponding to each pixel (u, v) in is back-projected into the camera coordinate system and then projected back into the world coordinate system. The formula is as follows:
[0124]
[0125] Among them, K i is the intrinsic parameter of the i-th camera, R i and T i is the camera extrinsic parameter, is the 3D position of the point in the camera coordinate system, p i Transform the point from the camera coordinate system back to the 3D coordinate system of the world coordinate system. Finally, all F i The average fusion method is used to map it to a unified 3D grid, which is recorded as The formula is as follows:
[0126]
[0127] Among them, Warp is the mapping function.
[0128] Optionally, the multi-view fusion network uses Transformer as the basic network to fuse the above aligned features. The formula is as follows:
[0129]
[0130] Among them, φ fusion is the Transformer network, F fused is the fused feature vector, and c′ is the number of channels of the 3D spatial feature obtained after fusion of multiple perspectives.
[0131] Optionally, use the PointNet decoder to generate a point cloud from the fused features:
[0132]
[0133] Among them, φ decode is the PointNet decoder, and P is the final point cloud.
[0134] Optionally, the 3D reconstruction model uses three loss functions to supervise the 3D reconstruction process. Specifically, it includes chamfer loss, alignment loss, and surface smoothness loss, and the formulas are as follows:
[0135]
[0136] L total =λ1L CD +λ2L align +λ3L smooth
[0137] Where G is the true value of the point cloud, p and q are the true point cloud set and the predicted point cloud set respectively. and is the feature after mapping from different perspectives, n i is the normal vector of the adjacent point, E is the edge set of the point cloud, L total is the total loss, λ1, λ2, and λ3 are the proportions of the three losses respectively.
[0138] Chamfer loss L CD Measuring the distance between the reconstructed point cloud and the real point cloud, reducing this loss can make the reconstructed point cloud closer to the real shape and optimize the overall geometric structure. align Constrain the reconstruction results to align with the multi-view images, improve the consistency and accuracy of the model under different perspectives, and solve the multi-view matching problem. smooth Make the reconstructed surface smoother, suppress noise and unreasonable bumps, enhance the visual rationality of the model surface, and conform to the surface characteristics of real objects. The three work together to achieve geometric accuracy, multi-view Figure 1 Optimize consistency and surface quality to improve the realism and reliability of 3D point cloud models.
[0139] Optionally, the disease diagnosis module diagnoses the diseased area based on the preprocessed human infrared thermal image and the patient's complaint and outputs a set of pixel points in the area, and then masks the diseased area using image masking technology.
[0140] Optionally, the mapping and matching module includes a 3D-2D projection module, a posture alignment module and a 2D-3D matching module; the 3D-2D projection module is used to project the vertices of the three-dimensional point cloud image to the infrared collector view; the posture alignment module is used to calculate the posture transformation matrix to align the posture of the three-dimensional point cloud image to the human posture space of the human infrared thermal map; the 2D-3D matching module is used to match the 2D points of the diseased area to the 3D projection points, and output the matched three-dimensional target area to obtain a three-dimensional infrared thermal map.
[0141] Optionally, the process of projecting the vertices of the three-dimensional point cloud image to the infrared thermal image of the human body is as follows: for each vertex in the three-dimensional point cloud image Calculate its projection coordinates in the i-th infrared camera view The formula is as follows:
[0142]
[0143] Among them, f is the focal length of the camera, dx and dy are the width and height of the actual physical size of a single pixel, c x and c y Represents the principal point of the camera in the x and y axis directions, and T i IR are the rotation matrix and translation matrix of the extrinsic parameters of the i-th infrared camera, v cam is the coordinate of the vertex in the three-dimensional model after being transformed into the camera coordinate system, v cam (1) v cam (2) and v cam (3) represent the values of the camera coordinates in the X, Y and Z axis directions respectively.
[0144] Furthermore, the process of the posture alignment module is as follows: using an iterative optimization algorithm to minimize the projected joint points With infrared joints The distance is as follows:
[0145]
[0146] The above optimization problem is solved by the existing related solution algorithm to obtain the posture transformation matrix (R pose ,T pose ), the posture of the 3D point cloud image can be aligned to the human posture space of the human infrared thermal map.
[0147] Furthermore, the 2D-3D matching module matches the 2D points of the diseased area to the 3D projection points as follows: given the target point of the diseased area in the infrared thermal image Calculate the projected coordinates of all visible vertices and Pixel distance, find the vertex with the smallest distance, the formula is as follows:
[0148]
[0149] In an optional embodiment of the present application, the 2D-3D matching module matches the 2D points of the diseased area to the 3D projection points as follows: First express it as homogeneous coordinates Using the posture transformation matrix (R pose ,T pose ) and camera intrinsic parameters Perform the inverse transform, the formula is as follows:
[0150]
[0151] Where V is the corresponding point in the 3D point cloud image. Ultimately, the diseased area in the infrared image is mapped to the 3D point cloud image.
[0152] The self-sensing robotic arm can perform two actions: pressing and hammering. Specifically, the self-sensing robotic arm first hammers the affected area and records the area where the patient feels pain. Then, it presses the above area and marks the points where the patient still feels pain, thus identifying the specific pain points.
[0153] The self-sensing manipulator has tactile sensors on its fingers and ulnar side, which can sense the force of pressing or hammering to avoid harming the patient; the self-sensing manipulator can move to the patient's diseased area according to the calculation results of the mapping matching module, and automatically calculate the motion trajectory according to the current task (hammering, pressing); specifically, any existing relevant path planning algorithm can be used, as long as it can realize the detection action of the diseased area, and this embodiment does not impose any restrictions on this.
[0154] The voice module collects the pain information fed back by the patient based on the detection action, and then determines the pain area based on the pain information. For example, when a detection action is performed on a certain point in the patient's diseased area, if the patient reports pain, it means that the point is a pain point; all pain points are clustered based on 3 cm to form a point set. For a point set with only 1 point, a circle with the point as the center and a radius of 1 cm is selected as the pain area; for a point set with 2 points, a line segment with the two points as endpoints and a width of 2 cm is selected as the pain area; for a point set with 3 or more points, the convex hull algorithm is used to calculate the minimum side enclosing polygon of the point set, and the polygon is used as the pain area. Reference Figure 8 .
[0155] The disease diagnosis module outputs diagnosis results based on the pain area, such as Figure 9 As shown. In this embodiment, the diagnosis result includes the disease type and the recommended treatment plan, wherein the treatment plan may include treatment methods and treatment points, such as moxibustion, acupuncture and other treatment methods and treatment points. Specifically, the multimodal large language model in the disease diagnosis module adds multiple rounds of question-answering process training in the pre-training stage. The first round of input of the large model is the patient's chief complaint and heat map, and the output is the patient's disease type and disease area. The second round of input of the large model is the patient's physical sign detection results, and the output is the treatment method and treatment point. In particular, the large model will determine the treatment method based on the medical knowledge base according to the disease type and physical sign detection results, and the selection of treatment points is related to the treatment method. When moxibustion treatment is used, the multimodal large model selects the treatment point according to the moxibustion thermodynamic model and the pain area to ensure that the pain area receives heat therapy evenly; when acupuncture treatment is used, the pain point is directly selected as the treatment point.
[0156] In this embodiment, the large model determines the disease type and the diseased area (which directly exists in the infrared thermal image) according to the infrared thermal image and the chief complaint, the sign detection module determines the pain point or the pain area according to the diseased area, and finally the large model determines the treatment point by comprehensively considering the disease type and the pain point / pain area. For example, the patient's chief complaint is "I have had right shoulder pain for 3 months", and the infrared thermal image shows that the right shoulder is higher in temperature. Then the multi-modal large model determines that the patient has shoulder periarthritis and points out the diseased area. The sign detection module detects the patient's shoulder area to obtain the pain area result as shown in FIG. 8. Figure 8
[0157] Based on the infrared thermal image assisted diagnosis system provided in this embodiment, by using the multi-modal large model, the image preprocessing, feature extraction and disease probability prediction can be completed within a few seconds; by combining the multi-modal technology, the infrared thermal image and the patient's clinical symptoms are cross-analyzed, which can more accurately distinguish rheumatoid arthritis from ordinary arthritis, and can effectively reduce the diagnosis rate.
[0158] Based on the infrared thermal image assisted diagnosis system provided in this embodiment, by using the multi-modal large model, the image preprocessing, feature extraction and disease probability prediction can be completed within a few seconds; by combining the multi-modal technology, the infrared thermal image and the patient's clinical symptoms are cross-analyzed, which can more accurately distinguish rheumatoid arthritis from ordinary arthritis, and can effectively reduce the diagnosis rate.
[0159] Reference Figure 10 A multi-modal large model based infrared thermal image assisted diagnosis method, comprising:
[0160] S101, collecting the chief complaint information described by the patient, the human infrared thermal image and the multi-angle human visible light image;
[0161] S102, constructing a three-dimensional point cloud image based on the multi-angle human visible light image;
[0162] S103, preprocessing the human infrared thermal image;
[0163] S104, determining the diseased area according to the preprocessed human infrared thermal image and the chief complaint information;
[0164] S105, mapping the diseased area to the three-dimensional point cloud image to obtain a three-dimensional infrared thermal image;
[0165] S106, detecting the diseased area;
[0166] S107, receiving the pain information fed back by the patient based on the detection action, and determining the pain area based on the pain information;
[0167] S108, outputting a diagnosis result based on the pain area.
[0168] Through the foregoing detailed description of the system, a person skilled in the art can clearly understand the specific implementation process of the method in the embodiment. For the sake of brevity of the description, the specific implementation process is not described here.
[0169] In order to better execute the program of the method described above, the embodiment of the present application further provides a diagnostic device, as shown in the figure, which comprises a processor, a memory and a communication bus for realizing the communication connection between the processor and the memory. Figure 11
[0170] The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function, and instructions for implementing the infrared thermal image assisted diagnosis method provided by the above embodiments, etc.; the data storage area can store data involved in the infrared thermal image assisted diagnosis method provided by the above embodiments, etc.
[0171] Optionally, the memory is a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compressed disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disc storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program codes in the form of instructions or data structures and capable of being accessed by a computer, but is not limited to this. The memory exists independently and is connected to the processor through the communication bus, or the memory is integrated with the processor.
[0172] The processor may include one or more processing cores. The processor calls the data stored in the memory by running or executing the instructions, programs, code sets or instruction sets stored in the memory, performs the various functions of the present application and processes data. The processor can be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller and a microprocessor. It is understandable that for different devices, the electronic device for realizing the above-mentioned processor function can also be other, and the embodiments of the present application are not specifically limited.
[0173] The communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, and so on. For ease of illustration, the figure shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0174] In an optional embodiment, the diagnostic device may further include a communication interface (not shown) for communication between the diagnostic device and other devices.
[0175] The present application provides a computer-readable storage medium, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code. The computer-readable storage medium stores a computer program capable of being loaded by a processor and executing the infrared thermal image-assisted diagnosis method of the above-described embodiment.
[0176] An embodiment of the present application also provides a computer program product, which includes a computer program tangibly contained on a readable medium, wherein the computer program includes program code for executing any infrared thermal image-assisted diagnosis method in the embodiments of the present application. The computer program can be downloaded and installed over the network, and / or installed from a removable medium (such as a disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc.).
[0177] The above embodiments are merely a detailed introduction to the technical solutions of the present application. However, the description of the above embodiments is only intended to help understand the method and core concept of the present application and should not be construed as limiting the present application. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application.
Claims
1. An infrared thermal image-assisted diagnosis system based on a multimodal large model, characterized in that: The infrared thermal image auxiliary diagnosis system includes: Infrared collector, used to collect infrared thermal images of the human body; Visible light optical camera, used to collect visible light images of the human body from multiple angles; A voice module is used to obtain the patient's chief complaint and pain information; The vital signs detection module is used to detect the diseased area; A large multimodal model, including a preprocessing module, a 3D reconstruction module, a mapping and matching module, and a disease diagnosis module; The preprocessing module is used to preprocess the infrared thermal image of the human body; The three-dimensional reconstruction module is used to construct a three-dimensional point cloud image based on visible light images of the human body collected from multiple angles; The disease diagnosis module is used to determine the diseased area based on the pre-processed human infrared thermal image and the main complaint information; The mapping and matching module is used to map the diseased area to a three-dimensional point cloud image to obtain a three-dimensional infrared thermal map; The disease diagnosis module is further configured to receive pain information collected by the voice module and fed back by the patient based on the detection action, thereby determining the pain area based on the pain information; and output a diagnosis result based on the pain area.
2. The infrared thermal image-assisted diagnosis system according to claim 1, characterized in that: The pre-processing module includes a view selection module, an image restoration module and an image enhancement module; The view selection module is used to select one or more infrared thermal images of the human body from the front and rear data captured by the infrared thermal imaging scanner according to the main complaint information; The image restoration module is used to determine the temperature difference between the affected side and the normal side of the patient based on the main complaint information, and restore the human body infrared thermal image selected by the view selection module into an RGB pseudo-color infrared image based on the central temperature and temperature difference; The image enhancement module is used to convert the RGB pseudo-color infrared image into an HSV image, enhance the hue and saturation channels, and then convert it into an RGB image.
3. The infrared thermal image-assisted diagnosis system according to claim 2, characterized in that: The view selection module consists of a pre-trained first BERT model and a text classification model; the pre-trained first BERT model converts the main complaint information into a first embedding vector; the text classification model predicts the human body infrared thermal image to be selected based on the first embedding vector.
4. The infrared thermal image-assisted diagnosis system according to claim 2, characterized in that: The image restoration module includes a grayscale image restorer, a human posture recognition model, a pre-trained second BERT model and a pseudo-color image generation network; The grayscale image restorer is used to map the temperature data in the selected human body infrared thermal image to the RGB pixel space 0 to 255 to generate a grayscale image; The human posture recognition model is used to identify the key points of the human body in the grayscale image, including the left upper arm aa, left forearm ab, left hand ac, right upper arm ba, right forearm bb, right hand bc, left thigh ca, left calf cb, left foot cc, right thigh da, right calf db, right foot dc, left chest ea, right chest eb, left abdomen ec, right abdomen ed, left buttock ee, right buttock ef, left brain fa, and right brain fb; The pre-trained second BERT model is used to encode the chief complaint information and generate a second embedding vector to guide the localization of the affected side area; The pseudo-color image generation network is used to generate temperature data into RGB pseudo-color images with the goal of maximizing the color difference between the affected side and the normal side.
5. The infrared thermal image-assisted diagnosis system according to claim 4, characterized in that: The pseudo-color image generation network is set with two learnable variables, namely core temperature (CT) and temperature width (TW); during training, the pseudo-color image generation network adopts self-supervised learning to maximize the color difference between the affected side and the normal side, and adjusts the values of core temperature and temperature width.
6. The infrared thermal image-assisted diagnosis system according to claim 2, characterized in that: The image enhancement module includes: Set two learnable parameters Hz, Ho and Sz, So for hue and saturation respectively, where H represents hue, S represents saturation, z represents scaling factor, and o represents offset; input a set of RGB pseudo-color images and convert the RGB pseudo-color images into HSV images using the following formula: H'=H*H z +H o S'=S*S z +S o Where H' and S' represent the enhanced hue and saturation respectively; Finally, the HSV image is restored to an RGB image.
7. The infrared thermal image-assisted diagnosis system according to any one of claims 1 to 6, characterized in that: The 3D reconstruction model includes a feature extraction network, a feature alignment module, a multi-view fusion network and a PointNet decoder; The feature extraction network is used to extract the 2D semantic features of the visible light image of the human body at each viewing angle and generate the corresponding feature vector; The feature alignment module is used to ensure that multi-view features are spatially consistent by minimizing the alignment loss, thereby achieving cross-view consistency of multi-view fusion features. The multi-view fusion network fuses the image features of multiple views into a unified 3D human body description through feature encoding; The PointNet decoder is used to decode the target human body mesh to obtain a three-dimensional point cloud image that can represent the human body 3D model.
8. The infrared thermal image-assisted diagnosis system according to any one of claims 1 to 6, characterized in that: The mapping matching module includes a 3D-2D projection module, a posture alignment module and a 2D-3D matching module; The 3D-2D projection module is used to project the vertices of the 3D point cloud image to the infrared collector view; The posture alignment module is used to calculate the posture transformation matrix to align the posture of the 3D point cloud image to the human posture space of the human infrared thermal image; The 2D-3D matching module is used to match the 2D points of the diseased area to the 3D projection points to obtain a three-dimensional infrared thermal map.
9. An infrared thermal image-assisted diagnosis method based on a multimodal large model, characterized in that: include: Collect the patient's chief complaint information, human infrared thermal images, and multi-angle visible light images of the human body; Construct a 3D point cloud image based on visible light images of the human body from multiple angles; Preprocessing of human body infrared thermal images; Determine the diseased area based on the pre-processed human infrared thermal image and the chief complaint information; Mapping the diseased area to a three-dimensional point cloud image to obtain a three-dimensional infrared thermal map; Perform detection actions on the diseased area; receiving pain information fed back by the patient based on the detection action, and determining a pain area based on the pain information; Outputs diagnosis results based on the pain area.
Citation Information
Patent Citations
Thermal imaging physiological detection system
CN108294731A
Multifunctional early examination comprehensive diagnostic instrument
CN108742703A
Automatic palpation method and device, electronic equipment and storage medium
CN113598714A
Tongue diagnosis basic information collection method and system
CN119732653A
Intelligent body temperature real-time monitoring system and method
CN119915383A