Image processing method, computer aided diagnosis method, and method for training image processing model

By extracting multi-scale feature information and deep learning training, the medical image recognition model is optimized, the problem of low recognition accuracy is solved, and efficient image recognition and multi-dimensional detection result generation are achieved.

WO2025185337A1PCT designated stage Publication Date: 2025-09-11ALIBABA (CHINA) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/070596
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2025-01-03
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

The recognition accuracy of existing medical image recognition models is low, and model training relies on a large amount of labeled data from experienced professional doctors, resulting in poor training results.

Method used

By receiving multiple target images, using image processing models to extract multi-scale feature information, generating detection annotation information, detection category information and detection guidance text, and combining deep learning and transfer training technology, the training process of the image processing model is optimized.

Benefits of technology

The detection accuracy of the image recognition model is improved, and multi-dimensional detection results are generated, including the location information, category information and guidance text of abnormal objects, which enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070596_12092025_PF_FP_ABST
    Figure CN2025070596_12092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are an image processing method, and a method for training an image processing model. The image processing method comprises: receiving an image processing task, wherein the image processing task carries a plurality of target images corresponding to a target detection region, and the image processing task is used for detecting whether there is an abnormal object in the target detection region; and inputting the plurality of target images into an image processing model, so as to obtain a detection result corresponding to the target detection region, wherein the detection result corresponding to the target detection region is generated on the basis of multi-scale feature information corresponding to the plurality of target images, and the detection result comprises detection annotation information, detection category information and detection guidance text. Multi-scale feature information corresponding to target images is obtained, thereby improving the accuracy of a subsequently generated detection result. The detection result comprises position information of an abnormal object, information of an object to be subjected to detection, and guidance text, so that the detection result is enriched, and multi-dimensional detection information is provided for a user, thereby improving the usage experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing methods, computer-aided diagnosis methods, and image processing model training methods

[0001] This disclosure claims priority to Chinese patent application number 202410257868.3 filed with the Patent Office of China on March 6, 2024, entitled “Image processing method, training method of image processing model”, the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The present disclosure relates to the field of computer technology, and in particular to an image processing method. Background Art

[0003] As people's living standards improve, more and more people pay attention to their own health. Tumors are one of the major factors affecting human health. Identifying tumors in medical images requires professional doctors to identify them based on their experience. Due to the limitations of doctors' experience, image recognition and analysis with the help of medical images has become an important topic.

[0004] Artificial intelligence systems have demonstrated tremendous potential. Using large models to identify medical images has made significant progress in computer-aided diagnosis (CAD) tasks. However, current models for medical image recognition and analysis suffer from low accuracy. Furthermore, training these models requires a large amount of labeled data, which in turn requires the expertise of experienced professional physicians. Consequently, model training effectiveness is limited. Therefore, improving the accuracy of image recognition models has become a pressing challenge for researchers. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide an image processing method. One or more embodiments of the present disclosure also provide an image processing method, a computer-aided diagnosis method for cancer, a training method for an image processing model, a computer-aided diagnosis method, an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0006] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:

[0007] receiving an image processing task, wherein the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;

[0008] The multiple target images are input into an image processing model to obtain detection results corresponding to the target detection area, wherein the image processing model generates detection results corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection results include detection annotation information, detection category information and detection guidance text.

[0009] According to a second aspect of an embodiment of the present disclosure, a computer-aided diagnosis method for cancer is provided, comprising:

[0010] receiving a computed tomography (CT) image processing task, wherein the CT image processing task carries a plurality of CT images corresponding to a target detection area, and the CT image processing task is used to detect whether a tumor exists in the target detection area;

[0011] The multiple CT images are input into a CT image processing model to obtain detection results corresponding to the target detection area, wherein the CT image processing model generates detection results corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection results include detection annotation information, detection category information and detection guidance text.

[0012] According to a third aspect of an embodiment of the present disclosure, a method for training an image processing model is provided, which is applied to a cloud-side device and includes:

[0013] Obtaining a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image;

[0014] Inputting the sample image and the sample guidance text into an image processing model to obtain predicted labeling information, predicted category information, and text loss value output by the image processing model;

[0015] Calculating a model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information;

[0016] Adjusting the model parameters of the image processing model according to the model loss value and the text loss value, and continuing to train the image processing model until a model training stop condition is reached, thereby obtaining the model parameters of the image processing model;

[0017] The model parameters of the image processing model are sent to the end-side device.

[0018] According to a fourth aspect of an embodiment of the present disclosure, a computer-aided diagnosis method is provided, comprising:

[0019] receiving a CT image processing task sent by a user, wherein the CT image processing task carries multiple CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area;

[0020] Inputting the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information, and detection guidance text;

[0021] Send the detection result corresponding to the target detection area to the user.

[0022] According to a fifth aspect of an embodiment of the present disclosure, there is provided a computing device, including:

[0023] memory and processor;

[0024] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned image processing method, CT image processing method or image processing model training method are implemented.

[0025] According to the sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.

[0026] According to the seventh aspect of the embodiments of the present disclosure, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.

[0027] An image processing method provided by one embodiment of the present disclosure uses an image processing model to obtain multiple scale feature information corresponding to a target image, thereby improving the accuracy of subsequent detection results. The generated detection results include the location of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results and providing users with multi-dimensional detection information, thereby enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] FIG1 is an architecture diagram of an image processing system provided by one embodiment of the present disclosure;

[0029] FIG2 is a flowchart of an image processing method provided by one embodiment of the present disclosure;

[0030] FIG3 is a schematic diagram of the structure of an image processing model provided by an embodiment of the present disclosure;

[0031] FIG4 is a flowchart of a CT image processing method provided by one embodiment of the present disclosure;

[0032] FIG5 is a flowchart of a method for training an image processing model provided by one embodiment of the present disclosure;

[0033] FIG6 is a flowchart of another image processing method provided by an embodiment of the present disclosure;

[0034] FIG7 is a flowchart of an image processing method applied to an esophageal cancer detection scenario provided by one embodiment of the present disclosure;

[0035] FIG8 is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure;

[0036] FIG9 is a schematic structural diagram of a CT image processing device provided by one embodiment of the present disclosure;

[0037] FIG10 is a schematic structural diagram of another image processing device provided by an embodiment of the present disclosure;

[0038] FIG11 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.

[0040] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0041] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0042] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0043] First, the terms involved in one or more embodiments of the present disclosure are explained.

[0044] CT (Computed Tomography): Computerized tomography uses precisely collimated X-ray beams, gamma rays, ultrasound waves, etc., together with highly sensitive detectors to perform cross-sectional scans around a certain part of the human body one by one. It has the characteristics of fast scanning time and clear images, and can be used to detect a variety of diseases.

[0045] CAD (computer aided diagnosis): Computer-aided diagnosis refers to the use of imaging, medical image processing technology and other possible physiological and biochemical means, combined with computer analysis and calculation, to assist in the discovery of lesions and improve the accuracy of diagnosis.

[0046] EC (esophageal cancer): Esophageal cancer is a highly lethal cancer with a low 5-year survival rate according to incomplete statistics. However, early detection of resectable / curable esophageal cancer can significantly reduce the mortality rate. Lymph node metastasis is a common and typical symptom.

[0047] As people's living standards improve, more and more people pay attention to their own health. Tumors are one of the major factors affecting human health. Identifying tumors in medical images requires professional doctors to identify them based on their experience. Due to the limitations of doctors' experience, image recognition and analysis with the help of medical images has become an important topic.

[0048] Artificial intelligence systems have demonstrated tremendous potential. Using large models to identify medical images has made significant progress in computer-aided diagnosis (CAD) tasks. However, current models for medical image recognition and analysis suffer from low accuracy. Most AI systems rely heavily on tumor grade annotation, which requires the expertise of experienced radiologists. Furthermore, clinical reports also contain rich descriptive information, which current CAD systems are unable to effectively utilize.

[0049] Based on this, an image processing method is provided in the present disclosure. The present disclosure also involves a CT image processing method, a training method for an image processing model, an image processing device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0050] Referring to FIG1 , FIG1 shows an architecture diagram of an image processing system provided by an embodiment of the present disclosure. The image processing system may include a client 100 and a server 200;

[0051] The client 100 is configured to send an image processing task to the server 200, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is configured to detect whether there is an abnormal object in the target detection area;

[0052] The server 200 is configured to input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection area, wherein the image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, the detection result including detection annotation information, detection category information, and detection guidance text; and send the detection result to the client 100;

[0053] The client 100 is also used to receive the detection results sent by the server 200.

[0054] Applying the solution of the embodiment of the present disclosure, an image processing task is received, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there are abnormal objects in the target detection area; the multiple target images are input into an image processing model to obtain a detection result corresponding to the target detection area, wherein the image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection result includes detection annotation information, detection category information and detection guidance text.

[0055] In this way, in the process of processing multiple target images using the image processing model, multi-scale feature information of multiple target images is extracted, and detection annotation information, detection category information and detection guidance text are generated based on the multi-scale feature information.

[0056] The image processing system may include multiple clients 100 and a server 200. The clients 100 may be referred to as client-side devices, and the server 200 may be referred to as cloud-side devices. Multiple clients 100 may establish communication connections through the server 200. In image processing scenarios, the server 200 provides image processing services to multiple clients 100. Multiple clients 100 may function as either senders or receivers, communicating through the server 200.

[0057] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100. In an image processing scenario, a user can publish a data stream to the server 200 through the client 100, and the server 200 generates detection results based on the data stream and pushes the detection results to other clients with which communication has been established.

[0058] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.

[0059] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client 100 can be based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0060] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that provide background training to support models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server that is integrated with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0061] It is worth noting that the image processing methods provided in the embodiments of the present disclosure are generally executed by the server. However, in other embodiments of the present disclosure, the client may also have similar functions to the server to execute the image processing methods provided in the embodiments of the present disclosure. In other embodiments, the image processing methods provided in the embodiments of the present disclosure may also be executed jointly by the client and the server.

[0062] Referring to FIG. 2 , FIG. 2 shows a flow chart of an image processing method provided by an embodiment of the present disclosure, which specifically includes the following steps:

[0063] Step 202: Receive an image processing task, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.

[0064] In actual applications, the image processing tasks sent by users can be received through the server or the client.

[0065] Specifically, an image processing task is a task for detecting the presence of abnormal objects within a target detection area. The image processing task carries multiple target images corresponding to the target detection area. Furthermore, the target detection area can be understood as a subarea used to detect the presence of abnormal objects, and the abnormal object specifically refers to a foreign object within the target detection area. For example, the abnormal object can be a tumor in the human body.

[0066] The target detection area can be any organ in the human body, such as the liver, lungs, stomach, esophagus, etc. By predicting whether there is an abnormal object in the target detection area, the status of the object to be detected can be further determined based on the prediction results, thereby providing assistance in locating the abnormal object and providing precise treatment.

[0067] The object to be detected can be understood as the object described in the target detection area. For example, if the target image is a CT image of Zhang San's stomach, then the target detection area is the stomach, the CT image is the target image, and the object to be detected is Zhang San. The image processing task is used to detect whether there is a tumor in the stomach area. In practical applications, the object to be detected can be a human or other living organism, and this is not limited to this in one or more specific embodiments provided in this disclosure.

[0068] It should be noted that in one or more embodiments of the present disclosure, the image processing task can be applied to the recognition of various types of medical images, and the determination of whether there are abnormal objects in the target detection area of ​​the medical image based on the image features. For example, in the application scenario of gastric cancer detection, the presence of a tumor in the stomach area can be predicted based on the medical image of the stomach area, thereby helping doctors to accurately locate the abnormal area; in the application scenario of esophageal cancer detection, the presence of a tumor in the esophagus can be predicted based on the medical image of the esophagus area, thereby helping doctors to accurately locate the abnormal area and facilitate subsequent treatment.

[0069] For example, in the scenario of esophageal cancer detection, the target image acquired is an image of the esophagus. Specifically, the multiple target images acquired are CT images of the esophagus. These multiple target images can form a 3D image of the esophageal region. The abnormal object can be understood as a tumor in the esophagus. Multiple CT images corresponding to the esophageal region are acquired and image detection processing is performed on the multiple CT images to detect whether a malignant tumor is present in the esophagus.

[0070] In practical applications, the abnormal object may be a certain type of cell, a certain type of tissue structure, etc., for example, a malignant tumor, a benign tumor, a hyperplastic tissue, etc. This is not limited in one or more embodiments provided in the present disclosure.

[0071] By receiving the image processing task, multiple target images corresponding to the target detection area carried in the image processing task can be used as input to detect whether there is an abnormal object in the target detection area.

[0072] In a specific embodiment provided by the present disclosure, before receiving the image processing task, the method further includes:

[0073] receiving an image segmentation task, wherein the image segmentation task carries a plurality of initial images corresponding to a target detection area, and the image segmentation task is used to extract a target image corresponding to the target detection area;

[0074] Each initial image is input into a pre-trained image segmentation model to obtain a target image corresponding to each initial image output by the image segmentation model.

[0075] In the embodiments provided in this disclosure, the target image can be understood as a close-up image corresponding to the target detection area. However, in actual applications, the multiple images received may include other areas in addition to the target detection area, and these other areas may affect the target detection area. Therefore, the method provided in this disclosure also processes the image.

[0076] Specifically, an image segmentation task is first obtained for multiple initial images, where the initial images include both target detection regions and regions that affect the target detection regions. The image segmentation task is used to extract the target images corresponding to the target detection regions from each of the initial images.

[0077] Multiple initial images are fed into a pre-trained image segmentation model for processing. The image segmentation model is trained to identify the target detection region in the initial image, extract the target detection region from the initial image, and generate a target image corresponding to the target detection region.

[0078] Taking CT images as an example, the method provided herein first acquires an initial CT image. This initial CT image is a plain scan CT image that meets image quality requirements. The initial CT image can be obtained from multiple CT scanners or from a single CT scanner, a fact not limited in this disclosure. After obtaining the initial CT images, the format of each initial CT image must be standardized. The image segmentation model can be 3DUNet.

[0079] 3DUNet is a deep learning architecture for image segmentation in three-dimensional space. It is an extension of the U-Net model, which was originally designed for semantic segmentation of two-dimensional biomedical images and is widely recognized for its excellent performance and high accuracy in segmenting small objects. 3DUNet applies this concept to three-dimensional datasets, such as medical images (such as CT and MRI scans), which are useful in many medical fields because these data often contain rich three-dimensional structural information.

[0080] 3DUNet includes an encoder-decoder structure for capturing global context information, restoring lost spatial details and generating accurate pixel-level segmentation labels. The 3D convolution kernel in the 3DUNet structure replaces the 2D convolution kernel in the network, and can simultaneously process the features of the input data in three dimensions (length, width, and height). At the same time, 3DUNet can also effectively fuse three-dimensional features at different levels, which helps to extract complex shape and structure information. This framework has a good effect in medical image segmentation methods. In the method provided in the present disclosure, the initial CT image is cut using the preprocessing strategy of the 3DUNet structure to obtain the target image corresponding to the target detection area for subsequent processing.

[0081] Step 204: Input the multiple target images into the image processing model to obtain the detection results corresponding to the target detection area, wherein the image processing model generates the detection results corresponding to the target detection area based on the multi-scale feature information corresponding to the multiple target images, and the detection results include detection annotation information, detection category information and detection guidance text.

[0082] In actual applications, after receiving an image processing task, multiple target images carried by the image processing task can be obtained from the image processing task, and the multiple target images can be input into the image processing model to obtain the detection results corresponding to the target detection area output by the image processing model. The detection results specifically include detection annotation information, detection category information and detection guidance text.

[0083] The detection annotation information specifically refers to the area where the abnormal object is marked in the target image. If there is no abnormal object in the target image, the detection annotation information is empty.

[0084] Detection category information specifically refers to the category information of the object to be detected corresponding to the target detection area for detection of the target image. If the target detection area contains an abnormal object, the detection category information is abnormal; if the target detection area does not contain an abnormal object, the detection category information is normal.

[0085] The detection guidance text specifically refers to the guidance text for the object to be detected, which is output based on the detection annotation information, detection category information, etc. of the target image, regarding the position, status, and other information of the abnormal object. It should be noted that the detection guidance text provided by the embodiment of the present disclosure is a guidance text generated by adding the detected abnormal object position, status, and other information to the guidance text template. If the detection annotation information determines that there is an abnormal object, the specific location of the abnormal object and the category information of the object to be detected will be given in the detection annotation information. For example, the detection guidance text can be "The patient has a tumor in the upper part of the esophagus, and the patient has esophageal cancer." For another example, the detection guidance text can be "No tumor was detected in the patient's stomach, and the patient was not detected with gastric cancer," and so on.

[0086] The multi-scale feature information of each target image is extracted within the image processing model, and then the detection results corresponding to the target detection area are extracted based on the multi-scale feature information.

[0087] Specifically, the image processing model includes a multi-scale feature extraction module, a feature fusion module, and a feature processing module;

[0088] Inputting the multiple target images into an image processing model to obtain detection results corresponding to the target detection area includes S2042-S2046:

[0089] S2042: Input the multiple target images into the multi-scale feature extraction module to obtain at least one scale feature information.

[0090] Among them, the multi-scale feature extraction module specifically refers to using different scale information to extract target feature information of a 3D model composed of multiple target images. For example, for a 3D model with a size of W*H*D, it is input into the multi-scale feature extraction module. The multi-scale feature extraction module can be understood as the backbone feature extraction network of 3DUNet to obtain a multi-scale feature map F = {F0, F1, F2, ... FS}, where S is the number of scale layers. In one embodiment provided in the present disclosure, S = 5, that is, the multi-scale feature map F = {F0, F1, F2, F3, F4, F5}. For the feature map corresponding to the i-th scale layer

[0091] Referring to Figure 3, Figure 3 shows a structural schematic diagram of an image processing model provided by an embodiment of the present disclosure. As shown in Figure 3, multiple target images are input into the multi-scale feature extraction module of the image processing model. After 5 stages of downsampling processing, multiple initial feature information of different scales is obtained. Then, the initial feature information of multiple different scales is convolved or deconvolved-convolved to obtain scale feature maps corresponding to each scale, thereby obtaining 6 multi-scale feature maps.

[0092] S2044: Input the feature information of each scale into the feature fusion module to obtain feature fusion information.

[0093] After obtaining multiple scale feature information, the multiple scale feature information is input into the feature fusion layer for feature fusion, and the multiple scale feature information is unified into the same scale and fused to obtain feature fusion information.

[0094] Referring to Figure 3, multiple scale feature maps {F0, F1, F2, F3, F4, F5} are input into the feature fusion layer. In the feature fusion layer, the multiple scale feature maps are unified into the same scale for feature fusion to obtain fused feature information Fa.

[0095] S2046: Input the feature fusion information into the feature processing module to obtain a detection result corresponding to the target detection area.

[0096] After obtaining the feature fusion information, the feature fusion information can be input into the feature processing module, where the feature fusion information is processed to generate a detection result corresponding to the target detection area. In the embodiments provided in the present disclosure, the detection result specifically includes detection annotation information, detection category information, and detection guidance text.

[0097] In a specific embodiment provided by the present disclosure, the feature processing module includes an abnormal object segmentation unit, an abnormal object classification unit and a guidance text generation unit;

[0098] Inputting the feature fusion information into the feature processing module to obtain the detection result corresponding to the target detection area includes:

[0099] Inputting the feature fusion information into the abnormal object segmentation unit to obtain detection labeling information;

[0100] Inputting the feature fusion information into the abnormal object classification unit to obtain detection category information;

[0101] Inputting the feature fusion information into the guidance text generation unit to obtain a detection guidance text;

[0102] A detection result corresponding to the target detection area is generated according to the detection label information, detection category information, and detection guidance text.

[0103] In practical applications, the feature processing module is used to process feature fusion information, extract features from the feature fusion information, and decode the extracted features to generate corresponding detection results. The feature processing module specifically includes an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit.

[0104] Among them, the abnormal object segmentation unit is used to mark the area of ​​the abnormal object in the target image according to the feature fusion information, and segment the abnormal object corresponding to the target detection area in the target image; the abnormal object classification unit is used to determine whether there is an abnormal object in the target detection area according to the feature fusion information, and determine the classification information for the object to be detected; the guidance text generation unit is used to generate guidance text for indicating information such as the position and status of the abnormal object to be detected according to the feature fusion information.

[0105] The final detection result is generated according to the detection annotation information, detection category information and detection guidance text output by the abnormal object segmentation unit, the abnormal object classification unit and the guidance text generation unit respectively.

[0106] Furthermore, the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit;

[0107] Inputting the feature fusion information into the guidance text generation unit to obtain a detection guidance text includes:

[0108] Inputting the feature fusion information into the abnormal position text subunit to obtain position guidance information;

[0109] Inputting the feature fusion information into the abnormal result text subunit to obtain result guidance information;

[0110] Generate a detection guidance text according to the position guidance information and the result guidance information.

[0111] In practical applications, the detection guidance text specifically includes guidance information on the abnormal location and the status of the object to be detected. Therefore, the guidance text generation unit includes an abnormal location text subunit and an abnormal result text subunit. The abnormal location text subunit is used to generate the location information of the abnormal object in the target detection area based on the feature fusion information, and the abnormal result text subunit is used to generate the status information of the object to be detected.

[0112] For example, when inspecting a user's stomach CT image to detect abnormalities in the stomach and generate detection guidance text, the abnormal location text subunit processes the feature fusion information and determines the location guidance information as "the middle of the stomach." The abnormal result text subunit also processes the feature fusion information and determines the result guidance information as "the user has stomach cancer." Based on the location guidance information and the result guidance information, the final detection guidance text is generated: "The user has a tumor in the middle of the stomach and has stomach cancer."

[0113] The image processing method provided by the embodiment of the present disclosure inputs multiple target images of the target detection area of ​​the object to be detected into the image processing model. In the image processing model, multiple scale features are extracted from the 3D images corresponding to the multiple target images. After the multiple scale features are fused, annotation information, category information, and guidance text are generated respectively, thereby finally generating a detection result. The accuracy of the subsequent generation of detection results is improved by the multiple scale feature information. The generated detection results include the location information of the abnormal object, the information of the object to be detected, and the guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.

[0114] With the continuous development of computer technology, deep learning has gradually been applied to various medical computer-assisted diagnosis tasks. Deep learning models rely on large-scale, accurately labeled training samples and sample labels as training data for model training. Currently, during the model training process, the training data for image processing models is images that include abnormal objects in the target detection area. The processing effect of images that do not include abnormal objects in the target detection area is poor. Based on this, in a specific embodiment provided by the present disclosure, the image processing model is trained by the following steps:

[0115] Obtaining a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image;

[0116] Inputting the sample image and the sample guidance text into an image processing model to obtain predicted labeling information, predicted category information, and text loss value output by the image processing model;

[0117] Calculating a model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information;

[0118] Adjust the model parameters of the image processing model according to the model loss value and the text loss value, and continue to train the image processing model until a model training stop condition is reached.

[0119] It should be noted that the training method of the image processing model provided in the present invention uses the ideas of supervised training and transfer training. First, an initial model is trained using labeled training data, and then the backbone network in the initial model is migrated to the image processing model, and further training is performed with another batch of training data to obtain the final image processing model.

[0120] Sample images specifically refer to images used for model training. In practical applications, sample images include both images with and without abnormal objects. Sample annotation information corresponding to sample images specifically refers to the annotation information for abnormal objects in the sample images; sample category information specifically refers to the status information of the object to be detected corresponding to the sample images; and sample guidance text specifically refers to the text that indicates the location of abnormal objects and the status of the object to be detected in the sample images. Sample category information can be obtained from the sample guidance text or can be separate sample category information.

[0121] Taking the training of auxiliary diagnostic image processing models in the field of computer-aided diagnosis as an example, during the training process, images of healthy people are introduced instead of only using images of tumor patients, thereby avoiding the image processing model from predicting incorrect results when faced with diverse tumor-free images during actual application.

[0122] After obtaining the sample images, they are fed into the image processing model (at this point, the model is still untrained). The image processing model generates prediction detection results based on each sample image. The prediction detection results include predicted label information, predicted category information, and predicted guidance text.

[0123] In the method provided in the present disclosure, the model structure of the image processing model, such as the image processing model in the above steps, also includes a multi-scale feature extraction module, a feature fusion module, and a feature processing module. The data processing process of the sample image in the untrained image processing model is the same as the data processing process of the image processing model in the above embodiment. Regarding the data processing process of the sample image in the untrained image processing model, please refer to the data processing process of the target image in the image processing model above, which will not be repeated here.

[0124] After obtaining the predicted detection results of the sample image, the model loss value can be calculated based on the predicted detection results and the sample detection results. In the method provided in the present disclosure, there are many methods for calculating the model loss value, such as cross entropy loss function, maximum loss function, average loss function, etc. In the present disclosure, the specific method of the loss function is not limited, and it is subject to actual application.

[0125] In one or more specific embodiments of the present disclosure, the sample detection result includes sample labeling information and sample category information, and the predicted detection result includes predicted labeling information and predicted category information. Calculating the model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information includes:

[0126] Calculating a first loss value according to the sample labeling information and the predicted labeling information;

[0127] Calculating a second loss value according to the sample category information and the predicted category information;

[0128] A model loss value is calculated based on the first loss value and the second loss value.

[0129] Specifically, in one or more specific embodiments of the present disclosure, sample detection results include sample labeling information and sample category information, and predicted detection results include predicted labeling information and predicted category information. Technicians hope that the predicted detection results are consistent with the sample detection results, thereby improving the accuracy of image processing model predictions.

[0130] The first loss value is calculated using sample labeling information and predicted labeling information, and the second loss value is calculated using sample category information and predicted category information. At the same time, the predicted guidance text and predicted labeling information can also be verified with each other, further improving the accuracy of model prediction.

[0131] The first loss value and the second loss value are then fused to obtain a model loss value. Specifically, the first loss value and the second loss value are added to obtain the model loss value.

[0132] In addition, it should be noted that in the training method of the image processing model provided in this disclosure, the sample guidance text is processed within the image processing model to obtain the text loss value. That is, the sample guidance text needs to be input into the image processing model. The image processing model includes a multi-scale feature extraction module, a feature fusion module, a feature processing module, and a text feature extraction module;

[0133] Inputting the sample image and the sample guidance text into an image processing model, and obtaining predicted labeling information, predicted category information, and text loss value output by the image processing model, including:

[0134] Inputting the sample image into the multi-scale feature extraction module to obtain at least one scale feature information;

[0135] Inputting the feature information of each scale into the feature fusion module to obtain feature fusion information;

[0136] Inputting the sample guidance text into the text feature extraction module to obtain sample text feature information;

[0137] The feature fusion information and the sample text feature information are input into the feature processing module to obtain the predicted labeling information, predicted category information and text loss value corresponding to the sample image.

[0138] The processing methods of the multi-scale feature extraction module and feature fusion module in the image processing model can be found in the above steps and will not be repeated here.

[0139] During the training process of the image processing model, the model also includes a text feature extraction module. During the training phase, the sample guidance text is input into the text feature extraction module to obtain sample text feature information. Simultaneously, during the training phase, the feature fusion information and sample text feature information are input into the feature processing module. After processing in the feature processing module, predicted label information, predicted category information, and text loss value are obtained.

[0140] After obtaining the text loss value and the model loss value, the model parameters of the image processing model are adjusted according to the text loss value and the model loss value. Specifically, the text loss value and the model loss value are added to obtain a new loss value, and the model parameters of the image processing model are adjusted by backpropagation according to the new loss value.

[0141] In a specific embodiment provided by the present disclosure, the feature processing module includes an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit;

[0142] Inputting the feature fusion information and the sample text feature information into the feature processing module to obtain predicted labeling information, predicted category information, and text loss value corresponding to the sample image, including:

[0143] Inputting the feature fusion information into the abnormal object segmentation unit to obtain detection labeling information;

[0144] Inputting the feature fusion information into the abnormal object classification unit to obtain detection category information;

[0145] Inputting the feature fusion information into the guidance text generation unit to obtain predicted text feature information;

[0146] A text loss value is obtained by calculation based on the predicted text feature information and the sample text feature information.

[0147] Specifically, the feature processing module in the training phase includes an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit. The abnormal object segmentation unit is used to generate detection annotation information based on feature fusion information; the abnormal object classification unit is used to generate detection category information based on feature fusion information; and the guidance text generation unit is used to generate predicted text feature information based on feature fusion information, specifically generating a predicted text feature vector, and then calculating the text loss value using the predicted text feature information and the sample text feature information.

[0148] Furthermore, the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit, and the sample text feature information includes sample position feature information and sample result feature information;

[0149] Inputting the feature fusion information into the guidance text generation unit to obtain predicted text feature information includes:

[0150] Inputting the feature fusion information into the abnormal position text subunit to obtain predicted position guidance information features;

[0151] Inputting the feature fusion information into the abnormal result text subunit to obtain prediction result guidance information features;

[0152] Accordingly, calculating and obtaining a text loss value according to the predicted text feature information and the sample text feature information includes:

[0153] The text loss value is calculated based on the sample position feature information, the sample result feature information, the predicted position guidance information feature and the predicted result guidance information feature.

[0154] In practical applications, the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit. In the process of generating predicted guidance text based on feature fusion information, the guidance text generation unit first generates predicted text feature information. Furthermore, the predicted text feature information is composed of predicted position guidance information features and predicted result guidance information features. The sample text feature information includes sample position feature information and sample result feature information. A text loss value is calculated based on the sample position feature information, sample result feature information, predicted position guidance information features, and predicted result guidance information features.

[0155] In practical applications, the sample guidance text specifically includes two parts: sample location guidance information and sample result guidance information. Sample location guidance information specifically refers to the location information of abnormal objects within the target detection area, while sample result guidance information specifically refers to the status information of the object to be detected. For example, sample location guidance information may include "a tumor is found in the middle of the esophagus," while sample result guidance information may include "the patient suffers from esophageal cancer." Sample location feature information specifically refers to the feature vector corresponding to the sample location guidance information, while sample result feature information specifically refers to the feature vector corresponding to the sample result guidance information.

[0156] The sample guidance text is input into a text feature extraction module. This text feature extraction module specifically refers to a module that can convert the sample guidance text into corresponding text feature information, such as the encoder of the Transformer model, the text encoder of the CLIP model, etc. Furthermore, the sample position guidance information and sample result guidance information in the sample guidance text are input into the feature extraction module to obtain the sample position guidance information features corresponding to the sample position guidance information and the sample result guidance information features corresponding to the sample result guidance information.

[0157] The image processing model includes a guidance text generation unit, which includes an abnormal position text sub-unit and an abnormal result text sub-unit. During the processing process, the image processing model will obtain the predicted position guidance information features output by the abnormal position text sub-unit and the predicted result guidance information features output by the abnormal result text sub-unit.

[0158] The position loss value is calculated based on the sample position guidance information feature and the predicted position guidance information feature, the result loss value is calculated based on the sample result guidance information feature and the predicted result guidance information feature, and then the text loss value is determined based on the position loss value and the result loss value.

[0159] The sample guidance text is processed by a text feature extraction module. The text feature extraction module can use the sample position guidance information and sample result guidance information of the object to be detected in the sample guidance text to supervise the image processing model, making full use of the existing sample guidance text and improving the training speed and prediction accuracy of the image processing model.

[0160] In actual applications, there is also the problem of incomplete sample image labeling information. In actual applications, only some images are set with annotation information and sample guidance text, while most images only have sample guidance text and no annotation information. For example, for CT images, only some CT images have tumor areas annotated by experienced doctors, while the majority of CT images do not have tumor areas annotated. In actual applications, the proportion of CT images with tumor areas annotated may be only 20%-30%, while the proportion of CT images without tumor areas annotated is 70%-80%. The sample guidance text can be understood as the clinical report corresponding to the CT image.

[0161] In order to make full use of the sample text, in another specific embodiment provided by the present disclosure, the sample image includes a first sample image and a second sample image, wherein the first sample image is marked with sample annotation information, and the second sample image is not marked with sample annotation information;

[0162] Obtaining a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image, including:

[0163] Obtaining a sample image and a sample guidance text corresponding to the sample image, and extracting sample category information from the sample guidance text;

[0164] Training an image annotation model according to a first sample image, sample annotation information corresponding to the first sample image, and sample guidance text to obtain an image annotation model for generating annotation information;

[0165] The second sample is input into the image annotation model to obtain predicted annotation information output by the image annotation model, and the predicted annotation information is used as sample annotation information of the second sample image.

[0166] The first sample image is a sample image with sample annotation information, and the second sample image is a sample image without sample annotation information. Both the first sample image and the second sample image have corresponding sample guidance texts.

[0167] The sample category information corresponding to the object to be detected can be identified and extracted in each sample guidance text. After obtaining the sample category information corresponding to each sample image, an image annotation model can be trained using the first sample image with sample annotation information to generate sample annotation information for the second sample image.

[0168] Specifically, since each sample image includes corresponding sample guidance text, the first sample image has sample annotation information, while the second sample image does not. An image annotation model is pre-trained using the first sample image and the sample annotation information and sample guidance text corresponding to the first sample image. The image annotation model includes an abnormal object segmentation unit and a guidance text generation unit.

[0169] The first sample image is input into the image annotation model to obtain predicted annotation information output by the abnormal object segmentation unit and predicted guidance text output by the guidance text generation unit. The predicted annotation information, sample annotation information, predicted guidance text, and sample guidance text are used to calculate a model loss value for the image annotation model. The model parameters of the image annotation model are adjusted based on the model loss value until a model training stopping condition is met, thereby obtaining a trained image annotation model that can annotate images that are not labeled with image annotation information.

[0170] The second sample image is input into the trained image annotation model. The image annotation model can generate predicted annotation information for the second sample image, and use the predicted annotation information as the sample annotation information corresponding to the second sample image.

[0171] During the training of the image annotation model, the image annotation model includes an abnormal object segmentation unit and a guidance text generation unit. The image processing model in the above steps also includes an abnormal object segmentation unit and a guidance text generation unit. In order to further save computing resources, the method further includes:

[0172] The abnormal object segmentation unit and the guidance text generation unit in the image annotation model are used as the abnormal object segmentation unit and the guidance text generation unit of the image processing model.

[0173] The image annotation model is used as a teacher model. During the training process of the image annotation model, the first sample image, the sample annotation information corresponding to the first sample image, and the sample guidance text are utilized. The abnormal object segmentation unit and guidance text generation unit in the obtained image annotation model have been trained and parameterized, and have corresponding data processing capabilities. The abnormal object segmentation unit and guidance text generation unit can be directly migrated to the image processing model, and the abnormal object segmentation unit and guidance text generation unit in the image annotation model can be used as the abnormal object segmentation unit and guidance text generation unit of the image processing model.

[0174] Furthermore, in the process of training the image annotation model, the image annotation model can be set to also include a multi-scale feature extraction module and a feature fusion module, and continue to be trained during the training process of the image annotation model. Then it is migrated to the image processing model, so that the multi-scale feature extraction module, feature fusion module, abnormal object segmentation unit and guidance text generation unit in the feature processing module in the image processing model are preliminarily trained during the training process of the image annotation model. When the image processing model is trained, it is trained again, or no longer trained, thereby improving the model training efficiency of the image processing model. In actual applications, the image processing model includes a text feature extraction module in the model training stage, and may not include the text feature extraction module in the model application stage of the image processing model.

[0175] The training method of the image processing model provided by the embodiment of the present disclosure utilizes a text feature extraction module to obtain corresponding sample text feature information from the sample guidance text through the text feature extraction module, and uses the sample text feature information to supervise the training of the image processing model, thereby improving the efficiency and accuracy of the image processing model training.

[0176] In addition, to address the situation where the labeling information of sample images is incomplete, a semi-supervised image labeling model training method is adopted. An image labeling model is trained using some of the labeled sample images, and the image labeling model is used to label the unlabeled sample images, thereby enriching the number of sample images with labeled information.

[0177] Finally, the model structure in the image annotation model is migrated to the image processing model, and the image processing model is further trained using the model structure in the trained image annotation model, which improves the training speed of the image processing model and avoids the waste of computing resources.

[0178] 4 , which shows a flow chart of a CT image processing method provided by an embodiment of the present disclosure, specifically comprising the following steps:

[0179] Step 402: Receive a CT image processing task, wherein the CT image processing task carries multiple CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area.

[0180] The abnormal object may be a tumor. The CT image processing method provided in this embodiment is a computer-aided diagnosis method for cancer, which is used to detect whether a tumor exists in a CT image corresponding to a target detection area.

[0181] Step 404: Input the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information and detection guidance text.

[0182] It should be noted that the implementation of steps 402 to 404 is the same as the implementation of steps 202 to 204 described above, and will not be described in detail in this embodiment of the present disclosure.

[0183] For example, taking the target detection area as the stomach, a CT image processing task is received. The CT image processing task includes multiple CT images corresponding to the target user's stomach. The multiple CT images can form a 3D image of the target user's stomach. The CT image processing task is used to detect whether there is a malignant tumor in the target user's stomach.

[0184] In applying the method of the embodiment of the present disclosure, the CT image processing model is the image processing model in the above embodiment, and the model structure of the CT image processing model is the same as the structure of the image processing model in the above embodiment, which will not be repeated here. By inputting multiple CT images corresponding to the stomach into the CT image processing model, the detection results corresponding to the stomach output by the image processing model can be obtained, thereby realizing automatic detection of whether there is a malignant tumor in the stomach. Specifically, the CT image processing model generates multiple scale feature information based on the multiple CT images, and then generates tumor annotation information corresponding to the stomach, the user's disease status (suffering from gastric cancer or not), user detection text (tumor location in the user's stomach, whether the user is sick, etc.) and other information based on the multiple scale feature information.

[0185] The CT image processing method provided by one or more embodiments of the present disclosure extracts multiple scale features from 3D images corresponding to multiple CT images within an image processing model. After fusing these multiple scale features, the method then generates annotation information, category information, and guidance text, ultimately generating a detection result. This multiple scale feature information improves the accuracy of subsequent detection results. The generated detection results include location information of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.

[0186] 5 , which shows a flow chart of a method for training an image processing model according to an embodiment of the present disclosure, which is applied to a cloud-side device and specifically includes the following steps:

[0187] Step 502: Obtain a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image.

[0188] Step 504: Input the sample image and the sample guidance text into the image processing model to obtain the predicted labeling information, predicted category information and text loss value output by the image processing model.

[0189] Step 506: Calculate a model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information.

[0190] Step 508: Adjust the model parameters of the image processing model according to the model loss value and the text loss value, and continue to train the image processing model until the model training stop condition is reached, thereby obtaining the model parameters of the image processing model.

[0191] Step 510: Send the model parameters of the image processing model to the terminal device.

[0192] It should be noted that steps 502 to 508 are implemented in the same manner as the training method of the above-mentioned image processing model, and will not be described in detail in this embodiment of the present disclosure.

[0193] In practical applications, model training requires large amounts of data and significant computing resources, which edge devices may not have the necessary processing capabilities. Therefore, model training can be performed on cloud-side devices. After obtaining the model parameters of the image processing model, the cloud-side device can also send the model parameters to the edge device. The edge device can then locally construct an image processing model based on the model parameters and further use the image processing model to perform image processing.

[0194] The method provided by the embodiment of the present disclosure utilizes a text feature extraction module during the training of an image processing model, obtains corresponding sample text feature information from the sample guidance text through the text feature extraction module, and uses the sample text feature information to supervise the training of the image processing model, thereby improving the efficiency and accuracy of the image processing model training.

[0195] In addition, to address the situation where the labeling information of sample images is incomplete, a semi-supervised image labeling model training method is adopted. An image labeling model is trained using some of the labeled sample images, and the image labeling model is used to label the unlabeled sample images, thereby enriching the number of sample images with labeled information.

[0196] Finally, the model structure in the image annotation model is migrated to the image processing model, and the image processing model is further trained using the model structure in the trained image annotation model, which improves the training speed of the image processing model and avoids the waste of computing resources.

[0197] 6 , which shows a flow chart of an image processing method provided by an embodiment of the present disclosure, specifically comprising the following steps:

[0198] Step 602: Receive an image processing task sent by a user, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.

[0199] Step 604: Input the multiple target images into the image processing model to obtain the detection results corresponding to the target detection area, wherein the image processing model generates the detection results corresponding to the target detection area based on the multi-scale feature information corresponding to the multiple target images, and the detection results include detection annotation information, detection category information and detection guidance text.

[0200] Step 606: Send the detection result corresponding to the target detection area to the user.

[0201] It should be noted that the specific implementation of steps 602 to 604 is the same as the implementation of steps 202 to 204 described above, and will not be described in detail in the embodiment of the present disclosure.

[0202] In this embodiment, an image processing request sent by a user is received, and the image processing request includes an image processing task. After the image processing method of the above embodiment is completed and the detection result is obtained, the detection result needs to be returned to the user so that the user can perform corresponding subsequent processing based on the detection result.

[0203] The image processing method provided by one or more embodiments of the present disclosure uses an image processing model to obtain multiple scale feature information corresponding to a target image, thereby improving the accuracy of subsequent detection results. The generated detection results include the location information of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results, providing users with multi-dimensional detection information, and enhancing the user experience.

[0204] One embodiment of the present disclosure also provides a computer-aided diagnosis method, comprising: receiving a CT image processing task sent by a user, wherein the CT image processing task carries multiple CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object within the target detection area; inputting the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information, and detection guidance text; and sending the detection result corresponding to the target detection area to the user. Specifically, the specific implementation of this method is the same as the implementation of steps 202-204 above, and will not be repeated in this embodiment of the present disclosure.

[0205] The following further illustrates the image processing method provided by the present disclosure using the application of the image processing method in the esophageal cancer detection scenario as an example, in conjunction with FIG7 . FIG7 shows a flowchart of an image processing method provided by one embodiment of the present disclosure for the esophageal cancer detection scenario, specifically comprising the following steps:

[0206] Step 702: Receive a CT image processing task, wherein the CT image processing task carries multiple CT images corresponding to the esophagus, and the CT image processing task is used to detect whether there is a tumor in the esophagus.

[0207] Step 704: Input the multiple CT images into a CT image processing model to obtain a detection result corresponding to the esophagus, wherein the CT image processing model generates a detection result corresponding to the esophagus based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information and detection guidance text.

[0208] In this embodiment, whether there is a tumor in the esophagus is used as an example for explanation, and the CT image processing model is trained in advance.

[0209] Specifically, the collected dataset includes esophageal cancer screening data from 1,617 patients, including their corresponding CT images and test reports. Of these, 946 patients had esophageal cancer, while 671 did not. Experienced physicians were invited to annotate the tumors in the CT images of the 30% of patients diagnosed with esophageal cancer and make final decisions for each patient based on the test reports.

[0210] First, based on the 30% of CT images marked with tumor locations and their corresponding detection reports, an image annotation model is trained. The image annotation model includes a multi-scale feature extraction module, a feature fusion module, and a first feature processing module, among which the first feature processing module includes an abnormal object segmentation unit and a guidance text generation unit.

[0211] After obtaining the image annotation model, the image annotation model is used to annotate the remaining CT images. The annotation information corresponding to the CT images of users who do not suffer from esophageal cancer is "empty".

[0212] After the annotation is completed, the multi-scale feature extraction module, feature fusion module, and first feature processing module in the image annotation model are extracted, and a new abnormal object classification unit is introduced. The abnormal object classification unit, the abnormal object segmentation unit, and the guidance text generation unit are combined to form a second feature processing module, and the CT image processing model is constructed by the multi-scale feature extraction module, the feature fusion module, and the second feature extraction module.

[0213] The patient status corresponding to each patient (i.e., whether the patient has esophageal cancer) is obtained from the test report as sample category information, the image annotation information is used as sample annotation information, and the location guidance information and result guidance information are extracted from the test report as sample guidance text.

[0214] The sample category information, sample annotation information, sample guidance text and CT images are used as training samples to train the CT image processing model until the model training stopping condition of the CT image processing model is reached.

[0215] The trained CT image processing model can be used to detect esophageal cancer. By inputting the CT image of a new patient into the CT image processing model for prediction, the corresponding detection annotation information, detection category information and detection guidance text of the patient can be obtained.

[0216] Corresponding to the above method embodiment, the present disclosure also provides an image processing device embodiment. FIG8 shows a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure. As shown in FIG8 , the device includes:

[0217] A receiving module 802 is configured to receive an image processing task, wherein the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;

[0218] The detection module 804 is configured to input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection area, wherein the image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection result includes detection annotation information, detection category information and detection guidance text.

[0219] Optionally, the image processing model includes a multi-scale feature extraction module, a feature fusion module, and a feature processing module;

[0220] Accordingly, the detection module 804 is further configured to:

[0221] Inputting the plurality of target images into the multi-scale feature extraction module to obtain at least one scale feature information;

[0222] Inputting the feature information of each scale into the feature fusion module to obtain feature fusion information;

[0223] The feature fusion information is input into the feature processing module to obtain the detection result corresponding to the target detection area.

[0224] Optionally, the feature processing module includes an abnormal object segmentation unit, an abnormal object classification unit and a guidance text generation unit;

[0225] Accordingly, the detection module 804 is further configured to:

[0226] Inputting the feature fusion information into the abnormal object segmentation unit to obtain detection labeling information;

[0227] Inputting the feature fusion information into the abnormal object classification unit to obtain detection category information;

[0228] Inputting the feature fusion information into the guidance text generation unit to obtain a detection guidance text;

[0229] A detection result corresponding to the target detection area is generated according to the detection label information, detection category information, and detection guidance text.

[0230] Optionally, the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit;

[0231] Accordingly, the detection module 804 is further configured to:

[0232] Inputting the feature fusion information into the abnormal position text subunit to obtain position guidance information;

[0233] Inputting the feature fusion information into the abnormal result text subunit to obtain result guidance information;

[0234] Generate a detection guidance text according to the position guidance information and the result guidance information.

[0235] Optionally, the device further includes a segmentation module configured to:

[0236] receiving an image segmentation task, wherein the image segmentation task carries a plurality of initial images corresponding to a target detection area, and the image segmentation task is used to extract a target image corresponding to the target detection area;

[0237] Each initial image is input into a pre-trained image segmentation model to obtain a target image corresponding to each initial image output by the image segmentation model.

[0238] Optionally, the device further includes a training module configured to:

[0239] Obtaining a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image;

[0240] Inputting the sample image and the sample guidance text into an image processing model to obtain predicted labeling information, predicted category information, and text loss value output by the image processing model;

[0241] Calculating a model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information;

[0242] Adjust the model parameters of the image processing model according to the model loss value and the text loss value, and continue to train the image processing model until a model training stop condition is reached.

[0243] Optionally, the training module is further configured to:

[0244] Calculating a first loss value according to the sample labeling information and the predicted labeling information;

[0245] Calculating a second loss value according to the sample category information and the predicted category information;

[0246] A model loss value is calculated based on the first loss value and the second loss value.

[0247] Optionally, the image processing model includes a multi-scale feature extraction module, a feature fusion module, a feature processing module, and a text feature extraction module;

[0248] The training module is further configured to:

[0249] Inputting the sample image and the sample guidance text into an image processing model, and obtaining predicted labeling information, predicted category information, and text loss value output by the image processing model, including:

[0250] Inputting the sample image into the multi-scale feature extraction module to obtain at least one scale feature information;

[0251] Inputting the feature information of each scale into the feature fusion module to obtain feature fusion information;

[0252] Inputting the sample guidance text into the text feature extraction module to obtain sample text feature information;

[0253] The feature fusion information and the sample text feature information are input into the feature processing module to obtain the predicted labeling information, predicted category information and text loss value corresponding to the sample image.

[0254] Optionally, the feature processing module includes an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit;

[0255] The training module is further configured to:

[0256] Inputting the feature fusion information into the abnormal object segmentation unit to obtain detection labeling information;

[0257] Inputting the feature fusion information into the abnormal object classification unit to obtain detection category information;

[0258] Inputting the feature fusion information into the guidance text generation unit to obtain predicted text feature information;

[0259] A text loss value is obtained by calculation based on the predicted text feature information and the sample text feature information.

[0260] Optionally, the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit, and the sample text feature information includes sample position feature information and sample result feature information;

[0261] The training module is further configured to:

[0262] Inputting the feature fusion information into the abnormal position text subunit to obtain predicted position guidance information features;

[0263] Inputting the feature fusion information into the abnormal result text subunit to obtain prediction result guidance information features;

[0264] The text loss value is calculated based on the sample position feature information, the sample result feature information, the predicted position guidance information feature and the predicted result guidance information feature.

[0265] Optionally, the sample image includes a first sample image and a second sample image, wherein the first sample image is marked with sample annotation information, and the second sample image is not marked with sample annotation information;

[0266] The training module is further configured to:

[0267] Obtaining a sample image and a sample guidance text corresponding to the sample image, and extracting sample category information from the sample guidance text;

[0268] Training an image annotation model according to a first sample image, sample annotation information corresponding to the first sample image, and sample guidance text to obtain an image annotation model for generating annotation information;

[0269] The second sample is input into the image annotation model to obtain predicted annotation information output by the image annotation model, and the predicted annotation information is used as sample annotation information of the second sample image.

[0270] Optionally, the image annotation model includes an abnormal object segmentation unit and a guidance text generation unit;

[0271] The training module is further configured to:

[0272] The abnormal object segmentation unit and the guidance text generation unit in the image annotation model are used as the abnormal object segmentation unit and the guidance text generation unit of the image processing model.

[0273] The image processing device provided by the embodiment of the present disclosure inputs multiple target images of the target detection area of ​​the object to be detected into the image processing model. In the image processing model, multiple scale features are extracted from the 3D images corresponding to the multiple target images. After the multiple scale features are fused, annotation information, category information, and guidance text are generated respectively, thereby finally generating a detection result. The accuracy of the subsequent generation of detection results is improved by the multiple scale feature information. The generated detection results include the location information of the abnormal object, the information of the object to be detected, and the guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.

[0274] In the process of training the image processing model, the text feature extraction module is used to obtain the corresponding sample text feature information of the sample guidance text through the text feature extraction module, and the sample text feature information is used to supervise the training of the image processing model, thereby improving the efficiency and accuracy of the image processing model training.

[0275] In addition, to address the situation where the labeling information of sample images is incomplete, a semi-supervised image labeling model training method is adopted. An image labeling model is trained using some of the labeled sample images, and the image labeling model is used to label the unlabeled sample images, thereby enriching the number of sample images with labeled information.

[0276] Finally, the model structure in the image annotation model is migrated to the image processing model, and the image processing model is further trained using the model structure in the trained image annotation model, which improves the training speed of the image processing model and avoids the waste of computing resources.

[0277] The above is a schematic diagram of an image processing device according to this embodiment. It should be noted that the technical solution of the image processing device and the technical solution of the above-mentioned image processing method are based on the same concept. For details not described in detail in the technical solution of the image processing device, please refer to the description of the technical solution of the above-mentioned image processing method.

[0278] Corresponding to the above method embodiment, the present disclosure also provides an embodiment of a CT image processing device. FIG9 shows a schematic structural diagram of a CT image processing device provided by an embodiment of the present disclosure. As shown in FIG9 , the device includes:

[0279] A receiving module 902 is configured to receive a CT image processing task, wherein the CT image processing task carries a plurality of CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area;

[0280] The detection module 904 is configured to input the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information and detection guidance text.

[0281] The CT image processing device provided by one or more embodiments of the present disclosure extracts multiple scale features from 3D images corresponding to multiple CT images within an image processing model. After fusing these multiple scale features, the device generates annotation information, category information, and guidance text, ultimately generating a detection result. The use of multiple scale feature information improves the accuracy of subsequent detection results. The generated detection results include the location information of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.

[0282] The above is a schematic diagram of a CT image processing device according to this embodiment. It should be noted that the technical solution of the CT image processing device and the technical solution of the aforementioned CT image processing method are based on the same concept. For details not described in detail in the technical solution of the CT image processing device, please refer to the description of the technical solution of the aforementioned CT image processing method.

[0283] Corresponding to the above method embodiment, the present disclosure also provides an image processing device embodiment. FIG10 shows a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure. As shown in FIG10 , the device includes:

[0284] The receiving module 1002 is configured to receive an image processing task sent by a user, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.

[0285] The detection module 1004 is configured to input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection area, wherein the image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection result includes detection annotation information, detection category information and detection guidance text.

[0286] The sending module 1006 is configured to send the detection result corresponding to the target detection area to the user.

[0287] The image processing device provided by one or more embodiments of the present disclosure extracts multiple scale features from 3D images corresponding to multiple images within an image processing model. After fusing these multiple scale features, the device then generates annotation information, category information, and guidance text, ultimately generating a detection result. This multiple scale feature information improves the accuracy of the subsequent detection results. The generated detection results include the location information of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.

[0288] The above is a schematic diagram of an image processing device according to this embodiment. It should be noted that the technical solution of the image processing device and the technical solution of the above-mentioned image processing method are based on the same concept. For details not described in detail in the technical solution of the image processing device, please refer to the description of the technical solution of the above-mentioned image processing method.

[0289] Figure 11 shows a block diagram of a computing device 1100 according to one embodiment of the present disclosure. Components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0290] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0291] In one embodiment of the present disclosure, the aforementioned components of the computing device 1100 and other components not shown in FIG11 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG11 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.

[0292] Computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1100 may also be a mobile or stationary server.

[0293] Among them, the processor 1120 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.

[0294] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method.

[0295] An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.

[0296] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method.

[0297] An embodiment of the present disclosure further provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.

[0298] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product is based on the same concept as the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method.

[0299] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0300] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.

[0301] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.

[0302] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0303] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. An image processing method, comprising: receiving an image processing task, wherein the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area; The multiple target images are input into an image processing model to obtain detection results corresponding to the target detection area, wherein the image processing model generates detection results corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection results include detection annotation information, detection category information and detection guidance text.

2. The method according to claim 1, wherein the image processing model comprises a multi-scale feature extraction module, a feature fusion module, and a feature processing module; Inputting the plurality of target images into an image processing model to obtain detection results corresponding to the target detection areas includes: Inputting the plurality of target images into the multi-scale feature extraction module to obtain at least one scale feature information; Inputting the feature information of each scale into the feature fusion module to obtain feature fusion information; The feature fusion information is input into the feature processing module to obtain the detection result corresponding to the target detection area.

3. The method according to claim 2, wherein the feature processing module comprises an abnormal object segmentation unit, an abnormal object classification unit and a guidance text generation unit; Inputting the feature fusion information into the feature processing module to obtain the detection result corresponding to the target detection area includes: Inputting the feature fusion information into the abnormal object segmentation unit to obtain detection labeling information; Inputting the feature fusion information into the abnormal object classification unit to obtain detection category information; Inputting the feature fusion information into the guidance text generation unit to obtain a detection guidance text; A detection result corresponding to the target detection area is generated according to the detection label information, detection category information, and detection guidance text.

4. The method according to claim 3, wherein the guidance text generation unit comprises an abnormal position text subunit and an abnormal result text subunit; Inputting the feature fusion information into the guidance text generation unit to obtain the detection guidance text includes: Inputting the feature fusion information into the abnormal position text subunit to obtain position guidance information; Inputting the feature fusion information into the abnormal result text subunit to obtain result guidance information; Generate a detection guidance text according to the position guidance information and the result guidance information.

5. The method according to any one of claims 1 to 4, further comprising, before receiving the image processing task: receiving an image segmentation task, wherein the image segmentation task carries a plurality of initial images corresponding to a target detection area, and the image segmentation task is used to extract a target image corresponding to the target detection area; Each initial image is input into a pre-trained image segmentation model to obtain a target image corresponding to each initial image output by the image segmentation model.

6. The method according to any one of claims 1 to 5, wherein the image processing model is trained by the following steps: Obtaining a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image; Inputting the sample image and the sample guidance text into an image processing model to obtain predicted labeling information, predicted category information, and text loss value output by the image processing model; Calculating a model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information; Adjust the model parameters of the image processing model according to the model loss value and the text loss value, and continue to train the image processing model until a model training stop condition is reached.

7. The method according to claim 6, wherein the model loss value is calculated based on the sample label information, the sample category information, the predicted label information, and the predicted category information, comprising: Calculating a first loss value according to the sample labeling information and the predicted labeling information; Calculating a second loss value according to the sample category information and the predicted category information; A model loss value is calculated based on the first loss value and the second loss value.

8. The method according to claim 6, wherein the image processing model comprises a multi-scale feature extraction module, a feature fusion module, a feature processing module, and a text feature extraction module; Inputting the sample image and the sample guidance text into an image processing model, and obtaining predicted labeling information, predicted category information, and text loss value output by the image processing model, including: Inputting the sample image into the multi-scale feature extraction module to obtain at least one scale feature information; Inputting the feature information of each scale into the feature fusion module to obtain feature fusion information; Inputting the sample guidance text into the text feature extraction module to obtain sample text feature information; The feature fusion information and the sample text feature information are input into the feature processing module to obtain the predicted labeling information, predicted category information and text loss value corresponding to the sample image.

9. The method according to claim 8, wherein the feature processing module comprises an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit; Inputting the feature fusion information and the sample text feature information into the feature processing module to obtain predicted labeling information, predicted category information, and text loss value corresponding to the sample image, including: Inputting the feature fusion information into the abnormal object segmentation unit to obtain detection labeling information; Inputting the feature fusion information into the abnormal object classification unit to obtain detection category information; Inputting the feature fusion information into the guidance text generation unit to obtain predicted text feature information; A text loss value is calculated based on the predicted text feature information and the sample text feature information.

10. The method according to claim 9, wherein the guidance text generation unit comprises an abnormal position text subunit and an abnormal result text subunit, and the sample text feature information comprises sample position feature information and sample result feature information; Inputting the feature fusion information into the guidance text generation unit to obtain predicted text feature information includes: Inputting the feature fusion information into the abnormal position text subunit to obtain predicted position guidance information features; Inputting the feature fusion information into the abnormal result text subunit to obtain prediction result guidance information features; Accordingly, calculating and obtaining a text loss value according to the predicted text feature information and the sample text feature information includes: The text loss value is calculated based on the sample position feature information, the sample result feature information, the predicted position guidance information feature and the predicted result guidance information feature.

11. The method according to claim 6, wherein the sample image comprises a first sample image and a second sample image, wherein: The first sample image is marked with sample annotation information, and the second sample image is not marked with sample annotation information; Obtaining a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image, including: Obtaining a sample image and a sample guidance text corresponding to the sample image, and extracting sample category information from the sample guidance text; Training an image annotation model according to a first sample image, sample annotation information corresponding to the first sample image, and sample guidance text to obtain an image annotation model for generating annotation information; The second sample is input into the image annotation model to obtain predicted annotation information output by the image annotation model, and the predicted annotation information is used as sample annotation information of the second sample image.

12. The method according to claim 11, wherein the image annotation model includes an abnormal object segmentation unit and a guidance text generation unit, and the method further comprises: The abnormal object segmentation unit and the guidance text generation unit in the image annotation model are used as the abnormal object segmentation unit and the guidance text generation unit of the image processing model.

13. A computer-aided diagnosis method for cancer, comprising: receiving a computed tomography (CT) image processing task, wherein the CT image processing task carries a plurality of CT images corresponding to a target detection area, and the CT image processing task is used to detect whether a tumor exists in the target detection area; The multiple CT images are input into a CT image processing model to obtain detection results corresponding to the target detection area, wherein the CT image processing model generates detection results corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection results include detection annotation information, detection category information and detection guidance text.

14. A training method for an image processing model, applied to a cloud-side device, comprising: Obtaining a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image; Inputting the sample image and the sample guidance text into an image processing model to obtain predicted labeling information, predicted category information, and text loss value output by the image processing model; Calculating a model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information; Adjusting the model parameters of the image processing model according to the model loss value and the text loss value, and continuing to train the image processing model until a model training stop condition is reached, thereby obtaining the model parameters of the image processing model; The model parameters of the image processing model are sent to the end-side device.

15. A computer-aided diagnosis method comprising: receiving a CT image processing task sent by a user, wherein the CT image processing task carries multiple CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area; Inputting the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information, and detection guidance text; Send the detection result corresponding to the target detection area to the user.

16. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 15 are implemented.

17. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 15.

18. A computer program product comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Image segmentation processing method and device, storage medium and electronic equipment

    CN114972368A

  • Medical image abnormal signal intensity detection method

    CN115115576A

  • Image processing method, computer readable storage medium and computer equipment

    CN116206331A

  • Image processing method and training method of image processing model

    CN116993663A

  • Image processing method and training method of image processing model

    CN117853490A